Service health degree assessment method, electronic equipment and readable storage medium
By automatically identifying the causal relationship of status monitoring data in a distributed system and evaluating the health of the distributed system, the problems of weak generalization ability and great influence of human factors in the existing technology are solved, and a more accurate and reliable health assessment is achieved.
Patent Information
- Application Number
- CN202410123142.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-29
- Publication Date
- 2025-07-29
AI Technical Summary
The existing service health measurement methods based on set thresholds have weak generalization capabilities in distributed systems, making it difficult to effectively monitor and evaluate service health, and are greatly affected by human factors.
By obtaining the status monitoring data of the distributed system, using the service health assessment model to determine the target operating status information and its causal relationship, automatically identify the associated operating status information, and evaluate the health of the system based on the associated data, without artificially setting evaluation rules and thresholds.
It improves the accuracy and reliability of service health assessment, reduces the influence of human factors, has good generalization ability, and can fully discover problems and alert root causes.
Smart Images

Figure CN120386680A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of system service health assessment, and in particular, to a service health assessment method, an electronic device, and a readable storage medium. Background Art
[0002] As a measurement model for service health, the service health assessment model is widely used in systems based on distributed architectures such as distributed software systems. At present, most service health assessment methods score the alarms generated by services according to pre-set assessment rules, and determine the service health of the distributed system by comparing the scoring results with the set thresholds. The selection of the assessment rules and thresholds for alarms usually depends on empirical values. With the continuous growth of the scale and complexity of distributed software systems, the generalization ability of related solutions is weak, and it is difficult to effectively monitor and evaluate the service health of distributed systems. Summary of the Invention
[0003] The embodiments of the present application provide a service health assessment method, an electronic device, and a readable storage medium, which are used to at least solve the problem that the existing service health measurement method based on set thresholds has weak generalization ability and is difficult to effectively monitor and evaluate the service health of distributed systems.
[0004] In a first aspect, the embodiments of the present application provide a service health assessment method, including:
[0005] In response to the collected status monitoring data of the distributed system, obtain the target operating status information that matches the status monitoring data among the multiple operating status information of the service health assessment model;
[0006] Based on the causal relationship between the respective operating status information among the multiple operating status information, determine the associated operating status information that has a direct causal relationship or an indirect causal relationship with the target operating status information;
[0007] Determine the associated data in the status monitoring data that matches the associated operating status information, and determine the service health of the distributed system based on the associated data.
[0008] In a second aspect, the embodiments of the present application provide an electronic device, including a processor and a memory, where the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the method described in the first aspect above are implemented.
[0009] In a third aspect, the embodiments of the present application provide a readable storage medium, where a program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect above are implemented.
[0010] In an embodiment of the present application, in response to the state monitoring data of the distributed system collected, the target operating state information that matches the state monitoring data is obtained from multiple operating state information of the service health assessment model; based on the causal relationship between each operating state information among the multiple operating state information, the associated operating state information that has a direct causal relationship or an indirect causal relationship with the target operating state information is determined; the associated data that matches the associated operating state information is determined from the state monitoring data, and the service health of the distributed system is determined based on the associated data. In this way, when the state monitoring data (such as metrics, alarms, logs, etc.) of the distributed system is collected, according to the causal relationship between each operating state information, the associated operating state information that has a direct causal relationship or an indirect causal relationship with the state monitoring data can be determined, and then the service health of the distributed system is determined according to the associated data that matches the associated operating state information. This method can evaluate the service health of the distributed system without artificially setting evaluation rules and corresponding thresholds for each state monitoring data of the distributed system, has good generalization ability, and can reduce the influence of human factors, improving the accuracy and reliability of service health assessment.
[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Brief Description of the Drawings
[0012] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0013] Figure 1 It shows a schematic flowchart of the service health assessment method provided by the embodiment of the present application;
[0014] Figure 2 It shows an example diagram of the service health assessment model provided by the embodiment of the present application;
[0015] Figure 3 It shows a schematic flowchart of the method for constructing the initial service health assessment model provided by the embodiment of the present application;
[0016] Figure 4a It shows an example diagram corresponding to the initial service health assessment model before update provided by the embodiment of the present application;
[0017] Figure 4b It shows an example diagram corresponding to the initial service health assessment model after update provided by the embodiment of the present application;
[0018] Figure 5The figure shows a schematic flowchart of a method for constructing a service health assessment model provided by an embodiment of the present application;
[0019] Figure 6 The figure shows a schematic structural diagram of a service health assessment system provided by an embodiment of the present application;
[0020] Figure 7 The figure shows a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0021] Here, exemplary embodiments will be described in detail, and examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0022] As a measurement model for service health, the service health assessment model is widely used in systems based on distributed architectures such as distributed software systems. Taking a distributed software system (Distributed Software Systems) as an example, it is a system that executes tasks on a multiprocessor architecture interconnected by a communication network. When the distributed software system is in a production environment, whether it can accurately reflect its healthy or faulty state is an important indicator for measuring whether a service health assessment model is reasonable.
[0023] At present, most service health assessment models are based on the relevant expert experience of distributed software systems. From the perspective of system operation and maintenance, services are divided into "available" and "unavailable" states, and certain weights are assigned to the alarms generated by the services according to the alarm levels. Then, the service health is calculated based on the alarm levels and weights, and the service state is obtained by comparing the service health with a set threshold. For example, the service health of a distributed software system can be evaluated by the following formula:
[0024] H = 100 - ∑L * k1 - ∑O * k2 - ∑P * k3;
[0025] Where H is the service health, L is the alarm level, O is the object level, P is the performance level, and k1, k2, and k3 are the weights corresponding to the alarm level L, object level O, and performance level P, respectively.
[0026] Specifically, the alarm information generated by the distributed software system can be divided into an alarm level L, an object level O, and a performance level P according to three different dimensions. Here, it is assumed that an alarm occurs for a certain metric of the current object in the distributed software system. The alarm level L can be used to measure the severity of the impact of the alarm on the system, the object level O can be used to measure the severity of the impact of the metric on the system, and the performance level P can be used to measure the performance of the current metric. Calculate the service health H according to the above formula, and artificially set a threshold based on experience. When the service health H is greater than the threshold, the distributed software system is available; when the service health H is less than the threshold, the distributed software system is unavailable.
[0027] However, with the continuous growth of the scale and the increase in complexity of the distributed software system, it has become increasingly difficult to effectively monitor and evaluate the service health status of the distributed software system. The above threshold-based service health evaluation method has weak generalization ability and cannot meet the environmental requirements of the current complex distributed software system.
[0028] In view of the problems existing in the above distributed system during the service health evaluation process, the embodiments of the present application provide a service health evaluation method. When the state monitoring data (such as metrics, alarms, logs, etc.) of the distributed system is collected, the method obtains the target operating state information that matches the state monitoring data among the multiple operating state information of the service health evaluation model; based on the causal relationship between each operating state information among the multiple operating state information, determine the associated operating state information that has a direct or indirect causal relationship with the target operating state information; determine the associated data in the state monitoring data that matches the associated operating state information, and determine the service health of the distributed system based on the associated data. In this way, without artificially setting evaluation rules and thresholds for the various state monitoring data of the distributed system, the service health of the distributed system can be evaluated. It has good generalization ability, can reduce the influence of human factors, and improve the accuracy and reliability of service health evaluation.
[0029] Please refer to Figure 1 , Figure 1The figure shows a schematic flowchart of the service health assessment method provided by an embodiment of the present application. The execution subject of this method can be a terminal device or a server. Among them, the terminal device can be a device such as a personal computer, or a mobile terminal device such as a mobile phone or a tablet computer. The terminal device can be the terminal device used by the user. The server can be an independent server or a server cluster composed of multiple servers. Moreover, the server can be the background server of a certain service or the background server of an application (such as a service health assessment system, an availability prediction system, etc.). In the embodiments of the present application, the case where the execution subject is a server is taken as an example for illustration. For the case of the terminal device, it can be processed according to the following relevant content and will not be elaborated here. As shown in the figure, the service health assessment method 100 may include the following steps:
[0030] S101: In response to the collected status monitoring data of the distributed system, obtain the target running status information that matches the status monitoring data among the multiple running status information of the service health assessment model.
[0031] Among them, the distributed system can be a system based on a distributed architecture. The distributed architecture is a design method that distributes different components of the system on multiple independent computer nodes. The distributed system satisfies one or several communication protocols. The most commonly used communication protocol for distributed systems at present is the http / https protocol for sending and requesting network data. Taking the distributed software system as an example, the distributed software system is a software system in which multiple computer nodes cooperate. It usually includes multiple different components, such as servers, databases, load balancers, caches, message queues, etc. These components can be distributed on different nodes to achieve distributed functions, such as data storage, processing, and transmission.
[0032] The status monitoring data can be index type data such as the resource utilization rate, network latency, and packet loss rate of the distributed system, can be log type data such as "the NameNode process heap memory is insufficient, and frequent Full GCs occur, resulting in master-slave switching", or can be alarm type data such as "the NameNode node is unavailable".
[0033] In the embodiments of the present application, when the status monitoring data of the distributed system (such as resource utilization rate, network latency, alarms, logs, etc.) is collected, the target running status information that matches the status monitoring data is searched among the multiple running status information of the pre-constructed service health assessment model. Specifically, the similarity between the status monitoring data and the multiple running status information can be calculated, and the target running status information that matches the status monitoring data among the multiple running status information can be determined according to the similarity.
[0034] S102: Determine, based on the causal relationships among the various operating status information in the multiple operating status information, the associated operating status information that has a direct or indirect causal relationship with the target operating status information.
[0035] Among them, the service health assessment model includes multiple operating status information and the causal relationships among the various operating status information; a direct causal relationship means that there is a direct and clear causal connection between two operating status information. For example, "unreasonable NameNode heap memory configuration" directly causes "NameNode heap memory shortage and frequent GC"; an indirect causal relationship means that there is an indirect connection or influence relationship between two or more operating status information. Continuing with the example of "unreasonable NameNode heap memory configuration", "unreasonable NameNode heap memory configuration" directly causes "NameNode heap memory shortage and frequent GC", "NameNode heap memory shortage and frequent GC" directly causes "NameNode node unavailable", and there is an indirect causal relationship between "unreasonable NameNode heap memory configuration" and "NameNode node unavailable".
[0036] In the embodiment of the present application, according to the target operating status information obtained in step S101 above and the service health assessment model, the associated operating status information that has a direct or indirect causal relationship with the target operating status information can be determined. As Figure 2 shown, when the warning information "NameNode node unavailable" generated by the distributed system is collected, then according to the multiple operating status information in the service health assessment model, the parent node information is searched. If there is a log message "NameNode process heap memory shortage and frequent Full GC, resulting in master-slave switchover", then the configuration information is searched. If it is found that the configuration is unreasonable, then at this time, "unreasonable NameNode heap memory configuration" is the root cause of this warning. At the same time, the child nodes are recursively searched to detect whether the "HDFS service is available". If it is determined that the "HDFS service is unavailable", then the next child node is searched. Finally, the associated operating status information that has a direct or indirect causal relationship with the target operating status information can be obtained.
[0037] S103: Determine the associated data in the status monitoring data that matches the associated operating status information, and determine the service health of the distributed system based on the associated data.
[0038] In the embodiment of the present application, according to the associated operation status information determined in step S102 above, the associated data in the status monitoring data of the distributed system that matches the associated operation status information is determined, and the service health of the distributed system is determined based on the associated data. For example, the rationality of the NameNode heap memory configuration can be determined according to the relevant data of the NameNode heap memory configuration; the availability of the HDFS service can be determined according to the relevant data of the HDFS service, so as to determine the availability of multiple related services based on the associated data, and then obtain the service health of the distributed system.
[0039] Compared with the service health evaluation method based on thresholds, the health evaluation method provided by the embodiment of the present application can evaluate the service health of the distributed system without artificially setting evaluation rules and corresponding thresholds for the status monitoring data of each item of the distributed system. It has good generalization ability, can reduce the influence of human factors, and improve the accuracy and reliability of service health evaluation. At the same time, it can more comprehensively discover problems in the status monitoring data, the causal relationship between problems and alarms, the root cause of alarms, etc., which is convenient for users to take corresponding measures in time to repair the faults of the distributed system.
[0040] In a possible implementation manner, in step S102 above, before obtaining the target operation status information in the multiple operation status information of the service health evaluation model that matches the status monitoring data, it further includes:
[0041] Step 1021: Based on the historical operation data of the distributed system, determine the first operation status information in the historical operation data, and the second operation status information that has a causal relationship with the first operation status information.
[0042] In the embodiment of the present application, the historical operation data of the distributed system within a preset time period can be obtained. For example, the historical operation data of the distributed system within the most recent week can be obtained. Here, the preset time period can be set according to actual needs. Through the historical operation data of the distributed system, the first operation status information in the historical operation data and the second operation status information that has a causal relationship with the first operation status information are determined.
[0043] Step 1022: Predict the third operation status information other than the second operation status information that has a causal relationship with the first operation status information according to the pre-trained target policy network.
[0044] In the embodiments of the present application, after obtaining the first operation state information in the historical operation data, the third operation state information other than the second operation state information that has a causal relationship with the first operation state information is predicted according to the pre-trained target policy network. In this way, the causal relationships between the operation state information in the service health assessment model can be complemented by the target policy network.
[0045] Step 1023: Generate the service health assessment model according to the first operation state information, the second operation state information, the third operation state information, and the causal relationships between the operation state information.
[0046] In the embodiments of the present application, the service health assessment model is generated according to the first operation state information, the second operation state information, the third operation state information, and the causal relationships between the operation state information obtained in the above steps 1021 and 1022. Since in the embodiments of the present application, in addition to determining the second operation state information that has a causal relationship with the first operation state information through the historical operation data of the distributed system, the third operation state information that has a causal relationship with the first operation state information is predicted through the target policy network, the causal relationships between the operation state information in the generated service health assessment model are more comprehensive, which is conducive to improving the accuracy and reliability of the service health assessment.
[0047] In a possible implementation manner, in the above step 1021, based on the historical operation data of the distributed system, determining the first operation state information in the historical operation data and the second operation state information that has a causal relationship with the first operation state information includes:
[0048] Obtain the structured and semi-structured operation state information data from the historical operation data of the distributed system; determine multiple candidate operation state information in the structured and semi-structured operation state information data based on the preset data extraction rules; perform feature association processing on the multiple candidate operation state information based on the feature information of the multiple candidate operation state information to obtain the first operation state information and the second operation state information that has a causal relationship with the first operation state information.
[0049] In the embodiments of the present application, taking a distributed software system as an example, the operation information of the distributed software system is stored in different databases. A data collection layer is constructed to connect to the data sources of the distributed software system. The data sources are such as (PG / MySQL / Elastic Search) or Rest API, etc. The expert experience is organized into semi-structured data, and various semi-structured and structured operation status information data collected by the data collection layer are used as data inputs. Based on preset data extraction rules, multiple candidate operation status information in the structured and semi-structured operation status information data are determined. Among them, the candidate operation status information may include features such as knowledge entities (such as index information like resource utilization rate, network latency, etc., log information, alarm information, etc.), knowledge relationships (such as time relationship, space relationship, causal relationship, etc.), and semantic attributes (such as the type to which the knowledge entity belongs). Based on the feature information of the multiple candidate operation status information, feature association processing is performed on the multiple candidate operation status information to obtain the first operation status information and the second operation status information that has a causal relationship with the first operation status information. Among them, the feature association processing may include entity alignment, entity disambiguation, entity association, etc. processing.
[0050] In a possible implementation manner, in the above step 1023, according to the first operation status information, the second operation status information, the third operation status information, and the causal relationships between the respective operation status information, generating the service health assessment model includes:
[0051] According to the first operation status information, the second operation status information, and the causal relationship between the first operation status information and the second operation status information, generating an initial service health assessment model, and evaluating the initial service health assessment model to obtain a first evaluation index;
[0052] According to the third operation status information and the causal relationship between the first operation status information and the third operation status information, updating the initial service health assessment model, and evaluating the updated initial service health assessment model to obtain a second evaluation index;
[0053] In the case where the deviation amount between the second evaluation index and the first evaluation index is less than a preset threshold, using the updated initial service health assessment model as the service health assessment model; in the case where the deviation amount is not less than the preset threshold, using the initial service health model as the service health assessment model.
[0054] In the embodiments of the present application, such as Figure 3As shown in the figure, the initial service health assessment model can be constructed through the following steps: 301: Organize data. Collect the historical operation data of the distributed system, and obtain the structured and semi-structured operation status information data from the historical operation data of the distributed system; 302: Knowledge extraction. Perform knowledge extraction on the historical operation data to obtain feature information such as knowledge entities, knowledge relationships, and semantic attributes; 303: Knowledge cleaning. Based on the feature information of multiple candidate operation status information, perform feature association processing such as entity alignment, entity disambiguation, and entity association on the candidate operation status information to obtain the first operation status information in the historical operation data, and the second operation status information that has a causal relationship with the first operation status information; According to the first operation status information, the second operation status information, and the causal relationship between the first operation status information and the second operation status information, generate the initial service health assessment model; Furthermore, S304: Evaluate the initial service health assessment model to obtain the first evaluation index. Among them, the first evaluation index may include accuracy and coverage rate.
[0055] Furthermore, according to the third operation status information and the causal relationship between the first operation status information and the third operation status information, update the above initial service health assessment model, and evaluate the updated initial service health assessment model to obtain the second evaluation index. Based on the historical operation data of the distributed system, obtain the Figure 4a initial service health assessment model shown in the figure, including the first operation status information, the second operation status information, and the causal relationship between the first operation status information and the second operation status information. Furthermore, according to the third operation status information and the causal relationship between the first operation status information and the third operation status information, update the initial service health assessment model to obtain the Figure 4b updated initial service health assessment model shown in the figure.
[0056] By comparing the first evaluation index and the second evaluation index, the improvement effect of the updated initial service health model can be measured. When the deviation between the second evaluation index and the first evaluation index is less than the preset threshold, the model tends to be stable and finally converges. At this time, the updated initial service health assessment model can be used as the service health assessment model; when the deviation is not less than the preset threshold, the model is still unstable at this time, and the model needs to be optimized again. At this time, the initial service health model is used as the service health assessment model.
[0057] In a possible implementation manner, in the above step S304, the initial service health assessment model can be evaluated in the following manner:
[0058] Obtain multiple pre-annotated sample data, and process the multiple sample data through the initial service health assessment model to obtain an availability assessment label corresponding to each sample data in the multiple sample data; the sample data carries an availability true label.
[0059] By comparing the availability assessment label corresponding to each sample data with the availability true label, determine the accuracy rate and coverage rate of the initial service health assessment model, and obtain the evaluation index corresponding to the initial service health assessment model according to the accuracy rate and coverage rate.
[0060] In the embodiment of the present application, first construct a test set, which includes multiple pre-annotated sample data, and the sample data carries an availability true label. For example, select 30 faulty operation behaviors (negative examples) and 20 normal operation behaviors (positive examples), and construct log data, metric data, alarm data, etc. and inject them into the distributed software system. Then process the multiple sample data through the initial service health assessment model to obtain an availability assessment label corresponding to each sample data in the multiple sample data; by comparing the availability assessment label corresponding to each sample data with the availability true label, determine the accuracy rate and coverage rate of the initial service health assessment model. For example, if the initial service health assessment model identifies the number of true positive examples TP, the number of false positive examples FP, the number of false negative examples FN, and the number of true negative examples TN, then the accuracy rate Accuracy can be determined by the following formula:
[0061]
[0062] Determine the coverage rate Coverage by the following formula:
[0063]
[0064] Obtain the evaluation index corresponding to the initial service health assessment model according to the accuracy rate and coverage rate. For example, the sum of the accuracy rate and coverage rate can be used as the evaluation index corresponding to the initial service health assessment model, or the average value of the accuracy rate and coverage rate can be used as the evaluation index corresponding to the initial service health assessment model, which is not specifically limited here.
[0065] In a possible implementation manner, when the deviation amount between the second evaluation index and the first evaluation index is less than a preset threshold, it further includes:
[0066] Obtain the monitoring data of the service health assessment model under multiple preset performance metrics; determine the performance prediction value of the service health assessment model according to the monitoring data under each performance metric; in the case that the performance prediction value meets the preset metric threshold, use the updated initial service health assessment model as the service health assessment model; in the case that the performance prediction value does not meet the preset metric threshold, use the initial service health model as the service health assessment model.
[0067] In the embodiments of the present application, a comprehensive evaluation is performed on the constructed service health model. The performance prediction criteria for a distributed system generally include latency, traffic, errors, saturation, etc. Among them, latency refers to the time required for a service to process a certain request. It is important to distinguish between successful requests and failed requests. For example, an HTTP 500 error caused by a database connection loss or other backend problems may have a very low latency; traffic refers to a measurement of the system load demand using a certain high-level pointer in the system. For a web server, this metric is usually the number of HTTP requests per second, and is also classified according to the request type (static requests and dynamic requests); errors refer to the rate of request failures, either explicitly failed, implicitly failed, or failed due to policy reasons; saturation can be understood as how "full" the service capacity is, and is usually a measurement of a specific metric of the most restricted resource in the system currently (in a memory-constrained system, it is memory; in an I / O-constrained system, it is I / O).
[0068] The metrics in the above performance prediction criteria can be used as the performance metrics for evaluating the service health assessment model. According to the monitoring data of the service health assessment model under the above multiple preset performance metrics, determine the performance prediction value of the service health assessment model (such as the time to process a request, the required traffic, the rate of request failures, the saturation of service capacity, etc.); in the case that the performance prediction value meets the preset metric threshold, use the updated initial service health assessment model as the service health assessment model; in the case that the performance prediction value does not meet the preset metric threshold, use the initial service health model as the service health assessment model.
[0069] In a possible implementation manner, in step 1022 above, after predicting, according to the pre-trained target policy network, the third running state information other than the second running state information that has a causal relationship with the first running state information, it further includes:
[0070] Obtain the observation parameters corresponding to the third running state information at multiple historical time points, and determine the adaptive threshold corresponding to the third running state information according to the observation parameters corresponding to the multiple historical time points;
[0071] In step 1023 above, generating the service health assessment model according to the first operating state information, the second operating state information, the third operating state information, and the causal relationships between the respective operating state information includes:
[0072] Generating the service health assessment model according to the first operating state information, the second operating state information, the third operating state information with the adaptive threshold, and the causal relationships between the respective operating state information.
[0073] In the embodiment of the present application, since the third operating state information is predicted according to the target policy network, after obtaining the third operating state information that has a causal relationship with the first operating state information, it is also possible to obtain the observation parameters corresponding to the third operating state information at multiple historical time points, and determine the adaptive threshold corresponding to the third operating state information according to the observation parameters corresponding to the multiple historical time points. Furthermore, the service health assessment model is generated according to the first operating state information, the second operating state information, the third operating state information with the adaptive threshold, and the causal relationships between the respective operating state information.
[0074] Based on this situation, in step S102 above, determining the associated operating state information that has a direct or indirect causal relationship with the target operating state information based on the causal relationships between the respective operating state information in the multiple operating state information includes:
[0075] Obtaining the first associated operating state information that has a direct causal relationship with the target operating state information; if it is determined that the first associated operating state information is abnormal according to the adaptive threshold corresponding to the first associated operating state information, obtaining the second associated operating state information that has a direct causal relationship with the first associated operating state information; if it is determined that the second associated operating state information is abnormal when according to the adaptive threshold corresponding to the second associated operating state information, obtaining the third associated operating state information that has a direct causal relationship with the second associated operating state information, and so on, until all the associated operating state information that has a direct or indirect causal relationship with the target operating state information is determined.
[0076] In an embodiment of the present application, continuing with the example of the alarm information "NameNode node is unavailable" generated by the distributed system, according to multiple running state information in the service health assessment model, find the child node "HDFS is unavailable" of "NameNode node is unavailable", and determine whether there is an abnormality in the child node according to the adaptive threshold corresponding to "HDFS is unavailable". If the child node is abnormal, that is, HDFS is unavailable, then continue to find the next child node "HBase is unavailable", and then determine whether the next child node has an abnormality according to the adaptive threshold corresponding to "HBase is unavailable", and so on, until all associated running state information having a direct causal relationship or an indirect causal relationship with the target running state information is determined.
[0077] Among them, the adaptive threshold corresponding to each running state information can be determined by the following method:
[0078] Construct a threshold time series model according to the autoregressive model and the moving average model; by inputting the observation parameters corresponding to each running state information at multiple historical time points into the threshold time series model, the adaptive threshold corresponding to each running state information is obtained.
[0079] In an embodiment of the present application, if the running state information is index type data such as resource utilization rate, network delay, network packet loss rate, etc., the adaptive threshold corresponding to each running state information can be obtained by inputting the observation parameters corresponding to each running state information at multiple historical time points into the threshold time series model. The threshold time series model can be constructed according to the autoregressive model and the moving average model. The expression of the autoregressive model is as follows:
[0080]
[0081] Among them, X t is the observation parameter corresponding to the current time point t, c is a constant, are the coefficients of the observation parameters corresponding to the lag time points respectively, and ε t is the white noise error.
[0082] The expression of the moving average model is as follows:
[0083] X t = μ + ε t + θ1ε t-1 + θ2ε t-2 + θ q ε t-q ;
[0084] Among them, μ is a constant, and ε t represents the white noise error, and θ1, θ2,..., θ q represent the coefficients corresponding to the lag errors.
[0085] Then the expression of the final threshold time series model is as follows:
[0086]
[0087] Where X t is the observation parameter corresponding to the current time point t, which is used to describe the relationship between the current observation parameter and the observation parameters corresponding to the past p time points. θ1, θ2,..., θ q is used to describe the relationship between the current observation parameter and the errors of the observation parameters corresponding to the past q time points. ε t is the error term at time point t, and c is a constant term. Usually, both p and q are positive integers between [1, 3]. The values of p and q can be selected by the exhaustive method to fit the model, and the relevant parameters can be determined by maximum likelihood estimation θ1, θ2,..., θ q , and finally the adaptive thresholds corresponding to each operation state information are obtained.
[0088] The target policy network in the above step 1022 is a neural network-based model used to learn and represent the action policy that the agent should take in a given state. The policy network here is a key concept in reinforcement learning. The reinforcement learning system mainly consists of the following two parts:
[0089] The first is the external environment ε, which constrains the dynamic interaction between the agent and the initial service health assessment model. This external environment is modeled as a Markov (MDP) decision process and is described by a tuple <S, A, P, R>. Among them, S represents the continuous state space, A = {a1, a2... a n} represents a series of selectable action spaces, P(S t+1 = s'|S t = s, A t = a) is the transition probability matrix, and R(s, a) is the reward function. In the formula of the transition probability matrix, S t represents that the state at the current moment is s, and A t represents that the action selected by the agent at the current moment is a. The meaning of its formula is that at the current moment t, the environment state is s. At this time, the agent makes an action a according to this state, and the probability that the next state of the system is s'.
[0090] The second part is the RL agent, which is described as a policy network π θ (s, a) = p(a|s; θ), which maps the state vector to a stochastic policy and updates the parameters θ of the neural network using the gradient descent method. This policy gradient is a stochastic policy that can avoid falling into local optimal solutions.
[0091] In a possible implementation, in the above step 1022, the target policy network can be trained in the following manner:
[0092] Use a preset path finding method to determine the relationship paths between the candidate running state information in the historical running data; pre-train a pre-constructed policy network based on the relationship paths, determine the rewards of the policy network, and update the network parameters of the policy network with maximizing the rewards as the optimization goal to obtain a pre-trained policy network; re-train the pre-trained policy network through multiple reward functions to obtain the target policy network.
[0093] In the embodiments of the present application, the external environment of the RL agent consists of three parts: action, state, and reward:
[0094] a. Action. For two given candidate running state information, denoted here as the entity pair (e s , e t ) and the relationship r, it is desired that the agent can find the path with the richest information between the entity pairs. The agent initially starts from the head entity e s , and at each time step step, finds an edge as the path through the policy network until reaching the tail entity e t . To ensure the consistency of the output dimension in the policy network, the size of the action space is defined as the number of all relationships.
[0095] b. State. In the initial service health model, entities and relationships are discrete information, and TransE and TransR based on the translation model can be used for modeling to describe the entities and relationships in the initial service health model. The agent first obtains the current state input, and then makes a decision to select an action to transfer from the current entity to the next entity. The state input can be the performance of a certain metric, such as a network delay of 1000 ms and a network packet loss rate of 2%; or a certain log, link, etc. The state vector at time step t is set as follows: s t = (e t , e target - e t ), where e t represents the vector corresponding to the entity where the agent is currently located. At the first time step, e t = e source , where e source is the source node and e target is the target node.
[0096] c. Reward. Multiple reward functions can be set, and the multiple reward functions include a path accuracy reward function, a path validity reward function, and a path diversity reward function.
[0097] Accuracy: In the modeling environment, the dimension of the action space is very large, that is, the number of incorrect paths is much larger than the number of correct paths, and the number of incorrect decision sequences increases exponentially with the growth of the path. To overcome this problem, the reward function r is set GLOBAL :
[0098]
[0099] That is, if the agent reaches the target entity e after a series of decisions target , a reward of 1 is given, otherwise -1.
[0100] Path validity: For the relationship reasoning task, it is generally considered that the reasoning of short paths is more reliable than that of long paths. Therefore, by setting the reward function r EFFICENCY the path p found by the agent is made as short as possible. The reward function r EFFICENCY is defined as follows:
[0101]
[0102] That is, a reward is given according to the length length(p) of the relationship path.
[0103] Path diversity: Using positive samples to train the agent to find paths, the agent is more inclined to find paths with similar semantics and syntax, and these paths often contain redundant information. To ensure the diversity of the paths found by the agent, the cosine similarity between the current path and the existing paths is used to set the diversity reward function r DIVERISITY :
[0104]
[0105] where represents a relationship chain, which is the current path; F represents the number of existing paths; p i represents the existing path.
[0106] Furthermore, supervised policy training is first carried out. The relationship paths between the candidate running state information in the historical running data are determined using a preset path finding method. For each relationship path, a subset of all positive samples is used to pre-train the pre-constructed policy network. The positive samples (e source , e target ) start from two nodes respectively and perform bidirectional BFS (Bidirectional Breadth-First Search) to find the correct path.
[0107] Here, bidirectional BFS is a graph search algorithm that starts searching from both the start point and the end point simultaneously to find the shortest path. In the bidirectional BFS algorithm, two queues or sets are used to record the reachable nodes starting from the start point and the end point respectively, and in each round, a node is alternately selected from the two queues or sets for expansion. If a node is visited from both directions, it means there is a path between these two nodes.
[0108] For each path p with a sequence of (r1, r2... r n ), Monte Carlo gradient descent is used to maximize the cumulative reward, and its loss function J(θ) is:
[0109]
[0110] where, represents the reward obtained when the state at time t is s t , and the action taken is α t .
[0111] Update the network parameters θ of the policy network with the goal of maximizing the reward to obtain the pre-trained policy network.
[0112] Further retrain the reward. To find the inference path through the reward function, the pre-trained policy network with supervised training is retrained using the reward. The inference of an entity pair (e source , e target ) is a sequence that starts from the head entity e source . The agent selects an edge / relation according to the random policy π(α|s), where π(α|s) represents the probability distribution of all possible actions. This relation link may point to a new entity, and at these failed time steps, the agent will receive a negative reward value. At these time steps, the agent will stay in the current state. Since the agent follows a random policy, the agent will not keep repeating wrong decisions. To improve the training efficiency, the maximum number of time steps for the agent to find the path will not exceed max_length. If the agent does not find the target entity e target within max_length time steps, stop executing this sequence. After each sequence is executed, update the policy network gradient formula:
[0113]
[0114] where, R total = λ1r GLOBAL + λ2r EFFICIENCY + λ3r DIVERSITY .
[0115] When the loss function J(θ) reaches the maximum value, the model training is completed, and the target policy network is obtained.
[0116] Figure 5 The flowchart shows the method for constructing a service health assessment model provided by an embodiment of the present application. As shown in the figure, the method may include the following steps:
[0117] Step 501: Data collection: The operation information of the distributed software system is stored in different databases. Connect to the data sources of the distributed software system and collect the historical operation data of the distributed software system.
[0118] Step 502: Construct an initial service health assessment model: Obtain structured and semi-structured operation status information data from the historical operation data of the distributed software system; Determine multiple candidate operation status information in the structured and semi-structured operation status information data based on preset data extraction rules; Perform knowledge cleaning and knowledge evaluation on the multiple candidate operation status information based on the feature information of the multiple candidate operation status information, and construct an initial service health assessment model.
[0119] Step 503: Use the target policy network in the reinforcement learning algorithm to optimize the initial service health assessment model and construct a model generation engine to generate a service health assessment model.
[0120] Step 504: Comprehensively evaluate the generated service health assessment model. If the evaluation passes, store the service health assessment model to obtain a service health assessment system.
[0121] The above service health assessment method can be implemented through software programs such as matlab and OpenCV. Figure 6 The structural diagram shows a service health assessment system provided by an embodiment of the present application. As shown in the figure, the system may include:
[0122] A data collection layer 610, which is used to collect the operation data of the distributed software system and measure the service health of the distributed system. It can connect to different data sources such as Prometheus, MySQL, ElasticSearch, and PostgreSQL.
[0123] A model construction layer 620, which is used to construct and improve the initial service health assessment model, including a knowledge acquisition module, a knowledge fusion module, a knowledge evaluation module, and a knowledge reasoning module.
[0124] A model optimization layer 630, which is used to optimize and adjust the initial service health assessment model, including a reinforcement learning model module and a model generation engine module.
[0125] The model generation layer 640 is used to intelligently generate a service health assessment model based on the initial service health assessment model and the multi-objective optimization model, including: a model comprehensive evaluation module and a service health model module.
[0126] In the embodiment of the present application, by applying the reinforcement learning algorithm to the generation process of the service health assessment model, an operation and maintenance knowledge base is constructed, and based on the knowledge reasoning technology, the service health of the distributed system can be analyzed efficiently and reliably, so as to reasonably plan the operation and maintenance strategy and reduce the operation and maintenance cost.
[0127] Figure 7 The figure shows a schematic hardware structure diagram of an electronic device provided by the embodiment of the present application. Referring to this figure, at the hardware level, the electronic device includes a processor, and optionally, an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.
[0128] The processor, the network interface, and the memory can be interconnected through the internal bus. The internal bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a bidirectional arrow is used in this figure to represent it, but it does not mean that there is only one bus or one type of bus.
[0129] The memory stores a program. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include a memory and a non-volatile memory, and provides instructions and data to the processor.
[0130] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, and forms a device for locating target users at the logical level. The processor executes the program stored in the memory and specifically executes: Figure 1 or Figure 5 The method disclosed in the shown embodiment realizes the functions and beneficial effects of the various methods described in the foregoing method embodiments, which will not be elaborated here.
[0131] The above is as in the present application Figure 1or Figure 5 The method disclosed in the illustrated embodiment can be applied in a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or by instructions in the form of software. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0132] The computer device can also execute the various methods described in the foregoing method embodiments and achieve the functions and beneficial effects of the various methods described in the foregoing method embodiments, which will not be elaborated here.
[0133] Of course, in addition to the software implementation, the electronic device of the present application does not exclude other implementation manners, such as a logic device or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and may also be hardware or a logic device.
[0134] The embodiments of the present application also propose a computer-readable storage medium. The computer-readable medium stores one or more programs. When the one or more programs are executed by an electronic device including a plurality of application programs, the electronic device is caused to execute Figure 1 or Figure 5 the method disclosed in the illustrated embodiment and achieve the functions and beneficial effects of the various methods described in the foregoing method embodiments, which will not be elaborated here.
[0135] Among them, the computer-readable storage medium includes a read-only memory (ROM for short), a random access memory (RAM for short), a magnetic disk, an optical disc, etc.
[0136] Furthermore, the embodiment of the present application also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the following process is implemented: Figure 1 or Figure 5 The method disclosed in the shown embodiment realizes the functions and beneficial effects of each method described in the foregoing method embodiments, and will not be elaborated herein.
[0137] The embodiment of the present application can be applied to various scenarios of electronic device cooperation or interconnection, including: cooperation and interconnection between a mobile phone and a laptop / tablet computer; cooperation and interconnection between a mobile terminal and a smart TV / display; cooperation and interconnection between a mobile phone or a tablet computer and an in-vehicle entertainment system; cooperation and interconnection between a mobile terminal and a smart conference system, etc. Thus, it can meet the diverse scenario requirements of users in smart home, smart office, smart travel, etc.
[0138] In summary, the above are only the preferred embodiments of the present application, and do not limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0139] The system, device, module or unit illustrated in the above embodiment can be specifically implemented by a computer chip or an entity, or by a product with a certain function. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0140] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0141] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0142] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.
Claims
1. A method for evaluating service health, characterized in that, Including: In response to the collected status monitoring data of the distributed system, obtain the target operating status information that matches the status monitoring data among multiple operating status information of the service health assessment model; Based on the causal relationships among the respective operating status information in the multiple operating status information, determine the associated operating status information that has a direct causal relationship or an indirect causal relationship with the target operating status information; Determine the associated data in the status monitoring data that matches the associated operating status information, and determine the service health of the distributed system based on the associated data.
2. The method according to claim 1, characterized in that Before obtaining the target operating status information that matches the status monitoring data among the multiple operating status information of the service health assessment model, it further includes: Based on the historical operating data of the distributed system, determine the first operating status information in the historical operating data and the second operating status information that has a causal relationship with the first operating status information; Predict, according to the pre-trained target policy network, the third operating status information other than the second operating status information that has a causal relationship with the first operating status information; Generate the service health assessment model according to the first operating status information, the second operating status information, the third operating status information, and the causal relationships among the respective operating status information.
3. The method according to claim 2, wherein The generating the service health assessment model according to the first operating status information, the second operating status information, the third operating status information, and the causal relationships among the respective operating status information includes: Generate an initial service health assessment model according to the first operating status information, the second operating status information, and the causal relationship between the first operating status information and the second operating status information, and evaluate the initial service health assessment model to obtain a first evaluation index; Update the initial service health assessment model according to the third operating status information and the causal relationship between the first operating status information and the third operating status information, and evaluate the updated initial service health assessment model to obtain a second evaluation index; In the case where the deviation amount between the second evaluation index and the first evaluation index is less than a preset threshold, use the updated initial service health assessment model as the service health assessment model; in the case where the deviation amount is not less than the preset threshold, use the initial service health model as the service health assessment model.
4. The method according to claim 3, wherein Evaluate the initial service health assessment model in the following manner: Obtain multiple pre-labeled sample data, process the multiple sample data through the initial service health assessment model to obtain an availability evaluation label corresponding to each sample data in the multiple sample data; The sample data carries an availability true label; By comparing the availability evaluation label corresponding to each sample data with the availability true label, determine the accuracy rate and coverage rate of the initial service health assessment model, and obtain the evaluation index corresponding to the initial service health assessment model according to the accuracy rate and coverage rate.
5. The method according to claim 3, wherein When the deviation amount between the second evaluation index and the first evaluation index is less than a preset threshold, it further includes: Obtaining monitoring data of the service health evaluation model under multiple preset performance indicators; Determining a performance prediction value of the service health evaluation model according to the monitoring data under each performance indicator; When the performance prediction value meets the preset index threshold, using the updated initial service health evaluation model as the service health evaluation model; when the performance prediction value does not meet the preset index threshold, using the initial service health model as the service health evaluation model.
6. The method according to claim 2, wherein After predicting, according to the pre-trained target policy network, third operating state information other than the second operating state information that has a causal relationship with the first operating state information, it further includes: Obtaining observation parameters corresponding to the third operating state information at multiple historical time points, and determining an adaptive threshold corresponding to the third operating state information according to the observation parameters corresponding to the multiple historical time points; The generating the service health evaluation model according to the first operating state information, the second operating state information, the third operating state information, and the causal relationship between each operating state information includes: Generating the service health evaluation model according to the first operating state information, the second operating state information, the third operating state information with the adaptive threshold, and the causal relationship between each operating state information.
7. The method according to claim 1, characterized in that, The determining, based on the causal relationship between each operating state information in the multiple operating state information, the associated operating state information that has a direct or indirect causal relationship with the target operating state information includes: Obtaining first associated operating state information that has a direct causal relationship with the target operating state information; If it is determined that the first associated operating state information is abnormal according to the adaptive threshold corresponding to the first associated operating state information, obtaining second associated operating state information that has a direct causal relationship with the first associated operating state information; if it is determined that the second associated operating state information is abnormal according to the adaptive threshold corresponding to the second associated operating state information, obtaining third associated operating state information that has a direct causal relationship with the second associated operating state information, and so on, until all the associated operating state information that has a direct or indirect causal relationship with the target operating state information is determined.
8. The method according to claim 6 or claim 7, characterized in that, Determining the adaptive threshold corresponding to each operating state information in the following manner: Constructing a threshold time series model according to an autoregressive model and a moving average model; Obtaining the adaptive threshold corresponding to each operating state information by inputting the observation parameters corresponding to each operating state information at multiple historical time points into the threshold time series model.
9. The method according to claim 2, characterized in that The determining, based on the historical operation data of the distributed system, the first operating state information in the historical operation data, and the second operating state information that has a causal relationship with the first operating state information includes: Obtain structured and semi-structured operation status information data from the historical operation data of the distributed system; Determine multiple candidate operation status information in the structured and semi-structured operation status information data based on preset data extraction rules; Perform feature correlation processing on the multiple candidate operation status information based on the feature information of the multiple candidate operation status information to obtain the first operation status information and the second operation status information that has a causal relationship with the first operation status information.
10. The method according to claim 2, wherein Train the target policy network in the following manner: Use a preset path finding method to determine the relationship paths between the various candidate operation status information in the historical operation data; Perform pre-training on a pre-constructed policy network based on the relationship paths, determine the rewards of the policy network, and update the network parameters of the policy network with maximizing the rewards as the optimization goal to obtain a pre-trained policy network; Re-train the pre-trained policy network through multiple reward functions to obtain the target policy network.
11. The method according to claim 10, wherein, The multiple reward functions include a path accuracy reward function, a path effectiveness reward function, and a path diversity reward function; the re-training of the pre-trained policy network through the multiple reward functions to obtain the target policy network includes: Use the path accuracy reward function to give a reward when the target candidate operation status information in each candidate operation status information is found according to the relationship path, and update the pre-trained policy network; Use the path effectiveness reward function to give a reward according to the length of the relationship path, and update the pre-trained policy network; Use the path diversity reward function to give a reward according to the similarity between the relationship path and the existing relationship paths, and update the pre-trained policy network; Determine the target policy network according to the updated pre-trained policy network.
12. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.
13. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.