Model capability assessment method and device, equipment and storage medium
By acquiring abnormal observation datasets under fault scenarios, analyzing and predicting fault information using the fault diagnosis model to be evaluated, and introducing a dynamic weighting mechanism, the problem of inaccurate model evaluation in existing technologies is solved, and a more accurate model capability evaluation is achieved.
Patent Information
- Application Number
- CN202511372740.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2026-01-02
AI Technical Summary
Existing model evaluation methods cannot accurately assess a model's ability to locate fault events in a real production environment, and their reliance on manually set weighting coefficients makes them susceptible to the influence of experience and understanding, leading to inaccurate evaluation results.
By acquiring anomaly observation datasets under fault scenarios, the fault diagnosis model to be evaluated is used to analyze and predict fault information. Based on the predicted fault information and the target fault information, the root cause analysis effect value is determined. A dynamic weighting mechanism is introduced to calculate the capability score by combining the importance and representativeness differences of the fault scenarios.
It improves the accuracy and reliability of model capability assessment, avoids the bias of subjective human evaluation and single fault scenario assessment, and is closer to the actual fault diagnosis needs.
Smart Images

Figure CN121256291A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of model evaluation technology, and in particular to a method, apparatus, device and storage medium for evaluating model capabilities. Background Technology
[0002] For operation and maintenance models, it is not only necessary to evaluate the model's general capabilities in language understanding, text generation quality, and logical reasoning, but more importantly, as a model for a specific vertical domain, it is necessary to evaluate whether the model can accurately locate fault events that occur in the actual production environment.
[0003] However, existing evaluation methods primarily assess model capability through option matching consistency (whether the output options match the standard options). This may limit model performance because real-world problems may not be multiple-choice questions; the model may need to generate answers rather than making choices. This could affect the realism of the evaluation, as it cannot be guaranteed whether the model's output options are derived from random guesses or genuine understanding.
[0004] Furthermore, when using multi-dimensional datasets to evaluate model capabilities, existing evaluation methods mainly rely on evaluators to set the weight coefficients corresponding to the evaluation results of each dataset. This approach is easily affected by the evaluators' experience and understanding of the datasets. When the weight coefficients are unreasonable, it will affect the authenticity of the evaluation and reduce the accuracy of the evaluation results. Summary of the Invention
[0005] This invention provides a model capability assessment method, apparatus, device, and storage medium to improve the accuracy of model capability assessment.
[0006] In a first aspect, embodiments of the present invention provide a model capability assessment method, the method comprising:
[0007] Obtain an anomaly observation dataset for each fault scenario, including at least one fault event for each fault scenario;
[0008] The abnormal observation dataset is analyzed using the fault diagnosis model to be evaluated to determine the predicted fault information;
[0009] Based on the predicted fault information and the target fault information, determine the root cause analysis effect value of the fault scenario;
[0010] The capability score of the fault diagnosis model to be evaluated is determined based on the root cause analysis effect value and the corresponding target weight of each fault scenario in the fault scenario set; the target weight is determined based on the number of fault events corresponding to each fault scenario and the nonlinear adjustment factor fault scenario corresponding to each number of fault events.
[0011] Secondly, embodiments of the present invention provide a model capability evaluation device, the device comprising:
[0012] The acquisition module is used to acquire anomaly observation datasets under fault scenarios, with each fault scenario including at least one fault event.
[0013] The prediction module is used to analyze the abnormal observation dataset through the fault diagnosis model to be evaluated, and to determine the predicted fault information.
[0014] The effect value determination module is used to determine the root cause analysis effect value of the fault scenario based on the predicted fault information and the target fault information.
[0015] The capability score determination module is used to determine the capability score of the fault diagnosis model to be evaluated based on the root cause analysis effect value and the corresponding target weight of each fault scenario in the fault scenario set; the target weight is determined based on the number of fault events corresponding to each fault scenario and the nonlinear adjustment factor fault scenario corresponding to each number of fault events.
[0016] Thirdly, embodiments of the present invention provide an electronic device, the electronic device comprising:
[0017] At least one processor;
[0018] and a memory communicatively connected to the at least one processor;
[0019] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the model capability evaluation method according to any embodiment of the present invention.
[0020] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing computer instructions that are used to cause a processor to execute the model capability evaluation method described in any embodiment of the present invention.
[0021] The technical solution of this invention involves acquiring anomaly observation datasets under fault scenarios, where each fault scenario includes at least one fault event; analyzing the anomaly observation datasets using the fault diagnosis model to be evaluated to determine predicted fault information; determining the root cause analysis effect value of the fault scenario based on the predicted fault information and target fault information; and determining the capability score of the fault diagnosis model to be evaluated based on the root cause analysis effect values of each fault scenario in the fault scenario set and their corresponding target weights. This method determines the root cause analysis effect value of the fault scenario based on the predicted fault information and actual target fault information of the fault diagnosis model to be evaluated relative to the anomaly observation dataset. It also introduces a dynamic weighting mechanism, considering the importance and representativeness differences of different fault scenarios, and aggregates the root cause analysis effect values of multiple fault scenarios. This overcomes the limitations of traditional averaging methods, determining a capability score that objectively and accurately reflects the comprehensive diagnostic capability of the fault diagnosis model to be evaluated. This improves the authenticity and credibility of the capability score, making it closer to actual fault diagnosis needs and effectively avoiding biases from subjective human evaluation or single fault scenario assessment.
[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 A flowchart of a model capability assessment method provided in an embodiment of the present invention;
[0025] Figure 2 This is a schematic diagram of the structure of a model capability assessment device provided in an embodiment of the present invention;
[0026] Figure 3 A schematic diagram of an electronic device that can be used to implement embodiments of the present invention is shown. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] It should be noted that one application scenario of this invention is as follows: when optimizing the shortcomings of self-developed models or when purchasing external operation and maintenance models, it is necessary to evaluate the capabilities of each model to achieve horizontal comparison and selection. However, when evaluating the capabilities of models, traditional manual evaluation methods are costly, inefficient, and difficult to handle massive amounts of data. In addition, although humans are better at judging fluency and grammatical correctness, they are not good at assessing factual accuracy and require sufficient professional knowledge, thus reducing the accuracy and objectivity of the evaluation.
[0030] Meanwhile, existing automated evaluation methods primarily assess model capability through option matching consistency (whether the output options match the standard options). However, since real-world problems may not be in multiple-choice format, the model needs to generate answers rather than making choices. This can affect the realism of the evaluation, as it cannot be guaranteed whether the model's output options are derived from random guesses or genuine understanding.
[0031] Furthermore, when using multi-dimensional datasets to evaluate model capabilities, existing evaluation methods mainly rely on evaluators to set the weight coefficients corresponding to the evaluation results of each dataset. This approach is easily affected by the evaluators' experience and understanding of the datasets. When the weight coefficients are unreasonable, it will affect the authenticity of the evaluation and reduce the accuracy of the evaluation results.
[0032] Based on this, embodiments of the present invention provide a method for evaluating model capabilities. Figure 1This is a flowchart of a model capability assessment method provided by an embodiment of the present invention. The embodiment of the present invention can be applied to scenarios for assessing model capabilities, especially scenarios for assessing the fault event location capability of the operation and maintenance model (such as the fault diagnosis model) of a business system. The method can be executed by a model capability assessment device, which can be implemented in the form of software and / or hardware, and optionally, by an electronic device, preferably a mobile terminal, desktop computer, laptop computer, or server.
[0033] like Figure 1 As shown, the model capability assessment method provided in this embodiment of the invention may specifically include:
[0034] S101. Obtain the abnormal observation dataset under the fault scenario. Each fault scenario includes at least one fault event.
[0035] A failure scenario can be considered as a situation in which a failure event occurs due to certain reasons in a specific system, device, or environment. Failure scenarios may change with time, operating conditions, or external environment. Different systems or devices may have different failure scenarios, and even the same system may have different failure scenarios.
[0036] A fault scenario can include one or more fault events. A fault event can be considered a specific point of failure or fault phenomenon that occurs within the fault scenario, causing the system or device to deviate from its normal operating state. Fault events are specific, observable fault phenomena, which can typically be detected through monitoring tools, logging systems, or other monitoring methods, and can be used to pinpoint specific components, locations, or points in time based on the detection results. For example, fault events can include CPU failures, memory failures, and Pod failures.
[0037] An anomaly observation dataset can be understood as a dataset automatically constructed based on observable data related to the abnormal state occurring in the system and target fault information when a fault event occurs. Observable data can be generated based on various abnormal system operation data generated by the system, reflecting the system's operating status, performance indicators, and fault information. Target fault information can be understood as the actual fault information corresponding to the actual fault event, which can be determined based on experience and set rules, or based on a pre-trained model. Target fault information can include the actual input fault event, fault description (such as fault event, fault location, etc.), solutions, and preventive measures.
[0038] In this embodiment, the abnormal observation dataset under the fault scenario can be obtained by obtaining the real-time constructed abnormal observation dataset from the set access interface, access address or database.
[0039] S102. Analyze the abnormal observation dataset using the fault diagnosis model to be evaluated to determine the predicted fault information.
[0040] The fault diagnosis model to be evaluated can be considered a fault diagnosis model with the capability assessment requirement. This model can be used to diagnose fault events occurring in the business system. Predicted fault information can be understood as the fault information diagnosed by the fault diagnosis model to be evaluated based on anomaly observation datasets. Predicted fault information may include predicted fault descriptions (such as fault events, fault locations, etc.), related evidence (such as abnormal time, scope of impact, etc.), root cause conclusions (such as "database connection pool leakage"), solutions (such as expanding the connection pool), and preventative measures.
[0041] In this embodiment, the abnormal observation dataset is input into the fault diagnosis model to be evaluated. The fault diagnosis model to be evaluated processes the abnormal observation data in the abnormal observation dataset and searches for historical cases and other knowledge that match the abnormal observation data in the pre-built knowledge base. Causal reasoning is applied to generate root cause hypotheses, and then the reliability of the conclusion is confirmed through secondary verification (such as data back lookup and threshold verification). After confirming the reliability of the conclusion, the predicted fault information is generated and output according to the preset unified template.
[0042] For example, the anomaly observation dataset may contain metric data, log data, and tracing data. The fault diagnosis model to be evaluated can process the anomaly observation data in the anomaly observation dataset in the following ways: process log data, parse log semantics, and identify error patterns, such as timeouts and connection failures; process metric data, detect performance inflection points, such as CPU spikes and response latency; process tracing data, locate service dependency failure paths, such as microservice call interruptions.
[0043] S103. Based on the predicted fault information and the target fault information, determine the root cause analysis effect value of the fault scenario.
[0044] The root cause analysis result value can be considered as an indicator reflecting the accuracy and reliability of the fault diagnosis model under evaluation in fault diagnosis. The larger the root cause analysis result value, the more accurate the predicted fault information output by the fault diagnosis model under evaluation, and the better the root cause analysis result.
[0045] In this embodiment, the root cause analysis effect value of the fault scenario can be determined based on the similarity between the predicted fault information and the target fault information; alternatively, it can be determined based on the number of predicted fault events in the predicted fault information and the number of target fault events in the target fault information. The target fault information can be obtained from the anomaly observation dataset or from the fault event input.
[0046] In an optional embodiment, the root cause analysis effect value of the fault scenario can be determined by dividing the number of predicted fault events in the predicted fault information by the number of target fault events in the target fault information, and determining the quotient as the root cause analysis effect value.
[0047] The predicted number of fault events can be understood as the number of faults diagnosed by the fault diagnosis model to be evaluated; the target number of fault events can be considered as the number of fault events that actually occur in the business system under the fault scenario.
[0048] S104. Based on the root cause analysis effect value and corresponding target weight of each fault scenario in the fault scenario set, determine the capability score of the fault diagnosis model to be evaluated.
[0049] The target weights are determined based on the number of fault events corresponding to each fault scenario and the nonlinear adjustment factor fault scenario corresponding to each number of fault events. The number of fault events can be considered as the total number of fault events existing in a fault scenario.
[0050] In this embodiment, the target weight can be considered as the weight corresponding to each root cause analysis effect value, which is used to adjust the degree of influence of the root cause analysis effect value on the ability score of the fault diagnosis model to be evaluated under each fault scenario.
[0051] Understandably, a single fault scenario is insufficient for a more accurate assessment of the capabilities of the fault diagnosis model under evaluation and to avoid errors. Therefore, it is necessary to simulate various fault scenarios, including those with single fault events and those with mixed types of fault events, to obtain anomaly observation datasets for each scenario. Based on the above steps, the root cause analysis performance value for each fault scenario is determined. However, different fault scenarios have varying degrees of complexity. The more fault events in a scenario, the more complex the overall business system environment becomes. This necessitates that the fault diagnosis model under evaluation possess fault diagnosis capabilities in complex fault scenarios. Therefore, when determining the capability score of the fault diagnosis model under evaluation, more weight needs to be allocated to the root cause analysis performance value for complex fault scenarios to minimize the impact of differences in the complexity of different fault scenarios and obtain a more objective and accurate capability score.
[0052] However, using the number of fault events as the basis for weight allocation implies that the more fault events a fault scenario contains, the higher its weight in the statistics should be. But considering that as the number of fault events increases, the environment itself becomes increasingly chaotic and complex, and fault events can influence each other, the abnormal observation data collected in the business system environment will inherently interfere with each other. This makes it difficult for the fault diagnosis model to analyze all fault events. Therefore, the weight allocation cannot simply increase linearly with the number of fault events in the fault scenario; a non-linear adjustment factor needs to be introduced to weaken the impact of the number of fault events.
[0053] In an optional embodiment, the target weight is the ratio of the first quantity influence value corresponding to the fault scenario to the total quantity influence value corresponding to all fault scenarios; the first quantity influence value is the product of the number of fault events in the fault scenario and the nonlinear adjustment factor corresponding to the number of fault events; the total quantity influence value is the sum of the products of the number of fault events corresponding to each fault scenario and the nonlinear adjustment factor corresponding to the number of fault events.
[0054] Optionally, the capability score of the fault diagnosis model to be evaluated can be determined by: determining the product of the root cause analysis effect value and the corresponding target weight for each fault scenario, and summing up the product values to obtain the capability score.
[0055] For example, determining the capability score of the fault diagnosis model to be evaluated. The specific method can be expressed as:
[0056]
[0057] in, This represents the total number of fault scenarios. For the first The target weights corresponding to each fault scenario; For the first The root cause analysis effect value of each failure scenario; For the first The number of failure events in each failure scenario; For the first Number of fault events in each fault scenario The corresponding nonlinear adjustment factor; For the first The number of fault events in each fault scenario; For the first Number of fault events in each fault scenario The corresponding nonlinear adjustment factor.
[0058] The model capability assessment method provided in this invention obtains anomaly observation datasets under fault scenarios, where each fault scenario includes at least one fault event. The method analyzes the anomaly observation datasets using the fault diagnosis model to be evaluated to determine predicted fault information. Based on the predicted fault information and target fault information, the root cause analysis effect value of the fault scenario is determined. Finally, based on the root cause analysis effect values of each fault scenario in the fault scenario set and their corresponding target weights, the capability score of the fault diagnosis model to be evaluated is determined. This method determines the root cause analysis effect value of the fault scenario based on the predicted fault information and actual target fault information of the fault diagnosis model to be evaluated relative to the anomaly observation dataset. It also introduces a dynamic weighting mechanism, considering the importance and representativeness differences of different fault scenarios, and aggregates the root cause analysis effect values of multiple fault scenarios. This overcomes the limitations of traditional averaging methods, determining a capability score that objectively and accurately reflects the comprehensive diagnostic capability of the fault diagnosis model to be evaluated. This improves the authenticity and credibility of the capability score, making it closer to actual fault diagnosis needs and effectively avoiding biases from subjective human evaluation or single fault scenario assessments.
[0059] As a first optional embodiment of the present invention, based on the above embodiments, the acquisition of the abnormal observation dataset under the fault scenario can be specified as the following steps:
[0060] a1) Input the fault events included in the fault scenario into the target business system, and collect system operation data under the fault scenario from the target business system.
[0061] The target business system can be understood as a business system with fault diagnosis requirements. System operation data can be understood as the data generated by the business system during operation.
[0062] In this embodiment, the way to input the fault events included in the fault scenario into the target business system can be: to arrange multiple fault events to form a fault scenario, and to conduct a chaotic experiment on the entire system with one fault event or a mixture of multiple fault events to simulate the complex situation in the real production environment.
[0063] Specifically, the first step is to orchestrate fault scenarios: one method is to create new fault scenarios from templates, using pre-set fault scenarios in the scenario library; the other is to create custom fault scenarios, orchestrating various fault events. Next, environment parameters are filled in: after orchestrating the fault scenarios, different environment parameters need to be filled in for different fault scenarios. For example, the fault scenario corresponding to a pod deletion fault event requires environment parameters such as the pod's partition and name, while the fault scenario corresponding to a database fault event requires environment parameters such as the database connection address and password.
[0064] Then, the fault events included in the fault scenarios are input into the target business system. For example, a fault injection client can be used that supports both physical host and Kubernetes chaos experiment scenarios and allows traffic injection through business simulation. Physical host fault scenarios support injecting various fault events into CPU, memory, network, and disk, such as using Go's goroutine technology to adjust CPU load, using functions like `new` and / or `make` to allocate memory in code blocks, and using `tc` technology to simulate network packet corruption. Kubernetes fault scenarios, by calling the k8s API Server interface to execute k8s commands, can inject various fault events into k8s resources such as pods, nodes, and containers, such as pod unavailability, node disk filling, and container deletion. The fault injection client can include functional modules such as an exercise center, fault management, business simulation, security verification, configuration management, and system management.
[0065] During the process of inputting fault events into the target business system, different types of fault events are handled by dedicated applications. For example, fault events in a Kubernetes environment are handled by the `chaos-operator` by executing relevant instructions through calls to the Kubernetes API Server; fault events on physical hosts are handled directly by the `chaos-agent` deployed on the target nodes. In terms of deployment architecture, the `chaos-operator` runs as a single instance within the Kubernetes cluster and requires the necessary Kubernetes privileges to perform its functions; while the `chaos-agent`, responsible for operating host resources, must be present on all target nodes. On Kubernetes nodes, `chaos-agent` is typically deployed as a DaemonSet; on bare metal servers, it needs to be directly installed and deployed in binary form. Ultimately, fault events in the infrastructure and various resources of the target business system are input by either the `chaos-operator` or the `chaos-agent`.
[0066] The method for collecting system operation data under the fault scenario from the target business system can be as follows: load data collection tools, such as eBPF program, cBPF program and AF_Packet program, into the kernel, and collect system operation data information of the target business system under the fault scenario, such as Packet data, Socket data and Function data, from the business applications, calls of the target business system, and the network card corresponding to the target business system.
[0067] b1) Determine the observable data under the fault scenario based on the system operation data, and determine the resource tag information of the observable data according to the preset resource tags and identification rules.
[0068] Resource tags can be understood as tags corresponding to resource entities (such as a specific instance of a specific service), such as service name, hostname, etc. Resource tag information can be understood as information composed of resource tags corresponding to observable data.
[0069] In an optional embodiment, observable data may include metric data, log data, and tracing data. Accordingly, the method for determining observable data under the fault scenario based on the system operation data can be as follows: Identify the application protocol in the system operation data using protocol parsing technology to obtain the corresponding Span information. Span information can be understood as a basic unit in distributed tracing, representing the execution process of an operation. Span information may include the operation name, start time, and end time. Based on the Span information, tracing data is formed through association, aggregation, and other processing. Performance metrics, such as request response time, error rate, and throughput, are extracted from the tracing data. Resource usage metrics, such as CPU utilization, memory utilization, network bandwidth, network layer throughput, latency, and anomalies, are extracted from the system operation data (Packet data and Socket data) to form metric data. Log information, such as error logs and warning logs, is extracted from the tracing data and performance metric data. For Packet data, TCP / UDPFlow can also be generated based on packet aggregation to generate flow logs. Log data is formed based on the extracted log information and the generated flow logs.
[0070] For example, resource labels may include: 1) Kubernetes custom labels: workload labels, replica set labels, and pod labels; common labels include owner, commitId, version, env, and group. 2) Kubernetes resource information: cluster, node, namespace, service, ingress, replica set, pod, workload (deployment / statefulset / daemonset), etc. 3) Cloud resource information: region, availability zone, cloud server, VPC, subnet, router, security group, NAT gateway, load balancer, peering connection, RDS, and Redis, etc.
[0071] In this embodiment, corresponding resource tags can be mapped to corresponding fields in observable data to obtain resource tag information of the observable data. For example, the mapping rules of resource tags to corresponding fields in observable data are shown in Table 1.
[0072] Table 1
[0073]
[0074] c1) Determine the target anomaly determination conditions corresponding to the observable data according to the pre-constructed anomaly determination logic, and determine the observable data that meets the target anomaly determination conditions as the initial anomaly observation data.
[0075] The anomaly detection logic can be understood as the logic used to determine whether there are data anomalies in the observable data. The target anomaly detection condition can be considered as the condition corresponding to the observable data that can be used to determine whether there are data anomalies. The initial anomaly observation data can be considered as independent observable data with data anomalies.
[0076] In this embodiment, the target anomaly determination condition corresponding to the observable data is determined according to the anomaly determination logic. It is then determined whether the observable data meets the target anomaly determination condition. If the target anomaly determination condition is not met, the observable data is determined as normal observation data. If the target anomaly determination condition is met, the observable data is determined as initial anomaly observation data.
[0077] As one implementation method, the construction steps of the exception determination logic can be further specified as follows:
[0078] c11) Define the data type of the data to be processed according to the business scenario, and use the data type as the first-level classification basis.
[0079] For example, the data types of the data to be processed can be defined according to the business scenario. The data types can be divided into three categories: 1) Metric data: structured numerical time-series data (such as CPU utilization, interface QPS and database connection count); 2) Log data: semi-structured / unstructured text logs (such as application logs, server logs, including timestamps, log levels and message content); 3) Trace data: call chain tracing data of distributed systems (such as OpenTelemetry Span data, including TraceID, SpanID, parent SpanID, time consumption, status code and service node).
[0080] c12) Based on the business scenario and business experience, define the abnormal feature type under each data type, and use the abnormal feature type as the basis for the second-level classification.
[0081] For example, the exception feature types under each data type are defined according to the business scenario and business experience, as shown in Table 2.
[0082] Table 2
[0083]
[0084] c13) Based on the business scenario and the data characteristics of the data to be processed, define the anomaly determination conditions under each anomaly feature type.
[0085] The anomaly detection criteria may include detection rules and parameters / thresholds.
[0086] For example, based on the business scenario and the data characteristics of the data to be processed, the exception judgment conditions under each exception feature type are defined, as shown in Table 3.
[0087] Table 3
[0088]
[0089] c14) Based on the first-level classification criteria, the second-level classification criteria, and the anomaly determination conditions, anomaly determination logic is constructed.
[0090] In this embodiment, the first-level classification criteria, the second-level classification criteria, and the anomaly determination conditions can be determined as first-level logic, second-level logic, and third-level logic, respectively, and the anomaly determination logic can be constructed in the order of logic level from first level to third level.
[0091] Correspondingly, the initial anomaly observation data determination method based on anomaly determination logic can be as follows: acquire observable data, determine the data type of observable data, locate the anomaly feature type according to the data type, match the corresponding anomaly determination conditions according to the anomaly feature type, determine the anomaly determination conditions as the target anomaly determination conditions, determine the observable data that meets the target anomaly determination conditions as the initial anomaly observation data and output it.
[0092] For example, if the data type of the observable data is determined to be indicator data, then the abnormal feature type is located and indicator anomaly detection is performed, including sudden increase / decrease detection, deviation from baseline 3σ detection, and golden signal anomaly detection. The target anomaly judgment conditions (judgment rules combined with parameters / thresholds) corresponding to each abnormal feature type are determined, and then it is determined whether the observable data meets any of the target anomaly judgment conditions. If it does, the observable data is determined as the initial abnormal observation data and output. If it does not meet all the target anomaly judgment conditions, the observable data is determined as normal observation data and output.
[0093] The above-described technical solution in this embodiment, through a three-dimensional hierarchical design, transforms abstract abnormal features into executable judgment rules, thus forming a systematic organization for the anomaly judgment of different types of observable data.
[0094] d1) Determine the target anomaly observation data based on the initial anomaly observation data and the resource tag information corresponding to the initial anomaly observation data.
[0095] It is understandable that system operation data in a fault scenario is usually multiple, therefore, both observable data and initial anomaly observation data can be multiple. Thus, in this embodiment, the initial anomaly observation data can be aggregated, and the resource tag information corresponding to each initial anomaly observation data can be obtained. Because different types of observable data are associated with unified resource tags (such as pod_cluster, pod_ns, etc.), different types of initial anomaly observation data (such as metric data and tracking data) can be linked through this common resource tag information. Based on the resource tag information, the resource corresponding to the initial anomaly observation data (such as the problematic service instance) can be directly located. Therefore, initial anomaly observation data with the same resource tag information can be extracted, associated, and integrated to form comprehensive target anomaly observation data.
[0096] As one implementation, the observable data includes indicator data, log data, and tracking data. Correspondingly, the initial anomaly observation data includes indicator data, log data, and tracking data. Therefore, the step of determining the target anomaly observation data based on the initial anomaly observation data and the resource tag information corresponding to the initial anomaly observation data can be further optimized into the following steps:
[0097] d11) Obtain the first indicator data and the target resource label information of the first indicator data contained in the initial anomaly observation data, and use the first indicator data as the target indicator data.
[0098] The first indicator data can be considered as the data of the indicator type in the initial anomaly observation data. The target resource tag information can be understood as the resource tag information corresponding to the first indicator data.
[0099] In this embodiment, indicator data is extracted from each initial anomaly observation data and recorded as first indicator data. Target resource tag information corresponding to each first indicator data is then identified and extracted from the corresponding fields of each first indicator data.
[0100] d12) Determine the timestamp information of the target indicator data, and determine the target time window based on the timestamp information and the preset time range.
[0101] In this embodiment, the timestamp of the identified target indicator data is used as a baseline anchor point. Based on a preset time range, a reasonable time range before and after this time point is defined as the target time window. For example, the timestamp information can be the collection time of the target indicator data.
[0102] d13) Based on the target indicator data, the target resource tag information, and the target time window, determine the target tracking data from the first tracking data included in the initial anomaly observation data.
[0103] The first tracking data can be understood as the data of type tracking in the initial anomaly observation data. The target tracking data can be considered as the first tracking data that is directly related to the target indicator data and where anomalies occur within the same target time window (effective time range).
[0104] In this embodiment, the characteristics of the target indicator data can be analyzed, such as high latency and error rate. By analyzing the characteristics of the target indicator data, it is possible to identify which indicators are problematic, and based on the target resource tag information, it is possible to determine where the problem occurs in the target business system (e.g., which resource, which service instance).
[0105] Since the resource tag information in the indicator data and tracking data is determined based on unified resource tags, the target tracking data within the corresponding location and time period can be directly found in the first tracking data using the target resource tag information (location) and the target time window (time period). This enables the correlation of different types of initial anomaly observation data based on common resource tag information. For example, the first tracking data can be the first tracking data containing slow requests or error responses.
[0106] d14) Based on the target tracking data and the target time window, determine the target log data from the first log data contained in the initial anomaly observation data.
[0107] The first log data can be considered as log-type data within the initial anomaly observation data. The target log data can be understood as the first log data associated with the target tracking data and within the same target time window.
[0108] In this embodiment, a key tracking identifier (such as trace_id) of the target tracking data is obtained. The key tracking identifier is used as a precise search key to retrieve the first log data that is associated with the same precise search key and is within the target time window. This first log data is then identified as the target log data.
[0109] d15) The target indicator data, the target tracking data, and the target tracking data are respectively used as target anomaly observation data.
[0110] The technical solution described in this embodiment determines a reasonable target time window by using the time point when the target indicator data becomes abnormal as a reference point, and performs precise time synchronization on abnormal data with different sampling frequencies and transmission delays. Within the target time window, it queries target tracking data that has the same target resource tag information as the target indicator data, and it also queries target log data that has the same key tracking identifier as the target tracking data. The target indicator data, target tracking data, and target tracking data are used as target abnormal observation data, respectively. This provides support for the subsequent construction of an abnormal observation dataset with consistent resource tag information, aligned time windows, and associated abnormal observation data of different types. This ensures that the tracking data and log data analyzed later occur within a valid time range that is directly related to the occurrence of abnormality in the indicator data, avoiding interference from noisy data and laying the foundation for subsequent deep correlation analysis.
[0111] e1) Determine the abnormal observation dataset based on the target abnormal observation data and the target fault information corresponding to the fault event.
[0112] It is understandable that modules or devices that construct fault scenarios and input fault events into the target business system can determine the target fault information corresponding to the fault events based on preset templates or rules, according to the fault scenarios and fault events.
[0113] In this embodiment, the scattered anomaly observation data of each target are integrated by associating and binding them with the target index data. Then, the integration results are aggregated and the target fault information corresponding to the fault event is obtained. The target fault information is associated with the aggregated dataset to form an anomaly observation dataset.
[0114] The above-described technical solution in this embodiment efficiently collects and determines multi-dimensional observable data by actively injecting controllable fault events into a real environment. Based on the anomaly judgment logic, it automatically and accurately identifies the initial anomaly observation data, and automatically and accurately associates and binds the scattered target anomaly observation data belonging to the unified resource instance based on resource tag information. Finally, it combines the accurate target fault information to form an anomaly observation dataset, realizing the automated and high-quality construction of the dataset. This provides strong support for the subsequent accurate evaluation of the deep correlation analysis and fault diagnosis capabilities of the fault diagnosis model to be evaluated, and can provide a more suitable and realistic training sample set for the training and optimization of the fault diagnosis to be evaluated.
[0115] Figure 2 This is a schematic diagram of the structure of a model capability assessment device provided in an embodiment of the present invention. Figure 2As shown, the device includes: an acquisition module 21, a prediction module 22, an effect value determination module 23, and an ability score determination module 24, wherein,
[0116] The acquisition module 21 is used to acquire an abnormal observation dataset under fault scenarios, where each fault scenario includes at least one fault event;
[0117] Prediction module 22 is used to analyze the abnormal observation dataset through the fault diagnosis model to be evaluated to determine the predicted fault information;
[0118] The effect value determination module 23 is used to determine the root cause analysis effect value of the fault scenario based on the predicted fault information and the target fault information.
[0119] The capability score determination module 24 is used to determine the capability score of the fault diagnosis model to be evaluated based on the root cause analysis effect value and the corresponding target weight of each fault scenario in the fault scenario set; the target weight is determined based on the number of fault events corresponding to each fault scenario and the nonlinear adjustment factor fault scenario corresponding to each number of fault events.
[0120] The model capability assessment device provided in this invention acquires anomaly observation datasets under fault scenarios, each fault scenario including at least one fault event; analyzes the anomaly observation datasets using the fault diagnosis model to be evaluated to determine predicted fault information; determines the root cause analysis effect value of the fault scenario based on the predicted fault information and target fault information; and determines the capability score of the fault diagnosis model to be evaluated based on the root cause analysis effect values of each fault scenario in the fault scenario set and their corresponding target weights. Using this device, the root cause analysis effect value of the fault scenario is determined based on the predicted fault information and actual target fault information of the fault diagnosis model to be evaluated relative to the anomaly observation dataset. A dynamic weighting mechanism is introduced to integrate and consider the differences in importance and representativeness of different fault scenarios, and the root cause analysis effect values of multiple fault scenarios are aggregated for analysis. This overcomes the limitations of traditional averaging methods, determining a capability score that objectively and accurately reflects the comprehensive diagnostic capability of the fault diagnosis model to be evaluated. This improves the authenticity and credibility of the capability score, making it closer to actual fault diagnosis needs and effectively avoiding biases from subjective human evaluation or single fault scenario assessment.
[0121] Furthermore, the acquisition module 21 may specifically include:
[0122] The acquisition unit is used to input the fault events included in the fault scenario into the target business system and acquire system operation data under the fault scenario from the target business system.
[0123] The observation data and tag information determination unit is used to determine the observable data under the fault scenario based on the system operation data, and to determine the resource tag information of the observable data according to the preset resource tags and identification rules;
[0124] An anomaly determination unit is used to determine the target anomaly determination conditions corresponding to the observable data according to the pre-constructed anomaly determination logic, and to determine the observable data that meets the target anomaly determination conditions as the initial anomaly observation data.
[0125] An abnormal data determination unit is used to determine target abnormal observation data based on the initial abnormal observation data and the resource tag information corresponding to the initial abnormal observation data;
[0126] An abnormal observation dataset determination unit is used to determine the abnormal observation dataset based on the target abnormal observation data and the target fault information corresponding to the fault event.
[0127] Furthermore, the acquisition module 21 may further include a logic construction unit, which may specifically be used for:
[0128] Define the data type of the data to be processed according to the business scenario, and use the data type as the first-level classification criterion;
[0129] Based on the business scenarios and business experience, define the abnormal feature types under each data type, and use the abnormal feature types as the basis for the second-level classification.
[0130] Based on the business scenario and the data characteristics of the data to be processed, define the anomaly determination conditions for each anomaly feature type;
[0131] Anomaly determination logic is constructed based on the first-level classification criteria, the second-level classification criteria, and the anomaly determination conditions.
[0132] Furthermore, the observable data includes indicator data, log data, and tracking data.
[0133] Furthermore, the abnormal data determination unit can specifically be used for:
[0134] Obtain the first indicator data and the target resource tag information of the first indicator data contained in the initial anomaly observation data, and use the first indicator data as the target indicator data.
[0135] Determine the timestamp information of the target indicator data, and determine the target time window based on the timestamp information and a preset time range;
[0136] Based on the target indicator data, the target resource tag information, and the target time window, target tracking data is determined from the first tracking data contained in the initial anomaly observation data;
[0137] Based on the target tracking data and the target time window, the target log data is determined from the first log data contained in the initial anomaly observation data;
[0138] The target indicator data, the target tracking data, and the target tracking data are respectively used as target anomaly observation data.
[0139] Furthermore, the effect value determination module 23 can specifically be used for:
[0140] The quotient is calculated by dividing the number of predicted fault events in the predicted fault information by the number of target fault events in the target fault information, and the quotient value is determined as the root cause analysis effect value.
[0141] Furthermore, the target weight is the ratio of the first quantity influence value corresponding to the fault scenario to the total quantity influence value corresponding to all fault scenarios;
[0142] The first quantity influence value is the product of the number of fault events in the fault scenario and the nonlinear adjustment factor corresponding to the number of fault events;
[0143] The total quantity impact value is the sum of the product of the number of fault events corresponding to each fault scenario and the nonlinear adjustment factor corresponding to the number of fault events.
[0144] The model capability assessment device provided in the embodiments of the present invention can execute the model capability assessment method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method.
[0145] Figure 3 A schematic diagram of an electronic device 30 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0146] like Figure 3As shown, the electronic device 30 includes at least one processor 31 and a memory, such as a read-only memory (ROM) 32 or a random access memory (RAM) 33, communicatively connected to the at least one processor 31. The memory stores computer programs executable by the at least one processor. The processor 31 can perform various appropriate actions and processes based on the computer program stored in the ROM 32 or loaded from storage unit 38 into the RAM 33. The RAM 33 can also store various programs and data required for the operation of the electronic device 30. The processor 31, ROM 32, and RAM 33 are interconnected via a bus 34. An input / output (I / O) interface 35 is also connected to the bus 34.
[0147] Multiple components in electronic device 30 are connected to I / O interface 35, including: input unit 36, such as keyboard, mouse, etc.; output unit 37, such as various types of monitors, speakers, etc.; storage unit 38, such as disk, optical disk, etc.; and communication unit 39, such as network card, modem, wireless transceiver, etc. Communication unit 39 allows electronic device 30 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0148] Processor 31 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 31 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 31 performs the various methods and processes described above, such as model capability evaluation methods.
[0149] In some embodiments, the model capability assessment method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 38. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 30 via ROM 32 and / or communication unit 39. When the computer program is loaded into RAM 33 and executed by processor 31, one or more steps of the model capability assessment method described above may be performed. Alternatively, in other embodiments, processor 31 may be configured to perform the model capability assessment method by any other suitable means (e.g., by means of firmware).
[0150] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0151] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0152] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0153] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0154] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0155] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0156] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0157] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A model capability evaluation method characterized by, The method comprises: acquiring an abnormal observation data set under a fault scenario, each fault scenario comprising at least one fault event; analyzing the abnormal observation data set by a fault diagnosis model to be evaluated to determine predicted fault information; determining a root cause analysis effect value of the fault scenario according to the predicted fault information and target fault information; determining a capability score of the fault diagnosis model to be evaluated according to the root cause analysis effect value of each fault scenario in a fault scenario set and a corresponding target weight; The target weight is determined according to the number of fault events corresponding to each fault scenario and the nonlinear adjustment factor corresponding to each fault event.
2. The method of claim 1, wherein, The abnormal observation data set under the fault scenario is acquired by: inputting the fault event included in the fault scenario into a target business system, and collecting system running data under the fault scenario from the target business system; determining observable data under the fault scenario according to the system running data, and determining resource tag information of the observable data according to a pre-set resource tag and identification rule; determining a target abnormal judgment condition corresponding to the observable data according to a pre-constructed abnormal judgment logic, and determining the observable data satisfying the target abnormal judgment condition as initial abnormal observation data; determining target abnormal observation data according to the initial abnormal observation data and the resource tag information corresponding to the initial abnormal observation data; determining the abnormal observation data set according to the target abnormal observation data and the target fault information corresponding to the fault event.
3. The method of claim 2, wherein, The construction steps of the abnormal judgment logic comprise: defining the data type of the data to be processed according to the business scenario, and taking the data type as the first classification basis; defining the abnormal feature type under each data type according to the business scenario and business experience, and taking the abnormal feature type as the second classification basis; defining the abnormal judgment condition under each abnormal feature type according to the business scenario and the data characteristics of the data to be processed; based on the first classification basis, the second classification basis and the abnormal judgment condition, the abnormal judgment logic is constructed.
4. The method of claim 2, wherein, The observable data comprises index data, log data and tracking data.
5. The method of claim 4, wherein, The target abnormal observation data is determined according to the initial abnormal observation data and the resource tag information corresponding to the initial abnormal observation data, comprising: acquiring first index data contained in the initial abnormal observation data and target resource tag information of the first index data, and taking the first index data as target index data; determining the timestamp information of the target index data, and determining a target time window according to the timestamp information and a pre-set time range; determining target tracking data from the first tracking data contained in the initial abnormal observation data according to the target index data, the target resource tag information and the target time window; determining target log data from the first log data contained in the initial abnormal observation data according to the target tracking data and the target time window; The target index data, the target tracking data and the target tracking data are respectively taken as target abnormal observation data.
6. The method of claim 1, wherein, The root cause analysis effect value of the fault scene is determined according to the predicted fault information and target fault information, including: The predicted fault event quantity in the predicted fault information is multiplied by the target fault event quantity in the target fault information, and the quotient value is determined as the root cause analysis effect value.
7. The method of claim 1, wherein, The target weight is a ratio of a first quantity influence value corresponding to the fault scene and a total quantity influence value corresponding to all fault scenes. The first quantity influence value is a product value of the fault event quantity of the fault scene and a nonlinear adjustment factor corresponding to the fault event quantity. The total quantity influence value is a sum of product values of the fault event quantities corresponding to each fault scene and the nonlinear adjustment factor corresponding to the fault event quantity.
8. A model capability evaluation apparatus characterized by comprising: It comprises: An acquisition module is configured to acquire an abnormal observation data set under a fault scene, each fault scene including at least one fault event; A prediction module is configured to analyze the abnormal observation data set by a fault diagnosis model to be evaluated to determine predicted fault information; An effect value determination module is configured to determine a root cause analysis effect value of the fault scene according to the predicted fault information and target fault information; An ability score determination module is configured to determine an ability score of the fault diagnosis model to be evaluated according to the root cause analysis effect value of each fault scene in a fault scene set and a corresponding target weight. The target weight is determined according to the fault event quantity corresponding to each fault scene and the nonlinear adjustment factor corresponding to each fault event quantity.
9. An electronic device, comprising: The electronic device comprises: At least one processor; and a memory connected in communication with the at least one processor; The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the model capability evaluation method of any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to execute the model capability evaluation method of any one of claims 1-7 when executed.