Abnormal attribution method and device, equipment, storage medium and product
By constructing a knowledge graph and Markov decision process model, anomaly attribution text is automatically generated, which solves the problem of low automation of anomaly attribution in existing technologies, realizes rapid location of problem points, and improves system stability.
Patent Information
- Application Number
- CN202510813188.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies have a low degree of automation in the platform anomaly attribution process and rely on the personal experience of operation and maintenance personnel, resulting in time-consuming and inefficient anomaly attribution.
Build a knowledge graph corresponding to operation and maintenance data, obtain target exception requests and request link data, build an exception attribution method based on the Markov decision process model, and automatically generate attribution text to quickly locate problem points.
It improves the automation and efficiency of anomaly attribution, and enhances system stability and fault location speed.
Smart Images

Figure CN120705505A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and more specifically, to an anomaly attribution method, apparatus, device, storage medium, and product. Background Art
[0002] General-purpose platform products often support multiple different business lines simultaneously, and their service stability risks have a significant amplification effect - a single point of failure may cause large-scale impacts across businesses. Therefore, early detection, rapid location and repair of problems are extremely important for platform stability.
[0003] Currently, platform anomalies are primarily attributed through the following methods: First, multi-dimensional data collection is performed. Combined with complex monitoring rules, the collected operational data is tested in real time or on a scheduled basis to determine whether alarm conditions are met. If the alarm conditions are met, an alarm text is generated and the alarm message in the text is notified to the on-duty personnel via SMS, WeChat, or smart calls.
[0004] However, the above method only supports the notification of alarm messages. The subsequent attribution of anomalies in the alarm messages relies heavily on the personal experience of the operation and maintenance personnel, resulting in problems such as low automation and time-consuming anomaly attribution. Summary of the Invention
[0005] The main purpose of this application is to provide an anomaly attribution method, device, equipment, storage medium and product that can automatically process the generated attribution text to quickly locate the problem point, improve the efficiency of anomaly attribution, and enhance system stability.
[0006] To achieve the above objectives, in a first aspect, the present application provides an abnormality attribution method, comprising:
[0007] Build a knowledge graph corresponding to operation and maintenance data;
[0008] Obtain target exception request and request link data;
[0009] Construct a Markov decision process model based on the knowledge graph and request link data corresponding to the operation and maintenance data;
[0010] Based on the Markov decision process model and the target abnormal request, a target attribution text of the target abnormal request is determined.
[0011] In a second aspect, an embodiment of the present application provides an abnormality attribution device, including:
[0012] Graph construction module, used to build the knowledge graph corresponding to operation and maintenance data;
[0013] Data acquisition module, used to obtain target abnormal request and request link data;
[0014] The model building module is used to build a Markov decision process model based on the knowledge graph corresponding to the operation and maintenance data and the request link data;
[0015] The exception attribution module is used to determine a target attribution text of a target exception request based on a Markov decision process model and the target exception request.
[0016] In a third aspect, an embodiment of the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above methods when executing the computer program.
[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0018] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which implements the steps of any of the above methods when executed by a processor.
[0019] The embodiments of the present application provide an exception attribution method, apparatus, device, storage medium, and product, including: first constructing a knowledge graph corresponding to the operation and maintenance data, then obtaining target exception requests and request link data, and then constructing a Markov decision process model based on the knowledge graph and request link data corresponding to the operation and maintenance data, thereby determining the target attribution text of the target exception request based on the Markov decision process model and the target exception request. The present application analyzes the target exception request through the Markov decision process model, automatically generates attribution text, and further processes the generated attribution text to quickly locate the problem point, improve the efficiency of exception attribution, and enhance system stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings that constitute part of this application are used to provide a further understanding of this application and make other features, objects and advantages of this application more apparent. The illustrative embodiment drawings of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0021] Figure 1 This is a schematic diagram of the structure of an anomaly attribution system provided in an embodiment of the present application;
[0022] Figure 2 This is a schematic diagram of the structure of a knowledge graph construction module provided in an embodiment of the present application;
[0023] Figure 3 This is a schematic diagram of the structure of a request link acquisition module provided in an embodiment of the present application;
[0024] Figure 4 This is a schematic diagram of the structure of an abnormality attribution module provided in an embodiment of the present application;
[0025] Figure 5 This is a structural diagram of an alarm merging module provided in an embodiment of the present application;
[0026] Figure 6 This is an application scenario diagram of an abnormality attribution method provided in an embodiment of the present application;
[0027] Figure 7 This is a flow chart of an abnormality attribution method provided in an embodiment of the present application;
[0028] Figure 8 This is a schematic diagram of the structure of an abnormality attribution device provided in an embodiment of the present application;
[0029] Figure 9 It is a schematic diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0031] The terms "first," "second," "third," "fourth," and so on (if any) in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in orders other than those illustrated or described herein.
[0032] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0033] It should be understood that in the various embodiments of the present application, the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0034] It should be understood that in this application, "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or apparatuses.
[0035] It should be understood that in this application, "multiple" means two or more. "And / or" is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "Contains A, B and C", "Contains A, B, C" means that A, B, and C are all included, "Contains A, B or C" means that one of A, B, and C is included, and "Contains A, B and / or C" means that any one, any two, or any three of A, B, and C are included.
[0036] It should be understood that, in this application, "B corresponding to A," "B corresponding to A," "A corresponds to B," or "B corresponds to A" means that B is associated with A and B can be determined based on A. Determining B based on A does not mean determining B based solely on A; B can also be determined based on A and / or other information. A and B match when the similarity between A and B is greater than or equal to a preset threshold.
[0037] Depending on the context, "if" as used herein may be interpreted as "when" or "when" or "in response to determining" or "in response to detecting."
[0038] The data involved in this application may be data authorized by the tester or fully authorized by all parties. The collection, dissemination, and use of the data shall comply with the relevant laws, regulations, and standards of the relevant countries and regions. The implementation methods / examples of this application may be combined with each other.
[0039] The following specific embodiments are used to describe the technical solution of the present application in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0040] First, to facilitate understanding of the solution of this application, the terms in this solution are explained as follows:
[0041] RCA (Root Cause Analysis): A systematic problem-solving approach designed to identify the underlying causes (rather than the symptoms) of a problem or incident. The core idea is to analyze the cause-and-effect chain layer by layer to identify the source of the problem, thereby developing effective corrective measures to prevent recurrence.
[0042] TARCA (Traceable Automated Root Cause Analysis, intelligent attribution system):
[0043] A root cause analysis system that combines automation technology and traceability. It uses machine learning, data mining, or causal reasoning algorithms to automatically analyze abnormal events in complex systems (such as industrial equipment and IT infrastructure) and generate explainable root cause reports.
[0044] MDP (Markov Decision Process): A mathematical framework for describing sequential decision problems. Based on the Markov property (i.e., future states depend only on the current state, not the past state), it optimizes long-term cumulative rewards through policies (such as dynamic programming and reinforcement learning).
[0045] Next, the present application solution will be described through specific embodiments with reference to the accompanying drawings.
[0046] General-purpose platform products often support multiple different business lines simultaneously, and their service stability risks have a significant amplification effect - a single point of failure may cause large-scale impacts across businesses. Therefore, early detection, rapid location and repair of problems are extremely important for platform stability.
[0047] Currently, platform anomalies are mainly attributed through the following methods, including the following two methods.
[0048] Method 1: The industry's conventional O&M approach involves building a monitoring, alerting, and downgrade system. This system collects data to establish monitoring indicators across various dimensions. This system then uses an alerting mechanism to promptly notify on-duty personnel of any issues. Ultimately, the impact of the issue is mitigated through downgrade and rollback measures. After the issue is identified and addressed, the system is re-iterated and released.
[0049] Method 2 involves multi-dimensional data collection. Combined with complex monitoring rules, the collected O&M data is tested in real time or on a scheduled basis to determine whether alarm conditions are met. If the alarm conditions are met, an alarm text is generated and the alarm message in the alarm text is notified to the on-duty personnel via SMS, WeChat, smart calls, etc.
[0050] However, a complete problem-handling process should include three steps: problem occurrence, problem location, and problem resolution. However, the aforementioned methods are relatively ineffective in problem location. They only support notification of alarm messages, and subsequent attribution of anomalies in alarm messages relies heavily on the individual experience of operations and maintenance personnel, resulting in low automation and time-consuming anomaly attribution.
[0051] To solve the above problems, this application proposes an abnormality attribution method.
[0052] See also Figure 1 , Figure 1 A schematic diagram of the structure of an abnormality attribution system provided in an embodiment of the present application.
[0053] The audio and sound effect processing system of the present application includes a knowledge graph construction module, a request link collection module, an anomaly attribution module and an alarm merging module.
[0054] Knowledge graph construction module: Using metadata-driven dynamic topology modeling technology, a service-resource global knowledge graph is constructed. Figure 2 As shown, metadata analysis is first performed on the operation and maintenance data, and then a knowledge graph is constructed. The nodes in the knowledge graph are configured with corresponding status data, such as component status and component dependency status, which can be refreshed in real time. The knowledge graph constructed in this application fully depicts the bidirectional mapping relationship between resources and the PageRank algorithm, and quantifies the weight of each link, thereby achieving intelligent topological sorting of service call chains and accurate analysis of fault propagation paths.
[0055] Request link collection module: through Java Agent probe (ie Figure 3 The Agent probe shown in the figure samples key indicators of each service (such as service A and service B) to obtain request link data, and adopts a heartbeat trigger mechanism to periodically upload the request link data in the cache to the monitoring server in batches.
[0056] Abnormal attribution module: such as Figure 4 As shown in Figure 1, this module integrates request link data, resource middleware status data in the knowledge graph, and application server status data, achieving minute-level multi-source data alignment and constructing unified feature value data. It then performs multiple modeling of the state space, action space, and reward function of the request based on the Markov decision process (MDP).
[0057] Alarm merging module: Figure 5 As shown in FIG, this module uses a Markov decision process model to infer the target abnormal request, outputs the initial attribution text, and merges the initial attribution text to generate the target attribution text.
[0058] See also Figure 6 , Figure 6 A schematic diagram of an application scenario of an anomaly attribution method provided in an embodiment of the present application.
[0059] The server 602 communicates with the client 604. The server 602 receives the operation and maintenance data, then constructs a knowledge graph corresponding to the operation and maintenance data, obtains the target abnormal request and request link data, and then constructs a Markov decision process model based on the knowledge graph corresponding to the operation and maintenance data and the request link data. Based on the Markov decision process model and the target abnormal request, the target attribution text of the target abnormal request is determined.
[0060] The client 604 and the server 602 can communicate through any communication method, including but not limited to network communication, and the above-mentioned network can include but not limited to: wired network, wireless network, wherein the wired network includes: local area network, metropolitan area network and wide area network, and the wireless network includes: Bluetooth, WIFI, and other networks that realize wireless communication. The client 604 includes but is not limited to at least one of the following: mobile phone (such as Android phone, iOS phone, etc.), laptop computer, tablet computer, PDA, mobile Internet device (Mobile Internet Device, MID), PAD, desktop computer, smart TV, etc. The server 602 can be an on-site server or a remote server, wherein both the on-site server and the remote server can be implemented as independent servers or a service cluster composed of multiple servers. The above is only an example, and no limitation is made to this in this embodiment.
[0061] See also Figure 7 , Figure 7 This is a flow chart of an abnormality attribution method provided in an embodiment of the present application. Figure 7 As shown, this method is applied to Figure 6 The server 602 shown includes the following steps:
[0062] Step S701: Construct a knowledge graph corresponding to the operation and maintenance data.
[0063] To construct the knowledge graph corresponding to the operation and maintenance data, you need to first obtain the operation and maintenance data, then perform metadata analysis on the operation and maintenance data to obtain multidimensional metadata, and then construct the knowledge graph corresponding to the operation and maintenance data based on the multidimensional metadata.
[0064] Among them, based on the multidimensional metadata, a knowledge graph corresponding to the operation and maintenance data is constructed, including: cleaning the multidimensional metadata to obtain the cleaned multidimensional metadata; performing entity recognition, relationship extraction and attribute extraction on the cleaned multidimensional metadata to obtain service instances, service dependencies and resource parameters corresponding to the service instances; using service instances as nodes and service dependencies as edges to construct a knowledge graph corresponding to the operation and maintenance data, wherein the knowledge graph corresponding to the operation and maintenance data maps the correspondence between service instances and the resource parameters corresponding to the service instances.
[0065] For example, an e-commerce company needs to obtain the operation and maintenance data of its e-commerce platform. This data comes from a wide range of sources, including server log files, which record the server's operating status, the processing of user requests, and other information; monitoring system data, such as server CPU usage, memory usage, network bandwidth, and other indicators; and application configuration files, which contain application service parameter settings and other content. For example, from the server log, you can obtain the time records of users visiting different pages, the response time of the request, and the error messages that occur; from the monitoring system, you can get specific values such as the server CPU utilization percentage per minute and the remaining memory space; from the configuration file, you can see the database connection parameters, the port number of the application service, and other settings.
[0066] Then metadata analysis is performed on the collected e-commerce operation and maintenance data to extract multidimensional metadata.
[0067] Time-dimensional metadata: This metadata is extracted from timestamps in server logs and monitoring data. For example, for user access records in server logs, the precise time of each access is recorded, including the year, month, day, hour, minute, and second. This allows us to understand the temporal distribution of operational data and analyze operational status trends over different time periods. For example, during an e-commerce promotion, time-dimensional metadata can clearly show the dramatic changes in server traffic and resource usage before and after the event.
[0068] Spatial metadata: This identifies the physical location or network topology of servers and application services. For example, a company's servers may be distributed across different data centers, each responsible for different business modules. For an e-commerce application, the front-end display service may be in one data center, while the back-end database service is in another. Spatial metadata can be used to understand the data transmission path within the network and the interactions between servers in different locations, helping to analyze potential issues such as network latency.
[0069] Content-dimensional metadata: This analyzes the specific content types involved in operational data. Examples include request types in server logs (page access requests, API call requests, etc.), resource types in monitoring data (CPU, memory, disk, etc.), and parameter categories in configuration files (database parameters, network parameters, etc.). For e-commerce applications, content-dimensional metadata can help differentiate operational data corresponding to different business functions, such as access requests to product display pages and API calls for adding items to a shopping cart, enabling more targeted operational analysis.
[0070] After obtaining the multidimensional metadata, you need to clean it first to obtain the cleaned multidimensional metadata. The following is the specific cleaning method:
[0071] Data deduplication: Duplicate request records may exist in server logs, especially in high-concurrency scenarios. For example, a user rapidly clicking the same link multiple times may generate nearly identical request records. Duplicate records need to be removed using a unique request identifier (such as a combination of the requesting IP address, the requesting timestamp accurate to the millisecond, and the requested resource path), and only one copy is required.
[0072] Correcting Erroneous Data: Monitoring data may contain erroneous values due to sensor failure or data transmission issues. For example, CPU usage may appear negative or exceed 100%, which is clearly unreasonable. Correcting these erroneous values can be done based on the normal range of historical data and the context of the data. Negative values can be corrected to 0. Values exceeding 100%, if they occur only occasionally and only slightly, may be due to data accuracy issues and can be capped at 100%. However, if values exceeding 100% are frequent and significant, further investigation of the monitoring system's malfunction is necessary.
[0073] Data completion: Some parameters in the configuration file may not be fully documented due to omission or default settings. For example, the timeout parameters for some application services may not be clearly stated. In this case, data completion is required based on industry standards or the application's default configuration to ensure metadata integrity.
[0074] After cleaning the multidimensional metadata, it is necessary to perform entity recognition, relationship extraction and attribute extraction on the cleaned multidimensional metadata. Specifically,
[0075] 1. Entity Recognition
[0076] Service instance identification: In e-commerce operations, service instances include various business services on the e-commerce platform, such as product display services, user authentication services, and order processing services. Service instances can be identified by analyzing the service name and port number in the configuration file, as well as the request processing service identifier in the server log. For example, if you see the service name "product-display-service" in the configuration file and the corresponding port number is 8080, and if you also see a large number of requests being routed to port 8080 and processed by "product-display-service" in the server log, you can confirm that this is a service instance.
[0077] Resource parameter identification: Resource parameters include server hardware resources (such as CPU, memory, and disk) and service software resources (such as database connection pool size and thread pool size). Server hardware resource parameters can be identified from monitoring data, such as Server A's 8 CPU cores and 32GB of memory. Service software resource parameters can be identified from configuration files, such as the maximum number of connections in the order processing service's database connection pool being 50.
[0078] 2. Relationship Extraction
[0079] Service dependency extraction: Service dependencies are extracted by analyzing the request processing flow in server logs and the service call relationships in configuration files. For example, in an e-commerce application, a user's order must first be verified by the user authentication service, which then calls the order processing service to create the order. The order processing service, in turn, calls the inventory management service to check product availability. Based on these call sequences and dependency relationships, it can be determined that the user authentication service depends on the order processing service, which in turn depends on the inventory management service.
[0080] 3. Attribute extraction
[0081] Extracting service instance attributes: In addition to the service instance itself, its attributes also need to be extracted. For example, the service instance's version number, deployment location (server or data center), and startup time can be obtained from the configuration file or the server's deployment management tool. The service instance's startup time can be determined from the server log by viewing the timestamp of the log record generated when the service started.
[0082] Resource parameter attribute extraction: For resource parameters, their attributes need to be extracted, such as resource usage thresholds and update frequencies. For example, a monitoring system might update data on a server's CPU usage every minute and have a usage threshold (such as 80%). When CPU usage exceeds this threshold, an alert is triggered. These attributes can be extracted from the business's operational and maintenance policy documents or the monitoring system's configuration parameters.
[0083] Finally, the identified service instances are treated as nodes in the knowledge graph. For example, node A is the user authentication service instance, node B is the order processing service instance, and node C is the inventory management service instance. Then, based on the service dependencies extracted earlier, edges are established between the nodes. For example, a directed edge is established between the user authentication service instance (node A) and the order processing service instance (node B), indicating that the user authentication service depends on the order processing service. A directed edge is also established between the order processing service instance (node B) and the inventory management service instance (node C), indicating that the order processing service depends on the inventory management service. Furthermore, the knowledge graph should also reflect the correspondence between service instances and resource parameters. For example, the user authentication service instance (node A) corresponds to the CPU resources of server D (including parameters such as the number of cores) and to the database connection pool size parameters of the service itself. These correspondences are indicated in the knowledge graph using dashed lines or special markers. This allows operations personnel to intuitively view the dependencies between service instances and the resources used by the service instances, providing powerful support for subsequent operations analysis, troubleshooting, and performance optimization.
[0084] Step S702: Obtain target abnormal request and request link data.
[0085] Among them, obtaining the target exception request and request link data includes: obtaining the initial exception request; filtering the initial exception request to obtain the target exception request; obtaining the key indicators corresponding to the target exception data, and sampling the key indicators of each service through Java probes to obtain request link data.
[0086] Based on the above embodiment, obtaining the initial abnormal request requires first obtaining the monitoring system alarm record: the e-commerce platform's monitoring system monitors server performance indicators in real time, such as response time and throughput. When the response time of a service exceeds a set threshold (such as 2 seconds) or the throughput is lower than the normal value, the monitoring system automatically records the relevant request as the initial abnormal request. For example, during a promotional event, the monitoring system recorded multiple requests for slow loading of product detail pages. These requests contained information such as user account, product ID, and request time.
[0087] The e-commerce platform's server logs record every request processing step in detail. The operations team uses log analysis tools to set filtering rules, such as error codes (e.g., 500 indicates an internal server error) and request status (failure status), to scan the log files. For example, they filter out purchase requests with a 500 error code and obtain information such as the request path and parameters, which they use as the initial abnormal request.
[0088] At the same time, the customer service department receives daily user feedback regarding issues like failed purchases and slow page loads. Dedicated personnel organize and analyze these responses, extracting information related to the requests involved. For example, if a user describes a purchase page that never loads, the corresponding request might include specific steps and browser type, and then be included in the initial abnormal request set.
[0089] After obtaining the initial abnormal request, it is necessary to filter the initial abnormal request to obtain the target abnormal request. Specifically,
[0090] Deduplication: Monitoring systems, user feedback, and logging may capture multiple instances of the same abnormal request. For example, a purchase request that times out may trigger both a monitoring alert and a user complaint, which may also be recorded in the logs. We use the request's unique identifier (such as the request's order number, user account, and request time) to identify duplicate requests, remove them, and retain only one.
[0091] Excluding known normal requests: Some requests, while appearing abnormal based on certain metrics, may actually be normal. For example, during a system upgrade, some user-initiated requests may time out, but this is due to normal delays in the upgrade process, not a system failure. Based on a pre-defined list of rules, these known normal requests are removed from the initial abnormal requests.
[0092] Business rule-based screening: Develop appropriate business rules for screening based on the business characteristics and transaction processes of the e-commerce platform. For example, even if a timeout occurs, purchase requests for very small amounts (e.g., less than 1 yuan) may have little impact on the business. Setting a threshold for the amount can filter these requests out. Alternatively, test requests related to non-core business functions can be excluded from being considered abnormal requests.
[0093] Afterwards, key indicators corresponding to the target abnormal data are obtained, including the following indicators:
[0094] Performance metrics: For target abnormal requests, key performance-related metrics are extracted from the monitoring system and server logs, such as request response time, throughput, and resource usage (CPU usage, memory usage, etc.). For example, for a specific target abnormal product purchase request, the response time is 10 seconds, which is outside the normal range. At the same time, the server CPU usage is recorded as 95% and the memory usage is 85%.
[0095] Business indicators: Key business indicators are obtained from the business database and business logic of the e-commerce platform, such as order creation success rate, order payment success rate, and shopping cart addition success rate. For example, for orders corresponding to target abnormal requests, the order creation success rate is 70%, which is lower than the normal level.
[0096] Finally, Java probes are used to sample key indicators of each service.
[0097] Install and configure Java probes in each microservice of the e-commerce platform (such as the product service, order service, and payment service). These probes are tightly integrated with the service code. Using technologies such as bytecode enhancement, they monitor and collect data from key business code snippets without modifying the code logic. For example, in the product service's product details display method, configure a probe to collect information such as the display time and return result status.
[0098] Set a reasonable sampling frequency and sampling range based on business needs and system performance requirements. For services in key transaction processes, the sampling frequency can be set to once per second or more frequently to obtain detailed data; for some non-core services, the sampling frequency can be appropriately reduced, such as once every 10 seconds. During the sampling process, focus on key indicators related to target abnormal requests. For example, in payment services, for transactions using the same payment method as the target abnormal order (such as Alipay), collect information such as the time required for transaction verification steps and the response time for connecting to third-party payment interfaces.
[0099] When a target exception request occurs, the Java probe collects key metrics from each service in real time, including service processing time, call counts, and error counts. This data is sent to backend data processing and storage systems via a distributed message queue or directly written to a database, providing a foundation for subsequent request chain data analysis.
[0100] By analyzing collected key indicator data, particularly the call sequence and dependencies between services, we can identify the request chain for the target abnormal request across the entire e-commerce platform. For example, a target abnormal request begins with a user initiating a purchase request, first calling the product service to query product details, then calling the order service to create the order, which in turn calls the payment service to process the payment, and finally calls the inventory service to update the inventory status. This constitutes a complete request chain.
[0101] The key indicator data of each service is integrated according to the order of the request link to form a clear request link data report. This report details the key indicator values of each service node, such as the processing time, response status, resource usage, etc. of each service, as well as the call relationship and sequence between them. For example, the report shows that the target abnormal request takes 3 seconds for the product service, 4 seconds for the order service (of which 3 seconds are spent on calling the payment service), and 1 second for the inventory service. Through this data, operation and maintenance personnel can intuitively understand the performance bottlenecks of the entire request link, providing a strong basis for subsequent troubleshooting and performance optimization.
[0102] Step S703: Construct a Markov decision process model based on the knowledge graph and request link data corresponding to the operation and maintenance data.
[0103] To construct a Markov decision process model based on the knowledge graph and request link data corresponding to the operation and maintenance data, it is necessary to first calculate the state space based on the knowledge graph, then generate a set of candidate actions through historical operation and maintenance abnormal operation records, and construct an action space based on the candidate action set. Then, based on the state space and action space, construct a reward function, and thus construct a Markov decision process model based on the state space, action space and reward function.
[0104] Among them, based on the knowledge graph, the state space is calculated, including: obtaining state data corresponding to the nodes in the knowledge graph, wherein the state data includes component state and component dependency state; constructing component performance indicators from request link data; aligning the component state, component dependency state and component performance indicators corresponding to the nodes in the knowledge graph to obtain characteristic values; generating a state value list based on the characteristic values and error rates, and constructing the state space from the state value list.
[0105] Specifically, the status values in the status value list are calculated as follows:
[0106]
[0107] Among them, S i Represents the state value of the current component (i-th component), S j Represents the state value of the dependent component (j-th component), Prefi Represents the characteristic value of the i-th component, ErrorRate i represents the error rate of the i-th component, ω ij Indicates the dependency weights of the current component and the dependent components, d ij Indicates the path length between the current component and the dependent component, D represents the dependent component set, β and γ are adjustable coefficients, which are set according to specific circumstances and are not specifically limited here.
[0108] Exemplarily, based on the above embodiment, to calculate the state space based on the knowledge graph, it is necessary to first obtain the state data corresponding to the nodes in the knowledge graph.
[0109] In the knowledge graph of an e-commerce platform, nodes mainly represent various service instances, such as product display services, user authentication services, order processing services, payment services, etc. State data includes component status and component dependency status.
[0110] Component status, such as the product display service, can be classified as normal, stuck, or crashed. User authentication service status can be classified as normal, slow response, or high authentication failure rate. These statuses are summarized based on historical operation and maintenance data and monitoring indicators, reflecting the operational status of each service.
[0111] For example, if the product display service depends on the database service, if the database service is slow to respond, the product display service's component dependency status will be "dependency service response is slow." The order processing service depends on the inventory service and user authentication service. Its component dependency status will be determined by the status of the inventory service and user authentication service, such as "dependency on the inventory service is normal, dependency on the user authentication service is slow to respond."
[0112] Component performance metrics are constructed from request link data. Based on the request link data of e-commerce platforms, various component performance metrics can be constructed. For example, for product display services, performance metrics include page load time (average, maximum, etc.), number of requests per second, and percentage of error requests. These performance metrics can quantitatively evaluate the performance of each service and are calculated by analyzing information such as response time and request status in request link data.
[0113] Then, the component status, component dependency status, and component performance indicators corresponding to the nodes in the knowledge graph are aligned to obtain the characteristic value. Taking the product display service as an example, its component status (such as normal), component dependency status (such as normal database service), and component performance indicators (such as an average page loading time of 1 second and 1000 requests per second) are integrated. Certain algorithms or rules can be used, such as quantizing and encoding the component status, component dependency status, and component performance indicators separately, and then combining them into a vector form to obtain the characteristic value. This characteristic value can comprehensively reflect the operating characteristics of the product display service at the current moment, that is, it describes the service from multiple dimensions of status, dependency, and performance.
[0114] Finally, based on the eigenvalues and error rates, a list of state values is generated, and a state space is constructed from the list of state values. Specifically, it is assumed that the e-commerce platform has service instance records with multiple different eigenvalues in the historical operation and maintenance data, and the error rate corresponding to each eigenvalue (such as the frequency of service anomalies) is counted. For example, for eigenvalue A (representing a specific combination of component status, component dependency status, and component performance indicators), the corresponding error rate is 5%; the error rate corresponding to eigenvalue B is 20%, and so on. These eigenvalues and their corresponding error rates are combined into a list of state values, and the state value can be expressed as a tuple form of (eigenvalue, error rate). Then, all possible state values are combined to construct a state space. This state space covers various operating states that may occur in various services of the e-commerce platform and their corresponding error conditions.
[0115] To generate a set of candidate actions and construct an action space based on historical abnormal operation records, we first need to collect historical abnormal operation records. Within the e-commerce platform's historical operation records, we collect all operation records related to service anomalies. For example, in the past, operations personnel may have addressed slow loading of the product display service by restarting the product display service, optimizing database query statements, and increasing server resources. Similarly, for other services such as user authentication and order processing, we also collect corresponding abnormal operation records.
[0116] We then analyze these historical O&M anomaly records, extracting common and effective operations and generating a set of candidate actions. For example, candidate actions might include restarting services, adjusting service configuration parameters (such as adjusting the database connection pool size or modifying caching policies), expanding server resources (such as increasing memory and CPU), and updating service software versions. These actions represent effective measures that O&M personnel might take to resolve service anomalies and represent candidate solutions drawn from historical experience.
[0117] The action space is then constructed from the set of candidate actions obtained above. Specifically, each action has a corresponding parameter range and operation object. For example, the operation object of the action "restarting a service" is a specific service instance, such as a product display service or a user authentication service. The parameter range of the action "adjusting service configuration parameters" might be a database connection pool size between 1 and 100 seconds, a cache expiration time between 1 and 1000 seconds, etc. The action space defines all possible operation choices in the operation and maintenance decision-making process, providing a foundation for subsequent reward function construction and strategy selection.
[0118] Among them, based on the state space and action space, a reward function is constructed, including: forming a state-action pair through parameters in the state space and action space, and configuring corresponding weights for the state-action pair; obtaining historical action effectiveness scores and experience decay factors; based on the state-action pair, the corresponding weights configured for the state-action pair, the historical action effectiveness scores and the experience decay factors, a reward function is constructed.
[0119] For example, in the state space, suppose the state value is (eigenvalue C, error rate 10%), which corresponds to a situation where the product display service page loads slowly and has a high error rate. In the action space, candidate actions include restarting the service and optimizing database queries. Therefore, the state value (eigenvalue C, error rate 10%) and the action (restarting the service) are combined into a state-action pair ((eigenvalue C, error rate 10%), restarting the service). Similarly, other state-action pairs can be combined into ((eigenvalue C, error rate 10%), optimizing database queries). Weights are assigned to these state-action pairs based on historical O&M experience and statistical data. For example, if historical data shows that the probability of resolving the problem after restarting the service is high in the state (eigenvalue C, error rate 10%), then a higher weight, such as 0.8, can be assigned to the state-action pair ((eigenvalue C, error rate 10%), restarting the service). On the other hand, if optimizing database queries is relatively ineffective in this state, a lower weight, such as 0.3, might be assigned. These weights reflect the relative effectiveness of different state-action pairs in solving the operation and maintenance problem.
[0120] Then, the historical action effectiveness score and experience decay factor are obtained. The historical action effectiveness score is mainly obtained by analyzing the historical operation and maintenance records of the e-commerce platform, scoring the effectiveness of each action under different conditions. For example, under multiple similar conditions in the past (such as service anomalies with similar feature values and similar error rates), the number of times the restart service action successfully solved the problem accounted for 80% of the total execution times, so its historical action effectiveness score was 0.8; while the optimization database query action had a success rate of only 50% under the same circumstances, its score was 0.5. These scores are based on historical data statistics and can reflect the effectiveness of the actions in actual applications. Considering that past experience may gradually become inapplicable over time, the experience decay factor is introduced. For example, setting the experience decay factor to 0.9 means that when calculating rewards, the effectiveness of historical experience will be multiplied by 0.9 after each new operation and maintenance decision cycle. This allows the model to place more emphasis on recent experience while not completely ignoring historical experience.
[0121] Finally, based on the state-action pair, the weight corresponding to the state-action pair configuration, the historical action effectiveness score and the experience decay factor, the reward function R(s,a) is constructed, which is specifically expressed as follows:
[0122]
[0123] Among them, f k (s, a) represents the characteristic function of the k-th state-action pair (such as error rate change, performance fluctuation value), λ k represents the k-th feature weight (i.e., the weight corresponding to the k-th state-action pair configuration), g(s,a) represents the effectiveness score of the historical action, and ε represents the experience decay factor.
[0124] Through the above steps, a Markov decision process model was constructed based on the knowledge graph and request link data corresponding to the e-commerce platform's operation and maintenance data. This model helps operations personnel, when faced with service anomalies, select the optimal operation (such as restarting the service or adjusting the configuration) from the action space based on the current state (comprehensive characteristics such as the service's component status, component dependency status, and performance indicators). The model also evaluates the effectiveness of different operation strategies based on a reward function, thereby enabling intelligent operation and maintenance decision-making and improving the stability and reliability of the e-commerce platform.
[0125] This application can realize intelligent operation and maintenance decision-making and improve platform stability and operation and maintenance efficiency by constructing a Markov decision process model.
[0126] Step S704: Determine the target attribution text of the target abnormal request based on the Markov decision process model and the target abnormal request.
[0127] To determine the target attribution text of the target exception request based on the Markov decision process model and the target exception request, it is necessary to first input the request link data and knowledge graph corresponding to the target exception request into the Markov decision process model, output the function value, and then compare the function value with the preset threshold. If the function value and the preset threshold meet the conditions, multiple initial attribution texts are generated, and then the multiple initial attribution texts are merged to obtain the target attribution text of the target exception request.
[0128] In addition, if multiple initial attribution texts are preset attribution texts, where the preset attribution texts are attribution texts generated by the same type of problem within the same time span; construct a multidimensional feature vector, where each feature vector in the multidimensional feature vector is configured with a corresponding weight; calculate the similarity of multiple initial attribution texts based on the multidimensional feature vector, and obtain the initial attribution text with the highest similarity as the target attribution text.
[0129] For example, in an e-commerce platform, for target abnormal requests (for example, user feedback on frequent payment failures when purchasing a popular product), its request link data is collected. These data include detailed information on each service call link from browsing the product page, adding to the shopping cart, submitting the order to the payment process, such as product display service, shopping cart service, order service, payment service, and interaction with the database and cache. At the same time, the knowledge graph corresponding to the operation and maintenance data is obtained, which contains information such as the dependency relationship between various service instances of the e-commerce platform (such as product services, user services, payment services, etc.), resource parameters (such as server CPU, memory usage, etc.), and historical operation and maintenance abnormal operation records.
[0130] The request link data and knowledge graph are then fed into the previously constructed Markov decision process model. The model calculates the state space (including characteristic values such as component states, component dependencies, and component performance metrics for each service, as well as error rates), the action space (a set of candidate O&M actions), and the reward function (which comprehensively considers the weights of state-action pairs, historical action effectiveness scores, and experience decay factors). The model then outputs a function value that reflects the expected reward for taking different O&M actions given the current target exception request. This function value represents an evaluation of the effectiveness of each possible action in resolving the exception request.
[0131] The function value is then compared to a preset threshold. For example, if the model outputs a function value of 0.75, which meets the medium-priority processing threshold (0.5-0.8), the process proceeds to the next step. If the function value is 0.9, it meets the high-priority processing threshold and also proceeds to the next step, but perhaps with a higher priority. If the function value is 0.4, attribution text is temporarily not generated, and further data collection or model adjustments may be required. The preset threshold is set based on specific circumstances and is not specifically defined here.
[0132] In one case, when the function value meets the preset threshold conditions, multiple initial attribution texts are generated based on the expected reward levels of each action calculated by the model. For example: Initial attribution text 1: "The slow response of the payment service resulted in payment failure. It is recommended to optimize the connection performance between the payment service and the third-party payment interface." (The corresponding action is to optimize the connection performance of the payment service, and the function value is higher). Initial attribution text 2: "The order service resource usage is too high, affecting the order submission and payment process. It is recommended to increase the server resources of the order service." (The corresponding action is to expand the order service resources, and the function value is second). Initial attribution text 3: "The database query efficiency is low, affecting the data reading of multiple services, resulting in a jam in the payment process." (The corresponding action is to optimize the database query, and the function value is relatively low but still within the threshold range). Merging the above texts will result in the target attribution text.
[0133] In another case, when the preset attribution text refers to the attribution text generated by the same type of problem within the same time span, for example, during a promotional activity on an e-commerce platform, if multiple service anomalies have occurred due to tight server resources before, and attribution text such as "Insufficient server resources affect the operation of multiple services" has been generated, then the currently generated initial attribution text with similar content can be identified as the preset attribution text.
[0134] Analyze the multiple generated initial attribution texts to see if any of them contain preset attribution texts. If so, construct a multi-dimensional feature vector to calculate the similarity of the initial attribution texts. The dimensions of the feature vector can include issue type (such as service response, resource usage, database performance, etc.), impact scope (number and type of services involved), historical frequency of occurrence, etc. For example:
[0135] Feature dimension 1: Problem type (service response problem weight is 0.4, resource usage problem weight is 0.3, database performance problem weight is 0.3, etc.); Feature dimension 2: Impact scope (impact on core services has a high weight, impact on edge services has a low weight); Feature dimension 3: Historical frequency of occurrence (frequently occurring problems have a higher weight).
[0136] Then, based on the multi-dimensional feature vector, the similarity between the multiple initial attribution texts and the preset attribution text is calculated. For example, a calculation method such as a cosine similarity algorithm is used. Assume that the similarity between initial attribution text 1 and the preset attribution text is 80%, the similarity between initial attribution text 2 and the preset attribution text is 60%, and the similarity between initial attribution text 3 and the preset attribution text is 70%. The initial attribution text with the highest similarity (in this example, initial attribution text 1) is obtained as the target attribution text.
[0137] An embodiment of the present application provides an exception attribution method, comprising: first constructing a knowledge graph corresponding to operation and maintenance data, then obtaining target exception requests and request link data, and then constructing a Markov decision process model based on the knowledge graph corresponding to the operation and maintenance data and the request link data, thereby determining a target attribution text for the target exception request based on the Markov decision process model and the target exception request. The present application analyzes the target exception request using the Markov decision process model, automatically generates attribution text, and further processes the generated attribution text to quickly locate problem points, improve exception attribution efficiency, and enhance system stability.
[0138] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0139] The following are device embodiments of the present application. For details not fully described therein, please refer to the corresponding method embodiments described above.
[0140] Figure 8 A schematic diagram of the structure of an anomaly attribution device provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown. An anomaly attribution device includes a graph construction module 801, a data acquisition module 802, a model construction module 803, and an anomaly attribution module 804, as follows:
[0141] Graph construction module 801, used to construct a knowledge graph corresponding to operation and maintenance data;
[0142] The data acquisition module 802 is used to obtain target abnormal request and request link data;
[0143] Model construction module 803, used to construct a Markov decision process model based on the knowledge graph corresponding to the operation and maintenance data and the request link data;
[0144] The exception attribution module 804 is configured to determine a target attribution text of the target exception request based on the Markov decision process model and the target exception request.
[0145] In one embodiment, the graph construction module 801 is further used to obtain operation and maintenance data;
[0146] Perform metadata analysis on operation and maintenance data to obtain multi-dimensional metadata;
[0147] Based on multi-dimensional metadata, a knowledge graph corresponding to operation and maintenance data is constructed.
[0148] In one embodiment, the graph construction module 801 is further configured to clean the multidimensional metadata to obtain cleaned multidimensional metadata;
[0149] Perform entity recognition, relationship extraction, and attribute extraction on the cleaned multidimensional metadata to obtain service instances, service dependencies, and resource parameters corresponding to the service instances;
[0150] A knowledge graph corresponding to the operation and maintenance data is constructed with service instances as nodes and service dependencies as edges. The knowledge graph corresponding to the operation and maintenance data maps the correspondence between service instances and the resource parameters corresponding to the service instances.
[0151] In one embodiment, the data acquisition module 802 is further configured to acquire an initial abnormal request;
[0152] Filter the initial abnormal request to obtain the target abnormal request;
[0153] Obtain the key indicators corresponding to the target abnormal data, and sample the key indicators of each service through Java probes to obtain request link data.
[0154] In one embodiment, the model building module 803 is further configured to calculate the state space based on the knowledge graph;
[0155] Generate a set of candidate actions through historical operation and maintenance abnormal operation records, and build an action space based on the candidate action set;
[0156] Construct a reward function based on the state space and action space;
[0157] Based on the state space, action space and reward function, a Markov decision process model is constructed.
[0158] In one embodiment, the model building module 803 is further configured to obtain state data corresponding to a node in the knowledge graph, wherein the state data includes component state and component dependency state;
[0159] Build component performance metrics from request link data;
[0160] Align the component status, component dependency status, and component performance indicators corresponding to the nodes in the knowledge graph to obtain feature values;
[0161] Based on the eigenvalues and error rates, a state value list is generated, and a state space is constructed from the state value list.
[0162] In one embodiment, the model building module 803 is further configured to form state-action pairs using parameters in the state space and the action space, and configure corresponding weights for the state-action pairs;
[0163] Get historical action effectiveness scores and experience decay factors;
[0164] Construct a reward function based on state-action pairs, weights corresponding to state-action pair configurations, historical action effectiveness scores, and experience decay factors.
[0165] In one embodiment, the anomaly attribution module 804 is further configured to input the request link data and the knowledge graph corresponding to the target abnormal request into the Markov decision process model and output a function value;
[0166] Compare the function value with a preset threshold;
[0167] If the function value and the preset threshold meet the conditions, multiple initial attribution texts are generated;
[0168] Merge multiple initial attribution texts to obtain the target attribution text of the target abnormal request.
[0169] In one embodiment, the device further includes a text acquisition module, the text acquisition module being configured to: if the plurality of initial attribution texts are preset attribution texts, wherein the preset attribution texts are attribution texts generated by the same type of question within the same time span;
[0170] Constructing a multi-dimensional feature vector, wherein each feature vector in the multi-dimensional feature vector is configured with a corresponding weight;
[0171] The similarity of multiple initial attribution texts is calculated based on the multi-dimensional feature vector, and the initial attribution text with the highest similarity is obtained as the target attribution text.
[0172] An embodiment of the present application provides an exception attribution device that first constructs a knowledge graph corresponding to operation and maintenance data, then obtains target exception requests and request link data, and then constructs a Markov decision process model based on the knowledge graph and request link data corresponding to the operation and maintenance data. This device then determines a target attribution text for the target exception request based on the Markov decision process model and the target exception request. The present application analyzes the target exception request using the Markov decision process model, automatically generates attribution text, and further processes the generated attribution text to quickly locate problem points, improve exception attribution efficiency, and enhance system stability.
[0173] This application Figure 9A schematic diagram of a computer device is provided. Figure 9 As shown, the computer device 9 of this embodiment includes: a processor 901, a memory 902, and steps in the embodiment of the abnormality attribution method stored in the memory 902 and executable on the processor 901, such as Figure 7 Alternatively, when the processor 901 executes the computer program 903, the functions of the modules / units in the above-mentioned embodiments of the abnormality attribution device are realized, for example Figure 8 Functionality of modules / units 801 to 804 shown.
[0174] The present application also provides a readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, it is used to implement the anomaly attribution method provided by the various embodiments described above.
[0175] Among them, the readable storage medium can be a computer storage medium or a communication medium. Communication media include any medium that facilitates the transmission of computer programs from one place to another. Computer storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer. For example, a readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application-specific integrated circuit (ASIC). In addition, the ASIC can be located in a user device. Of course, the processor and the readable storage medium can also exist in a communication device as discrete components. The readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0176] The present application also provides a program product, comprising execution instructions stored in a readable storage medium. At least one processor of a device can read the execution instructions from the readable storage medium, and at least one processor can execute the execution instructions to cause the device to implement the anomaly attribution methods provided in the various embodiments described above.
[0177] In the embodiments of the above-mentioned devices, it should be understood that the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in this application may be directly implemented as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0178] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An abnormality attribution method, characterized in that: include: Build a knowledge graph corresponding to operation and maintenance data; Obtain target abnormal request and request link data; Constructing a Markov decision process model based on the knowledge graph corresponding to the operation and maintenance data and the request link data; Based on the Markov decision process model and the target abnormal request, a target attribution text of the target abnormal request is determined.
2. The abnormality attribution method according to claim 1, characterized in that: The construction of the knowledge graph corresponding to the operation and maintenance data includes: Obtaining the operation and maintenance data; Performing metadata analysis on the operation and maintenance data to obtain multidimensional metadata; Based on the multidimensional metadata, a knowledge graph corresponding to the operation and maintenance data is constructed.
3. The abnormality attribution method according to claim 2, characterized in that: The step of constructing a knowledge graph corresponding to the operation and maintenance data based on the multidimensional metadata includes: Cleaning the multidimensional metadata to obtain cleaned multidimensional metadata; Performing entity recognition, relationship extraction, and attribute extraction on the cleaned multidimensional metadata to obtain service instances, service dependencies, and resource parameters corresponding to the service instances; A knowledge graph corresponding to the operation and maintenance data is constructed with the service instance as a node and the service dependency as an edge, wherein the knowledge graph corresponding to the operation and maintenance data maps the correspondence between the service instance and the resource parameters corresponding to the service instance.
4. The abnormality attribution method according to claim 1, characterized in that: The obtaining of target abnormal request and request link data includes: Get the initial exception request; Filtering the initial abnormal request to obtain a target abnormal request; The key indicators corresponding to the target abnormal data are obtained, and the key indicators of each service are sampled through a Java probe to obtain the request link data.
5. The abnormality attribution method according to claim 1, characterized in that: The constructing of a Markov decision process model based on the knowledge graph corresponding to the operation and maintenance data and the request link data includes: Calculating a state space based on the knowledge graph; Generate a set of candidate actions through historical operation and maintenance abnormal operation records, and construct an action space based on the candidate action set; Constructing a reward function based on the state space and the action space; The Markov decision process model is constructed based on the state space, the action space and the reward function.
6. The abnormality attribution method according to claim 5, characterized in that: The calculating of the state space based on the knowledge graph includes: Obtaining state data corresponding to a node in the knowledge graph, wherein the state data includes component state and component dependency state; constructing component performance indicators from the request link data; Aligning component states, component dependency states, and component performance indicators corresponding to nodes in the knowledge graph to obtain feature values; A state value list is generated based on the feature value and the error rate, and the state space is constructed from the state value list.
7. An abnormality attribution device, characterized in that: include: Graph construction module, used to build the knowledge graph corresponding to operation and maintenance data; Data acquisition module, used to obtain target abnormal request and request link data; A model building module, configured to build a Markov decision process model based on the knowledge graph corresponding to the operation and maintenance data and the request link data; The exception attribution module is used to determine a target attribution text of the target exception request based on the Markov decision process model and the target exception request.
8. A computer device, characterized in that: comprising a memory, and one or more processors communicatively connected to the memory; The memory stores instructions that can be executed by the one or more processors, and the instructions are executed by the one or more processors to enable the one or more processors to implement the abnormality attribution method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The method comprises a program or an instruction, which, when executed on a computer, implements the abnormality attribution method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the abnormality attribution method according to any one of claims 1 to 6.