Model training method and business fault reason determination method

By using reinforcement learning model training methods on the business platform to build a knowledge graph, the problem of low positioning accuracy in large-scale business platform operation and maintenance is solved, and more efficient fault positioning and business stability are achieved.

CN120045375APending Publication Date: 2025-05-27CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510207202.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the operation and maintenance of large-scale business platforms, existing fault diagnosis methods are difficult to effectively process production data with huge data volume, complex and variable causes, resulting in low accuracy of the fault cause positioning.

Method used

The reinforcement learning model training method is adopted to build a knowledge graph by obtaining the historical production data of the business platform, and to determine the cause of business failure using relationship selection, entity selection and fact extraction modules.

Benefits of technology

It improves the accuracy of the location of the root cause of failure, reduces business interruption time, and enhances operation and maintenance efficiency and business stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045375A_ABST
    Figure CN120045375A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method and a business fault reason determination method. The model training method comprises the following steps: constructing a knowledge graph based on historical production data; traversing entities in the knowledge graph, and determining a target relationship corresponding to the current entity; determining a target entity having a target relationship with the current entity; extracting a fact triple of the current entity, and updating the knowledge graph according to the fact triple; obtaining a reward obtained by each module in the reinforcement learning model, and determining a loss function of each module according to the reward; and updating parameters of each module in the reinforcement learning model until the loss function satisfies a preset convergence condition, and obtaining a trained reinforcement learning model. The technical problem of low fault root cause positioning accuracy in large-scale service platform operation and maintenance due to the fact that production data with huge data volume, complex fault causes and variability in the operation and maintenance process is difficult to effectively process in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of communication devices, and more specifically, to a model training method and a method for determining the cause of a service failure. Background Art

[0002] In the current rapid development of informatization and digitalization, the operation and maintenance management of business platforms faces unprecedented challenges. With the wide application of technologies such as cloud computing, big data, and the Internet of Things, the scale and complexity of business platforms have increased rapidly, and the interdependent relationships among various services and components have become increasingly complex. This complexity has led to the diversification and unpredictability of fault causes. Even under the management of professional operation and maintenance teams, the rapid location and resolution of faults have become extremely difficult.

[0003] In the field of business platform operation and maintenance, fault diagnosis mainly relies on the observation and analysis of fault phenomena and the processing of relevant data of the business platform. However, related fault diagnosis methods, such as rule-based expert systems and statistical analysis, have obvious limitations. On the one hand, the rule system requires a large number of rules to be manually written. Facing the constantly changing business environment and fault types, the cost of maintaining and updating rules is extremely high, and it is difficult to cover all possible fault scenarios. On the other hand, although statistical analysis methods can process a large amount of data, they often lack in-depth logical reasoning ability and are difficult to dig out the real cause of faults from the data, especially in the complex event chain before the occurrence of faults.

[0004] In addition, with the explosive growth of the data volume of business platforms, simply counting the warning messages before a fault and setting thresholds can no longer meet the needs of efficient fault location. This threshold-based method ignores the correlation between warning messages and the complexity of the fault occurrence mechanism, often resulting in low accuracy of fault location. Especially for the fault diagnosis of multi-hop relationships, that is, faults caused by a series of interrelated events or performance changes, related technologies are often difficult to effectively handle, easily causing misjudgment or omission of fault causes.

[0005] In summary, related fault diagnosis and root cause location methods have key problems such as low location accuracy and poor processing efficiency when dealing with complex fault scenarios in the operation and maintenance of large-scale business platforms.

[0006] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0007] This application provides a model training method and a method for determining the cause of a service failure, so as to at least solve the technical problem of low accuracy of fault root cause location in the operation and maintenance of large-scale business platforms caused by the fact that related technologies are difficult to effectively process the huge amount of production data with complex and changeable fault causes during the operation and maintenance process.

[0008] According to one aspect of the present application, a model training method is provided, including: obtaining historical production data of a business platform and constructing a knowledge graph based on the historical production data, wherein the entities in the knowledge graph include at least one of the following: devices, time, personnel, and fault phenomena; traversing the entities in the knowledge graph by using a relationship selection module in a reinforcement learning model to determine a target relationship corresponding to the current entity, where the target relationship is the relationship with the largest relevance index among the multiple relationships corresponding to the current entity; determining a target entity having a target relationship with the current entity by using an entity selection module in the reinforcement learning model; extracting a fact triple of the current entity from an external corpus by using a fact extraction module in the reinforcement learning model and updating the knowledge graph according to the fact triple, where the fact triple includes: the current entity, a relationship, and other entities or attributes having a relationship with the current entity; obtaining the rewards obtained by each module in the reinforcement learning model and determining a loss function for each module according to the rewards, where the reward is the reward obtained by each module according to its own state, selecting an action through its own policy, executing the selected action, and the reward is used to evaluate the influence degree of the action on the accuracy of determining the cause of a business fault; updating the parameters of each module in the reinforcement learning model until the loss function meets a preset convergence condition to obtain a trained reinforcement learning model, where the reinforcement learning model is used to determine the cause of a business fault in production data.

[0009] Optionally, the fact extraction module determines the state, action, and reward through the following method: calculating the attention weight between the current entity and other entities having a target relationship with the current entity through a self-attention neural network; determining a target vector representation according to a linear transformation matrix and the attention weight, where the linear transformation matrix is used to align different attention weights to the same space or scale; determining a set of sentences corresponding to the current entity in the external corpus, where each sentence in the set of sentences includes context information related to the current entity; determining the state of the fact extraction module according to the current entity, the set of sentences corresponding to the current entity, and the target vector representation; determining the following as the action of the fact extraction module: predicting the scores of each set of sentences related to the target relationship and determining the set of sentences with the highest score; determining the reward mechanism of the relationship selection module as the reward mechanism of the fact extraction module.

[0010] Optionally, the relationship selection module determines the state, action, and reward through the following method: determining the state of the relationship selection module according to the current entity, the fact triples extracted by the fact extraction module, the query entity, the query relationship, and the query target; determining the following as the action of the relationship selection module: predicting the probability distribution of all relationships existing for the current entity and determining the relationship with the highest probability value; determining the following as the reward of the relationship selection module: whether the relationship with the highest probability value is determined as the target relationship corresponding to the current entity.

[0011] Optionally, predicting the probability distribution of all relationships existing for the current entity includes: determining all relationships selected at the i-th time step according to the relationship selected at the (i - 1)-th time step and the action determined at the (i - 1)-th time step, where i is a positive integer greater than 1; processing all relationships selected at the i-th time step using a multi-layer perceptron layer in a neural network to obtain the probability distribution of all relationships selected at the i-th time step, where the neural network is a feed-forward network including a rectified linear unit non-linear layer.

[0012] Optionally, the entity selection module determines the state, action, and reward through the following method: determining the state of the entity selection module according to the current entity, the fact triples extracted by the fact extraction module, the query entity, the query relationship, and the query target; determining the following as the action of the entity selection module: predicting the probability distribution of all entities having a target relationship with the current entity and determining the entity with the highest probability value; determining the following as the reward of the entity selection module: whether the entity with the highest probability value is determined as the target entity corresponding to the current entity.

[0013] Optionally, determining the following as the reward of the entity selection module: whether the entity with the highest probability value is determined as the target entity corresponding to the current entity includes: obtaining a first reward value when the entity with the highest probability value is determined as the target entity corresponding to the current entity; obtaining a second reward value when an entity other than the entity with the highest probability value is determined as the target entity corresponding to the current entity; obtaining a third reward value when a preset entity is determined as the target entity corresponding to the current entity, where the preset entity is used to indicate that the entity selection module fails to find a corresponding entity within a preset number of steps, the first reward value is greater than the third reward value, and the third reward value is greater than the second reward value.

[0014] Optionally, constructing a knowledge graph based on production data includes: determining, in the schema layer of the knowledge graph, a schema for representing the relationships between upper ontology data, and receiving, in the schema layer, the attributes of the upper ontology data and the association relationships between the upper ontology data, where the upper ontology data includes at least one of the following keywords: service, host, metric, fault, and phenomenon; adding, based on the schema layer, the entities and relationships in the production data to the data layer of the knowledge graph.

[0015] According to another aspect of the present application, a method for determining the cause of a service failure is further provided, including: obtaining service failure data of a service platform in a knowledge graph, and extracting triple data from the service failure data, wherein the entities in the knowledge graph include at least one of the following: device, time, person, failure phenomenon, and the triple data includes: an entity, a relationship, and another entity or attribute corresponding to the entity through the relationship; using a trained reinforcement learning model to determine the cause of the service failure corresponding to the triple data, wherein the reinforcement learning model is obtained by training through the above model training method.

[0016] Optionally, using a trained reinforcement learning model to determine the cause of the service failure corresponding to the triple data includes: using a relationship selection module to select a target relationship based on a state, where the target relationship represents a possible connection between the current entity and an adjacent node; using an entity selection module to determine a target entity based on the target relationship selected by the relationship selection module; using a fact extraction module to extract fact triples of the current entity from an external corpus and update the knowledge graph according to the fact triples; updating the state according to the target relationship and the target entity selected by the relationship selection module and the entity selection module, and recording an inference path until the entities and relationships in the state indicated by the inference path meet a preset matching target.

[0017] According to another aspect of the present application, there is also provided a model training device, including: a first acquisition module, configured to acquire historical production data of a business platform and construct a knowledge graph based on the historical production data, wherein the entities in the knowledge graph include at least one of the following: devices, time, personnel, and fault phenomena; a first determination module, configured to traverse the entities in the knowledge graph by using a relationship selection module in a reinforcement learning model to determine a target relationship corresponding to the current entity, wherein the target relationship is the relationship with the largest relevance index to the current entity among multiple relationships corresponding to the current entity; a second determination module, configured to use an entity selection module in the reinforcement learning model to determine a target entity having a target relationship with the current entity; an extraction module, configured to extract a fact triple of the current entity from an external corpus by using a fact extraction module in the reinforcement learning model and update the knowledge graph according to the fact triple, wherein the fact triple includes: the current entity, a relationship, and another entity or attribute having a relationship with the current entity; a second acquisition module, configured to acquire the rewards obtained by each module in the reinforcement learning model and determine a loss function for each module according to the rewards, wherein the reward is the reward obtained by each module according to its own state, selecting an action through its own policy, executing the selected action, and the reward is used to evaluate the influence degree of the action on the accuracy of determining the cause of a business failure; a training module, configured to update the parameters of each module in the reinforcement learning model until the loss function meets a preset convergence condition, and obtain a trained reinforcement learning model, wherein the reinforcement learning model is used to determine the cause of a business failure in production data.

[0018] According to another aspect of the present application, there is also provided a non-volatile storage medium, where the storage medium includes a stored program, and when the program runs, it controls the device where the storage medium is located to execute the above model training method.

[0019] According to another aspect of the present application, there is also provided an electronic device, including: a memory and a processor, where the processor is configured to run a program stored in the memory, and when the program runs, it executes the above model training method.

[0020] According to another aspect of the present application, there is also provided a computer program, where when the computer program is executed by a processor, it implements the above model training method.

[0021] According to another aspect of the present application, there is also provided a computer program product, where the computer program product includes a non-volatile computer-readable storage medium, and the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above model training method.

[0022] In this application, historical production data of the business platform is obtained, and a knowledge graph based on the historical production data is constructed. Among them, the entities in the knowledge graph include at least one of the following: equipment, time, personnel, and fault phenomena. The relationship selection module in the reinforcement learning model traverses the entities in the knowledge graph to determine the target relationship corresponding to the current entity, where the target relationship is the relationship with the largest relevance index among the multiple relationships corresponding to the current entity. The entity selection module in the reinforcement learning model determines the target entity that has a target relationship with the current entity. The fact extraction module in the reinforcement learning model extracts the fact triples of the current entity from the external corpus and updates the knowledge graph according to the fact triples, where the fact triples include: the current entity, the relationship, and other entities or attributes that have a relationship with the current entity. The rewards obtained by each module in the reinforcement learning model are acquired, and the loss function of each module is determined according to the rewards. The reward is the reward obtained by each module based on its own state, selecting actions through its respective policies, executing the selected actions, and the reward is used to evaluate the influence degree of the action on the accuracy of determining the cause of the business failure. The parameters of each module in the reinforcement learning model are updated until the loss function meets the preset convergence condition, and the trained reinforcement learning model is obtained. The way in which the reinforcement learning model is used to determine the cause of the business failure in the production data achieves the purpose of improving the accuracy of fault root cause location, thus realizing the technical effect of reducing the business interruption time, and further solving the technical problem of low accuracy of fault root cause location in the large-scale business platform operation and maintenance due to the difficulty of effectively processing the huge amount of production data, complex and variable fault causes in the operation and maintenance process in the related technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0024] Figure 1 is a flowchart of a model training method according to an embodiment of the present application;

[0025] Figure 2 is a structural diagram of a model training device according to an embodiment of the present application;

[0026] Figure 3 is a hardware structure block diagram of a computer terminal of a model training method according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0029] According to the embodiments of this application, a method embodiment of a model training method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.

[0030] Figure 1 is a flowchart of a model training method according to the embodiments of this application. As Figure 1 shown, the method includes the following steps:

[0031] Step S101, obtain the historical production data of the business platform, and construct a knowledge graph based on the historical production data, where the entities in the knowledge graph include at least one of the following: equipment, time, personnel, and fault phenomenon.

[0032] For example, production data is obtained from the operation records of a large cloud service platform, including the operation status of devices, the operation records of maintenance personnel, the description of fault phenomena, user feedback, etc. within the past year. The above data is preprocessed, including data cleaning, format unification, and entity recognition, and converted into structured triples, such as ('Device A', 'Abnormal operation', 'Time 1'), ('Person B', 'Execute', 'Operation X'), etc., for constructing a knowledge graph. The entity types in the knowledge graph include devices, time, personnel, fault phenomena, etc., and the relationships between nodes are defined according to the relationships in the triples, such as the relationship between devices and operation status, the relationship between personnel and operations, the relationship between time and events, etc. Through the above method, the knowledge graph not only contains the historical production data of the service platform but also can reflect the complex associations between entities.

[0033] Step S102: Use the relationship selection module in the reinforcement learning model to traverse the entities in the knowledge graph to determine the target relationship corresponding to the current entity, where the target relationship is the relationship with the largest relevance index to the current entity among multiple relationships corresponding to the current entity.

[0034] The relationship selection module is initialized and loaded into the model. At the beginning, the module starts from a starting entity, for example, starting from Device A, and determines the next relationship selection by traversing the outgoing edges of Device A. Each relationship corresponds to a relevance index, which reflects the importance of the relationship in indicating the root cause of the fault. The module selects the relationship corresponding to the maximum value as the target relationship, such as 'Abnormal operation', by comparing the relevance indices of each relationship. This process is carried out under the guidance and control of the model, and the model evaluates the importance of the relationship based on historical data and the current state, thus guiding the further in-depth fault location.

[0035] Step S103: Use the entity selection module in the reinforcement learning model to determine the target entity that has a target relationship with the current entity.

[0036] After determining the target relationship, the entity selection module is responsible for selecting the target entity that has a target relationship with the current entity from the knowledge graph. For example, if the current entity is Device A and the target relationship is 'Abnormal operation', the entity selection module will select the entity that is most likely to cause the abnormal operation from all entities related to Device A, such as Software S or Operator P. This process is also driven by the model, and the model predicts the most likely target entity according to the characteristics of the relationship, the attributes of the entity, and historical fault data, so as to gradually approach the root cause of the fault.

[0037] Step S104: Use the fact extraction module in the reinforcement learning model to extract the fact triples of the current entity from the external corpus, and update the knowledge graph according to the fact triples. The fact triples include: the current entity, the relationship, and other entities or attributes related to the current entity.

[0038] The fact extraction module extracts the fact triples related to the current entity from the external corpus, such as ('Device A', 'installed Software S', 'Time T'), ('Person P', 'performed Operation X', 'Time T'), etc. These fact triples contain not only static information such as device specifications and software versions, but also dynamic information such as operation logs and user feedback. By dynamically adding these triples to the knowledge graph, the content of the graph can be updated in real time to ensure the comprehensiveness and timeliness of the graph information, thereby improving the accuracy and efficiency of fault location.

[0039] Step S105: Obtain the rewards obtained by each module in the reinforcement learning model, and determine the loss function of each module according to the rewards. The reward is the reward obtained by each module based on its own state, selecting an action through its own policy, executing the selected action, and the reward is used to evaluate the influence degree of the action on the accuracy of determining the cause of the business failure.

[0040] During the model training process, the relationship selection module, the entity selection module, and the fact extraction module will all obtain corresponding rewards according to the effects of their actions. The calculation of the reward is based on the influence degree of the action on the accuracy of determining the cause of the business failure. For example, if the entity selection module selects the correct fault source entity, the model will obtain a positive reward; if an irrelevant or incorrect entity is selected, a negative reward will be obtained. The reward mechanisms of the relationship selection module and the fact extraction module are similar, based on whether their actions are helpful for fault location.

[0041] Step S106: Update the parameters of each module in the reinforcement learning model until the loss function meets the preset convergence condition, and obtain the trained reinforcement learning model. The reinforcement learning model is used to determine the cause of the business failure in the production data.

[0042] The model will update its parameters according to the rewards obtained by each module to minimize the loss function. The loss function reflects the difference between the model prediction result and the actual root cause of the failure. Through iterative optimization, the model can adjust its strategy and parameters to obtain more accurate fault location. The iterative process uses the policy gradient method in reinforcement learning, such as the REINFORCE algorithm, and the ADAM optimizer to select and update the parameters of the module. When the training reaches the preset convergence condition, the model is considered to have completed the training and can be used for the actual fault root cause location task, significantly improving the location speed and accuracy.

[0043] Through the specific implementation of the above steps, the fault root cause location system based on the knowledge graph can quickly and accurately locate the fault root cause in a complex business environment, effectively improving the operation and maintenance efficiency and business stability, and providing strong technical support for the efficient operation and maintenance of the cloud business platform.

[0044] The following Figure 1 illustrates and explains the steps shown by way of example.

[0045] According to some optional embodiments of the present application, the fact extraction module determines the state, action, and reward through the following method: calculating the attention weights between the current entity and other entities having a target relationship with the current entity through a self-attention neural network; determining a target vector representation according to a linear transformation matrix and the attention weights, wherein the linear transformation matrix is used to align different attention weights to the same space or scale; determining a set of sentences corresponding to the current entity in an external corpus, wherein each sentence in the set of sentences includes context information related to the current entity; determining the state of the fact extraction module according to the current entity, the set of sentences corresponding to the current entity, and the target vector representation; determining the following as the action of the fact extraction module: predicting the scores of each set of sentences related to the target relationship and determining the set of sentences with the highest score; determining the reward mechanism of the relationship selection module as the reward mechanism of the fact extraction module.

[0046] In the above embodiment, the fact extraction module is used to provide the most relevant facts of the entity e in the fault knowledge graph at the i-th time step, and extract them from the external corpus in the form of (e i , r i , e i , e i+1 ). Among them, r is the relationship, and e i+1 is another entity having a relationship with e i . The above relevant facts will be temporarily added to the fault knowledge graph. In order to generate the embedding at the sentence pack level Transformer is used as a sentence encoder to obtain a distributed representation. Therefore, the fact extraction module plays an important role in expanding the inference path. A quadruple (S F , A F , R F , π F ) is used to define the fact extraction module, where S F is the state space, A F is the action space, and π F is the policy function.

[0047] The state of the fact extraction module is determined through the following method.

[0048] If the relationship selection module reaches ei , the fact extraction module captures relevant sentences from the external corpus and indicates the most promising outgoing edges of entity e i . If the entity selection module reaches r i , the fact extraction module captures information from the external corpus and indicates the most promising entity e based on (e i , r i ). i+1 .

[0049] The state of the fact extraction module is defined where S F represents the entire state space including all reasonable combinations of entities and equivalent sentence packs, and a tti represents the GAT embedding.

[0050] It should be noted that the purpose of introducing the GAT embedding a tti is to help the fact extraction module focus on fact triples from the sentence pack based on the current entity e i . The node-level attention mechanism is defined as follows:

[0051]

[0052] where W represents the linear transformation matrix and N s represents the number of neighbors of entity e outside the knowledge graph i . α ij is the attention weight between the i-th entity and the j-th entity calculated by a single-layer self-attention neural network:

[0053] a i,j = LeakyReLU(μ T [We i ; We j )

[0054] where μ represents the learnable weight vector shared by all entities.

[0055] The action of the fact extraction module is determined by the following method.

[0056] The purpose of the fact extraction module is to select the inference-related triples existing in the external corpus. For each fact (e , r i , r i , e i+1 ) captured from the action can be obtained and the action space at the i-th step is represented as A F represents the entire action space including all facts extracted from the external corpus. The policy network is defined as:

[0057]

[0058] Among them, represents the learnable weight. The purpose of the fact extraction module is to predict each sentence pack i related to the relationship r The sentence pack with the highest score will be selected.

[0059] The reward of the fact extraction module is determined by the following method. When the fact extraction module completes fact extraction and infers the correct target, the path is considered to be effective and positive for the inference process. The fact extraction module shares the same reward mechanism with the relationship selection module and the entity selection module at each time step.

[0060] According to some other alternative embodiments of the present application, the relationship selection module determines the state, action, and reward by the following method: According to the current entity, the fact triple extracted by the fact extraction module, the query entity, the query relationship, and the query target, determine the state of the relationship selection module; Determine the following as the action of the relationship selection module: Predict the probability distribution of all relationships existing for the current entity, and determine the relationship with the highest probability value; Determine the following as the reward of the relationship selection module: Whether the relationship with the highest probability value is determined as the target relationship corresponding to the current entity.

[0061] Specifically, predicting the probability distribution of all relationships existing for the current entity includes: According to the relationship selected at the (i - 1)-th time step and the action determined at the (i - 1)-th time step, determine all the relationships selected at the i-th time step, where i is a positive integer greater than 1; Use the multi-layer perceptron layer in the neural network to process all the relationships selected at the i-th time step to obtain the probability distribution of all the relationships selected at the i-th time step, where the neural network is a feed-forward network including a rectified linear unit non-linear layer.

[0062] In the above embodiment, the relationship selection module is defined as a quadruple (S R , A R , R R , π R ). During the training process, the object entity has been pre-labeled. The state of the relationship selection module is determined by the following method. The state at the i-th step is defined as a tuple (e i , e ifa , e q , r q , e t ), where is the entity where the relationship selection module is currently located, e ifa represents the extracted fact, e q is the query entity, r q represents the query relationship, et Denotes the query target. At the beginning, e ifa is equal to e q , and the starting state is represented as (e i , e ifa , e q , r q , e t ), and the final state is (e i ′, e ifa ′, e q ′, r q ′, e t ′). After the module takes an action, it will move to the next state.

[0063] The action of the relationship selection module is determined by the following method.

[0064] For the state of the action space is the set of out-edges of entity e i . In order to select the most promising relationship at each time step, a policy network is used to guide the selection of an action from all available actions . The historical embedding is calculated through an LSTM network:

[0065]

[0066] where, is all the relationships selected at the i-th time step, is the relationship selected at the (i - 1)-th time step, is the action determined at the (i - 1)-th time step. d is the dimension of the relationship / entity embedding, and the embedding of the action at time step i - 1 are concatenated together.

[0067] Subsequently, it is input into a feed-forward network with a ReLU non-linear layer, and the MLP layer generates the probability distribution of all available relationships after softmax

[0068]

[0069] where, represents the trainable weights of the policy network, determines the possibility that the relationship selection module takes the corresponding action at time step i. According to sample an action The policy network of the relationship module is specifically

[0070]

[0071] As some alternative embodiments of the present application, the entity selection module determines the state, action, and reward through the following method: Based on the current entity, the fact triples extracted by the fact extraction module, the query entity, the query relationship, and the query target, determine the state of the entity selection module; Determine the following as the action of the entity selection module: Predict the probability distribution of all entities having the target relationship with the current entity, and determine the entity with the highest probability value; Determine the following as the reward of the entity selection module: Whether the entity with the highest probability value is determined as the target entity corresponding to the current entity.

[0072] In the above embodiment, the entity selection module is defined as a quadruple (S R , A R , R R , π R ), and the entity selection module is used to select an entity based on the relationship r i selected by the relationship selection module at time step i.

[0073] The state of the relationship selection module is determined through the following method: The state at the i-th step is defined as a tuple (r i-1 , e ifa , e q , r q , e t ), where ri-1 ∈ R is the last relationship. The starting state is represented as (e q , e ifa , e q , r q , e t ), and the final state is (e t ′, e ifa ′, e q ′, r q ′, e t ′) if the target entity is reached, otherwise ('STOP', e s , e t ) within the maximum length L. After taking an action, the module will move to the next state.

[0074] The action of the relationship selection module is determined through the following method: For the action space matrix of the state is the out-edge set of the last relationship r i-1 . Similar to the relationship selection module, a policy network is used to guide the selection of an action from all available actions. Specifically:

[0075]

[0076] Among them, all entities selected at the i-th time step are the entities selected at the (i - 1)-th time step is the action determined at the (i - 1)-th time step.

[0077] Among them are the weights of the policy network. determines the possibility of the action corresponding to the time step i. According to sample an action The historical dependence policy of the entity module is designed as When the entity selection module decides not to take any action at the time step i, a special action is selected. Therefore, Among them, the policy represents the probability distribution of all possible actions of the entity selection module. In the path inference task, the transition is deterministic because it is known in advance which nodes the environment will transition to when the entity selection module takes an action.

[0078] In some alternative embodiments of the present application, the following is determined as the reward of the entity selection module: whether the entity with the highest probability value is determined as the target entity corresponding to the current entity, including: in the case where the entity with the highest probability value is determined as the target entity corresponding to the current entity, a first reward value is obtained; in the case where an entity other than the entity with the highest probability value is determined as the target entity corresponding to the current entity, a second reward value is obtained; in the case where a preset entity is determined as the target entity corresponding to the current entity, a third reward value is obtained, where the preset entity is used to indicate that the entity selection module fails to find the corresponding entity within the preset number of steps, the first reward value is greater than the third reward value, and the third reward value is greater than the second reward value.

[0079] It should be noted that to solve the problem that the module cannot reach the answer entity within a limited number of steps, a synthetic "no answer" action is added in this embodiment, and this action leads to a special entity e NOANSWER . Therefore, a ternary reward mechanism is proposed in this embodiment, that is, if the module reaches the target entity e targ , it will receive a positive reward; if the module reaches the wrong entity, the module will get a negative reward. If the module reaches the no answer entity e NOANSWER , the module gets a neutral reward. Specifically:

[0080] As some other alternative embodiments of the present application, constructing a knowledge graph based on production data can be achieved through the following method: determining a schema for representing the relationships between upper ontology data in the schema layer of the knowledge graph, and receiving the attributes of the upper ontology data and the association relationships between the upper ontology data in the schema layer, where the upper ontology data includes at least one of the following keywords: business, host, metric, fault, and phenomenon; based on the schema layer, adding entities and relationships in the production data to the data layer of the knowledge graph.

[0081] In the above embodiment, the starting point of the schema layer design is to clarify the relationship types between key concepts in business platform fault location. For example, defining the "runs on" relationship between "business" and "host" means that the business runs on a specific host; the "monitors" relationship between "host" and "metric" means that the running state of the host is monitored by a series of metrics; the "is manifested as" relationship between "fault" and "phenomenon" means that the fault is manifested through specific phenomena; the "originates from" relationship between "fault" and "host" means that the fault originates from a specific host. After designing the relationship schema in the schema layer, next, receive the attributes of the upper ontology data and the relationships between the data. The attributes include the characteristic descriptions of entities, such as the CPU utilization rate of the host, the number of users of the business, etc. The association relationships are based on the above-defined relationship schema and describe the specific connections between entities in the upper ontology data. For example, a specific business runs on host A, and the CPU utilization rate of host A is extremely high, and these anomalies are manifested as an extended business response time, etc.

[0082] From the production data, specific entities such as business service B, host H, monitoring metric I, etc. are identified, and then these entities are mapped to the relationship patterns in the schema layer. For example, if the production data contains the information that business service B runs on host H, it is mapped to the "runs on" relationship; if the data contains the monitoring relationship between metric I and host H, it is mapped to the "monitors" relationship. The identified entities and relationships are instantiated into the data layer of the knowledge graph. Specifically, nodes can be created for each entity, and the node information can be filled according to its attribute values; at the same time, edges are created based on the relationships between entities, and the type and direction of the edges are based on the relationship patterns defined in the schema layer. For example, create the edges "Business B -> runs on -> Host H" and "Host H -> monitors -> Metric I", and the weight of the edges can be defined based on the numerical values of entity attributes or the severity of the faults. As the business platform runs and faults occur, new entities and relationships will continuously emerge. Therefore, the data layer needs to be dynamically updated according to the latest production data to ensure that the knowledge graph can reflect the actual state of the business platform. For example, when it is detected that the CPU utilization rate of host H suddenly increases and the response time of business B also increases accordingly, the system will automatically identify this new phenomenon and map it to the knowledge graph, creating new edges "Host H -> abnormal CPU utilization rate -> Phenomenon X" and "Business B -> increased response time -> Phenomenon X".

[0083] An embodiment of this application also provides a method for determining the cause of a business fault, including the following steps:

[0084] Step S21, obtain the business fault data of the business platform in the knowledge graph, and extract triple data from the business fault data. Among them, the entities in the knowledge graph include at least one of the following: device, time, person, fault phenomenon, and the triple data includes: entity, relationship, and another entity or attribute corresponding to the entity through the relationship.

[0085] Step S22, use the trained reinforcement learning model to determine the cause of the business fault corresponding to the triple data, where the reinforcement learning model is obtained by Figure 1 the model training method shown.

[0086] Specifically, at the beginning of fault location and diagnosis, the system initialization state includes the current entity, relationship, and historical path. The current entity can be the node where the fault is first detected, such as business service B. The relationship selection module, entity selection module, and fact extraction module are initialized and ready to perform reasoning based on the fault knowledge graph.

[0087] The relationship selection module first selects a potential relationship based on the current state and entity B, such as "runs on", which represents the connection between business service B and the host that supports it. Subsequently, the entity selection module selects the next entity node based on the "runs on" relationship, such as host H1. This process is based on the model's policy network, which determines the best choice by calculating the scores (or probability distributions) of different relationships and entities, gradually delving into the environment where the fault occurred.

[0088] The fact extraction module dynamically generates fact triples related to the current entity H1 from an external corpus, such as ('Host H1', 'CPU utilization', 'abnormally increased'). These triples not only include the status information of entity H1 but may also contain other entities or attributes related to fault location. The fact extraction module uses Transformer to obtain sentence-pack-level embeddings and combines the attention weights calculated by GAT to select the most relevant fact triples, which are temporarily added to the knowledge graph to support the subsequent reasoning process.

[0089] According to the selections of the relationship selection module and the entity selection module, update the current state and add the relationship and entity to the historical path. For example, the path is updated from ('Business service B') to ('Business service B', 'runs on', 'Host H1'). This process continues until the entities and relationships on the path match the root cause of the fault or the maximum number of reasoning steps is reached, at which point the reasoning process is considered complete.

[0090] In each step of the reasoning process, the system calculates a reward based on whether the selected relationship and entity contribute to locating the root cause of the fault. If the selection is correct, the reward is positive; if the selection is incorrect or irrelevant, the reward is negative. The calculation of the reward is based on the distance between the target entity and the current entity, as well as the contribution of the selected relationship and entity to fault location. These rewards are used to update the parameters of the model's policy network, and a reinforcement learning algorithm is used to optimize the selection of the reasoning path, making future reasoning processes more efficient and accurate.

[0091] When the reasoning process is complete, the system combines the entire reasoning path and the rewards to determine the final reasoning results, including the root cause entity of the fault, the logical path where the fault occurred, and relevant facts. This information is organized into a report and output to the operations and maintenance personnel to help them quickly understand and handle the fault, reducing business interruption time and operations and maintenance costs.

[0092] Through the above steps, the trained knowledge graph reasoning model can automatically locate the root cause of the fault from complex data, provide a decision-making basis for fault location, and greatly improve the operations and maintenance efficiency of the business platform and the accuracy of fault handling. In addition, the self-learning and adaptive capabilities of the model enable it to continuously optimize the reasoning strategy to adapt to changes in the business environment, becoming a powerful tool for the operations and maintenance management of the business platform.

[0093] To verify the effectiveness of the above method, in this embodiment, 20 types of relationships and 2,498 entities are sorted and counted from the actual operation and maintenance data of a certain business platform. The training set includes 25,487 triples, the validation set includes 8,767 triples, and the test set includes 3,000 triples. A comparative experiment is conducted on the knowledge graph proposed in this application and other classic root cause localization methods, and MAP is selected as the evaluation index. The following table shows the MAP scores and baseline models achieved by this application in the fact prediction task, including embedding-based methods and multi-hop reasoning methods. This application has significant improvements and higher accuracy.

[0094] Model MAP(%) TransD 50.4 TransR 52.2 TransH 55.4 PRA 63.7 DeepPath 64.1 Multi-hop 63.7 M-Walk 65.3 The model provided by this application 76.5

[0095] Figure 2 is a structural diagram of a model training device according to an embodiment of the present application, as Figure 2 shown. The device includes:

[0096] A first acquisition module 21, configured to acquire historical production data of a business platform and construct a knowledge graph based on the historical production data, where the entities in the knowledge graph include at least one of the following: devices, time, personnel, and fault phenomena.

[0097] A first determination module 22, configured to traverse the entities in the knowledge graph by using a relationship selection module in a reinforcement learning model to determine a target relationship corresponding to the current entity, where the target relationship is the relationship with the largest relevance index to the current entity among multiple relationships corresponding to the current entity.

[0098] A second determination module 23, configured to determine a target entity having a target relationship with the current entity by using an entity selection module in the reinforcement learning model.

[0099] An extraction module 24, configured to extract a fact triple of the current entity from an external corpus by using a fact extraction module in the reinforcement learning model and update the knowledge graph according to the fact triple, where the fact triple includes: the current entity, a relationship, and other entities or attributes having a relationship with the current entity.

[0100] A second acquisition module 25, configured to acquire rewards obtained by each module in the reinforcement learning model and determine a loss function for each module according to the rewards, where the rewards are the rewards obtained by each module according to its own state, selecting actions through its own policies, executing the selected actions, and the rewards are used to evaluate the influence degree of the actions on the accuracy of determining the cause of business failures.

[0101] A training module 26, configured to update the parameters of each module in the reinforcement learning model until the loss function meets a preset convergence condition, so as to obtain a trained reinforcement learning model, where the reinforcement learning model is used to determine the cause of business failures in production data.

[0102] Optionally, the fact extraction module determines the state, action, and reward through the following method: calculating the attention weights between the current entity and other entities having a target relationship with the current entity through a self-attention neural network; determining a target vector representation according to a linear transformation matrix and the attention weights, where the linear transformation matrix is used to align different attention weights to the same space or scale; determining a set of sentences corresponding to the current entity in an external corpus, where each sentence in the set of sentences includes context information related to the current entity; determining the state of the fact extraction module according to the current entity, the set of sentences corresponding to the current entity, and the target vector representation; determining the following as the action of the fact extraction module: predicting the scores of each set of sentences related to the target relationship and determining the set of sentences with the highest score; determining the reward mechanism of the relationship selection module as the reward mechanism of the fact extraction module.

[0103] Optionally, the relationship selection module determines the state, action, and reward through the following method: determining the state of the relationship selection module according to the current entity, the fact triples extracted by the fact extraction module, the query entity, the query relationship, and the query target; determining the following as the action of the relationship selection module: predicting the probability distribution of all relationships existing for the current entity and determining the relationship with the highest probability value; determining the following as the reward of the relationship selection module: whether the relationship with the highest probability value is determined as the target relationship corresponding to the current entity.

[0104] Optionally, predicting the probability distribution of all relationships existing for the current entity includes: determining all relationships selected at the i-th time step according to the relationship selected at the (i - 1)-th time step and the action determined at the (i - 1)-th time step, where i is a positive integer greater than 1; processing all relationships selected at the i-th time step through a multi-layer perceptron layer in a neural network to obtain the probability distribution of all relationships selected at the i-th time step, where the neural network is a feed-forward network including a rectified linear unit non-linear layer.

[0105] Optionally, the entity selection module determines the state, action, and reward through the following method: determining the state of the entity selection module according to the current entity, the fact triples extracted by the fact extraction module, the query entity, the query relationship, and the query target; determining the following as the action of the entity selection module: predicting the probability distribution of all entities having a target relationship with the current entity, and determining the entity with the highest probability value; determining the following as the reward of the entity selection module: whether the entity with the highest probability value is determined as the target entity corresponding to the current entity.

[0106] Optionally, determining the following as the reward of the entity selection module: whether the entity with the highest probability value is determined as the target entity corresponding to the current entity, including: in the case where the entity with the highest probability value is determined as the target entity corresponding to the current entity, obtaining a first reward value; in the case where an entity other than the entity with the highest probability value is determined as the target entity corresponding to the current entity, obtaining a second reward value; in the case where a preset entity is determined as the target entity corresponding to the current entity, obtaining a third reward value, where the preset entity is used to indicate that the entity selection module fails to find a corresponding entity within a preset number of steps, the first reward value is greater than the third reward value, and the third reward value is greater than the second reward value.

[0107] Optionally, constructing a knowledge graph based on production data, including: determining a schema for representing the relationships between upper ontology data in the schema layer of the knowledge graph, and receiving the attributes of the upper ontology data and the association relationships between the upper ontology data in the schema layer, where the upper ontology data includes at least one of the following keywords: business, host, metric, fault, and phenomenon; adding entities and relationships in the production data to the data layer of the knowledge graph based on the schema layer.

[0108] It should be noted that the above Figure 2 each module can be a program module (for example, a set of program instructions for implementing a specific function), or a hardware module. For the latter, it can be presented in the following forms, but not limited to: the manifestation form of each of the above modules is a processor, or the functions of each of the above modules are implemented by a processor.

[0109] It should be noted that Figure 2 the preferred implementation manners of the illustrated embodiments can be referred to Figure 1 the relevant descriptions of the illustrated embodiments, which will not be elaborated here.

[0110] Figure 3 shows a hardware structure block diagram of a computer terminal for implementing a model training method. As Figure 3As shown, the computer terminal 30 may include one or more processors 302 (shown as 302a, 302b, ……, 302n in the figure) (the processor 302 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 304 for storing data, and a transmission module 306 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 3 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 30 may further include more or fewer components than Figure 3 shown therein, or have a different configuration from Figure 3 that shown.

[0111] It should be noted that the above one or more processors 302 and / or other data processing circuits are generally referred to as "data processing circuits" herein. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 30. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistor terminal path connected to an interface).

[0112] The memory 304 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the model training method in the embodiments of the present application. The processor 302 executes various functional applications and data processing by running the software programs and modules stored in the memory 304, that is, implements the above-mentioned model training method. The memory 304 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 304 may further include a memory remotely set relative to the processor 302, and these remote memories can be connected to the computer terminal 30 through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0113] The transmission module 306 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the computer terminal 30. In one example, the transmission module 306 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission module 306 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0114] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computer terminal 30.

[0115] It should be noted here that in some alternative embodiments, the above Figure 3 illustrated computer terminal may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 3 is only an example of a specific concrete example and is intended to show the types of components that may exist in the above computer terminal.

[0116] It should be noted that Figure 3 the illustrated computer terminal is used to execute Figure 1 the illustrated model training method. Therefore, the relevant explanations in the execution method of the above commands also apply to this electronic device, which will not be elaborated here.

[0117] The embodiment of the present application also provides a non-volatile storage medium. The non-volatile storage medium includes a stored program, wherein when the program runs, it controls the device where the storage medium is located to execute the above model training method.

[0118] A program for a non-volatile storage medium to perform the following functions: obtaining historical production data of a business platform and constructing a knowledge graph based on the historical production data, where the entities in the knowledge graph include at least one of the following: devices, time, personnel, and fault phenomena; traversing the entities in the knowledge graph using a relationship selection module in a reinforcement learning model to determine the target relationship corresponding to the current entity, where the target relationship is the relationship with the largest relevance index among the multiple relationships corresponding to the current entity; determining the target entity having the target relationship with the current entity using an entity selection module in the reinforcement learning model; extracting the fact triple of the current entity from an external corpus using a fact extraction module in the reinforcement learning model and updating the knowledge graph according to the fact triple, where the fact triple includes: the current entity, the relationship, and the other entity or attribute having a relationship with the current entity; obtaining the rewards obtained by each module in the reinforcement learning model and determining the loss function of each module according to the rewards, where the reward is the reward obtained by each module according to its own state, selecting an action through its own policy, executing the selected action, and the reward is used to evaluate the influence degree of the action on the accuracy of determining the cause of the business fault; updating the parameters of each module in the reinforcement learning model until the loss function satisfies a preset convergence condition to obtain a trained reinforcement learning model, where the reinforcement learning model is used to determine the cause of the business fault in the production data.

[0119] An embodiment of the present application further provides an electronic device, including: a memory and a processor, where the processor is configured to run a program stored in the memory, and when the program runs, it executes the above model training method.

[0120] The processor is used to run a program that performs the following functions: obtaining historical production data of the business platform and constructing a knowledge graph based on the historical production data, where the entities in the knowledge graph include at least one of the following: devices, time, personnel, and fault phenomena; traversing the entities in the knowledge graph using the relationship selection module in the reinforcement learning model to determine the target relationship corresponding to the current entity, where the target relationship is the relationship with the largest relevance index among the multiple relationships corresponding to the current entity; determining the target entity that has a target relationship with the current entity using the entity selection module in the reinforcement learning model; extracting the fact triples of the current entity from the external corpus using the fact extraction module in the reinforcement learning model and updating the knowledge graph according to the fact triples, where the fact triples include: the current entity, the relationship, and other entities or attributes that have a relationship with the current entity; obtaining the rewards obtained by each module in the reinforcement learning model and determining the loss function of each module according to the rewards, where the rewards are the rewards obtained by each module according to its own state, selecting actions through their respective policies, executing the selected actions, and the rewards obtained based on the selected actions, and the rewards are used to evaluate the influence degree of the actions on the accuracy of determining the cause of the business failure; updating the parameters of each module in the reinforcement learning model until the loss function meets the preset convergence condition to obtain a trained reinforcement learning model, where the reinforcement learning model is used to determine the cause of the business failure in the production data.

[0121] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0122] In the above embodiments of the present application, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0123] In the above embodiments of the present application, the collected information is information and data authorized by the user or fully authorized by all parties, and the processing of the relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with relevant laws, regulations, and standards, take necessary protection measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0124] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0125] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0126] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0127] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0128] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A model training method, characterized in that: include: Acquire historical production data of the business platform, and construct a knowledge graph based on the historical production data, wherein the entities in the knowledge graph include at least one of the following: equipment, time, personnel, and fault phenomenon; Using the relationship selection module in the reinforcement learning model to traverse the entities in the knowledge graph, and determine the target relationship corresponding to the current entity, wherein the target relationship is the relationship with the largest correlation index with the current entity among the multiple relationships corresponding to the current entity; Determine a target entity having the target relationship with the current entity using an entity selection module in the reinforcement learning model; The fact extraction module in the reinforcement learning model is used to extract the fact triple of the current entity from the external corpus, and the knowledge graph is updated according to the fact triple, wherein the fact triple includes: the current entity, the relationship and other entities or attributes that have the relationship with the current entity; Obtaining rewards obtained by each module in the reinforcement learning model, and determining the loss function of each module according to the rewards, wherein the rewards are rewards obtained by each module selecting an action according to its own state and executing the selected action through its own strategy, and the rewards are used to evaluate the influence of the action on the accuracy of determining the cause of the business failure; The parameters of each module in the reinforcement learning model are updated until the loss function meets the preset convergence condition, thereby obtaining a trained reinforcement learning model, wherein the reinforcement learning model is used to determine the cause of the business failure in the production data.

2. The method according to claim 1, characterized in that The fact extraction module determines the state, action and reward by: Calculating the attention weights between the current entity and other entities having the target relationship with the current entity through a self-attention neural network; Determining a target vector representation based on a linear transformation matrix and the attention weights, wherein the linear transformation matrix is ​​used to align different attention weights to the same space or scale; Determine a sentence set corresponding to the current entity in the external corpus, wherein each sentence in the sentence set includes context information related to the current entity; Determining a state of the fact extraction module according to the current entity, a set of sentences corresponding to the current entity, and the target vector representation; Determining the following as actions of the fact extraction module: predicting a score for each set of sentences associated with the target relation, and determining a set of sentences with the highest scores; The reward mechanism of the relationship selection module is determined as the reward mechanism of the fact extraction module.

3. The method according to claim 1 or 2, characterized in that: The relationship selection module determines the state, action and reward by the following method: Determining the state of the relationship selection module according to the current entity, the fact triples extracted by the fact extraction module, the query entity, the query relationship and the query target; Determine the following as actions of the relationship selection module: predict the probability distribution of all relationships existing in the current entity, and determine the relationship with the highest probability value; The following content is determined as a reward for the relationship selection module: whether to determine the relationship with the highest probability value as the target relationship corresponding to the current entity.

4. The method according to claim 3, characterized in that Predict the probability distribution of all relationships existing in the current entity, including: Determine all the relationships selected at the i-th time step according to the relationships selected at the i-1th time step and the actions determined at the i-1th time step, where i is a positive integer greater than 1; A multi-layer perceptron layer in a neural network is used to process all relations selected by the i-th time step to obtain a probability distribution of all relations selected by the i-th time step, wherein the neural network is a feedforward network including a nonlinear layer of a rectified linear unit.

5. The method according to claim 1 or 2, characterized in that: The entity selection module determines the state, action and reward by the following method: Determining a state of the entity selection module according to the current entity, the fact triples extracted by the fact extraction module, the query entity, the query relationship, and the query target; Determining the following as actions of the entity selection module: predicting the probability distribution of all entities that have the target relationship with the current entity, and determining the entity with the highest probability value; The following content is determined as a reward for the entity selection module: whether to determine the entity with the highest probability value as the target entity corresponding to the current entity.

6. The method according to claim 5, characterized in that The following content is determined as a reward of the entity selection module: whether to determine the entity with the highest probability value as the target entity corresponding to the current entity, including: In a case where the entity with the highest probability value is determined to be the target entity corresponding to the current entity, a first reward value is obtained; In the case where the entity with the highest non-probability value is determined as the target entity corresponding to the current entity, a second reward value is obtained; When a preset entity is determined as the target entity corresponding to the current entity, a third reward value is obtained, wherein the preset entity is used to indicate that the entity selection module has not found a corresponding entity within a preset number of steps, the first reward value is greater than the third reward value, and the third reward value is greater than the second reward value.

7. The method according to claim 1, characterized in that Constructing a knowledge graph based on the production data, including: Determining, at the pattern layer of the knowledge graph, a pattern for representing the relationship between the upper-layer ontology data, and receiving, at the pattern layer, the attributes of the upper-layer ontology data and the association relationship between the upper-layer ontology data, wherein the upper-layer ontology data includes at least one of the following keywords: business, host, indicator, fault, and phenomenon; Based on the model layer, entities and relationships in the production data are added to the data layer of the knowledge graph.

8. A method for determining the cause of a service failure, characterized in that: include: Obtaining business fault data of a business platform in a knowledge graph, and extracting triple data from the business fault data, wherein the entity in the knowledge graph includes at least one of the following: equipment, time, personnel, and fault phenomenon, and the triple data includes: an entity, a relationship, and another entity or attribute corresponding to the entity through the relationship; A trained reinforcement learning model is used to determine the cause of the business failure corresponding to the triple data, wherein the reinforcement learning model is obtained by training using the model training method described in any one of claims 1 to 7.

9. The method according to claim 8, characterized in that The trained reinforcement learning model is used to determine the cause of the business failure corresponding to the triple data, including: A relation selection module is used to select a target relation based on the state, wherein the target relation represents a possible connection between the current entity and an adjacent node; Determine a target entity by using an entity selection module based on the target relationship selected by the relationship selection module; Using a fact extraction module to extract fact triples of the current entity from an external corpus, and updating the knowledge graph according to the fact triples; According to the target relationship and target entity selected by the relationship selection module and the entity selection module, the state is updated and the reasoning path is recorded until the reasoning path indicates that the entities and relationships in the state meet the preset matching target.

10. A model training device, characterized in that: include: A first acquisition module is used to acquire historical production data of the business platform and construct a knowledge graph based on the historical production data, wherein the entities in the knowledge graph include at least one of the following: equipment, time, personnel, and fault phenomenon; A first determination module is used to traverse the entities in the knowledge graph using a relationship selection module in a reinforcement learning model to determine a target relationship corresponding to a current entity, wherein the target relationship is a relationship with the largest correlation index with the current entity among multiple relationships corresponding to the current entity; A second determination module, configured to use the entity selection module in the reinforcement learning model to determine a target entity that has the target relationship with the current entity; An extraction module, configured to extract a fact triple of the current entity from an external corpus using the fact extraction module in the reinforcement learning model, and update the knowledge graph according to the fact triple, wherein the fact triple includes: the current entity, the relationship, and other entities or attributes that have the relationship with the current entity; A second acquisition module is used to obtain rewards obtained by each module in the reinforcement learning model, and determine the loss function of each module according to the rewards, wherein the rewards are rewards obtained by each module selecting an action according to its own state through its own strategy, executing the selected action, and based on the selected action, and the rewards are used to evaluate the influence of the action on the accuracy of determining the cause of the business failure; A training module is used to update the parameters of each module in the reinforcement learning model until the loss function meets the preset convergence condition, thereby obtaining a reinforcement learning model that has completed training, wherein the reinforcement learning model is used to determine the cause of the business failure in the production data.

11. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the model training method described in any one of claims 1 to 7.

12. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program executes the model training method described in any one of claims 1 to 7 when running.

13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the model training method described in any one of claims 1 to 7 is implemented.