Fault reasoning analysis method and system based on large language model
Through the fault reasoning analysis method based on the large language model, multi-source operation and maintenance data is obtained for semantic enhancement and the construction of a semantic association network, which solves the problems of low efficiency and poor accuracy in traditional methods and realizes automated and intelligent fault diagnosis and repair.
Patent Information
- Application Number
- CN202511211509.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Traditional system fault diagnosis methods rely on manual experience and preset rules, and have problems such as low efficiency, poor accuracy, inability to fully cover complex fault scenarios, and insufficient semantic correlation of multi-source data.
A fault reasoning analysis method based on a large language model is adopted to obtain multi-source operation and maintenance data, perform semantic enhancement and build a semantic association network, use the large language model to perform root cause reasoning analysis, and generate operation and maintenance decision instructions to trigger the automated repair process.
It improves the accuracy and comprehensiveness of fault reasoning, realizes the automation and intelligence of fault diagnosis and repair, reduces the cost and error rate of manual intervention, and improves the efficiency and reliability of system operation and maintenance.
Smart Images

Figure CN120745841A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of system operation and maintenance technology, and in particular to a fault reasoning analysis method and system based on a large language model. Background Art
[0002] In modern, complex system operations and maintenance scenarios, the systems to be diagnosed are often composed of numerous components that work together to achieve specific functions. As systems continue to expand in size and complexity, the frequency and complexity of system failures are also increasing significantly. Traditional system fault diagnosis methods rely primarily on manual experience and pre-set rule engines.
[0003] While manual diagnosis can leverage the deep knowledge and experience of professionals to address complex issues, it has significant limitations. Firstly, manual diagnosis is inefficient. Faced with massive amounts of multi-source O&M data, manual analysis is not only time-consuming and labor-intensive, but also prone to missing critical information. Secondly, manual diagnosis results are significantly influenced by personal experience and subjective factors. Different people may reach different conclusions on the same fault phenomenon, resulting in a lack of consistency and accuracy.
[0004] While pre-set rule engines can quickly handle known, rule-defined failure scenarios, they are often unable to address emerging, complex failure modes. Because the system's operating environment and component interactions are constantly changing, pre-set rules struggle to fully cover all possible failure scenarios, and rule updates and maintenance are expensive. Furthermore, existing methods fail to fully consider the semantic connections between different types of data when processing multi-source operation and maintenance data, making it impossible to comprehensively grasp the system's operating status and the root causes of failures. This results in inaccurate and incomplete fault reasoning and analysis. Summary of the Invention
[0005] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a fault reasoning and analysis method based on a large language model, the method comprising: Acquire a multi-source operation and maintenance data set of the system to be diagnosed, wherein the multi-source operation and maintenance data set includes structured performance indicator data, semi-structured service log data, and unstructured text description data sorted by timestamp; Performing semantic enhancement on the multi-source operation and maintenance data set to obtain a semantic data unit including entity semantic labels, relationship semantic labels, and attribute semantic labels, wherein the entity semantic labels correspond to system component entities, the relationship semantic labels correspond to interaction relationships between components, and the attribute semantic labels correspond to entity state attributes; Constructing a semantic association network based on the semantic data unit, wherein the nodes of the semantic association network are the entity semantic labels, the directed edges between the nodes are the relationship semantic labels, and the edge weights are the relationship strength parameters calculated based on the attribute semantic labels; Calling a pre-trained large language model to perform root cause reasoning analysis on the semantic association network to generate a set of candidate root causes sorted by confidence, wherein the candidate root cause set includes potential cause entities of the system anomaly and corresponding association path descriptions; An operation and maintenance decision instruction including an entity operation sequence and a priority ranking is generated according to the candidate root cause set, and the operation and maintenance decision instruction is sent to a system management terminal to trigger an automated repair process.
[0006] On the other hand, an embodiment of the present invention also provides a fault reasoning analysis system based on a large language model, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.
[0007] Based on the above aspects, the embodiment of the present invention obtains a multi-source operation and maintenance data set of the system to be diagnosed, which comprehensively covers structured performance indicator data, semi-structured service log data and unstructured text description data sorted by timestamps, and performs semantic enhancement on the multi-source operation and maintenance data set to obtain a semantic data unit containing entity semantic labels, relationship semantic labels and attribute semantic labels. It can accurately identify system component entities, interaction relationships between components and entity state attributes, and construct a semantic association network based on the semantic data unit, with entity semantic labels as nodes, relationship semantic labels as directed edges, and edge weights as relationship strength parameters calculated based on attribute semantic labels. It presents the complex association relationships between system components and helps to grasp the operating status of the system as a whole. Calling a pre-trained large language model to perform root cause reasoning analysis on the semantic association network, using the powerful language understanding and reasoning capabilities of the large language model to generate a set of candidate root causes sorted by confidence, which can accurately find the potential cause entities of system abnormalities and the corresponding association path descriptions, greatly improving the accuracy and comprehensiveness of fault reasoning. Based on the set of candidate root causes, operation and maintenance decision instructions containing entity operation sequences and priority rankings are generated and sent to the system management terminal to trigger the automated repair process, realizing the automation and intelligence of fault diagnosis and repair, significantly improving the efficiency and reliability of system operation and maintenance, and reducing the cost and error rate of manual intervention. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 It is a schematic diagram of the execution flow of the fault reasoning and analysis method based on the large language model provided by an embodiment of the present invention.
[0009] Figure 2 Schematic diagram of exemplary hardware and software components of a fault reasoning and analysis system based on a large language model provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0010] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flow chart of a fault reasoning and analysis method based on a large language model provided by an embodiment of the present invention. The fault reasoning and analysis method based on a large language model is introduced in detail below.
[0011] This embodiment takes the fault diagnosis of a distributed server cluster as an application scenario and describes in detail the specific implementation process of the fault reasoning and analysis method based on a large language model.
[0012] Step S110: Acquire a multi-source operation and maintenance data set of the system to be diagnosed, wherein the multi-source operation and maintenance data set includes structured performance indicator data, semi-structured service log data, and unstructured text description data sorted by timestamp.
[0013] In a distributed server cluster, the system to be diagnosed includes multiple physical servers, various application services deployed on the servers, network switches, storage devices, and other components. To fully understand the system's operating status, it is necessary to collect multi-source operation and maintenance data from these components.
[0014] Structured performance indicator data comes from built-in monitoring agents within each component. These monitoring agents collect data at fixed intervals (e.g., every 30 seconds) and store it in a pre-set format (e.g., JSON). This data includes information such as server CPU usage, memory usage, disk I / O rate, network bandwidth utilization; application service response time, number of concurrent connections, number of error requests; network switch port traffic and forwarding latency; storage device read / write speeds, and storage space utilization. All of this data is accurately timestamped and arranged in chronological order to form a structured performance indicator data set.
[0015] Semi-structured service log data is automatically generated by application services during operation and is typically stored in text files. Each log entry contains fields such as a timestamp, log level (such as INFO, WARNING, or ERROR), source service identifier, target service identifier, and interaction description. However, the field separation method may not be fixed, and some fields may be missing. For example, a log entry for an application service might read "[2025-08-03 10:00:05] ERROR: ServiceA->ServiceB: Connection timeout," which includes a timestamp, log level, source service identifier (ServiceA), target service identifier (ServiceB), and interaction description (Connection timeout).
[0016] Unstructured text description data comes from various channels, including descriptions of system anomalies recorded by operation and maintenance personnel during inspections (such as "the temperature of Server 1 in the server room rises abnormally, accompanied by unusual noises"), fault reporting information submitted by users (such as "when using the App, there is no response for a long time after submitting an order"), and communication records between technical support personnel and developers (such as "Service C crashes when processing a large number of concurrent requests and still frequently reports errors after restarting"), etc. The above data has no fixed format and exists in the form of natural language text.
[0017] When collecting this data, privacy-sensitive data (such as user account information that may be included in troubleshooting reports submitted by users) requires the use of privacy protection technologies. For example, user account information can be desensitized, replacing specific account numbers with meaningless identifiers; data encryption technology can be used to encrypt data during transmission and storage to prevent data leakage; and access control can be implemented to ensure that only authorized operation and maintenance personnel can access the relevant data.
[0018] Step S120: Perform semantic enhancement on the multi-source operation and maintenance data set to obtain a semantic data unit containing entity semantic labels, relationship semantic labels and attribute semantic labels. The entity semantic labels correspond to system component entities, the relationship semantic labels correspond to interaction relationships between components, and the attribute semantic labels correspond to entity state attributes.
[0019] Semantic enhancement of multi-source operation and maintenance data sets aims to transform the raw data into a semantic form that is easier to understand and process, facilitating the subsequent construction of a semantic association network. This process requires processing structured performance indicator data, semi-structured service log data, and unstructured text description data separately.
[0020] Step S121: Perform indicator semantic mapping processing on the structured performance indicator data, extract system component keywords in the indicator name, and generate entity semantic labels corresponding to the system component keywords based on a preset entity dictionary. The entity semantic labels include component name, component level and unique identification code.
[0021] The preset entity dictionary is pre-built and contains information about all possible system component entities in a distributed server cluster. Each entry in the preset entity dictionary corresponds to a system component entity and includes the component name (such as "Server1," "ServiceA," "Switch1," or "Storage1"), the component level (such as "Physical Device Layer - Server," "Application Service Layer - Business Service," "Network Device Layer - Switch," or "Storage Device Layer - Disk Array"), and a unique identifier (such as "PHY-SRV-001," "APP-SVC-001," "NET-SWT-001," or "STR-DEV-001").
[0022] When performing indicator semantic mapping on structured performance indicator data, the indicator name of each performance indicator data item is first analyzed. For example, in the indicator name "CPU usage of Server1", the system component keyword is "Server1"; in the indicator name "Response time of ServiceA", the system component keyword is "ServiceA".
[0023] The extracted system component keywords are then matched against a pre-set entity dictionary. If a match is successful, the corresponding component name, component level, and unique identifier are retrieved from the dictionary to generate an entity semantic label. For example, when "Server1" matches an entry in the pre-set entity dictionary, the generated entity semantic label is (Component Name: Server1, Component Level: Physical Device Layer - Server, Unique Identifier: PHY-SRV-001); the corresponding entity semantic label for "ServiceA" is (Component Name: ServiceA, Component Level: Application Service Layer - Business Service, Unique Identifier: APP-SVC-001).
[0024] Step S122: Perform log field structuring processing on the semi-structured service log data, parse the source component field, target component field and interaction operation field in the log entry, and generate a relationship semantic label representing the call direction and interaction type between components based on the verb meaning of the interaction operation field. The relationship semantic label includes a source entity identifier, a target entity identifier and a relationship type code.
[0025] When structuring semi-structured service log data into log fields, we first use regular expression matching to extract the source component field, target component field, and interaction field from each log entry. For example, for the log entry "[2025-08-03 10:00:05]ERROR:ServiceA->ServiceB:Connectiontimeout," the regular expression can be used to extract the source component field as "ServiceA," the target component field as "ServiceB," and the interaction field as "Connectiontimeout."
[0026] Next, the source component field and the target component field are matched against the preset entity dictionary to obtain the corresponding unique identification codes as the source entity identifier and target entity identifier. For example, if the unique identification code corresponding to "ServiceA" is "APP-SVC-001" and the unique identification code corresponding to "ServiceB" is "APP-SVC-002", the source entity identifier is "APP-SVC-001" and the target entity identifier is "APP-SVC-002".
[0027] Next, the verb meaning of the interaction operation field is analyzed to determine the interaction type between components. Preset interaction types include calls (e.g., "call," "invoke"), responses (e.g., "response," "reply"), connections (e.g., "connect," "link"), disconnections (e.g., "disconnect," "unlink"), and data transfers (e.g., "send," "receive"). Each interaction type corresponds to a relationship type code (e.g., "REL-001" for calls, "REL-002" for responses, "REL-003" for connections, "REL-004" for disconnections, and "REL-005" for data transfers).
[0028] For the interaction operation field "Connection timeout," its core verb meaning "connection failure" can be categorized as a connection-related interaction type, and the corresponding relationship type code is "REL-003." (Here, connection failure is also classified as a special case of the connection type to facilitate unified processing later.) Therefore, the generated relationship semantic label is (source entity identifier: APP-SVC-001, target entity identifier: APP-SVC-002, relationship type code: REL-003).
[0029] Step S123: Call a large language model to perform semantic parsing processing on the unstructured text description data, identify entity mentions, attribute descriptions and relationship assertions therein, match entity mentions with the entity semantic labels for consistency, and map attribute descriptions and relationship assertions into attribute semantic labels and supplementary relationship semantic labels containing attribute names, attribute values and timestamps, respectively.
[0030] Step S1231: performing text preprocessing on the unstructured text description data, removing repeated characters therein, and segmenting the data into multiple text segment units according to sentence boundaries, each of which is associated with a position index of the original unstructured text description data.
[0031] Unstructured text description data may contain various special symbols (such as "@," "#," and "$") and repeated characters (such as multiple consecutive spaces and line breaks), which can affect subsequent semantic parsing. During text preprocessing, a character filtering algorithm is first used to remove special symbols, and then a repeated character compression algorithm is used to reduce repeated characters to single characters.
[0032] For example, for the text description "The temperature of Server 1 in the server room rises abnormally, accompanied by an abnormal sound!!", after removing the special symbol "!" and repeated characters ",," the result is "The temperature of Server 1 in the server room rises abnormally, accompanied by an abnormal sound."
[0033] Next, the preprocessed text is segmented into multiple text segments based on sentence boundaries (such as periods, question marks, exclamation points, and commas). In the example above, this can be segmented into two text segments: "The temperature of Server 1 in the server room has risen abnormally" and "accompanied by an unusual sound." Each text segment is associated with a position index that records the start and end positions of the segment in the original unstructured text description data for subsequent tracing.
[0034] Step S1232: constructing a semantic parsing prompt template for each text segment unit, wherein the semantic parsing prompt template includes a task description, an entity type list, a relationship type list, and an attribute type list.
[0035] The task description of the semantic parsing prompt template is "Please analyze the following text fragment and identify the entity mentions, attribute descriptions, and relationship assertions therein." The entity type list includes entity types that may exist in a distributed server cluster, such as "server," "application service," "network switch," and "storage device." The relationship type list is consistent with the interaction types preset in step S122, such as "call," "response," and "connection." The attribute type list includes attributes that each entity may have, such as "temperature," "response time," "CPU usage," and "storage space."
[0036] For example, the semantic parsing prompt template constructed for the text fragment unit "The temperature of Server 1 in the server room rises abnormally" is: "Please analyze the following text fragment and identify the entity mentions, attribute descriptions and relationship assertions therein. Entity type list: server, application service, network switch, storage device; relationship type list: call, response, connection, disconnection, data transmission; attribute type list: temperature, response time, CPU usage, storage space. Text fragment: The temperature of Server 1 in the server room rises abnormally."
[0037] Step S1233: splice the text fragment unit and the semantic parsing prompt template into a model input text, call the large language model to perform sequence labeling processing on the model input text, and generate a labeling result sequence including entity labeling, attribute labeling and relationship labeling. The entity labeling includes the entity name and entity type, the attribute labeling includes the attribute name, attribute value and the entity to which it belongs, and the relationship labeling includes the source entity, the target entity and the relationship type.
[0038] The text snippet "Server 1 in the server room has an abnormally high temperature" is concatenated with the corresponding semantic parsing prompt template to form the model input text. A pre-trained large language model (such as a model based on the Transformer architecture) is used to process the model input text.
[0039] The large language model performs sequence labeling based on its understanding of the model's input text. For the above model input text, the generated entity labels are (Entity Name: Server1, Entity Type: Server); the attribute labels are (Attribute Name: Temperature, Attribute Value: Abnormally High, Entity: Server1). Since this text fragment does not involve any relationships between entities, the relationship labels are empty.
[0040] Step S1234: perform entity linking processing on the annotation result sequence, match the identified entity mentions with the standard entity names in the preset entity dictionary, and generate standardized entity identifiers. The preset entity dictionary contains the standard names and unique identification codes of all component entities in the system.
[0041] The entity mention "Server1" successfully matches the standard entity name "Server1" in the preset entity dictionary, and its unique identification code "PHY-SRV-001" is obtained from the dictionary as the standardized entity identifier.
[0042] Step S1235: Based on the standardized entity identifier, the attribute annotation is associated with the corresponding entity to generate an attribute semantic label including the entity identifier, attribute name, attribute value and extraction confidence.
[0043] The attribute annotation (attribute name: temperature, attribute value: abnormally high, entity: Server1) is associated with the standardized entity identifier "PHY-SRV-001" to generate an attribute semantic label. The large language model also provides a confidence score (e.g., 0.92) for the attribute extraction to assess the reliability of the extraction result. Therefore, the generated attribute semantic label is (entity identifier: PHY-SRV-001, attribute name: temperature, attribute value: abnormally high, extraction confidence: 0.92).
[0044] Step S1236: Replace the source entity and target entity in the relationship annotation with standardized entity identifiers, and generate a supplementary relationship semantic label including the source entity identifier, target entity identifier, relationship type, and extraction confidence.
[0045] If a text segment contains a relationship annotation, for example, "ServiceA called ServiceB," its relationship annotation is (source entity: ServiceA, target entity: ServiceB, relationship type: call). Replace the source entity "ServiceA" and target entity "ServiceB" with their corresponding standardized entity identifiers "APP-SVC-001" and "APP-SVC-002," and combine the extraction confidence (e.g., 0.95) to generate a supplementary relationship semantic label (source entity identifier: APP-SVC-001, target entity identifier: APP-SVC-002, relationship type: call, extraction confidence: 0.95).
[0046] Step S1237: Integrate the attribute semantic tag and the supplementary relationship semantic tag to generate a semantic parsing result set corresponding to the text segment unit, where each tag entry in the semantic parsing result set is associated with a position index of the text segment unit.
[0047] The attribute semantic tags generated in step S1235 and the supplementary relationship semantic tags generated in step S1236 are combined to form a semantic parsing result set. Each tag entry is associated with the position index of the text segment unit to facilitate tracing its source. For example, for the text segment unit "The temperature of Server1 in the server room has risen abnormally," its semantic parsing result set is [(entity identifier: PHY-SRV-001, attribute name: temperature, attribute value: abnormally risen, extraction confidence: 0.92, position index: 0-25)].
[0048] Step S124: Perform redundancy detection processing on the entity semantic labels, relationship semantic labels and the supplementary relationship semantic labels to remove duplicate label entries and low-confidence label entries. The low-confidence label entries are labels whose confidence scores output by the large language model are lower than the preset confidence threshold.
[0049] The preset confidence threshold can be set based on the actual application scenario, such as 0.8. During redundancy detection, the content of each tag entry is first compared. If all fields of two tag entries are identical (such as the component name, component hierarchy, and unique identifier code of the entity semantic tag; the source entity identifier, target entity identifier, and relationship type code of the relationship semantic tag; and the entity identifier, attribute name, and attribute value of the attribute semantic tag), they are considered duplicate tag entries, and only one of them is retained.
[0050] Then, the confidence score of the tag entry (for tag entries generated by the large language model, such as attribute semantic tags and supplementary relationship semantic tags) is checked. If the confidence score is lower than a preset confidence threshold (such as 0.8), the tag entry is determined to be low-confidence and removed. For example, if the extraction confidence score of an attribute semantic tag is 0.75, which is lower than 0.8, the tag entry is removed.
[0051] Step S125: Fusion of the de-redundant entity semantic labels, relationship semantic labels, and attribute semantic labels to generate a semantic data unit containing label type identification, associated data source, and timestamp information. Each label entry of the semantic data unit is associated with a data segment index in the original multi-source operation and maintenance data set, and the data segment index is used to trace the original data location corresponding to the label.
[0052] After removing redundancy, the entity, relationship, and attribute semantic labels need to be merged. Each label entry is assigned a label type identifier (e.g., "LABEL-ENT" for entity semantic labels, "LABEL-REL" for relationship semantic labels, and "LABEL-ATR" for attribute semantic labels), the associated data source (e.g., from structured performance indicator data, semi-structured service log data, or unstructured text description data), and timestamp information (obtained from the original data).
[0053] Each tag entry is also associated with a data segment index within the original multi-source O&M data set. This data segment index records the location of the original data corresponding to the tag entry within the multi-source O&M data set. For example, if an entity semantic tag originates from a metric such as "Server1's CPU usage" in the structured performance indicator data, its data segment index records the storage location of that metric within the structured performance indicator data set. This information is then used to generate semantic data units.
[0054] Step S130: constructing a semantic association network based on the semantic data unit, wherein the nodes of the semantic association network are the entity semantic labels, the directed edges between the nodes are the relationship semantic labels, and the edge weights are relationship strength parameters calculated based on the attribute semantic labels.
[0055] Semantic association networks can intuitively display the associations and strengths between system component entities. The construction process includes determining the node set and directed edge set, collecting node attribute features, calculating edge weights, and dynamically evolving the network.
[0056] Step S131: Extract all entity semantic tags from the semantic data unit as a network node set, assign a unique node identifier and corresponding component type attributes to each node, the component type attributes include hardware component type, software component type and external dependency type, and associate each node with the component name and unique identification code in the corresponding entity semantic tag.
[0057] Extract all entity semantic labels from the semantic data unit, remove duplicate entity semantic labels, and form a set of network nodes. Assign a unique node identifier to each node (the same encoding as the unique identifier in the entity semantic label can be used), and determine the component type attribute based on the component hierarchy in the entity semantic label.
[0058] For example, a node corresponding to an entity semantic label (Component Name: Server1, Component Level: Physical Device Layer - Server, Unique Identifier: PHY-SRV-001) has a unique node identifier of "PHY-SRV-001" and a component type attribute of hardware component type. A node corresponding to an entity semantic label (Component Name: ServiceA, Component Level: Application Service Layer - Business Service, Unique Identifier: APP-SVC-001) has a component type attribute of software component type. If a component depends on an external system (such as a third-party payment interface), its component type attribute is the external dependency type. Each node is associated with a corresponding component name and unique identifier to clearly identify the corresponding system component entity.
[0059] Step S132: Extract the relationship semantic tags in the semantic data unit to construct a set of directed edges between nodes. The directed edge set represents the direction of the interaction relationship between components. Each directed edge is associated with a corresponding relationship type identifier. The relationship type identifier includes a deployment relationship, a call relationship, and a dependency relationship. The source node and the target node of each directed edge correspond to the source entity identifier and the target entity identifier in the relationship semantic tag respectively.
[0060] Extract all relationship semantic labels from the semanticized data unit (including the relationship semantic labels generated in step S122 and the supplementary relationship semantic labels generated in step S1236). These labels will serve as directed edges between the nodes. The direction of each directed edge is determined by the source entity identifier and the target entity identifier. The node corresponding to the source entity identifier is the source node of the directed edge, and the node corresponding to the target entity identifier is the target node of the directed edge.
[0061] Each directed edge is associated with a corresponding relationship type identifier. These relationship type identifiers are determined during the relationship semantic label generation process, such as deployment relationships (e.g., an application service is deployed on a server), call relationships (e.g., Service A calls Service B), and dependency relationships (e.g., a service relies on the storage space provided by a storage device). For example, the directed edge corresponding to the relationship semantic label (source entity identifier: APP-SVC-001, target entity identifier: APP-SVC-002, relationship type code: REL-001) has a source node of APP-SVC-001, a target node of APP-SVC-002, and a relationship type identifier of a call relationship.
[0062] Step S133: Collect the attribute semantic tags in the semantic data unit as node attribute features, and the node attribute features include current state parameters, historical state parameters and threshold range parameters. The current state parameters are the attribute values corresponding to the latest timestamp, and the historical state parameters are the attribute value sequences within a preset time window.
[0063] Attribute semantic labels are extracted from semantic data units and associated with corresponding nodes according to entity identifiers, serving as attribute features for the nodes. The current state parameter refers to the attribute value corresponding to the latest timestamp in the attribute semantic label. For example, if the temperature attribute value corresponding to the latest timestamp in the attribute semantic label of a node (PHY-SRV-001) is "abnormally elevated," then the temperature parameter in the current state parameter for that node is "abnormally elevated."
[0064] A historical status parameter is a sequence of attribute values within a preset time window (such as the past hour). For example, for the temperature attribute of Server1, the attribute values recorded in chronological order by timestamp within the past hour might be "normal," "slightly increased," "significantly increased," or "abnormally increased," forming a historical status parameter sequence.
[0065] The threshold range parameter is the normal value range of the attribute. For example, the threshold range parameter of the server temperature may be "normal range: 20-35 degrees Celsius". When the current status parameter exceeds this range, it indicates that there may be an abnormality.
[0066] Step S134: Calculate the dynamic relationship weight of each directed edge based on the node attribute characteristics. The dynamic relationship weight is the product of the node's current state parameter and the relationship influence coefficient. The relationship influence coefficient is determined according to the preset weight matrix of the relationship type identifier. Different relationship type identifiers correspond to different initial values of the relationship influence coefficient.
[0067] Step S1341: extracting current state parameters from the node attribute characteristics as node activity indicators, wherein the node activity indicators include CPU utilization, memory occupancy and response time parameters. Nodes with different component type attributes correspond to different activity indicator combinations.
[0068] For hardware component nodes (such as servers), node activity indicators may include CPU utilization, memory usage, and disk I / O rate. For software component nodes (such as application services), node activity indicators may include response time, number of concurrent connections, and number of error requests. For externally dependent nodes, indicators may include interface call success rate, data transmission latency, etc. For example, if the current status parameters of a server node (PHY-SRV-001) show CPU utilization of "80%" and memory usage of "70%", its node activity indicators would be CPU utilization of 80% and memory usage of 70%.
[0069] Step S1342: querying a preset relationship influence coefficient matrix according to the relationship type identifier of the directed edge to obtain a basic influence coefficient corresponding to the relationship type identifier, wherein different relationship type identifiers in the relationship influence coefficient matrix correspond to basic influence coefficients of different numerical ranges.
[0070] The preset relationship impact coefficient matrix is pre-defined based on historical experience and system characteristics. The rows and columns in the matrix correspond to different relationship type identifiers, and the matrix elements are the basic impact coefficients corresponding to the relationship type identifier. For example, the basic impact coefficient corresponding to a call relationship may be between 0.6 and 0.8, the basic impact coefficient corresponding to a deployment relationship may be between 0.4 and 0.6, and the basic impact coefficient corresponding to a dependency relationship may be between 0.7 and 0.9. When the relationship type of a directed edge is a call relationship, querying this matrix will obtain the corresponding basic impact coefficient, such as 0.7.
[0071] Step S1343: Collect the threshold range parameters in the node attribute characteristics, calculate the ratio of the current state parameter to the upper threshold value as the state deviation, and the state deviation is the current state parameter divided by the upper threshold value, and the value range is 0 to 1.
[0072] For example, if the upper threshold value in the threshold range parameter is 100% and the CPU utilization in the current state parameter is 80%, the state deviation is 80% / 100% = 0.8. If the current state parameter is 120% (assuming it exceeds the threshold), the state deviation is 120% / 100% = 1.2. However, subsequent processing can limit this to the range of 0 to 1.
[0073] Step S1344: multiply the basic influence coefficient by the state deviation of the source node to obtain a preliminary relationship weight, which reflects the influence of the source node state on the relationship strength.
[0074] For example, if the basic influence coefficient is 0.7 and the state deviation of the source node is 0.8, the preliminary relationship weight is 0.7×0.8=0.56.
[0075] Step S1345: Collect the attribute semantic labels related to the directed edge in the semantic data unit, extract the change rate of the attribute value as a dynamic adjustment factor, and the change rate is the difference between the current timestamp attribute value and the previous timestamp attribute value divided by the time interval.
[0076] The attribute semantic labels associated with a directed edge are those that reflect changes in the strength of the relationship represented by the directed edge. For example, for a directed edge in which ServiceA calls ServiceB, the relevant attribute semantic label might be the frequency of ServiceA calling ServiceB. If the call frequency at the current timestamp is 100 times / minute, the call frequency at the previous timestamp was 80 times / minute, and the interval is 1 minute, then the rate of change is (100 - 80) / 1 = 20 times / minute.
[0077] Step S1346: Add the preliminary relationship weight and the dynamic adjustment factor to obtain the dynamic relationship weight. The numerical range of the dynamic relationship weight is limited to a preset interval through normalization processing. The normalization processing uses the maximum and minimum normalization method to map the weight value to the range of 0 to 1.
[0078] Assuming the initial relationship weight is 0.56 and the dynamic adjustment factor is 0.2 (after conversion, the call frequency change rate is converted to a value in the range of 0-1), the dynamic relationship weight is 0.56 + 0.2 = 0.76. If this value exceeds the preset range, it is mapped to the range of 0 to 1 through the maximum and minimum normalization method. For example, if the calculated result is 1.2, the maximum possible value is 1.5, and the minimum possible value is 0, the normalized value is 1.2 / 1.5 = 0.8.
[0079] Step S135: Integrate the node set, directed edge set and dynamic relationship weights into an initial semantic association network, and perform time slicing processing on the initial semantic association network according to the timestamp information of the semantic data unit to generate multi-period semantic association network snapshots arranged in chronological order, each semantic association network snapshot corresponds to the network status within a time window.
[0080] The node set determined in step S131, the directed edge set determined in step S132, and the dynamic relationship weights calculated in step S134 are integrated to form an initial semantic association network. This initial semantic association network is then time-sliced according to fixed time windows (e.g., one window every 5 minutes) based on the timestamp information of the semantic data units.
[0081] Each time window corresponds to a semantic association network snapshot. The snapshot contains the set of nodes, directed edges, and corresponding dynamic relationship weights within the time window, reflecting the association status between system component entities during that time period. For example, a snapshot from 10:00:00 to 10:05:00 contains the attribute characteristics of each node, the directed edges between nodes, and the dynamic relationship weights during that time period.
[0082] Step S136: Perform dynamic evolution analysis on the multi-period semantic association network snapshots, calculate the change in node attribute characteristics and the change rate of directed edge relationship weights between adjacent snapshots, and generate a semantic association network that includes the evolution trend of the network topology structure. The relationship weights in the semantic association network are updated in real time with the node attribute characteristics, and the update frequency is consistent with the collection frequency of the multi-source operation and maintenance data set.
[0083] For two adjacent semantically associated network snapshots (e.g., snapshots T1 and T2, where T2 occurs after T1), the change in each node's attribute feature is calculated as the attribute value at T2 minus the attribute value at T1. For example, if Server 1's temperature attribute value is "normal" at T1 and "slightly elevated" at T2, the change is "slightly elevated - normal" (calculated by converting textual descriptions into numerical values using pre-defined numerical conversion rules).
[0084] At the same time, the rate of change of the relationship weight of each directed edge is calculated: (relationship weight at time T2 - relationship weight at time T1) / (time interval T2 - T1). Using these changes and rates, we can analyze the evolution of the network topology, such as the changing trends of node attributes and the increase or decrease of relationship weights.
[0085] The generated semantic association network will update the relationship weights in real time according to the changes in node attribute characteristics. The update frequency is consistent with the collection frequency of multi-source operation and maintenance data sets (such as once every 30 seconds) to ensure that the network can reflect the latest status of the system in a timely manner.
[0086] Step S140: calling a pre-trained large language model to perform root cause reasoning analysis on the semantic association network to generate a set of candidate root causes sorted by confidence, wherein the candidate root cause set includes potential cause entities of the system anomaly and corresponding association path descriptions.
[0087] By analyzing the semantic association network through a large language model, we can discover the potential root causes and associated paths of system anomalies. This process includes the extraction of network topology data, the conversion of graph structure description text, and multi-round reasoning processing.
[0088] Step S141: extracting the network topology data at the current moment from the semantic association network, wherein the network topology data includes the current values of the node set, directed edge set, dynamic relationship weight and node attribute characteristics, and the current value is the node state parameter corresponding to the latest timestamp.
[0089] The current network topology data refers to the most recently updated content in the semantic association network, including information about all nodes (node identifiers, component type attributes, etc.), information about directed edges between nodes (source node, target node, relationship type identifier, etc.), the dynamic relationship weight of each directed edge, and the current value of each node's attribute characteristics (the state parameter corresponding to the latest timestamp). For example, the current node set includes PHY-SRV-001, APP-SVC-001, etc., and the directed edge includes a call relationship edge from APP-SVC-001 to APP-SVC-002 with a dynamic relationship weight of 0.7. The current attribute value of PHY-SRV-001 is "abnormally high temperature."
[0090] Step S142: Convert the network topology structure data into a graph structure description text that can be parsed by a large language model. The graph structure description text includes a natural language description of a node identifier-component type correspondence table, a directed edge relationship type list, and a dynamic relationship weight matrix. The natural language description uses a preset template to convert structured data into text paragraphs.
[0091] Preset templates are used to convert structured network topology data into natural language text. For example, a natural language description of a node identifier-component type table might read, "The component type of node PHY-SRV-001 is a hardware component type (server), and the component type of node APP-SVC-001 is a software component type (application service)..."; a description of a directed edge relationship type list might read, "There is a directed edge with a call relationship from node APP-SVC-001 to node APP-SVC-002, and a directed edge with a deployment relationship from node PHY-SRV-001 to node APP-SVC-001..."; and a description of a dynamic relationship weight matrix might read, "The dynamic relationship weight of the directed edge from APP-SVC-001 to APP-SVC-002 is 0.7, and the dynamic relationship weight of the directed edge from PHY-SRV-001 to APP-SVC-001 is 0.5...". These descriptions form a graph structure description text.
[0092] Step S143: calling a large language model to perform multiple rounds of reasoning on the graph structure description text, wherein the first round of reasoning generates an initial root cause hypothesis set, which includes potential abnormal nodes and corresponding abnormal propagation path descriptions.
[0093] A pre-trained large language model (such as a fine-tuned industry-specific large model) is called to process the text describing the graph structure. In the first round of inference, the model generates an initial set of root cause hypotheses based on its understanding of the text and the knowledge learned during training. For example, the model might infer that "the abnormally high temperature of node PHY-SRV-001 (Server1) is a potential abnormal node. This abnormality may propagate to node APP-SVC-001 (ServiceA) through the deployment relationship, resulting in a longer response time for ServiceA. This abnormality may then propagate to node APP-SVC-002 (ServiceB) through the call relationship, causing a connection timeout for ServiceB." This assumption is one of the initial root cause hypotheses.
[0094] Step S144: Based on the multi-period semantic association network snapshots of the semantic association network, perform temporal consistency verification on the abnormal propagation path in the initial root cause hypothesis set, and calculate the matching degree between the attribute feature change trend of each node in the abnormal propagation path during the abnormal period and the propagation path. The matching degree is the consistency score between the change direction of the node state parameter and the path propagation direction.
[0095] Step S1441: extracting a multi-period semantic association network snapshot of the abnormality occurrence period from the semantic association network, where the abnormality occurrence period is a preset time window before and after the system alarm triggering timestamp.
[0096] After a system alarm is triggered, the anomaly occurrence period is determined to be the preset time window before and after the alarm trigger timestamp (for example, if the alarm trigger time is 10:00:00 and the preset time window is 30 minutes before and after, the anomaly occurrence period is from 9:30:00 to 10:30:00). Multi-period semantic association network snapshots within this period are extracted from the semantic association network for subsequent temporal consistency verification.
[0097] Step S1442: For each abnormal propagation path sequence of the initial root cause hypothesis, extract the node identifiers in the abnormal propagation path in the propagation order and generate a node sequence list, where the first node in the node sequence list is the root cause entity and the last node is the abnormal manifestation entity.
[0098] For the anomaly propagation path described in the initial root cause hypothesis, node identifiers are extracted in propagation order. For example, if the anomaly propagation path is "PHY-SRV-001->APP-SVC-001->APP-SVC-002," the generated node sequence list is [PHY-SRV-001, APP-SVC-001, APP-SVC-002], where PHY-SRV-001 is the root cause entity and APP-SVC-002 is the anomaly manifestation entity.
[0099] Step S1443: For each node in the node sequence list, extract the attribute feature time series of the corresponding node during the abnormality occurrence period from the multi-period semantic association network snapshot, where the attribute feature time series is a set of state parameter values sorted by timestamp.
[0100] Extract the attribute feature time series of each node in the node sequence list from the multi-period semantic association network snapshot during the anomaly period. For example, for node PHY-SRV-001, the time series of its temperature attribute during the anomaly period might be [normal, slightly increased, significantly increased, abnormally increased], arranged in chronological order by timestamp.
[0101] Step S1444: Calculate the change trend slope of the attribute feature time series. The change trend slope is the slope value obtained by linear regression fitting. A positive value indicates that the state parameter increases, and a negative value indicates that it decreases.
[0102] Perform linear regression fitting on the attribute feature time series to obtain the slope of the change trend. For example, after converting the temperature attribute time series of PHY-SRV-001 into numerical values (e.g., normal = 1, slightly increased = 2, significantly increased = 3, abnormally increased = 4), a linear regression fit is performed and a positive slope is obtained, indicating an upward trend in temperature.
[0103] Step S1445: Based on the propagation direction of the abnormal propagation path, determine whether the change trend slope of the adjacent nodes meets the causal relationship expectations.
[0104] Anomaly propagation paths have a specific direction, such as from PHY-SRV-001 to APP-SVC-001 and then to APP-SVC-002. Causal relationships are expected to follow: if the attribute characteristics of the root cause node show an upward trend (such as rising temperature), the related attribute characteristics of the next node to which it propagates should also show an upward trend (such as longer response time). For example, if the temperature trend of PHY-SRV-001 has a positive slope, the response time trend of APP-SVC-001 is also expected to have a positive slope. If this is the case, then causal relationships are expected to follow.
[0105] Step S1446: Count the ratio of the number of adjacent node pairs that meet the causal relationship expectations in the abnormal propagation path to the total number of node pairs as the matching degree. The higher the matching degree, the stronger the temporal consistency of the abnormal propagation path.
[0106] For example, if there are three nodes in the node sequence list, forming two pairs of adjacent nodes, if one of the pairs meets the causal relationship expectation, the matching degree is 1 / 2 = 0.5; if both pairs meet the causal relationship expectation, the matching degree is 1.
[0107] Step S145: adjusting the confidence score of each hypothesis in the initial root cause hypothesis set according to the matching degree, retaining hypotheses with matching degrees higher than a preset matching degree threshold, and generating an intermediate root cause hypothesis set, wherein the confidence score is adjusted to a weighted sum of the original score and the matching degree.
[0108] The preset match threshold can be set to 0.6. For each hypothesis in the initial set of root cause hypotheses, its original confidence score (given by the first round of inference on the large language model) and the match are weighted and summed (for example, with weights of 0.7 and 0.3, respectively) to obtain an adjusted confidence score. If the match is higher than 0.6, the hypothesis is retained in the intermediate set of root cause hypotheses. For example, if an initial root cause hypothesis has an original confidence score of 0.8 and a match of 0.7, the adjusted confidence score after weighted summation is 0.8 × 0.7 + 0.7 × 0.3 = 0.56 + 0.21 = 0.77. Since the match of 0.7 is higher than the preset threshold of 0.6, the hypothesis is retained. If another hypothesis has a match of 0.5, which is lower than 0.6, it is excluded regardless of its original confidence score.
[0109] Step S146: Call the large language model to perform conflict detection processing on the intermediate root cause hypothesis set, identify and remove hypothesis entries with logical contradictions, where the logical contradiction refers to the existence of mutually exclusive propagation path descriptions for the same abnormal node, and the mutual exclusivity refers to opposite propagation directions or conflicting relationship types.
[0110] Step S1461: Convert each hypothesis entry in the intermediate root cause hypothesis set into a logical expression, where the logical expression includes a root cause entity, a propagation path condition, and an abnormal phenomenon description.
[0111] Each hypothesis entry contains information such as the root cause entity, propagation path, and abnormal phenomenon. This information is converted into a logical expression. For example, a hypothesis entry such as "Root cause entity PHY-SRV-001, through the path PHY-SRV-001->APP-SVC-001->APP-SVC-002, causes a connection timeout on APP-SVC-002" is converted into the logical expression "Root cause entity = PHY-SRV-001 ∧ Propagation path = PHY-SRV-001--APP-SVC-001--APP-SVC-002 ∧ Abnormal phenomenon = APP-SVC-002 connection timeout."
[0112] Step S1462: Construct a conflict detection prompt template, which includes a task description, a conflict definition, and a list of hypothesis items.
[0113] The task description is "Please check whether there are logical conflicts between the following hypothesis entries." A conflict is defined as "mutually exclusive propagation path descriptions for the same anomaly node, including those with opposite propagation directions or conflicting relationship types." The hypothesis entry list is the logical expression corresponding to all hypothesis entries in the intermediate root cause hypothesis set. For example, a conflict detection prompt template might be "Please check whether there are logical conflicts between the following hypothesis entries. Conflict definition: Mutually exclusive propagation path descriptions for the same anomaly node, including those with opposite propagation directions or conflicting relationship types. Hypothesis entry list: 1. Root cause entity = PHY-SRV-001 ∧ Propagation path = PHY-SRV-001--APP-SVC-001--APP-SVC-002 ∧ Anomaly = APP-SVC-002 connection timeout; 2. Root cause entity = APP-SVC-003 ∧ Propagation path = APP-SVC-003--APP-SVC-002 ∧ Anomaly = APP-SVC-002 connection timeout..."
[0114] Step S1463: splice the logical expression and the conflict detection prompt template into a model input text, call the large language model to perform conflict judgment on the model input text, and generate a conflict matrix between the hypothesis entries. The conflict matrix is a two-dimensional matrix, and the matrix elements indicate whether there is a conflict between the corresponding two hypothesis entries.
[0115] The logical expressions of all hypothesis entries are concatenated into model input text in the format of the conflict detection prompt template, and the large language model is called for processing. The model analyzes the logical relationships between the hypothesis entries to determine whether there is a conflict between any two hypothesis entries and generates a conflict matrix. The rows and columns of the conflict matrix correspond to the hypothesis entries. If the matrix element is 1, it means that there is a conflict between the corresponding two hypothesis entries; if it is 0, it means that there is no conflict. For example, for two hypothesis entries, if they both involve the abnormal node APP-SVC-002, but the propagation paths are in opposite directions, the corresponding element in the conflict matrix is 1.
[0116] Step S1464: constructing a conflict graph based on the conflict matrix, wherein the nodes of the conflict graph are hypothetical entries and the directed edges represent the conflict relationships between the entries.
[0117] Based on the conflict matrix, a conflict graph is constructed. Each node in the conflict graph represents a hypothesis entry. If two hypothesis entries conflict (i.e., the corresponding element in the conflict matrix is 1), a directed edge is added between the two nodes to indicate the conflict relationship. For example, if hypothesis entry 1 and hypothesis entry 2 conflict, a directed edge connects nodes 1 and 2 in the conflict graph.
[0118] Step S1465: performing maximum conflict-free subset selection on the conflict graph, and selecting a subset containing the largest number of hypothesis entries without mutual conflicts.
[0119] Analyze the conflict graph using a graph theory algorithm (such as a greedy algorithm) to select a maximal conflict-free subset. This maximal conflict-free subset contains the largest number of hypothetical entries, with no conflicts between any two entries. For example, if there are five hypothetical entry nodes in the conflict graph, where node 1 conflicts with nodes 2 and 3, node 2 conflicts with node 4, and node 3 conflicts with node 5, the algorithm may select a maximal conflict-free subset consisting of nodes 1, 4, and 5 (assuming no conflicts between them).
[0120] Step S1466: remove the hypothesis entries that are not selected into the maximum conflict-free subset from the intermediate root cause hypothesis set to generate a conflict-free intermediate root cause hypothesis set.
[0121] Based on the selection result of the maximum conflict-free subset, remove the unselected hypothesis entries from the intermediate root cause hypothesis set. For example, if the maximum conflict-free subset contains hypothesis entries 1, 4, and 5, then remove hypothesis entries 2 and 3 from the intermediate root cause hypothesis set to obtain the conflict-free intermediate root cause hypothesis set.
[0122] Step S1467: recalculate the confidence scores of the deconflicted hypothesis entries, increase the scores of the hypothesis entries verified to be conflict-free during the conflict detection process by a preset reward value, and improve their sorting priority.
[0123] The default reward value can be set to 0.1. For each deconflicted hypothesis, the reward value is increased based on its previously adjusted confidence score. For example, if the adjusted confidence score of a hypothesis is 0.77, an increase of 0.1 will bring it to 0.87, giving it a higher priority in subsequent sorting.
[0124] Step S147: Sort the intermediate root cause hypothesis set after conflict resolution from high to low according to the confidence score, and extract a preset number of hypothesis entries before sorting as a candidate root cause set, wherein the candidate root cause set includes a root cause entity identifier, an abnormal propagation path sequence and a confidence score, and the abnormal propagation path sequence is a list of node identifiers arranged in propagation order.
[0125] The preset number can be set based on actual needs, such as 5. After de-conflicting and recalculating the confidence scores, the hypotheses are sorted from high to low by score, and the top 5 hypotheses are selected as the candidate root cause set. Each candidate root cause includes a root cause entity identifier (such as PHY-SRV-001), an anomaly propagation path sequence (such as [PHY-SRV-001, APP-SVC-001, APP-SVC-002]), and a corresponding confidence score (such as 0.87).
[0126] Step S150: Generate an operation and maintenance decision instruction including an entity operation sequence and a priority ranking according to the candidate root cause set, and send the operation and maintenance decision instruction to the system management terminal to trigger an automated repair process.
[0127] Generating operational decision instructions involves converting a set of candidate root causes into specific operational steps to fix system anomalies. This process requires identifying target entities, querying operational flows, adjusting the order of operations, and setting priorities.
[0128] Step S151: parse the root cause entity identifier and the abnormal propagation path sequence corresponding to each candidate root cause in the candidate root cause set, determine the target entity set that needs to perform intervention operations and the dependency order between entities, and the dependency order is the reverse order of the abnormal propagation path sequence.
[0129] The root cause entity identifier and anomaly propagation path sequence for each candidate root cause are parsed to determine the target entity requiring intervention. For example, if the anomaly propagation path sequence for a candidate root cause is [PHY-SRV-001, APP-SVC-001, APP-SVC-002], the target entity set is {PHY-SRV-001, APP-SVC-001, APP-SVC-002}. The dependency order between entities is the reverse order of the anomaly propagation path sequence, i.e., APP-SVC-002 -- APP-SVC-001 -- PHY-SRV-001. This means that operations must be performed on APP-SVC-002 first, then on APP-SVC-001, and finally on PHY-SRV-001. This ensures that operations are performed sequentially from the end entity to the root cause entity, gradually eliminating the anomaly.
[0130] Step S152: Based on the target entity set, query the preset operation and maintenance operation manual to obtain the standard operating procedure corresponding to each target entity. The standard operating procedure includes an operation step sequence, precondition constraints and expected effect description. The precondition constraints are the entity state parameter range that must be met before performing the operation.
[0131] The pre-set operation and maintenance manual stores the standard operating procedures for each system component entity. For example, for the target entity PHY-SRV-001 (server), the standard operating procedure might be: the sequence of steps is [Check cooling fan status, clean dust inside the server, restart the server]; the precondition is "The server's current load rate is less than 30%"; and the expected result is "The server temperature returns to the normal range (20-35 degrees Celsius)." For APP-SVC-001 (application service), the standard operating procedure might be [Stop service, check configuration files, restart service]; the precondition is "No other services depend on this service to run"; and the expected result is "Service response time returns to normal range (<1 second)."
[0132] Step S153: Call the large language model to perform adaptability analysis on the standard operating process and the exception propagation path sequence, and adjust the execution order of the operating steps to match the reverse path of the exception propagation, so that the operations are executed in order from the root cause entity to the end entity.
[0133] For example, step S1531: converting the sequence of operation steps of the standard operating procedure into an operation directed graph, where the nodes of the operation directed graph are operation steps, and the directed edges represent the dependency relationship between the steps.
[0134] For example, the standard operating procedure sequence for PHY-SRV-001 is [Check cooling fan status, clean dust inside the server, and restart the server]. When converted into an operation directed graph, the nodes are Step 1 (Check cooling fan status), Step 2 (Clean dust inside the server), and Step 3 (Restart the server). The directed edges are Step 1-Step 2 (indicating that checking cooling fan status is required before cleaning dust), and Step 2-Step 3 (indicating that cleaning dust is required before restarting the server).
[0135] Step S1532: extracting the reverse path of the abnormal propagation path sequence as the target operation sequence, wherein the reverse path is the node sequence from the last node to the first node of the abnormal propagation path sequence.
[0136] The exception propagation path sequence is [PHY-SRV-001, APP-SVC-001, APP-SVC-002], and its reverse path is [APP-SVC-002, APP-SVC-001, PHY-SRV-001], that is, the target operation sequence is to perform the corresponding operations in the order of APP-SVC-002, APP-SVC-001, and PHY-SRV-001.
[0137] Step S1533: Associating a corresponding target entity identifier with each operation step in the operation directed graph, and determining the system component entity to which each step acts.
[0138] For example, in the operation directed graph of PHY-SRV-001, step 1, step 2, and step 3 are all associated with the target entity identifier PHY-SRV-001; the steps in the operation directed graph of APP-SVC-001 are all associated with the target entity identifier APP-SVC-001.
[0139] Step S1534: Match the target operation sequence with the node target entity identifier of the operation directed graph, identify the corresponding relationship between the operation steps and the target entity, and generate an entity-step mapping table.
[0140] The target operation sequence is [APP-SVC-002, APP-SVC-001, PHY-SRV-001]. After matching with the target entity identifiers of each operation directed graph, an entity-step mapping table is generated. For example, APP-SVC-002 corresponds to the step sequence in its operation directed graph, APP-SVC-001 corresponds to the step sequence in its operation directed graph, and PHY-SRV-001 corresponds to the step sequence in its operation directed graph.
[0141] Step S1535: constructing a compatibility analysis prompt template, wherein the compatibility analysis prompt template includes an operation directed graph description, a target operation sequence description, and adaptability adjustment requirements.
[0142] The operation directed graph description is a textual description of the operation directed graph corresponding to each target entity; the target operation sequence is described as "the target operation sequence is APP-SVC-002--APP-SVC-001--PHY-SRV-001"; the adaptability adjustment requirement is "adjust the execution order of the operation steps in each operation directed graph so that its overall execution order is consistent with the target operation sequence and does not violate the dependency relationship between the steps."
[0143] Step S1536: input the operation directed graph, target operation sequence and entity-step mapping table into the adaptability analysis prompt template, call the large language model to perform sequence planning processing on the adaptability analysis prompt template, and generate an adjusted operation step sequence.
[0144] Based on the input information, the large language model adjusts the order of steps in each operation directed graph to ensure that the overall operation sequence matches the target operation sequence. For example, the operation step sequence of APP-SVC-002 is executed first, followed by the operation step sequence of APP-SVC-001, and finally the operation step sequence of PHY-SRV-001, while ensuring that the step dependencies within each operation directed graph remain unchanged.
[0145] Step S1537: Verify the dependencies of the adjusted operation step sequence to check whether the step dependencies in the original operation directed graph are violated. If there is a step sequence that violates the dependencies, re-call the large language model for local adjustments.
[0146] For example, if in the adjusted operation step sequence, step 3 (restarting the server) of PHY-SRV-001 is executed before step 2 (cleaning the dust inside the server), this violates the dependency relationship in the original operation directed graph. The large language model needs to be re-called to locally adjust the order of these steps to ensure that step 2 is executed before step 3.
[0147] Step S1538: Calculate the matching degree between the entity sequence corresponding to the adjusted operation step sequence and the target operation sequence, where the matching degree is the ratio of the length of the continuous subsequences with consistent entity sequences to the total length of the target operation sequence.
[0148] The total length of the target operation sequence is 3 (containing 3 entities). If the entity order corresponding to the adjusted operation step sequence is [APP-SVC-002, APP-SVC-001, PHY-SRV-001], which is exactly the same as the target operation sequence, the match degree is 3 / 3 = 1. If the entity order is [APP-SVC-002, PHY-SRV-001, APP-SVC-001], then the continuous subsequence [APP-SVC-002] is consistent with the target sequence, has a length of 1, and the match degree is 1 / 3 ≈ 0.33.
[0149] Step S1539: If the matching degree reaches the preset matching degree threshold, the adjusted operation step sequence is determined to be the adapted operation process; if the matching degree does not reach the preset matching degree threshold, the constraints in the adaptability analysis prompt template are added, and the sequential planning process is re-executed until the matching degree meets the standard or the maximum number of retries is reached.
[0150] The preset matching threshold can be set to 0.8. If the matching degree is 1, which meets the threshold, the operation step sequence is determined to be the adapted operation flow. If the matching degree is 0.33, which does not meet the threshold, constraints are added to the adaptability analysis prompt template (e.g., "The operation steps of each entity must be arranged strictly in the order APP-SVC-002--APP-SVC-001--PHY-SRV-001"), and the large language model is re-invoked for sequence planning. If the matching degree still does not meet the standard after multiple retries (e.g., five times), the operation step sequence with the highest matching degree is used.
[0151] Step S154: Determine the priority weight of each target entity operation based on the confidence score of the candidate root cause set. The higher the confidence score, the greater the corresponding priority weight. The priority weight is used to determine the execution order of operations corresponding to different candidate root causes.
[0152] Normalize the confidence scores of each candidate root cause in the candidate root cause set to obtain the priority weight for each target entity operation. For example, if the confidence score of candidate root cause 1 is 0.87 and the confidence score of candidate root cause 2 is 0.75, the total score is 0.87 + 0.75 = 1.62. Therefore, the priority weight for the target entity operation corresponding to candidate root cause 1 is 0.87 / 1.62 (≈ 0.54), and the priority weight for candidate root cause 2 is 0.75 / 1.62 (≈ 0.46). The larger the priority weight, the sooner the corresponding operation is executed.
[0153] Step S155: Generate a preliminary operation and maintenance operation sequence based on the priority weight and the adjusted operation step sequence, wherein the preliminary operation and maintenance operation sequence includes an operation object identifier, an operation type, an execution time window and an expected state parameter, and the execution time window is determined according to the entity dependency order.
[0154] The operation object identifier is the unique identifier of the target entity (e.g., PHY-SRV-001). The operation type is the specific operation action (e.g., check, clean, restart, etc.). The execution time window is determined by the entity dependency order and the estimated duration of the operation steps. For example, the estimated duration of the APP-SVC-002 operation is 5 minutes, and its execution time window is 10:00:00-10:05:00. The APP-SVC-001 operation begins after the APP-SVC-002 operation completes, with an estimated duration of 3 minutes and an execution time window of 10:05:00-10:08:00. The PHY-SRV-001 operation begins after the APP-SVC-001 operation completes, with an estimated duration of 10 minutes and an execution time window of 10:08:00-10:18:00. The expected state parameter is the state that the entity should reach after the operation completes (e.g., the expected temperature of PHY-SRV-001 is 20-35 degrees Celsius).
[0155] Step S156: Conflict detection is performed on the preliminary operation and maintenance operation sequence to identify operation steps with the same operation objects or resource competition, and the execution order of the conflicting operations is adjusted based on priority weights.
[0156] If two steps in the initial O&M sequence target the same operation object (for example, simultaneously restarting and checking PHY-SRV-001), or if the steps require the same resource (such as a maintenance terminal), a conflict exists. The execution order of conflicting operations is adjusted based on their priority weights, with the operation with the higher priority weight executed first. For example, if the operation corresponding to Candidate Root Cause 1 competes for resources with the operation corresponding to Candidate Root Cause 2, and Candidate Root Cause 1 has a higher priority weight, the operation corresponding to Candidate Root Cause 1 is executed first.
[0157] Step S157: Integrate the operation and maintenance operation sequence and priority weight after conflict resolution to generate an operation and maintenance decision instruction including the operation sequence number, target entity identifier, operation instruction content, priority ranking and expected completion time. The priority ranking of the operation and maintenance decision instruction is positively correlated with the confidence score of the candidate root cause set, and the expected completion time is calculated based on the estimated execution time of the operation steps.
[0158] When integrating the post-conflict resolution operation sequence with the priority weight, each operation step is first assigned a unique operation sequence number. This operation sequence number increases in the order in which the operations are performed. For example, if the post-conflict resolution operation sequence is to cool down PHY-SRV-001, restart APP-SVC-001, and check the connection status of APP-SVC-002, the corresponding operation sequence numbers are 001, 002, and 003, respectively.
[0159] The target entity identifier is the unique identifier of the node targeted by the operation, such as PHY-SRV-001, APP-SVC-001, etc. The operation instruction content is a specific description of the operation, such as "Enable the cooling fan of PHY-SRV-001 to run at maximum power to reduce the temperature of the device," "Restart the APP-SVC-001 service," and "Check whether the connection between APP-SVC-002 and APP-SVC-001 has returned to normal."
[0160] Prioritization is determined by priority weight, with steps with higher priority weights placed higher in the ranking. Because priority weights are positively correlated with the confidence scores of the candidate root cause set, steps corresponding to candidate root causes with higher confidence scores are ranked higher in priority.
[0161] To calculate the expected completion time, you must first determine the estimated execution time for each operation step. This estimated execution time can be set based on historical operation records or empirical data. For example, the estimated execution time for cooling PHY-SRV-001 is 10 minutes, the estimated execution time for restarting APP-SVC-001 is 5 minutes, and the estimated execution time for checking the connection status of APP-SVC-002 is 3 minutes. Based on the start time of the first operation step, the expected completion time for the first operation is the start time plus 10 minutes, the expected completion time for the second operation is the expected completion time of the first operation plus 5 minutes, the expected completion time for the third operation is the expected completion time of the second operation plus 3 minutes, and so on, cumulatively calculating the time.
[0162] Through the above method, a complete operation and maintenance decision instruction is generated, which contains all the key information required to perform operation and maintenance operations and can guide the system management terminal to accurately and orderly execute the automated repair process.
[0163] Step S160: Send the operation and maintenance decision instruction to the system management terminal to trigger an automated repair process.
[0164] Before sending the operation and maintenance decision instructions to the system management terminal, the instructions need to be converted to a format that the system management terminal can recognize (such as XML format or a specific API call format). After the conversion is completed, the instructions are sent to the system management terminal via a secure communication protocol (such as HTTPS protocol).
[0165] After receiving the operation and maintenance decision-making instructions, the system management terminal parses the instructions and executes the corresponding operation steps in sequence according to the operation sequence number and priority. During the execution process, the system management terminal provides real-time feedback on the operation execution status (such as "Operation 001 Executing," "Operation 001 Completed," "Operation 002 Failed," etc.). If an operation fails, the system management terminal will retry according to the preset retry mechanism. If multiple retries fail, an alarm will be sent to the fault reasoning and analysis system for further processing.
[0166] Through the above process, a closed loop from fault reasoning analysis to automated repair is achieved, improving the efficiency and accuracy of distributed server cluster fault handling.
[0167] Furthermore, the method may also include step S210: a training step of pre-training a large language model.
[0168] Step S211: Collecting operation and maintenance domain corpus data, wherein the operation and maintenance domain corpus data includes historical fault handling records, system operation and maintenance manuals, technical documents, etc.
[0169] Collected historical troubleshooting records include descriptions of past failures, root cause analysis, handling procedures, and results. System operation and maintenance manuals detail the operation and maintenance specifications and troubleshooting methods for various system components. Technical documentation includes system architecture design documents, component interface documentation, and performance indicator specifications. This data needs to be screened and cleaned to remove invalid information (such as duplicate content and ambiguous descriptions) and sensitive information (such as data involving user privacy). Sensitive information is processed using data desensitization techniques, such as replacing key identifiers and obfuscating, to ensure data security.
[0170] Step S212: pre-processing the collected operation and maintenance domain corpus data, including word segmentation, part-of-speech tagging, entity recognition, and relationship extraction.
[0171] Word segmentation breaks down continuous text into independent words or lexical units for easier model processing. Part-of-speech tagging labels each word with its part of speech (e.g., noun, verb, adjective, etc.). Entity recognition identifies entities such as system components and fault types mentioned in the text. Relationship extraction extracts relationships between entities (e.g., "deployed on," "called," "caused by," etc.). For example, the text "ServiceA is deployed on Server1. When Server1 experiences a CPU overload, ServiceA's response may be delayed" is preprocessed and segmented to obtain "ServiceA / deployed / on / Server1 / . When / Server1 / experiences a CPU overload / , it will / cause / ServiceA / response / delay / ." The part-of-speech tagging is "ServiceA (noun) / deployed (verb) / on (preposition) / Server1 (noun) / on (particle) / ..."; entity recognition identifies entities such as "ServiceA," "Server1," "CPU overload," and "response delay." Relationship extraction obtains relationships such as "ServiceA-deployed-on-Server1," "Server1-experiences-CPU-overload," and "CPU overload-causes-response delay."
[0172] Step S213: constructing a pre-training task, wherein the pre-training task includes a masked language model task, a next sentence prediction task, and an operation and maintenance knowledge question answering task.
[0173] The masked language model task randomly masks some words in the corpus and allows the model to predict the masked words to enhance its understanding of contextual semantics. The next sentence prediction task allows the model to determine whether two sentences form a continuous context to enhance its understanding of the logical relationship between sentences. The operation and maintenance knowledge question-answering task constructs question-answer pairs based on corpus data in the operation and maintenance field (e.g., "Q: Which application services may be abnormal due to server CPU overload? A: It may cause response delays and decreased concurrent processing capabilities of application services deployed on this server"), allowing the model to generate answers based on the questions, thereby enhancing the model's ability to apply operation and maintenance knowledge.
[0174] Step S214: Based on the pre-processed operation and maintenance domain corpus data and the constructed pre-training task, pre-train the basic large language model and adjust the model parameters until the performance of the model on the pre-training task reaches the preset indicators.
[0175] The basic large language model can use a general Transformer architecture model. During pre-training, pre-processed corpus data is fed into the Transformer architecture model, the loss function is calculated according to the pre-training task requirements, and the model weight parameters are adjusted using the backpropagation algorithm. Pre-set metrics include prediction accuracy on the masked language model task, classification accuracy on the next sentence prediction task, and answer similarity on the operation and maintenance knowledge question-and-answer task. When the model's performance on these metrics reaches the preset threshold, pre-training is terminated, resulting in a pre-trained large language model suitable for the operation and maintenance domain.
[0176] Step S215: Use partially labeled fault reasoning samples to fine-tune the pre-trained large language model to further optimize the performance of the model on the fault reasoning task.
[0177] The labeled fault inference samples include semantic association network descriptions, correct root cause inference results, and operational decision recommendations. During fine-tuning, the semantic association network descriptions in the samples serve as model inputs, while the correct root cause inference results and operational decision recommendations serve as the expected outputs. The loss between the model output and the expected output is calculated, and the model parameters are adjusted using gradient descent. After fine-tuning, the model is evaluated. If the model meets the requirements for accuracy and recall on the fault inference task, it is selected as the final large language model for fault inference analysis.
[0178] Figure 2 A schematic diagram illustrates exemplary hardware and software components of a large language model-based fault reasoning and analysis system 100 that can implement the concepts of the present application, as provided in some embodiments of the present application. For example, the processor 120 can be used in the large language model-based fault reasoning and analysis system 100 to perform the functions of the present application.
[0179] For example, the fault reasoning analysis system 100 based on a large language model may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and storage media 140 in different forms, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the fault reasoning analysis system 100 based on a large language model may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The fault reasoning analysis system 100 based on a large language model also includes an I / O interface 150 between the computer and other input and output devices.
[0180] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When a processor executes the computer-executable instructions, the above-mentioned fault reasoning and analysis method based on the large language model is implemented.
[0181] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.
Claims
1. A fault reasoning analysis method based on a large language model, characterized in that: The method comprises: Acquire a multi-source operation and maintenance data set of the system to be diagnosed, wherein the multi-source operation and maintenance data set includes structured performance indicator data, semi-structured service log data, and unstructured text description data sorted by timestamp; Performing semantic enhancement on the multi-source operation and maintenance data set to obtain a semantic data unit including entity semantic labels, relationship semantic labels, and attribute semantic labels, wherein the entity semantic labels correspond to system component entities, the relationship semantic labels correspond to interaction relationships between components, and the attribute semantic labels correspond to entity state attributes; Constructing a semantic association network based on the semantic data unit, wherein the nodes of the semantic association network are the entity semantic labels, the directed edges between the nodes are the relationship semantic labels, and the edge weights are the relationship strength parameters calculated based on the attribute semantic labels; Calling a pre-trained large language model to perform root cause reasoning analysis on the semantic association network to generate a set of candidate root causes sorted by confidence, wherein the candidate root cause set includes potential cause entities of the system anomaly and corresponding association path descriptions; An operation and maintenance decision instruction including an entity operation sequence and a priority ranking is generated according to the candidate root cause set, and the operation and maintenance decision instruction is sent to a system management terminal to trigger an automated repair process.
2. The fault reasoning analysis method based on a large language model according to claim 1 is characterized in that: The semantic enhancement of the multi-source operation and maintenance data set to obtain a semantic data unit including entity semantic labels, relationship semantic labels, and attribute semantic labels includes: Performing indicator semantic mapping processing on the structured performance indicator data, extracting system component keywords from the indicator name, and generating entity semantic tags corresponding to the system component keywords based on a preset entity dictionary, wherein the entity semantic tags include the component name, component level, and unique identification code; Perform log field structuring on semi-structured service log data, parse the source component field, target component field, and interaction field in the log entry, and generate relationship semantic tags representing the call direction and interaction type between components based on the verb meaning of the interaction field. The relationship semantic tags contain the source entity identifier, target entity identifier, and relationship type code. Calling a large language model to perform semantic parsing on the unstructured text description data, identifying entity mentions, attribute descriptions, and relationship assertions therein, matching the entity mentions with the entity semantic labels for consistency, and mapping the attribute descriptions and relationship assertions into attribute semantic labels and supplementary relationship semantic labels containing attribute names, attribute values, and timestamps, respectively; Performing redundancy detection on the entity semantic labels, relationship semantic labels, and supplementary relationship semantic labels to remove duplicate label entries and low-confidence label entries, where the low-confidence label entries are labels whose confidence scores output by the large language model are lower than a preset confidence threshold; The de-redundant entity semantic labels, relationship semantic labels and attribute semantic labels are integrated to generate a semantic data unit containing label type identification, associated data source and timestamp information. Each label entry in the semantic data unit is associated with the data segment index in the original multi-source operation and maintenance data set, and the data segment index is used to trace the original data location corresponding to the label.
3. The fault reasoning analysis method based on a large language model according to claim 2 is characterized in that: The calling of a large language model to perform semantic parsing on the unstructured text description data to identify entity mentions, attribute descriptions, and relationship assertions therein includes: Performing text preprocessing on the unstructured text description data to remove repeated characters therein and segmenting the data into a plurality of text segment units according to sentence boundaries, wherein each text segment unit is associated with a position index of the original unstructured text description data; Constructing a semantic parsing prompt template for each text segment unit, wherein the semantic parsing prompt template includes a task description, an entity type list, a relationship type list, and an attribute type list; The text segment unit and the semantic parsing hint template are spliced into a model input text, and a large language model is called to perform sequence annotation processing on the model input text to generate an annotation result sequence including entity annotation, attribute annotation and relationship annotation, wherein the entity annotation includes the entity name and entity type, the attribute annotation includes the attribute name, attribute value and the entity to which it belongs, and the relationship annotation includes the source entity, the target entity and the relationship type; Performing entity linking on the annotation result sequence, matching the identified entity mentions with standard entity names in a preset entity dictionary to generate standardized entity identifiers, wherein the preset entity dictionary contains the standard names and unique identification codes of all component entities in the system; Based on the standardized entity identifier, attribute annotations are associated with the corresponding entities to generate attribute semantic labels containing entity identifiers, attribute names, attribute values, and extraction confidence; Replace the source entity and target entity in the relationship annotation with standardized entity identifiers, and generate a supplementary relationship semantic label containing the source entity identifier, target entity identifier, relationship type, and extraction confidence; The attribute semantic tag and the supplementary relationship semantic tag are integrated to generate a semantic parsing result set corresponding to the text segment unit, wherein each tag entry in the semantic parsing result set is associated with a position index of the text segment unit.
4. The fault reasoning analysis method based on a large language model according to claim 1, characterized in that: The step of constructing a semantic association network based on the semantic data unit includes: Extract all entity semantic tags from the semantic data unit as a set of network nodes, assign a unique node identifier and a corresponding component type attribute to each node, wherein the component type attribute includes a hardware component type, a software component type, and an external dependency type, and associate each node with the component name and unique identification code in the corresponding entity semantic tag; Extract the relationship semantic tags in the semantic data unit to construct a directed edge set between nodes, the directed edge set represents the direction of the interaction relationship between components, each directed edge is associated with a corresponding relationship type identifier, the relationship type identifier includes a deployment relationship, a call relationship and a dependency relationship, and the source node and target node of each directed edge correspond to the source entity identifier and the target entity identifier in the relationship semantic tag respectively; Collecting attribute semantic tags in the semantic data unit as node attribute features, wherein the node attribute features include current state parameters, historical state parameters, and threshold range parameters. The current state parameters are attribute values corresponding to the latest timestamp, and the historical state parameters are attribute value sequences within a preset time window. Calculate the dynamic relationship weight of each directed edge based on the node attribute characteristics, where the dynamic relationship weight is the product of the node's current state parameter and the relationship influence coefficient, and the relationship influence coefficient is determined according to a preset weight matrix of the relationship type identifier, with different relationship type identifiers corresponding to different initial values of the relationship influence coefficient; Integrating the node set, directed edge set, and dynamic relationship weights into an initial semantic association network, and performing time slicing processing on the initial semantic association network according to the timestamp information of the semantic data unit to generate multi-period semantic association network snapshots arranged in chronological order, each semantic association network snapshot corresponding to a network state within a time window; A dynamic evolution analysis is performed on the multi-period semantic association network snapshots, and the change amount of the node attribute characteristics and the change rate of the directed edge relationship weights between adjacent snapshots are calculated to generate a semantic association network containing the evolution trend of the network topology structure. The relationship weights in the semantic association network are updated in real time with the node attribute characteristics, and the update frequency is consistent with the collection frequency of the multi-source operation and maintenance data set.
5. The fault reasoning analysis method based on a large language model according to claim 4 is characterized in that: The calculating the dynamic relationship weight of each directed edge based on the node attribute characteristics includes: Extracting current state parameters from the node attribute features as node activity indicators, wherein the node activity indicators include CPU utilization, memory occupancy, and response time parameters. Nodes with different component type attributes correspond to different activity indicator combinations. Querying a preset relationship influence coefficient matrix according to the relationship type identifier of the directed edge to obtain a basic influence coefficient corresponding to the relationship type identifier, wherein different relationship type identifiers in the relationship influence coefficient matrix correspond to basic influence coefficients with different numerical ranges; Collecting the threshold range parameters in the node attribute characteristics, calculating the ratio of the current state parameter to the upper threshold as the state deviation, wherein the state deviation is the current state parameter divided by the upper threshold, and the value range is 0 to 1; Multiplying the basic influence coefficient by the state deviation of the source node to obtain a preliminary relationship weight, where the preliminary relationship weight reflects the influence of the source node state on the relationship strength; Collecting attribute semantic labels related to the directed edge in the semantic data unit, and extracting a change rate of the attribute value as a dynamic adjustment factor, wherein the change rate is a difference between a current timestamp attribute value and a previous timestamp attribute value divided by a time interval; The preliminary relationship weight is added to the dynamic adjustment factor to obtain the dynamic relationship weight. The numerical range of the dynamic relationship weight is limited to a preset interval through normalization processing. The normalization processing uses the maximum and minimum normalization method to map the weight value to the range of 0 to 1.
6. The fault reasoning analysis method based on a large language model according to claim 1, characterized in that: The calling of the pre-trained large language model to perform root cause reasoning analysis on the semantic association network to generate a set of candidate root causes sorted by confidence level includes: Extracting network topology data at a current moment from the semantic association network, the network topology data including current values of a node set, a directed edge set, dynamic relationship weights, and node attribute features, wherein the current values are node state parameters corresponding to the latest timestamp; Converting the network topology data into a graph structure description text that can be parsed by a large language model, wherein the graph structure description text includes a natural language description of a node identifier-component type correspondence table, a directed edge relationship type list, and a dynamic relationship weight matrix, wherein the natural language description uses a preset template to convert the structured data into a text paragraph; Calling a large language model to perform multiple rounds of reasoning on the graph structure description text, wherein the first round of reasoning generates an initial root cause hypothesis set, wherein the initial root cause hypothesis set includes potential abnormal nodes and corresponding abnormal propagation path descriptions; Based on the multi-period semantic association network snapshots of the semantic association network, a temporal consistency verification is performed on the abnormal propagation path in the initial root cause hypothesis set, and the matching degree between the attribute feature change trend of each node in the abnormal propagation path during the abnormal period and the propagation path is calculated. The matching degree is a consistency score between the change direction of the node state parameter and the path propagation direction; Adjusting the confidence score of each hypothesis in the initial root cause hypothesis set according to the matching degree, retaining hypotheses with matching degrees higher than a preset matching degree threshold, and generating an intermediate root cause hypothesis set, wherein the confidence score is adjusted to a weighted sum of the original score and the matching degree; Calling the large language model to perform conflict detection processing on the intermediate root cause hypothesis set, identifying and removing hypothesis entries with logical contradictions, wherein the logical contradiction refers to the existence of mutually exclusive propagation path descriptions for the same abnormal node, and the mutual exclusivity refers to opposite propagation directions or conflicting relationship types; The intermediate root cause hypothesis set after de-conflict is sorted from high to low according to the confidence score, and a preset number of hypothesis entries before sorting are extracted as the candidate root cause set. The candidate root cause set includes the root cause entity identifier, the abnormal propagation path sequence and the confidence score. The abnormal propagation path sequence is a list of node identifiers arranged in the propagation order.
7. The fault reasoning analysis method based on a large language model according to claim 6 is characterized in that: The multi-period semantic association network snapshot based on the semantic association network performs temporal consistency verification on the abnormal propagation path in the initial root cause hypothesis set, and calculates the matching degree between the attribute feature change trend of each node in the abnormal propagation path during the abnormal period and the propagation path, including: Extracting a multi-period semantic association network snapshot of the abnormality occurrence period from the semantic association network, wherein the abnormality occurrence period is a preset time window before and after the system alarm triggering timestamp; For each anomaly propagation path sequence of the initial root cause hypothesis, extract the node identifiers in the anomaly propagation path in the propagation order and generate a node sequence list, where the first node in the node sequence list is the root cause entity and the last node is the anomaly manifestation entity; For each node in the node sequence list, extract the attribute feature time series of the corresponding node during the abnormality occurrence period from the multi-period semantic association network snapshot, wherein the attribute feature time series is a set of state parameter values sorted by timestamp; Calculate the change trend slope of the attribute feature time series, where the change trend slope is the slope value obtained by linear regression fitting, a positive value indicates an increase in the state parameter, and a negative value indicates a decrease; According to the propagation direction of the abnormal propagation path, determine whether the change trend slope of the adjacent nodes meets the causal relationship expectations; The ratio of the number of adjacent node pairs that meet the causal relationship expectations in the abnormal propagation path to the total number of node pairs is counted as the matching degree. The higher the matching degree, the stronger the temporal consistency of the abnormal propagation path.
8. The fault reasoning analysis method based on a large language model according to claim 6, characterized in that: The calling of the large language model to perform conflict detection processing on the intermediate root cause hypothesis set to identify and remove hypothesis entries with logical contradictions includes: Convert each hypothesis entry in the intermediate root cause hypothesis set into a logical expression, wherein the logical expression includes a root cause entity, a propagation path condition, and an abnormal phenomenon description; Constructing a conflict detection prompt template, wherein the conflict detection prompt template includes a task description, a conflict definition, and a list of hypothesis items; splicing the logical expression and the conflict detection prompt template into a model input text, calling a large language model to perform conflict judgment on the model input text, and generating a conflict matrix between hypothesis entries, wherein the conflict matrix is a two-dimensional matrix, and the matrix elements represent whether there is a conflict between two corresponding hypothesis entries; Constructing a conflict graph based on the conflict matrix, wherein the nodes of the conflict graph are hypothetical entries and the directed edges represent the conflict relationships between the entries; Performing maximum conflict-free subset selection on the conflict graph, selecting a subset containing the largest number of hypothesis entries without mutual conflicts; Removing hypothesis entries that are not selected into the maximum conflict-free subset from the intermediate root cause hypothesis set to generate a conflict-free intermediate root cause hypothesis set; The confidence scores of the deconflicted hypothesis entries are recalculated, and the scores of the hypothesis entries that are verified to be conflict-free during the conflict detection process are increased by the preset reward value to improve their sorting priority.
9. The fault reasoning analysis method based on a large language model according to claim 1, characterized in that: Generating an operation and maintenance decision instruction including an entity operation sequence and a priority ranking according to the candidate root cause set includes: Parsing the root cause entity identifier and the anomaly propagation path sequence corresponding to each candidate root cause in the candidate root cause set, determining the target entity set requiring intervention and the dependency order between the entities, wherein the dependency order is the reverse order of the anomaly propagation path sequence; Based on the target entity set, a preset operation and maintenance manual is queried to obtain the standard operating procedure corresponding to each target entity. The standard operating procedure includes a sequence of operation steps, precondition constraints, and a description of expected effects. The precondition constraints are the range of entity state parameters that must be met before performing the operation; Calling a large language model to perform adaptability analysis on the standard operating procedure and the exception propagation path sequence, adjusting the execution order of the operation steps to match the reverse path of the exception propagation, so that the operations are executed in the order from the root cause entity to the end entity; Determine the priority weight of each target entity operation based on the confidence score of the candidate root cause set, where a higher confidence score corresponds to a greater priority weight, and the priority weight is used to determine the execution order of operations corresponding to different candidate root causes; generating a preliminary operation and maintenance operation sequence based on the priority weights and the adjusted operation step sequence, wherein the preliminary operation and maintenance operation sequence includes an operation object identifier, an operation type, an execution time window, and an expected state parameter, wherein the execution time window is determined according to an entity dependency order; Performing conflict detection on the preliminary operation and maintenance operation sequence, identifying operation steps with the same operation objects or resource competition, and adjusting the execution order of conflicting operations based on priority weights; The conflict-processed operation sequence and priority weight are integrated to generate an operation and maintenance decision instruction containing the operation sequence number, target entity identifier, operation instruction content, priority ranking, and expected completion time. The priority ranking of the operation and maintenance decision instruction is positively correlated with the confidence score of the candidate root cause set, and the expected completion time is calculated by accumulating the estimated execution time of the operation steps.
10. A fault reasoning analysis system based on a large language model, characterized in that: It includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the fault reasoning analysis method based on a large language model as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Enterprise global data analysis method based on knowledge graph and large language model
CN120218256A
Intelligent question and answer method based on knowledge graph
CN120297415A
Metro equipment fault intelligent diagnosis method and system assisted by large language model
CN120337106A
Multi-source data management method based on machine learning and related device
CN120429549A
Translation of natural language questions and requests to a structured query format
US20180095962A1
Cited By
Lock control terminal work log query method and system based on substation operation and maintenance
CN121092403A
Hybrid model fault early warning method and system based on time sequence prediction and fuzzy logic
CN121093053A
Equipment operation and maintenance intelligent analysis and control method and system based on fault root cause
CN121433954A
Industrial IoT-oriented equipment adaptive access method and system
CN121567791A
Self-adaptive disk intelligent operation and maintenance and self-healing method and system based on AI large model
CN121862186A