Methods, apparatus, equipment, and media for inferring the root causes of host status anomalies.

By analyzing abnormal host status alarms, combining real-time data and historical cases, and using a large language reasoning model for multimodal fusion, the problem of traditional fault diagnosis relying on human experience is solved, and automated and accurate fault root cause inference is achieved.

CN121547343BActive Publication Date: 2026-04-03TAIPING FINANCIAL SERVICE CENT (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-20
Publication Date
2026-04-03

Smart Images

  • Figure CN121547343B_ABST
    Figure CN121547343B_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, device, and medium for root cause inference based on host status anomalies. It includes: receiving anomaly alarm notifications from a target host system and parsing them to obtain structured alarm data containing host identifiers and alarm events; matching anomaly components in a relational database, generating a list of fault-related components based on their upstream and downstream components, and extracting the target host system architecture topology; calling monitoring interfaces and configuration management database systems to collect real-time operational data of the listed components, and generating semantic real-time status data by combining semantic mapping templates; searching for preliminary historical cases based on alarm events, and generating case data sorted by matching degree based on user feedback; performing multimodal processing on the above data and inputting it into a large language inference model, outputting root cause inference results sorted by probability, and displaying them on a user interface. This invention achieves rapid and accurate root cause localization of host status anomalies in complex systems, improving operational efficiency and reducing business losses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of system operation and maintenance, and particularly to a root cause speculation method, device, equipment and medium based on abnormal host status. Background Art

[0002] With the continuous iteration of information technology, the architectures of various business systems are rapidly developing towards the direction of multi-components, multi-links and high coupling, and the scale and complexity of the systems are constantly increasing. Under this background, the triggering scenarios of system failures are becoming more diverse, and the fault conduction paths are also more complex, which puts forward higher requirements for the timeliness and accuracy of fault troubleshooting. As the core operation and maintenance link to ensure business continuity, fault troubleshooting directly affects system downtime losses and service quality.

[0003] Traditional fault handling modes often highly rely on the personal experience and professional capabilities of operation and maintenance personnel. At the same time, there are differences in the working habits of different teams or operation and maintenance personnel, resulting in uneven efficiency and effects of fault handling, and it is impossible to guarantee the consistency and reliability of the handling process. In addition, in a high-pressure operation and maintenance environment, operation and maintenance personnel often prefer to take temporary repair measures to quickly restore the business, lacking sufficient time and resources to carry out in-depth root cause analysis, resulting in the failure to fundamentally solve the fault and the risk of repeated occurrence. Summary of the Invention

[0004] Based on this, the present invention provides a root cause speculation method, device, equipment and medium based on abnormal host status to solve the problems that the fault location of complex systems overly relies on manual experience, lacks a unified standard, has low troubleshooting efficiency and is difficult to deeply cure.

[0005] In the first aspect, an embodiment of the present invention provides a root cause speculation method based on abnormal host status, and the method includes:

[0006] When receiving a host status abnormal alarm notification of a target host system sent by an external monitoring system, parsing the host status abnormal alarm notification to obtain structured alarm data, where the structured alarm data includes a host identifier and an alarm item;

[0007] Obtaining abnormal components matching the host identifier in a relational database, generating a list of fault-related components according to all upstream and downstream related components associated with the abnormal components, and extracting an architecture topology relationship corresponding to the target host system from the relational database;

[0008] Invoking a monitoring interface and a configuration management database system to collect real-time operation data of all components in the list of fault-related components, and generating semantic real-time status data according to the real-time operation data and a preset semantic mapping template;

[0009] Based on the alarms, preliminary historical case data is searched in the historical fault knowledge base. At the same time, the user interface is monitored, and historical case data sorted by matching degree is generated according to the user feedback results of the user interface.

[0010] The structured alarm data, architecture topology, historical case data, and semantic real-time status data are fused using a multimodal method. The fusion result is then input into a large language inference model to obtain multiple root cause inference results ordered by probability, which are then displayed in the user interface.

[0011] Secondly, embodiments of the present invention provide a root cause estimation device based on host state anomalies, the device comprising:

[0012] The alarm data acquisition module is used to parse the alarm notification of abnormal host status of the target host system when it receives the alarm notification of abnormal host status sent by the external monitoring system, and obtain structured alarm data, wherein the structured alarm data includes host identifier and alarm item.

[0013] The fault association component determination module is used to obtain abnormal components that match the host identifier in the relational database, generate a fault association component list based on all upstream and downstream associated components associated with the abnormal components, and extract the architecture topology relationship corresponding to the target host system from the relational database.

[0014] The semantic transformation module is used to call the monitoring interface and the configuration management database system to collect real-time running data of all components in the fault-related component list, and generate semantic real-time status data based on the real-time running data and the preset semantic mapping template.

[0015] The historical case data generation module is used to search for preliminary historical case data in the historical fault knowledge base based on the alarm items, and at the same time monitor the user interaction interface to generate historical case data sorted by matching degree according to the user feedback results of the user interaction interface.

[0016] The fault root cause prediction result generation module is used to perform multimodal fusion processing on the structured alarm data, architecture topology relationship, historical case data and semantic real-time status data, and input the fusion result into the large language inference model to obtain multiple fault root cause prediction results sorted by probability and displayed in the user interface.

[0017] Thirdly, embodiments of the present invention provide an electronic device, the electronic device comprising:

[0018] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform a root cause inference method based on host state anomalies as described in any embodiment of the present invention.

[0019] Fourthly, a computer-readable storage medium is also provided, the computer-readable storage medium storing computer instructions, the computer instructions being used to cause a processor to execute and implement a root cause inference method based on host state anomalies as described in any embodiment of the present invention.

[0020] The technical solution of this invention accurately locates abnormal hosts through structured alarms and clarifies the analysis scope by associating upstream and downstream components; it integrates multi-source information such as real-time data, historical cases, and topological relationships, inputs it into the model through multimodal fusion, and optimizes matching based on user feedback to reduce single-dimensional bias; it automates alarm parsing, data collection, and case retrieval, replacing tedious manual processes; semantic data lowers the professional understanding threshold, and sorted historical cases can be quickly referenced, significantly shortening investigation time; the user interface receives supplementary feedback to fill information blind spots in monitoring alarms, and root cause results are displayed in probability sorting with supporting evidence, balancing the efficiency of automated reasoning with the need for manual decision verification; it relies on the architectural topology to sort component dependencies, covering cascading failures caused by multiple component linkages, not limited to a single abnormal component, and improving the applicability of fault analysis for complex business systems.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of a root cause inference method based on host state anomalies provided in Embodiment 1 of the present invention;

[0024] Figure 2 This is a flowchart of another root cause inference method based on host state anomalies provided in Embodiment 2 of the present invention;

[0025] Figure 3This is a schematic diagram of a root cause inference device based on host state anomalies provided in Embodiment 3 of the present invention;

[0026] Figure 4 This is a schematic diagram of the structure of an electronic device that implements a root cause prediction method based on host state anomalies according to an embodiment of the present invention. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] Example 1

[0030] Figure 1 This is a flowchart of a root cause prediction method based on host status anomalies provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where host status anomalies occur in complex business systems, requiring rapid and accurate location of the root cause of the fault and standardized processing procedures to reduce business losses. This method can be executed by a root cause prediction device based on host status anomalies. This device can be implemented in hardware and / or software and can be configured in an intelligent operation and maintenance platform. Figure 1 As shown, the method includes:

[0031] S110. When receiving a host status abnormality alarm notification from an external monitoring system, the host status abnormality alarm notification is parsed to obtain structured alarm data, wherein the structured alarm data includes a host identifier and alarm items.

[0032] An external monitoring system refers to an independently deployed technical system capable of collecting host operating status data and sending alarms. It acquires operational metrics of the target host system through standardized protocols and triggers alarm notifications when preset abnormal conditions are met. The target host system refers to a complete business support system comprised of multiple computing nodes and business function components deployed on them; it is the target of anomaly monitoring and root cause analysis. Host status anomaly alarm notifications are structured notifications generated by the external monitoring system based on collected operational metrics, containing core information related to the anomaly. They serve as the trigger signal for initiating the root cause analysis process. Structured alarm data is fixed-format data formed after parsing and standardizing host status anomaly alarm notifications. The core fields are the host identifier and the alarm item, adapting to subsequent automated data query and matching processes. The host identifier is the unique identifier of the computing node in the target host system, used to establish the mapping relationship between computing nodes and related business components. The alarm item is a specific technical description of the host status anomaly, including the associated components, anomaly type, and quantitative indicators, providing core evidence for anomaly localization.

[0033] S120. Obtain the abnormal component that matches the host identifier in the relational database, generate a list of fault-related components based on all upstream and downstream related components associated with the abnormal component, and extract the architecture topology relationship corresponding to the target host system from the relational database.

[0034] A relational database refers to a database system that uses a relational model to store static information about the target host system architecture. Its core storage includes "component node-business component binding relationships" and "upstream and downstream dependencies between business components," providing data support for component queries and topology extraction. Anomaly components are business function components bound to host identifiers and directly associated with alarm events; they are the core starting point for root cause analysis. Upstream and downstream related components are business components that form direct data interaction or function call links with anomaly components through dependency relationships, including upstream data-providing components and downstream function-calling components. The fault-related component list is a structured list integrating anomaly components and all upstream and downstream related components, clearly defining the scope and boundaries of subsequent data collection. The architecture topology relationship is a structured description of the functional dependencies and data flow links between fault-related components within the target host system, reflecting the technical relationship logic between components.

[0035] Using the host identifier in the structured alarm data as the search condition, the component information table in the relational database is queried to obtain all business components bound to the host, and the abnormal component is identified by combining the alarm items; using the unique identifier of the abnormal component as the search condition, the dependency table in the relational database is queried to obtain all its upstream and downstream related components, and the abnormal component and its upstream and downstream related components are integrated to generate a structured list of fault-related components; based on the target host system identifier and the list of fault-related components, the dependencies and data flow links between components are extracted from the relational database to form the architectural topology.

[0036] S130 calls the monitoring interface and configuration management database system to collect real-time running data of all components in the fault-related component list, and generates semantic real-time status data based on the real-time running data and preset semantic mapping template.

[0037] The monitoring interface refers to the standardized data interface provided by the external monitoring system, supporting batch querying of dynamic operating metrics by component identifier. In this embodiment, the configuration management database system is a dedicated system for storing static configuration information, primarily containing fixed attribute information such as the deployment location, hardware parameters, preset thresholds, and dependent resources of business components. Real-time operating data is component-level complete status data formed by integrating the dynamic operating metrics collected by the monitoring interface with the static configuration information extracted from the configuration management database system. The preset semantic mapping template is a predefined set of rules corresponding to "quantitative indicators - natural language descriptions," used to achieve semantic conversion of dynamic operating metrics. The rules are consistent with the alarm thresholds of the external monitoring system. Semantic real-time status data is component status description data formed through conversion using the preset semantic mapping template, possessing both machine-understandable and human-readable attributes.

[0038] S140. Based on the alarm, search for preliminary historical case data in the historical fault knowledge base, and at the same time monitor the user interaction interface, and generate historical case data sorted by matching degree according to the user feedback results of the user interaction interface.

[0039] The historical fault knowledge base is a structured knowledge base storing technical information on past fault cases. Its core includes technical elements such as fault phenomena, related components, root cause conclusions, and solutions, and supports keyword-based similarity retrieval. Preliminary historical case data is a collection of similar fault cases filtered from the historical fault knowledge base using alarm events as search keywords. The user interface is a visual technical interface that allows users to view alarm information, input supplementary fault descriptions, and browse analysis results, and it features natural language input. User feedback results refer to supplementary fault information entered by users through the interface that is not included in the alarm notification. Historical case data sorted by matching degree refers to a collection of historical fault cases sorted in descending order of matching degree after combining alarm events and user feedback results using a similarity algorithm.

[0040] Optionally, based on the alarm events, preliminary historical case data is searched in the historical fault knowledge base, while simultaneously monitoring the user interface and generating historical case data sorted by matching degree based on user feedback from the user interface. This may include:

[0041] Using the alarm items as initial search conditions, preliminary historical case data is generated in the historical fault knowledge base through text similarity calculation and vector retrieval technology. At the same time, the front-end interface is used to receive user feedback results returned by the user interaction interface in real time.

[0042] If no user feedback is received from the user interface within the preset time window, the initial historical case data will be maintained and historical case data sorted by matching degree will be formed.

[0043] If a natural language fault description, manually inputted by the user interface, is received within a preset time window, the natural language fault description and the alarm item are combined into a two-dimensional search condition.

[0044] Using the aforementioned dual-dimensional search criteria, new historical case data is obtained from the historical fault knowledge base through text similarity calculation and vector retrieval technology, and the new historical case data is sorted according to the matching degree to form historical case data.

[0045] Text similarity calculation and vector retrieval techniques are a combination of technologies used to measure the similarity between search criteria and fault descriptions in historical cases. Text similarity calculation quantifies the semantic overlap of text through algorithms; vector retrieval converts text into high-dimensional vectors and quickly locates similar cases using vector space distance. The combination of these two techniques improves search efficiency and accuracy. A preset time window refers to the maximum time the system pre-sets for waiting for user feedback, used to balance analysis timeliness and case matching accuracy. If no feedback is received within the time limit, the initial case is used by default, avoiding analysis delays due to waiting for feedback. Dual-dimensional search criteria, composed of alarm items and user feedback results, provide a more comprehensive coverage of fault characteristics compared to single alarm items, improving the relevance of case matching. Natural language fault descriptions are supplementary fault information expressed in natural language by the user through the interactive interface. This is the core content of user feedback results, used to supplement details not covered by alarm items.

[0046] Using alarm events as the sole search criteria, the historical fault knowledge base retrieval process is initiated. Alarm events are converted into text vectors, and vector retrieval technology is used to quickly filter candidate case sets. Then, text similarity calculation quantifies the semantic overlap between candidate cases and alarm events, generating preliminary historical case data. Simultaneously, the user interface is monitored in real-time via a front-end interface, awaiting user input and feedback to ensure timely capture of feedback information. If a natural language fault description is received from the user within a preset time window, its validity is verified. The natural language fault description and alarm event are semantically fused to form dual-dimensional search criteria, covering both machine alarms and human observations. The fused criteria are then cleaned to ensure compatibility with the historical case database's retrieval format. Using the dual-dimensional search criteria as new search keywords, the historical fault knowledge base retrieval interface is re-invoked, expanding the search dimensions to dual features. New historical case data that highly matches both criteria is selected and sorted in descending order of comprehensive similarity between the new cases and the dual-dimensional search criteria, ultimately forming historical case data sorted by matching degree.

[0047] S150. Perform multimodal fusion processing on the structured alarm data, architecture topology relationship, historical case data and semantic real-time status data, and input the fusion result into the large language inference model to obtain multiple fault root cause inference results sorted by probability and display them in the user interaction interface.

[0048] Multimodal fusion processing refers to the technical process of converting data of different types and formats, such as structured alarm data, architectural topology relationships, historical case data sorted by matching degree, and semantic real-time status data, into a unified format and establishing logical connections. The fusion result is unified format data formed after multimodal fusion processing, containing all the core technical information required for fault analysis and providing complete input for root cause reasoning. The large language reasoning model is a pre-trained large language model fine-tuned for fault root cause analysis scenarios, possessing the ability to understand fault context, associate historical cases, and infer root cause probabilities. Multiple fault root cause inference results sorted by probability refer to the set of multiple possible fault root causes output by the large language reasoning model based on the fusion result, arranged in descending order of the probability of occurrence of each root cause.

[0049] This invention, through structured alarms, accurately locates abnormal hosts and clarifies the analysis scope by associating upstream and downstream components; it integrates multi-source information such as real-time data, historical cases, and topological relationships, inputs it into the model through multimodal fusion, and optimizes matching based on user feedback to reduce single-dimensional bias; it automates alarm parsing, data collection, and case retrieval, replacing tedious manual processes; semantic data lowers the professional understanding threshold, and sorted historical cases can be quickly referenced, significantly shortening investigation time; the user interface receives supplementary feedback, filling information blind spots in monitoring alarms, and root cause results are displayed in probability sorting with supporting evidence, balancing the efficiency of automated reasoning with the need for manual decision verification; relying on architectural topological relationships to sort component dependencies, it can cover cascading failures caused by multiple component linkages, not limited to a single abnormal component, improving the applicability of fault analysis for complex business systems.

[0050] Example 2

[0051] Figure 2 This is a flowchart of another root cause prediction method based on host state anomalies provided in Embodiment 2 of the present invention. This embodiment is a refinement based on the above embodiment, and correspondingly, as follows: Figure 2 As shown, the method specifically includes:

[0052] S210. When receiving a host status abnormality alarm notification from an external monitoring system, the host status abnormality alarm notification is parsed to obtain structured alarm data, wherein the structured alarm data includes a host identifier and alarm items.

[0053] S220. Using the host identifier as the search keyword, query the component information table of the pre-built relational database, and obtain the corresponding abnormal component and the abnormal component identifier of the abnormal component through field matching.

[0054] The abnormal component identifier is a specific subset of component identifiers, specifically referring to the unique technical identifier corresponding to the abnormal component determined to be directly related to the host's abnormal state. It is completely consistent with the general component identifier fields in the component information table and dependency table of the relational database, with no format or definition differences, and is specifically identified only because it is an abnormal component. The component information table is one of the pre-built core data tables in the relational database, used to store the "host-component" binding relationship. Key fields include "host identifier," "component identifier," "component name," and "component type." Using the host identifier as the core search condition, a query operation is initiated in the component information table of the relational database. The search keyword "host identifier" is precisely matched with the host identifier field in the component information table to filter out the set of all business components bound to that host. Combined with alarm events, components directly related to the abnormality are filtered from the above component set and identified as the corresponding abnormal components. The component identifier corresponding to this abnormal component in the component information table is extracted. Because the component role is an abnormal component, this component identifier is specifically designated as the abnormal component identifier of the abnormal component.

[0055] S230. Based on the abnormal component identifier, query the dependency table of the relational database to determine all upstream and downstream related components of the abnormal component, generate a list of fault-related components containing the abnormal component, upstream and downstream related components and the abnormal component identifier, and extract the architecture topology relationship corresponding to the target host system from the relational database.

[0056] The dependency table is one of the core pre-built data tables in a relational database, used to store the upstream and downstream dependency logic between business components. Key fields include "component identifier," "upstream component identifier," "downstream component identifier," and "dependency type." Its core function is to query the directly related upstream and downstream components through the component identifier, providing dependency data for generating a list of fault-related components. Using the faulty component identifier of the faulty component as the search keyword, the dependency table of the relational database is queried, and the keyword is precisely matched with the component identifier field in the table. Based on the matching results, all components corresponding to the upstream component identifier field and all components corresponding to the downstream component identifier field are extracted, ensuring that all upstream and downstream related components are covered without omission.

[0057] Optionally, before retrieving the abnormal component matching the host identifier from the relational database, generating a list of fault-related components based on all upstream and downstream related components associated with the abnormal component, and extracting the architectural topology relationship corresponding to the target host system from the relational database, the method further includes:

[0058] Obtain the static architecture diagram of all systems to be managed within the local area network, and convert the static architecture diagram of each system into structured architecture data that stores the architecture topology relationship;

[0059] From the structured architecture data corresponding to each system to be managed, extract the basic attributes and relationships of the host and components, form the component information table corresponding to the current system to be managed, and store it in a relational database;

[0060] From the structured architecture data corresponding to each system to be managed, extract the upstream and downstream related components of each component, form the dependency relationship table corresponding to the current system to be managed, and store it in a relational database.

[0061] A static architecture diagram refers to the architectural design blueprint of the system to be managed within a local area network (LAN). It graphically presents the system's constituent elements, including the distribution of hosts and business components, as well as the relationships between hosts and components, and between components themselves. It is a static snapshot of the system architecture. Structured architecture data is machine-readable data generated by digitizing the static architecture diagram. It contains the attribute information of all entities in the diagram and the relationships between them, serving as the structured carrier of the static architecture diagram. The system to be managed comprises all business systems within the LAN that require anomaly monitoring and root cause analysis (the target host system is one of them). Each system corresponds to an independent static architecture diagram and structured architecture data, ensuring clear boundaries for data management. The basic attributes of hosts and components refer to the inherent characteristic information of hosts and components extracted from the structured architecture data. Host basic attributes include host identifiers, models, IP addresses, etc.; component basic attributes can include component identifiers, names, types, etc., serving as the core basis for unique entity identification.

[0062] Image recognition technology can be used to convert the "hosts, component entities" and their "deployment / call relationships" in the static architecture diagram into structured data, ultimately forming structured architecture data containing the architecture topology. By pre-converting the static architecture diagram into structured data and building component information tables and dependency tables, standardized and queryable basic data is provided for the "host-component-dependency" retrieval chain, ensuring the accuracy and efficiency of subsequent abnormal component location and related component query. This is a crucial data infrastructure construction step in the entire root cause inference method.

[0063] S240. Collect real-time operating indicator data for each component in the fault-related component list by calling the monitoring interface; wherein the real-time operating indicator data is a dynamic quantitative indicator generated by each component during operation.

[0064] Real-time operational metrics data is a set of quantitative indicators dynamically generated during component operation. It features real-time updates, reflecting the component's current dynamic operational status, and is collected through monitoring interfaces. Component configuration information is a set of static basic attributes for the component. Its core includes dependencies between upstream and downstream related components, as well as fixed attributes such as component deployment location, hardware parameters, preset operating thresholds, and dependent database addresses. These are stored in the configuration management database system and are inherent attributes of the component. Based on a list of fault-related components, only metric data for all components within the list is collected to avoid data redundancy from irrelevant components. Using the component identifier of each component in the list as a unique request parameter, monitoring interfaces are called in batches. The real-time operational metric data returned by the interface is bound one-to-one with the corresponding component identifier, ensuring that dynamic metrics are not misaligned with components.

[0065] S250. Extract the component configuration information of each component in the fault-related component list by calling the configuration management database system; wherein, the component configuration information is static basic information containing the dependency relationship between upstream and downstream related components.

[0066] Using the list of fault-related components as the search scope, only the configuration information of the components in the list is extracted to accurately match the collection requirements. Using the component identifier as the search condition, the query interface of the configuration management database system is called to accurately extract the static configuration information of each component.

[0067] S260. The component configuration information and real-time operation index data of each component are associated and integrated, and then standardized and structured after association and integration to obtain real-time operation data containing the component identifier, standardized dynamic quantitative index and standardized static basic data of each component.

[0068] Standardized dynamic quantitative indicators refer to structured quantitative data obtained by unifying the format, standardizing units, and removing redundancy from real-time operational indicator data, ensuring that the indicator format is unambiguous and comparable. Standardized static basic data refers to structured data obtained by standardizing the fields and unifying the expression of component configuration information, ensuring that the static information format is standardized and associative. Using the component identifier as the unique association key, the real-time operational indicator data of the same component is bound one-to-one with the component configuration information, forming a relationship of "component identifier → dynamic indicator + static configuration", ensuring the data integrity of each component.

[0069] S270. For each component in the real-time running data, the component identifier is used to match the corresponding target indicator threshold in the preset semantic mapping template, and the standardized dynamic quantitative indicator of the current component is compared with the target indicator threshold to obtain the threshold comparison result of the current component.

[0070] The target metric threshold refers to the baseline range set for the dynamic quantification metric of each component in the preset semantic mapping template. It is used to determine whether the real-time running status of the component is abnormal and is strongly correlated with the component type and functional characteristics. The threshold comparison result is the comparison conclusion between the component's current standardized dynamic quantification metric and the target metric threshold, which is the core basis for semantic transformation. Each component in the real-time running data is traversed, using the component identifier as the search key. The target metric threshold corresponding to that component is accurately matched in the preset semantic mapping template. The standardized dynamic quantification metric of the component is compared with the matched target metric threshold one by one, the quantification difference is calculated, and the abnormality of the status is determined.

[0071] S280. Based on the threshold comparison result, search for semantic description information corresponding to the current threshold comparison result in the preset semantic mapping template.

[0072] The semantic description information consists of natural language descriptions that correspond one-to-one with the threshold comparison results in a preset semantic mapping template, realizing the conversion of quantitative indicators from "machine language" to "human-readable language". Based on the threshold comparison results, the system searches for matching semantic description information in the preset semantic mapping template. If the component has multiple types of dynamic indicators, it matches the corresponding semantic description for the threshold comparison results of each type of indicator, ensuring that all abnormal / normal states are covered.

[0073] S290. Based on the dependency relationship between upstream and downstream related components in the standardized static basic data corresponding to the current component, search for topological influence association information corresponding to the threshold comparison result in the preset semantic mapping template.

[0074] The preset semantic mapping template refers to a predefined set of rules, which corely includes three types of mapping relationships: "component identifier → target indicator threshold", "threshold comparison result → semantic description information", and "threshold comparison result + dependency relationship → topological impact association information". These are stored in a structured configuration file and support dynamic adjustment based on component type. The topological impact association information is a description of the association impact generated in the preset semantic mapping template by combining the threshold comparison result with the dependency relationship between upstream and downstream related components, reflecting the transmission of anomalies.

[0075] S2100: Summarize the component identifier, threshold comparison results, semantic description information and topological influence association information corresponding to each component to obtain semantic real-time status data corresponding to the real-time running data.

[0076] Four types of information are integrated for each component: component identifier, threshold comparison results, semantic description information, and topological influence correlation information, forming semantic state data for a single component. The semantic state data of all components are arranged in the order of the fault-related component list to form semantic real-time state data covering all analysis objects, stored in structured text, which supports reading by large language models and can also be directly displayed on the user interface.

[0077] S2110. Based on the alarm, search for preliminary historical case data in the historical fault knowledge base, and at the same time monitor the user interaction interface, and generate historical case data sorted by matching degree according to the user feedback results of the user interaction interface.

[0078] S2120. Perform multimodal fusion processing on the structured alarm data, architecture topology relationship, historical case data and semantic real-time status data, and input the fusion result into the large language inference model to obtain multiple fault root cause inference results sorted by probability and display them in the user interaction interface.

[0079] Furthermore, before inputting the fusion result into the large language reasoning model, it may also include:

[0080] Historical failure cases bound to root cause conclusions are extracted from the historical failure knowledge base, system architecture topology data is extracted from the relational database, and historical real-time operation data and historical alarm sample data of components are obtained from the external monitoring system and aggregated to form pre-training data.

[0081] After standardizing the pre-training data, alarm sample data, system architecture topology relationship data, historical real-time operation data, and historical fault cases are used as pre-training input samples.

[0082] The root cause conclusions associated with the historical failure cases are used as pre-trained output samples.

[0083] The initial large language model is iteratively trained using the pre-trained input samples and pre-trained output samples until the root cause inference accuracy reaches a preset accuracy threshold, thus obtaining the large language inference model.

[0084] The pre-training data is a historical dataset used to train the initial large language model, integrating multi-dimensional fault-related information. Pre-training input samples refer to the feature set extracted from the pre-training data for model input, including alarm sample data, system architecture topology data, historical real-time operation data, and historical fault cases, simulating the input scenario for actual root cause inference. Pre-training output samples refer to the root cause conclusions corresponding one-to-one with the historical fault cases in the pre-training input samples, and are the target output for model training. The initial large language model is a general pre-trained language model that has not undergone fine-tuning in fault root cause analysis scenarios. It possesses basic text understanding and logical reasoning capabilities but lacks professional knowledge in the field of fault analysis. A preset accuracy threshold is the quantitative standard for terminating model training. It is determined by evaluating the matching degree between the model's root cause prediction results on the validation set and the actual root causes, ensuring that the model has practical inference accuracy.

[0085] The technical solution of this invention refines the overall solution, clarifying the technical implementation of key aspects such as component association and list generation, real-time data collection and conversion, and basic data preparation. Specifically: by standardizing component positioning, association retrieval, and underlying data table construction, the accuracy of component association and data reliability are ensured; by refining the data collection, integration, and semantic conversion processes, the data's support for reasoning is enhanced; and by clarifying the model training logic, the professional reasoning capabilities of the large language model are guaranteed, ultimately further strengthening the accuracy, automation level, and engineering feasibility of the solution.

[0086] Example 3

[0087] Figure 3 This is a schematic diagram of a root cause prediction device based on host state anomalies provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes:

[0088] The alarm data acquisition module 310 is used to parse the host status abnormality alarm notification sent by the external monitoring system to obtain structured alarm data when receiving the host status abnormality alarm notification of the target host system. The structured alarm data includes the host identifier and alarm items.

[0089] The fault association component determination module 320 is used to obtain abnormal components that match the host identifier in the relational database, generate a fault association component list based on all upstream and downstream associated components associated with the abnormal components, and extract the architecture topology relationship corresponding to the target host system from the relational database.

[0090] The semantic transformation module 330 is used to call the monitoring interface and the configuration management database system to collect real-time running data of all components in the fault-related component list, and generate semantic real-time status data based on the real-time running data and the preset semantic mapping template.

[0091] The historical case data generation module 340 is used to search for preliminary historical case data in the historical fault knowledge base based on the alarm items, and at the same time monitor the user interaction interface to generate historical case data sorted by matching degree according to the user feedback results of the user interaction interface.

[0092] The fault root cause prediction result generation module 350 is used to perform multimodal fusion processing on the structured alarm data, architecture topology relationship, historical case data and semantic real-time status data, and input the fusion result into the large language inference model to obtain multiple fault root cause prediction results sorted by probability and displayed in the user interaction interface.

[0093] In this embodiment of the invention, abnormal hosts are accurately located through structured alarms, and the analysis scope is clarified by associating upstream and downstream components; multi-source information such as real-time data, historical cases, and topological relationships are integrated and input into the model through multimodal fusion, and the matching is optimized by combining user feedback to reduce single-dimensional bias; alarm parsing, data collection, and case retrieval are completed automatically, replacing tedious manual processes; semantic data lowers the professional understanding threshold, and sorted historical cases can be quickly referenced, significantly shortening the investigation time; the user interface receives supplementary feedback to fill the information blind spots of monitoring alarms, and root cause results are displayed in probability sorting with supporting evidence, balancing the efficiency of automated reasoning with the need for manual decision verification; and component dependencies are sorted based on architectural topological relationships, which can cover chain failures caused by the linkage of multiple components, not limited to a single abnormal component, improving the applicability of fault analysis for complex business systems.

[0094] Optionally, based on the above embodiments, the fault association component determination module 320 may include:

[0095] An abnormal component matching unit is used to query a pre-built relational database component information table using the host identifier as the search keyword, and obtain the corresponding abnormal component and the abnormal component identifier of the abnormal component through field matching.

[0096] The dependency table query unit is used to query the dependency table of the relational database based on the abnormal component identifier, determine all upstream and downstream related components of the abnormal component, and generate a list of fault-related components containing the abnormal component, the dependency relationships between upstream and downstream related components, and the abnormal component identifier.

[0097] Optionally, based on the above embodiments, the semantic conversion module 330 may include:

[0098] The dynamic quantitative indicator acquisition unit is used to collect real-time operation indicator data of each component in the fault-related component list by calling the monitoring interface; wherein, the real-time operation indicator data is the dynamic quantitative indicator generated by each component during operation.

[0099] The static basic information extraction unit is used to extract the component configuration information of each component in the fault-related component list by calling the configuration management database system; wherein, the component configuration information is static basic information containing the dependency relationship between upstream and downstream related components;

[0100] The real-time running data generation unit is used to associate and integrate the component configuration information and real-time running indicator data of each component, and then perform standardization and structuring processing after association and integration to obtain real-time running data containing the component identifier, standardized dynamic quantitative indicators and standardized static basic data of each component.

[0101] Optionally, based on the above embodiments, the semantic conversion module 330 may further include:

[0102] The threshold comparison unit is used to match the corresponding target indicator threshold in a preset semantic mapping template for each component in the real-time running data using the component identifier, and compare the standardized dynamic quantitative indicator of the current component with the target indicator threshold to obtain the threshold comparison result of the current component.

[0103] The semantic description information lookup unit is used to search for semantic description information corresponding to the current threshold comparison result in a preset semantic mapping template based on the threshold comparison result.

[0104] The topology influence association information lookup unit is used to search for topology influence association information corresponding to the threshold comparison result in a preset semantic mapping template based on the dependency relationship between upstream and downstream related components in the standardized static basic data corresponding to the current component.

[0105] The semantic real-time status data generation unit is used to summarize the component identifier, threshold comparison results, semantic description information and topological influence association information corresponding to each component to obtain semantic real-time status data corresponding to the real-time running data.

[0106] Optionally, based on the above embodiments, the historical case data generation module 340 may include:

[0107] The historical case data preliminary retrieval unit is used to generate preliminary historical case data in the historical fault knowledge base by using the alarm items as initial retrieval conditions and text similarity calculation and vector retrieval technology. At the same time, it uses the front-end interface to receive user feedback results returned by the user interaction interface in real time.

[0108] The historical case data maintenance unit is used to maintain the initial historical case data and form historical case data sorted by matching degree if no user feedback results are received from the user interaction interface within a preset time window.

[0109] A dual-dimensional retrieval condition construction unit is used to combine the natural language fault description and the alarm item into a dual-dimensional retrieval condition if a natural language fault description returned by the user interface is received within a preset time window.

[0110] The historical case data re-retrieval unit is used to retrieve new historical case data from the historical fault knowledge base using text similarity calculation and vector retrieval technology based on the aforementioned dual-dimensional retrieval conditions, and then sort the new historical case data according to the matching degree to form historical case data.

[0111] Optionally, based on the above embodiments, it further includes a relational database pre-storage unit, used to obtain abnormal components matching the host identifier in the relational database, generate a list of fault-related components based on all upstream and downstream related components associated with the abnormal components, and obtain static architecture diagrams of all systems to be managed in the local area network before extracting the architecture topology relationship corresponding to the target host system from the relational database, and convert the static architecture diagram of each system into structured architecture data for storing architecture topology relationships;

[0112] From the structured architecture data corresponding to each system to be managed, extract the basic attributes and relationships of the host and components, form the component information table corresponding to the current system to be managed, and store it in a relational database;

[0113] Extract the dependencies and connections between components from the structured architecture data corresponding to each system to be managed, form a dependency table for the current system to be managed, and store it in a relational database.

[0114] Optionally, based on the above embodiments, it also includes a large language model pre-training unit, which is used to extract historical fault cases bound to root cause conclusions from the historical fault knowledge base, extract system architecture topology relationship data from the relational database, and obtain historical real-time operation data and historical alarm sample data of components from the external monitoring system before inputting the fusion results into the large language reasoning model, and aggregate them to form pre-training data.

[0115] After standardizing the pre-training data, alarm sample data, system architecture topology relationship data, historical real-time operation data, and historical fault cases are used as pre-training input samples.

[0116] The root cause conclusions associated with the historical failure cases are used as pre-trained output samples.

[0117] The initial large language model is iteratively trained using the pre-trained input samples and pre-trained output samples until the root cause inference accuracy reaches a preset accuracy threshold, thus obtaining the large language inference model.

[0118] The root cause inference device based on host state anomaly provided in the embodiments of the present invention can execute the root cause inference method based on host state anomaly provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0119] Example 4

[0120] Figure 4A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0121] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0122] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0123] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a root cause inference method based on host state anomalies.

[0124] That is, when receiving a host status abnormality alarm notification from an external monitoring system, the host status abnormality alarm notification is parsed to obtain structured alarm data, which includes a host identifier and alarm items.

[0125] Retrieve the abnormal component that matches the host identifier from the relational database, generate a list of fault-related components based on all upstream and downstream related components associated with the abnormal component, and extract the architecture topology relationship corresponding to the target host system from the relational database.

[0126] The system calls the monitoring interface and configuration management database system to collect real-time running data of all components in the fault-related component list, and generates semantic real-time status data based on the real-time running data and the preset semantic mapping template.

[0127] Based on the alarms, preliminary historical case data is searched in the historical fault knowledge base. At the same time, the user interface is monitored, and historical case data sorted by matching degree is generated according to the user feedback results of the user interface.

[0128] The structured alarm data, architecture topology, historical case data, and semantic real-time status data are fused using a multimodal method. The fusion result is then input into a large language inference model to obtain multiple root cause inference results ordered by probability, which are then displayed in the user interface.

[0129] In some embodiments, a root cause inference method based on host state anomalies may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the root cause inference method based on host state anomalies described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform a root cause inference method based on host state anomalies by any other suitable means (e.g., by means of firmware).

[0130] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0131] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0132] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0133] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0134] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0135] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0136] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0137] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A root cause prediction method based on host state anomalies, characterized in that, include: When receiving a host status anomaly alarm notification from an external monitoring system, the host status anomaly alarm notification is parsed to obtain structured alarm data. The structured alarm data includes a host identifier and an alarm item. The target host system refers to a business support system consisting of multiple computing nodes and business function components deployed on them. The host identifier is a unique identifier of a computing node in the target host system, used to establish a mapping relationship between computing nodes and associated business components. Retrieve the abnormal component that matches the host identifier from the relational database, generate a list of fault-related components based on all upstream and downstream related components associated with the abnormal component, and extract the architecture topology relationship corresponding to the target host system from the relational database. The system calls the monitoring interface and configuration management database system to collect real-time running data of all components in the fault-related component list, and generates semantic real-time status data based on the real-time running data and the preset semantic mapping template. Based on the alarms, preliminary historical case data is searched in the historical fault knowledge base. At the same time, the user interface is monitored, and historical case data sorted by matching degree is generated according to the user feedback results of the user interface. The structured alarm data, architecture topology, historical case data, and semantic real-time status data are fused in a multimodal manner, and the fusion result is input into a large language inference model to obtain multiple root cause inference results of faults ordered by probability, which are then displayed in the user interface. The system calls the monitoring interface and configuration management database system to collect real-time operational data for all components in the fault-related component list, including: By calling the monitoring interface, real-time operational indicator data of each component in the list of fault-related components is collected; wherein, the real-time operational indicator data is a dynamic quantitative indicator generated by each component during operation; The component configuration information of each component in the list of fault-related components is extracted by calling the configuration management database system; wherein, the component configuration information is static basic information containing the dependency relationship between upstream and downstream related components; The component configuration information and real-time operation indicator data of each component are linked and integrated, and then standardized and structured after linkage and integration to obtain real-time operation data containing the component identifier, standardized dynamic quantitative indicators and standardized static basic data of each component. Based on real-time operational data and preset semantic mapping templates, semantic real-time status data is generated, including: For each component in the real-time running data, the component identifier is used to match the corresponding target indicator threshold in the preset semantic mapping template, and the standardized dynamic quantitative indicator of the current component is compared with the target indicator threshold to obtain the threshold comparison result of the current component. Based on the threshold comparison result, search for semantic description information corresponding to the current threshold comparison result in the preset semantic mapping template; Based on the dependencies between upstream and downstream related components in the standardized static basic data corresponding to the current component, the topological influence association information corresponding to the threshold comparison result is searched in the preset semantic mapping template; By summarizing the component identifier, threshold comparison results, semantic description information, and topological influence association information corresponding to each component, semantic real-time status data corresponding to the real-time running data is obtained.

2. The method according to claim 1, characterized in that, Retrieve the abnormal component matching the host identifier from the relational database, and generate a list of fault-related components based on all upstream and downstream related components associated with the abnormal component, including: Using the host identifier as the search keyword, query the component information table of the pre-built relational database, and obtain the corresponding abnormal component and the abnormal component identifier of the abnormal component through field matching; Based on the abnormal component identifier, the dependency table of the relational database is queried to determine all upstream and downstream related components of the abnormal component, and a list of fault-related components containing the abnormal component, upstream and downstream related components, and the abnormal component identifier is generated.

3. The method according to claim 1, characterized in that, Based on the aforementioned alarm events, preliminary historical case data is retrieved from the historical fault knowledge base. Simultaneously, the user interface is monitored, and historical case data sorted by matching degree is generated based on user feedback from the user interface, including: Using the alarm items as initial search conditions, preliminary historical case data is generated in the historical fault knowledge base through text similarity calculation and vector retrieval technology. At the same time, the front-end interface is used to receive user feedback results returned by the user interaction interface in real time. If no user feedback is received from the user interface within the preset time window, the initial historical case data will be maintained and historical case data sorted by matching degree will be formed. If a natural language fault description, manually inputted by the user interface, is received within a preset time window, the natural language fault description and the alarm item are combined into a two-dimensional search condition. Using the aforementioned dual-dimensional search criteria, new historical case data is obtained from the historical fault knowledge base through text similarity calculation and vector retrieval technology, and the new historical case data is sorted according to the matching degree to form historical case data.

4. The method according to claim 2, characterized in that, Before retrieving the abnormal component matching the host identifier from the relational database, generating a list of fault-related components based on all upstream and downstream related components associated with the abnormal component, and extracting the architectural topology relationship corresponding to the target host system from the relational database, the process also includes: Obtain the static architecture diagram of all systems to be managed within the local area network, and convert the static architecture diagram of each system into structured architecture data that stores the architecture topology relationship; From the structured architecture data corresponding to each system to be managed, extract the basic attributes and relationships of the host and components, form the component information table corresponding to the current system to be managed, and store it in a relational database; From the structured architecture data corresponding to each system to be managed, extract the upstream and downstream related components of each component, form the dependency relationship table corresponding to the current system to be managed, and store it in a relational database.

5. The method according to any one of claims 1-4, characterized in that, Before inputting the fusion results into the large language reasoning model, the following steps are also included: Historical failure cases bound to root cause conclusions are extracted from the historical failure knowledge base, system architecture topology data is extracted from the relational database, and historical real-time operation data and historical alarm sample data of components are obtained from the external monitoring system and aggregated to form pre-training data. After standardizing the pre-training data, historical alarm sample data, system architecture topology relationship data, historical real-time operation data, and historical fault cases are used as pre-training input samples. The root cause conclusions associated with the historical failure cases are used as pre-trained output samples. The initial large language model is iteratively trained using the pre-trained input samples and pre-trained output samples until the root cause inference accuracy reaches a preset accuracy threshold, thus obtaining the large language inference model.

6. A root cause prediction device based on host state anomalies, characterized in that, include: The alarm data acquisition module is used to parse the abnormal host status alarm notification sent by the external monitoring system to obtain structured alarm data. The structured alarm data includes a host identifier and an alarm item. The target host system refers to a business support system consisting of multiple computing nodes and business function components deployed on them. The host identifier is a unique identifier of the computing node in the target host system, used to establish the mapping relationship between the computing node and the associated business components. The fault association component determination module is used to obtain abnormal components that match the host identifier in the relational database, generate a fault association component list based on all upstream and downstream associated components associated with the abnormal components, and extract the architecture topology relationship corresponding to the target host system from the relational database. The semantic transformation module is used to call the monitoring interface and the configuration management database system to collect real-time running data of all components in the fault-related component list, and generate semantic real-time status data based on the real-time running data and the preset semantic mapping template. The historical case data generation module is used to search for preliminary historical case data in the historical fault knowledge base based on the alarm items, and at the same time monitor the user interaction interface to generate historical case data sorted by matching degree according to the user feedback results of the user interaction interface. The fault root cause prediction result generation module is used to perform multimodal fusion processing on the structured alarm data, architecture topology relationship, historical case data and semantic real-time status data, and input the fusion result into the big language inference model to obtain multiple fault root cause prediction results sorted by probability, which are then displayed in the user interface. The semantic transformation module includes: The dynamic quantitative indicator acquisition unit is used to collect real-time operation indicator data of each component in the fault-related component list by calling the monitoring interface; wherein, the real-time operation indicator data is the dynamic quantitative indicator generated by each component during operation. The static basic information extraction unit is used to extract the component configuration information of each component in the fault-related component list by calling the configuration management database system; wherein, the component configuration information is static basic information containing the dependency relationship between upstream and downstream related components; The real-time running data generation unit is used to associate and integrate the component configuration information and real-time running indicator data of each component, and to perform standardization and structuring processing after association and integration to obtain real-time running data containing the component identifier, standardized dynamic quantitative indicators and standardized static basic data of each component. The semantic transformation module also includes: The threshold comparison unit is used to match the corresponding target indicator threshold in a preset semantic mapping template for each component in the real-time running data using the component identifier, and compare the standardized dynamic quantitative indicator of the current component with the target indicator threshold to obtain the threshold comparison result of the current component. The semantic description information lookup unit is used to search for semantic description information corresponding to the current threshold comparison result in a preset semantic mapping template based on the threshold comparison result. The topology influence association information lookup unit is used to search for topology influence association information corresponding to the threshold comparison result in a preset semantic mapping template based on the dependency relationship between upstream and downstream related components in the standardized static basic data corresponding to the current component. The semantic real-time status data generation unit is used to summarize the component identifier, threshold comparison results, semantic description information and topological influence association information corresponding to each component to obtain semantic real-time status data corresponding to the real-time running data.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform a root cause inference method based on host state anomalies according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute and implement the root cause inference method based on host state anomalies as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Alarm processing method and device of data center, electronic equipment and storage medium

    CN120750727A

  • Fault analysis method and device based on artificial intelligence, equipment and storage medium

    CN121099359A