Artificial intelligence-based fault resolution method, device, equipment and storage medium
Patent Information
- Application Number
- CN202610969788.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-01
AI Technical Summary
例如在金融技术领域多应用同时接入服务器,如银行应用程序、财险应用程序、车险应用程序、寿险应用程序和理财应用程序等等同时接入服务器,当服务器接收到同一业务故障触发多条技术告警,难以快速定位“业务影响”与“主告警”,导致告警异常处理冗余和技术告警无法直接映射到业务事件,导致告警处理效率低,从而严重影响用户的使用
[0009]This application provides an artificial intelligence-based fault solution method, apparatus, device, and storage medium. In response to receiving a multi-source heterogeneous alarm request, this application performs alarm noise reduction and aggregation on the multi-source heterogeneous alarm request according to topological dependencies to obtain a target heterogeneous alarm request; performs abnormal event analysis and business tag matching on the target heterogeneous alarm request based on a preset business attribution knowledge graph to obtain heterogeneous alarm data; identifies the root causes of anomalies in the heterogeneous alarm data using a preset anomaly root cause identification model, which is a pre-trained convergent neural network model; generates a fault solution based on the anomaly root cause information using the preset business attribution knowledge graph, and performs fault anomaly handling based on the fault solution. This application first performs alarm noise reduction and aggregation on multi-source heterogeneous alarm requests through topological dependency relationships to obtain target heterogeneous alarm requests, which can effectively avoid duplicate alarm information and omission of alarm information. A preset business attribution knowledge graph performs abnormal event analysis and business tag matching on the target heterogeneous alarm requests to obtain heterogeneous alarm data, which can effectively associate heterogeneous alarm data with business attribution, thereby improving the accuracy of root cause identification. A neural network model is used to identify abnormal root causes of heterogeneous alarm data, making the obtained abnormal root cause information more accurate. Finally, based on the preset business attribution knowledge graph, fault solutions are generated from the abnormal root cause information, which can accurately obtain fault solutions and perform fault anomaly handling according to the fault solutions, which can efficiently and accurately solve multi-source heterogeneous alarm faults.
Smart Images

Figure CN122679025A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a fault solution method, apparatus, device and storage medium based on artificial intelligence. Background Technology
[0002] With the rapid development of network interconnection, most industries have adopted a business layout based on servers, middleware, multi-cloud, and containers. This layout can effectively improve the efficiency of business operations. For example, in the financial technology field, multiple applications simultaneously connect to the server, such as banking applications, property insurance applications, auto insurance applications, life insurance applications, and wealth management applications. When the server receives multiple technical alarms triggered by the same business failure, it is difficult to quickly locate the "business impact" and the "primary alarm." This leads to redundant alarm handling and the inability to directly map technical alarms to business events, resulting in low alarm handling efficiency and seriously affecting user experience.
[0003] Therefore, improving the efficiency of troubleshooting alarm anomalies is an urgent problem to be solved. Summary of the Invention
[0004] The main purpose of this application is to provide a fault solution method, device, equipment and storage medium based on artificial intelligence, which aims to improve the efficiency and accuracy of fault resolution for alarm anomalies.
[0005] Firstly, this application provides a fault-solving method based on artificial intelligence, which includes the following steps: In response to receiving a multi-source heterogeneous alarm request, the multi-source heterogeneous alarm request is denoised and aggregated according to topological dependencies to obtain the target heterogeneous alarm request; Based on a preset business attribution knowledge graph, the target heterogeneous alarm request is analyzed for abnormal events and matched with business tags to obtain heterogeneous alarm data. The heterogeneous alarm data is subjected to anomaly root cause identification through a preset anomaly root cause identification model to obtain anomaly root cause information. The preset anomaly root cause identification model is a pre-trained convergent neural network model. Based on a preset business attribution knowledge graph, a fault solution is generated from the abnormal root cause information to obtain a fault solution, and fault and abnormal handling is performed based on the fault solution.
[0006] Secondly, this application also provides a fault-solving device, which includes an acquisition module, a generation module, and an exception handling module, wherein: The acquisition module is used to receive multi-source heterogeneous alarm requests; The generation module is used to perform alarm noise reduction and aggregation on the multi-source heterogeneous alarm requests according to the topological dependency relationship to obtain the target heterogeneous alarm request; The generation module is used to perform abnormal event analysis and business tag matching on the target heterogeneous alarm request based on a preset business attribution knowledge graph to obtain heterogeneous alarm data. The generation module is further configured to perform anomaly root cause identification on the heterogeneous alarm data through a preset anomaly root cause identification model to obtain anomaly root cause information. The preset anomaly root cause identification model is a pre-trained convergent neural network model. The generation module is also used to generate fault solutions from the abnormal root cause information based on a preset business attribution knowledge graph, thereby obtaining fault solutions. The exception handling module is used to handle fault exceptions according to the fault solution.
[0007] Thirdly, this application also provides a computer device, the computer device including a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the steps of the above-described artificial intelligence-based fault solution method.
[0008] Fourthly, this application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the above-described artificial intelligence-based fault-solving method.
[0009] This application provides an artificial intelligence-based fault solution method, apparatus, device, and storage medium. In response to receiving a multi-source heterogeneous alarm request, this application performs alarm noise reduction and aggregation on the multi-source heterogeneous alarm request according to topological dependencies to obtain a target heterogeneous alarm request; performs abnormal event analysis and business tag matching on the target heterogeneous alarm request based on a preset business attribution knowledge graph to obtain heterogeneous alarm data; identifies the root causes of anomalies in the heterogeneous alarm data using a preset anomaly root cause identification model, which is a pre-trained convergent neural network model; generates a fault solution based on the anomaly root cause information using the preset business attribution knowledge graph, and performs fault anomaly handling based on the fault solution. This application first performs alarm noise reduction and aggregation on multi-source heterogeneous alarm requests through topological dependency relationships to obtain target heterogeneous alarm requests, which can effectively avoid duplicate alarm information and omission of alarm information. A preset business attribution knowledge graph performs abnormal event analysis and business tag matching on the target heterogeneous alarm requests to obtain heterogeneous alarm data, which can effectively associate heterogeneous alarm data with business attribution, thereby improving the accuracy of root cause identification. A neural network model is used to identify abnormal root causes of heterogeneous alarm data, making the obtained abnormal root cause information more accurate. Finally, based on the preset business attribution knowledge graph, fault solutions are generated from the abnormal root cause information, which can accurately obtain fault solutions and perform fault anomaly handling according to the fault solutions, which can efficiently and accurately solve multi-source heterogeneous alarm faults. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A flowchart illustrating an artificial intelligence-based fault solution provided in an embodiment of this application; Figure 2 for Figure 1 A flowchart illustrating the sub-steps of an AI-based fault resolution method. Figure 3 for Figure 1 A flowchart illustrating the sub-steps of an AI-based fault resolution method. Figure 4 A schematic block diagram of a fault-solving device provided in an embodiment of this application; Figure 5 for Figure 4 A schematic block diagram of a submodule of the fault-solving device in the system; Figure 6 This is a schematic block diagram of the structure of a computer device provided in an embodiment of this application.
[0012] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0015] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0016] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0017] With the rapid development of network interconnection, most industries have adopted a business layout based on servers, middleware, multi-cloud, and containers. This layout can effectively improve the efficiency of business operations. For example, in the financial technology field, multiple applications simultaneously connect to the server, such as banking applications, property insurance applications, auto insurance applications, life insurance applications, and wealth management applications. When the server receives multiple technical alarms triggered by the same business failure, it is difficult to quickly locate the "business impact" and the "primary alarm." This leads to redundant alarm handling and the inability to directly map technical alarms to business events, resulting in low alarm handling efficiency and seriously affecting user experience.
[0018] To address the aforementioned issues, embodiments of this application provide an AI-based fault solution method, apparatus, device, and storage medium. This AI-based fault solution method includes, in response to receiving a multi-source heterogeneous alarm request, performing alarm noise reduction and aggregation on the multi-source heterogeneous alarm request based on topological dependencies to obtain a target heterogeneous alarm request; performing anomaly event analysis and business tag matching on the target heterogeneous alarm request based on a preset business attribution knowledge graph to obtain heterogeneous alarm data; identifying the root causes of anomalies in the heterogeneous alarm data using a preset anomaly root cause identification model, wherein the preset anomaly root cause identification model is a pre-trained convergent neural network model; generating a fault solution based on the anomaly root cause information using the preset business attribution knowledge graph to obtain a fault solution; and performing fault anomaly handling based on the fault solution.
[0019] This AI-based troubleshooting method can be applied to computer devices such as mobile phones, tablets, laptops, desktop computers, and personal digital assistants.
[0020] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0021] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an artificial intelligence-based fault solution provided for an embodiment of this application.
[0022] like Figure 1 As shown, the fault-solving method based on artificial intelligence includes steps S101 to S104.
[0023] Step S101: In response to receiving a multi-source heterogeneous alarm request, perform alarm noise reduction and aggregation on the multi-source heterogeneous alarm request according to the topological dependency relationship to obtain the target heterogeneous alarm request.
[0024] Among them, multi-source heterogeneous alarms include various platform layers connected to the computer, including but not limited to the infrastructure layer, platform layer, application layer, security layer, and monitoring layer. The infrastructure layer includes but is not limited to servers, networks, and storage; the platform layer includes but is not limited to databases, middleware, and cloud platforms; the application layer includes but is not limited to core business systems, channel systems, and customer service systems; the security layer includes but is not limited to firewalls, application firewalls, and intrusion detection systems; and the monitoring layer includes but is not limited to Zabbix, Prometheus, and ELK.
[0025] In some embodiments, log information from each platform layer is acquired, and alarm identification is performed on each log information. When an abnormal alarm is detected, the log information of the abnormal alarm is stored, and a multi-source heterogeneous alarm request is generated based on the log information of the abnormal alarm. By monitoring the log information of each platform, multi-source heterogeneous alarm requests can be effectively obtained.
[0026] In some embodiments, data cleaning and standardization are performed on multi-source heterogeneous alarm requests to generate pre-processed multi-source heterogeneous alarm requests. By performing data cleaning and standardization on multi-source heterogeneous alarm requests, the accuracy of abnormal alarm identification can be effectively improved.
[0027] It should be noted that this data cleaning includes, but is not limited to, deduplication and merging of alarm data, time / format standardization, terminology mapping, and completion or removal of abnormal fields. Format standardization includes, but is not limited to, converting heterogeneous data from different sources into a unified JSON structure to extract core fields such as time, source, level, content, IP / host, service, and metrics. Standardization processing includes, but is not limited to, unifying timestamps, mapping alarm levels to a unified standard, and uniformly naming service / applications according to business domains.
[0028] In some embodiments, similar alarms are aggregated from multi-source heterogeneous alarm requests according to a preset clustering algorithm to obtain candidate multi-source heterogeneous alarm requests; alarm noise reduction and aggregation are then performed on the candidate multi-source heterogeneous alarm requests based on topological dependencies to obtain the target heterogeneous alarm request. The preset clustering algorithm includes, but is not limited to, K-means clustering and DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering. By performing alarm noise reduction and aggregation on multi-source heterogeneous alarm requests, interference from invalid alarm information can be effectively reduced, and the identification of alarm anomalies can be improved.
[0029] It should be noted that this topology dependency is generated in advance based on the topology relationships generated by service calls, hosts, applications, and databases.
[0030] In some embodiments, the method of obtaining the target heterogeneous alarm request by performing alarm noise reduction and aggregation on candidate multi-source heterogeneous alarm requests according to topological dependencies can be as follows: erroneous alarm information of candidate multi-source heterogeneous alarm requests is eliminated according to topological dependencies, and chain alarms caused by the same fault are aggregated to obtain the target heterogeneous alarm request.
[0031] Step S102: Based on the preset business attribution knowledge graph, perform abnormal event analysis and business tag matching on the target heterogeneous alarm request to obtain heterogeneous alarm data.
[0032] The preset business attribution knowledge graph is a knowledge graph pre-constructed based on the business entity layer, business relationship layer, and business rule layer.
[0033] For example, in the financial technology field, the business entity layer includes physical (IT) entities, business entities, channel entities, and user entities. The IT entities include servers, databases, middleware, microservices, and gateways. The business entities include insurance application, underwriting, claims processing, payment, and policy maintenance. The channel entities include applications, official websites, bancassurance, and agent channels. The user entities include individual entities and enterprise entities. The business relationship layer establishes relationships between entities. For example, service A calls service B, the database supports the underwriting system, the claims system belongs to the property insurance business system, and gateway anomalies affect application channels. The business rule layer defines rules for faults and their impact on business. For example, database lock anomalies correspond to insurance timeouts, Redis timeouts correspond to policy query failures, Kafka backlogs correspond to message delays, and API 500 (Internal Server Error) corresponds to claims interface anomalies, etc. Computer devices extract data from the Configuration Management Database (CMDB), service registry, API gateway, APM links, configuration center, and historical fault tickets to construct a knowledge graph, resulting in a business attribution knowledge graph.
[0034] In some embodiments, such as Figure 2 As shown, step S102 includes sub-steps S1021 to S1023.
[0035] Sub-step S1021: Extract alarm anomaly events and match business domains for the target heterogeneous alarm request based on the preset business attribution knowledge graph to obtain heterogeneous alarm indicator information and business domain information.
[0036] The preset business attribution knowledge graph is used to extract services, hosts, error codes and indicators from the target heterogeneous alarm requests to obtain heterogeneous alarm indicator information. In addition, the preset business attribution knowledge graph is used to extract the corresponding insurance business domain, process and channel information from the target heterogeneous alarm requests to obtain business domain information.
[0037] Sub-step S1022: Generate alarm service tags by processing the heterogeneous alarm indicator information and the service domain information to obtain alarm service tags.
[0038] A pre-defined mapping table between heterogeneous alarm indicator information, business domain information, and business tags is obtained. At least one business tag matching the heterogeneous alarm indicator information and business domain information is then retrieved from this mapping table to obtain the alarm business tag. This mapping table is pre-established based on the heterogeneous alarm indicator information, business domain information, and business tags. The mapping table can be established according to actual circumstances, and this embodiment does not impose specific limitations on it. This mapping table allows for accurate retrieval of the alarm business tag corresponding to the heterogeneous alarm indicator information and business domain information.
[0039] Sub-step S1023: Perform heterogeneous alarm aggregation on the heterogeneous alarm indicator information, the service domain information and the alarm service label to obtain heterogeneous alarm data.
[0040] Heterogeneous alarm aggregation is performed on business domain information and alarm business tags of the same type to obtain heterogeneous alarm data including at least one sub-heterogeneous alarm data.
[0041] Step S103: The heterogeneous alarm data is subjected to abnormal root cause identification through a preset abnormal root cause identification model to obtain abnormal root cause information. The preset abnormal root cause identification model is a pre-trained convergent neural network model.
[0042] The anomaly root cause identification model is a pre-trained and converged neural network model. The pre-set anomaly root cause identification model includes a feature extraction layer, an anomaly classification layer, and a root cause output layer. The feature extraction layer is used to extract features from heterogeneous alarm data, the anomaly classification layer is used for fault feature matching, and the root cause output layer is used for anomaly root cause matching to output anomaly root cause information. The feature extraction layer can be TF-IDF, the anomaly classification layer can be a convolutional neural network model, and the root cause output layer includes a business attribution knowledge graph.
[0043] In some embodiments, a sample dataset is obtained, which includes multiple sample data, including heterogeneous alarm data and labeled anomaly root cause information; an initialized anomaly root cause identification model is obtained, and a sample data is selected from the sample dataset as target sample data, which includes heterogeneous alarm data and labeled anomaly root cause information; the heterogeneous alarm data of the target sample data is input into the initialized anomaly root cause identification model to identify anomalies root causes and obtain predicted anomaly root cause information; based on the labeled anomaly root cause information and the predicted anomaly root cause information, it is determined whether the anomaly root cause identification model has converged; if the anomaly root cause identification model has not converged, the model parameters of the anomaly root cause identification model are adjusted, and the step of selecting a sample data from the sample dataset as target sample data is continued until a converged anomaly root cause identification model is obtained.
[0044] It should be noted that the specific implementation method of inputting the heterogeneous alarm data of the target sample data into the initialized abnormal root cause identification model to identify abnormal root causes and obtain the predicted abnormal root cause information can be referred to in the following embodiments for the specific implementation method of identifying abnormal root causes of the heterogeneous alarm data through the preset abnormal root cause identification model to obtain abnormal root cause information.
[0045] In some embodiments, determining whether the anomaly root cause identification model has converged based on the labeled and predicted anomaly root cause information can be achieved by: determining the loss value of the anomaly root cause identification model based on the labeled and predicted anomaly root cause information; if the loss value is less than or equal to a preset loss value, the anomaly root cause identification model is determined to have converged; if the loss value is greater than the preset loss value, the anomaly root cause identification model is determined to have not converged. The preset loss value can be set according to actual circumstances, and this application does not specifically limit it.
[0046] In some embodiments, the loss value of the anomaly root cause identification model can be determined based on the labeled anomaly root cause information and the predicted anomaly root cause information as follows: calculate the similarity between the labeled anomaly root cause information and the predicted anomaly root cause information to obtain the current similarity; subtract the current similarity from the unit value to obtain the current loss value; obtain the historical loss value, which is the average total loss generated during model training before the current time point; calculate the average of the historical loss value and the current loss value to obtain the loss value of the anomaly root cause identification model. The method for calculating the similarity between the labeled anomaly root cause information and the predicted anomaly root cause information can be selected according to the actual situation, and this application embodiment does not specifically limit this method. For example, the cosine similarity between the labeled anomaly root cause information and the predicted anomaly root cause information can be calculated.
[0047] In some embodiments, such as Figure 3 As shown, step S103 includes sub-steps S1031 to S1033.
[0048] Sub-step S1031: Extract features from the heterogeneous alarm data through the feature extraction layer to obtain heterogeneous alarm feature vectors.
[0049] Feature extraction of heterogeneous alarm data is performed using a pre-defined TF-IDF layer to obtain heterogeneous alarm feature vectors. This pre-defined TF-IDF layer extracts features based on the positional weights of word groups. This method of feature extraction from heterogeneous alarm data using the pre-defined TF-IDF layer can accurately obtain heterogeneous alarm feature vectors.
[0050] It should be noted that the preset TF-IDF layer extracts adjacent word groups as features, increases the weight of words / word groups near the end of the text, and introduces inter-class distribution coefficients to measure the distribution differences of words in different fault categories. It assigns higher weights to words that appear frequently only in a certain type of fault, thereby improving the discriminative power of features and effectively improving the accuracy of feature vector extraction.
[0051] Sub-step S1032: The heterogeneous alarm feature vector is classified into business events and matched with business fault features through the anomaly classification layer to obtain anomaly fault information.
[0052] Heterogeneous alarm feature vectors are clustered to obtain multiple heterogeneous alarm feature class vectors. Anomaly identification is then performed on these heterogeneous alarm feature class vectors to obtain candidate anomaly information for each vector, including the anomaly category and its corresponding confidence level. The anomaly category with the highest confidence level is selected as the target anomaly, and its anomaly information is output. This anomaly classification layer, which performs business event classification and business fault feature matching on the heterogeneous alarm feature vectors, accurately obtains anomaly information.
[0053] Sub-step S1033: Perform abnormal root cause matching on the abnormal fault information through the root cause output layer to obtain the abnormal root cause information.
[0054] Based on preset link tracing information, fault location matching is performed on abnormal fault information to obtain candidate abnormal root cause information; based on preset fault tree, fault propagation path matching is performed on candidate abnormal root cause information to obtain abnormal root cause information.
[0055] The preset link tracing information and preset fault tree are obtained by pre-constructing link tracing information and fault tree based on historical fault cases.
[0056] In some embodiments, the method of obtaining candidate abnormal root cause information by performing fault location matching on abnormal fault information based on preset link tracing information can be as follows: the preset link tracing information is used to perform fault service location, interface level location and root cause range narrowing on abnormal fault information to obtain candidate abnormal root cause information.
[0057] In some embodiments, the method of obtaining the abnormal root cause information by matching the fault propagation path of the candidate abnormal root cause information based on the preset fault tree can be: verifying the fault propagation path of the candidate abnormal root cause information based on the preset fault tree to obtain the abnormal root cause information.
[0058] Step S104: Generate a fault solution from the abnormal root cause information based on the preset business attribution knowledge graph, obtain the fault solution, and perform fault and abnormal processing based on the fault solution.
[0059] In some embodiments, the abnormal root cause information is matched with similar fault cases based on a preset business attribution knowledge graph to obtain at least one similar fault case; the error information in the abnormal root cause information is corrected based on the abnormal root cause information of the at least one similar fault case to obtain the target abnormal root cause information; and the fault solution matched with the target abnormal root cause information is obtained. By correcting the deviation of the abnormal root cause information through similar fault cases, the accuracy of the abnormal root cause information can be effectively improved.
[0060] In some embodiments, obtaining the fault solution matching the target anomaly root cause information can be achieved by: obtaining a preset mapping table between anomaly root cause information and fault solutions, and querying the fault solution matching the target anomaly root cause information from the mapping table. This mapping table is pre-established based on anomaly root cause information and fault solutions, and can be established according to actual circumstances; this embodiment does not specifically limit its implementation. The fault solution can be accurately retrieved through this mapping table.
[0061] In some embodiments, fault anomaly handling based on fault solutions can accurately and efficiently complete fault anomaly handling.
[0062] For example, in the field of financial management technology, multiple applications simultaneously access a server. Upon receiving a multi-source heterogeneous financial management alarm request, the server performs alarm noise reduction and aggregation based on topological dependencies to obtain the target multi-source heterogeneous financial management alarm request. Based on a preset business attribution knowledge graph, the server performs anomaly event analysis and business tag matching on the target multi-source heterogeneous financial management alarm request to obtain heterogeneous alarm data. A preset anomaly root cause identification model is used to identify the root causes of anomalies in the heterogeneous alarm data to obtain anomaly root cause information. Based on the preset business attribution knowledge graph, a fault solution is generated from the anomaly root cause information to obtain a fault solution. Finally, fault anomaly handling is performed based on the fault solution to complete the processing of multi-source heterogeneous financial management alarms.
[0063] For example, in the field of healthcare technology, medical insurance service applications, hospital applications, life insurance applications, and medical insurance applications are simultaneously connected to a server. Upon receiving a multi-source heterogeneous healthcare alarm request, the server performs alarm noise reduction and aggregation on the multi-source heterogeneous healthcare alarm request based on topological dependencies to obtain the target healthcare heterogeneous alarm request. Based on a preset business attribution knowledge graph, the server performs anomaly event analysis and business tag matching on the target healthcare heterogeneous alarm request to obtain healthcare heterogeneous alarm data. A preset anomaly root cause identification model is used to identify the root causes of anomalies in the healthcare heterogeneous alarm data to obtain healthcare anomaly root cause information. Based on the preset business attribution knowledge graph, a fault solution is generated from the healthcare anomaly root cause information to obtain a fault solution, and fault anomaly processing is performed based on the fault solution to complete the processing of multi-source heterogeneous healthcare alarms.
[0064] The AI-based fault resolution method provided in the above embodiments, in response to receiving multi-source heterogeneous alarm requests, performs alarm noise reduction and aggregation on the multi-source heterogeneous alarm requests according to topological dependencies to obtain target heterogeneous alarm requests; performs abnormal event analysis and business tag matching on the target heterogeneous alarm requests based on a preset business attribution knowledge graph to obtain heterogeneous alarm data; identifies the root causes of anomalies in the heterogeneous alarm data through a preset anomaly root cause identification model, which is a pre-trained convergent neural network model; generates fault solutions based on the anomaly root cause information according to the preset business attribution knowledge graph, obtains fault solutions, and performs fault anomaly handling according to the fault solutions. This application first performs alarm noise reduction and aggregation on multi-source heterogeneous alarm requests through topological dependency relationships to obtain target heterogeneous alarm requests, which can effectively avoid duplicate alarm information and omission of alarm information. A preset business attribution knowledge graph performs abnormal event analysis and business tag matching on the target heterogeneous alarm requests to obtain heterogeneous alarm data, which can effectively associate heterogeneous alarm data with business attribution, thereby improving the accuracy of root cause identification. A neural network model is used to identify abnormal root causes of heterogeneous alarm data, making the obtained abnormal root cause information more accurate. Finally, based on the preset business attribution knowledge graph, fault solutions are generated from the abnormal root cause information, which can accurately obtain fault solutions and perform fault anomaly handling according to the fault solutions, which can efficiently and accurately solve multi-source heterogeneous alarm faults.
[0065] Please see Figure 4 , Figure 4 This is a schematic block diagram of a fault-solving device provided in an embodiment of this application.
[0066] like Figure 4As shown, the fault-solving device 200 includes an acquisition module 210, a generation module 220, and an exception handling module 230, wherein: The acquisition module 210 is used to receive multi-source heterogeneous alarm requests; The generation module 220 is used to perform alarm noise reduction and aggregation on the multi-source heterogeneous alarm requests according to the topological dependency relationship to obtain the target heterogeneous alarm request. The generation module 220 is also used to perform abnormal event analysis and business tag matching on the target heterogeneous alarm request based on a preset business attribution knowledge graph to obtain heterogeneous alarm data; The generation module 220 is further configured to perform anomaly root cause identification on the heterogeneous alarm data through a preset anomaly root cause identification model to obtain anomaly root cause information. The preset anomaly root cause identification model is a pre-trained convergent neural network model. The generation module 220 is also used to generate a fault solution based on the abnormal root cause information according to a preset business attribution knowledge graph, so as to obtain a fault solution. The exception handling module 230 is used to perform fault exception handling according to the fault solution.
[0067] In some embodiments, such as Figure 5 As shown, the generation module 220 includes a feature extraction module 221, a fault feature matching module 222, and an anomaly root cause matching module 223, wherein: The feature extraction module 221 is used to extract features from the heterogeneous alarm data through the feature extraction layer to obtain a heterogeneous alarm feature vector. The fault feature matching module 222 is used to perform business event classification and business fault feature matching on the heterogeneous alarm feature vector through the anomaly classification layer to obtain abnormal fault information. The abnormal root cause matching module 223 is used to perform abnormal root cause matching on the abnormal fault information through the root cause output layer to obtain the abnormal root cause information.
[0068] In some embodiments, the feature extraction module 221 is further configured to: The heterogeneous alarm data is subjected to feature extraction by a preset TF-IDF layer to obtain a heterogeneous alarm feature vector. The preset TF-IDF layer performs feature extraction based on the position weight of the word group.
[0069] In some embodiments, the fault feature matching module 222 is further configured to: Cluster the heterogeneous alarm feature vectors to obtain multiple heterogeneous alarm feature class vectors; Anomaly fault identification is performed on the multiple heterogeneous alarm feature class vectors to obtain candidate anomaly fault information corresponding to each heterogeneous alarm feature class vector. The candidate anomaly fault information includes anomaly fault category and corresponding confidence level. The anomaly category with the highest confidence level is selected as the target anomaly, and the anomaly information of the target anomaly is output.
[0070] In some embodiments, the anomaly root cause matching module 223 is further configured to: Based on preset link tracing information, the abnormal fault information is matched to obtain candidate abnormal root cause information. Based on a pre-defined fault tree, the candidate anomaly root cause information is matched with the fault propagation path to obtain the anomaly root cause information.
[0071] In some embodiments, the generation module 220 is further configured to: Based on the preset business attribution knowledge graph, the abnormal root cause information is matched with similar fault cases to obtain at least one similar fault case. Based on the abnormal root cause information of at least one similar fault case, the error information in the abnormal root cause information is corrected to obtain the target abnormal root cause information. Obtain the fault solution matched with the root cause information of the target anomaly.
[0072] In some embodiments, the generation module 220 is further used for Based on the preset business attribution knowledge graph, the target heterogeneous alarm request is subjected to alarm anomaly event extraction and business domain matching to obtain heterogeneous alarm indicator information and business domain information. Service tags are generated from the heterogeneous alarm indicator information and the service domain information to obtain alarm service tags; Heterogeneous alarm data is obtained by aggregating the heterogeneous alarm indicator information, the service domain information, and the alarm service tags.
[0073] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the above-mentioned fault-solving device can be referred to the corresponding process in the aforementioned embodiment of the fault-solving method based on artificial intelligence, and will not be repeated here.
[0074] Please see Figure 6 , Figure 6 This is a schematic block diagram of the structure of a computer device provided in an embodiment of this application.
[0075] like Figure 6As shown, the computer device 300 includes a processor 302 and a memory 303 connected via a system bus 301, wherein the memory 303 may include a storage medium and internal memory.
[0076] The storage medium may store a computer program. This computer program includes program instructions that, when executed, cause the processor to perform any artificial intelligence-based fault-solving method.
[0077] The processor 302 provides computing and control capabilities to support the operation of the entire computer device.
[0078] Internal memory provides an environment for the execution of computer programs stored in the storage medium. When these computer programs are executed by the processor, the processor can perform any kind of artificial intelligence-based fault-solving method.
[0079] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0080] It should be understood that processor 302 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, the general-purpose processor can be a microprocessor or any conventional processor.
[0081] In one embodiment, the processor 302 is configured to run a computer program stored in a memory to perform the following steps: In response to receiving a multi-source heterogeneous alarm request, the multi-source heterogeneous alarm request is denoised and aggregated according to topological dependencies to obtain the target heterogeneous alarm request; Based on a preset business attribution knowledge graph, the target heterogeneous alarm request is analyzed for abnormal events and matched with business tags to obtain heterogeneous alarm data. The heterogeneous alarm data is subjected to anomaly root cause identification through a preset anomaly root cause identification model to obtain anomaly root cause information. The preset anomaly root cause identification model is a pre-trained convergent neural network model. Based on a preset business attribution knowledge graph, a fault solution is generated from the abnormal root cause information to obtain a fault solution, and fault and abnormal handling is performed based on the fault solution.
[0082] In one embodiment, the preset anomaly root cause identification model includes a feature extraction layer, an anomaly classification layer, and a root cause output layer; when the processor 302 performs anomaly root cause identification on the heterogeneous alarm data through the preset anomaly root cause identification model to obtain anomaly root cause information, it is used to implement: The feature extraction layer extracts features from the heterogeneous alarm data to obtain a heterogeneous alarm feature vector. The anomaly classification layer performs business event classification and business fault feature matching on the heterogeneous alarm feature vector to obtain anomaly fault information. The abnormal fault information is obtained by performing abnormal root cause matching on the abnormal fault information through the root cause output layer.
[0083] In one embodiment, when the processor 302 performs feature extraction on the heterogeneous alarm data through the feature extraction layer to obtain a heterogeneous alarm feature vector, it is used to implement: The heterogeneous alarm data is subjected to feature extraction by a preset TF-IDF layer to obtain a heterogeneous alarm feature vector. The preset TF-IDF layer performs feature extraction based on the position weight of the word group.
[0084] In one embodiment, when the processor 302 performs business event classification and business fault feature matching on the heterogeneous alarm feature vector through the anomaly classification layer to obtain anomaly fault information, it is used to implement: Cluster the heterogeneous alarm feature vectors to obtain multiple heterogeneous alarm feature class vectors; Anomaly fault identification is performed on the multiple heterogeneous alarm feature class vectors to obtain candidate anomaly fault information corresponding to each heterogeneous alarm feature class vector. The candidate anomaly fault information includes anomaly fault category and corresponding confidence level. The anomaly category with the highest confidence level is selected as the target anomaly, and the anomaly information of the target anomaly is output.
[0085] In one embodiment, when the processor 302 performs abnormal root cause matching on the abnormal fault information through the root cause output layer to obtain the abnormal root cause information, it is configured to: Based on preset link tracing information, the abnormal fault information is matched to obtain candidate abnormal root cause information. Based on a pre-defined fault tree, the candidate anomaly root cause information is matched with the fault propagation path to obtain the anomaly root cause information.
[0086] In one embodiment, when the processor 302 generates a fault solution based on the abnormal root cause information according to a preset business attribution knowledge graph, it is configured to: Based on the preset business attribution knowledge graph, the abnormal root cause information is matched with similar fault cases to obtain at least one similar fault case. Based on the abnormal root cause information of at least one similar fault case, the error information in the abnormal root cause information is corrected to obtain the target abnormal root cause information. Obtain the fault solution matched with the root cause information of the target anomaly.
[0087] In one embodiment, when the processor 302 performs anomaly event parsing and business tag matching on the target heterogeneous alarm request based on a preset business attribution knowledge graph to obtain heterogeneous alarm data, it is used to: Based on the preset business attribution knowledge graph, the target heterogeneous alarm request is subjected to alarm anomaly event extraction and business domain matching to obtain heterogeneous alarm indicator information and business domain information. Service tags are generated from the heterogeneous alarm indicator information and the service domain information to obtain alarm service tags; Heterogeneous alarm data is obtained by aggregating the heterogeneous alarm indicator information, the service domain information, and the alarm service tags.
[0088] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the computer equipment described above can be referred to the corresponding process in the aforementioned embodiments of the fault solution method based on artificial intelligence, and will not be repeated here.
[0089] This application also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, and the method implemented when the program instructions are executed can refer to the various embodiments of the fault solution based on artificial intelligence in this application.
[0090] The computer-readable storage medium can be an internal storage unit of the computer device described in the foregoing embodiments, such as a hard disk or memory of the computer device. The computer-readable storage medium can be non-volatile or volatile. Alternatively, the computer-readable storage medium can be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.
[0091] Furthermore, the computer-readable storage medium may primarily include a program storage area and a data storage area, wherein the program storage area may store the operating system, at least one application required for a function, etc.; and the data storage area may store data created based on the use of blockchain nodes, etc.
[0092] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0093] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” are intended to include the plural forms.
[0094] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0095] It should also be understood that the term "and / or" as used in this specification refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. It should be noted that, herein, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0096] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. The above descriptions are merely specific implementations of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A fault-solving method based on artificial intelligence, characterized in that, include: In response to receiving a multi-source heterogeneous alarm request, the multi-source heterogeneous alarm request is denoised and aggregated according to topological dependencies to obtain the target heterogeneous alarm request; Based on a preset business attribution knowledge graph, the target heterogeneous alarm request is analyzed for abnormal events and matched with business tags to obtain heterogeneous alarm data. The heterogeneous alarm data is subjected to anomaly root cause identification through a preset anomaly root cause identification model to obtain anomaly root cause information. The preset anomaly root cause identification model is a pre-trained convergent neural network model. Based on a preset business attribution knowledge graph, a fault solution is generated from the abnormal root cause information to obtain a fault solution, and fault and abnormal handling is performed based on the fault solution.
2. The fault-solving method based on artificial intelligence as described in claim 1, characterized in that, The preset abnormal root cause identification model includes a feature extraction layer, an abnormal classification layer, and a root cause output layer; The step of identifying the root causes of heterogeneous alarm data using a preset root cause identification model to obtain root cause information includes: The feature extraction layer extracts features from the heterogeneous alarm data to obtain a heterogeneous alarm feature vector. The anomaly classification layer performs business event classification and business fault feature matching on the heterogeneous alarm feature vector to obtain anomaly fault information. The abnormal fault information is obtained by performing abnormal root cause matching on the abnormal fault information through the root cause output layer.
3. The fault-solving method based on artificial intelligence as described in claim 2, characterized in that, The step of extracting features from the heterogeneous alarm data through the feature extraction layer to obtain heterogeneous alarm feature vectors includes: The heterogeneous alarm data is subjected to feature extraction by a preset TF-IDF layer to obtain a heterogeneous alarm feature vector. The preset TF-IDF layer performs feature extraction based on the position weight of the word group.
4. The fault-solving method based on artificial intelligence as described in claim 2, characterized in that, The step of classifying the heterogeneous alarm feature vectors into business events and matching business fault features through the anomaly classification layer to obtain anomaly fault information includes: Cluster the heterogeneous alarm feature vectors to obtain multiple heterogeneous alarm feature class vectors; Anomaly fault identification is performed on the multiple heterogeneous alarm feature class vectors to obtain candidate anomaly fault information corresponding to each heterogeneous alarm feature class vector. The candidate anomaly fault information includes anomaly fault category and corresponding confidence level. The anomaly category with the highest confidence level is selected as the target anomaly, and the anomaly information of the target anomaly is output.
5. The fault-solving method based on artificial intelligence as described in claim 2, characterized in that, The step of performing abnormal root cause matching on the abnormal fault information through the root cause output layer to obtain the abnormal root cause information includes: Based on preset link tracing information, the abnormal fault information is matched to obtain candidate abnormal root cause information. Based on a pre-defined fault tree, the candidate anomaly root cause information is matched with the fault propagation path to obtain the anomaly root cause information.
6. The fault-solving method based on artificial intelligence as described in claim 1, characterized in that, The step of generating a fault solution based on the abnormal root cause information according to a preset business attribution knowledge graph, to obtain a fault solution, includes: Based on the preset business attribution knowledge graph, the abnormal root cause information is matched with similar fault cases to obtain at least one similar fault case. Based on the abnormal root cause information of at least one similar fault case, the error information in the abnormal root cause information is corrected to obtain the target abnormal root cause information. Obtain the fault solution matched with the root cause information of the target anomaly.
7. The fault-solving method based on artificial intelligence as described in claim 1, characterized in that, The method involves parsing abnormal events and matching business tags for the target heterogeneous alarm requests based on a preset business attribution knowledge graph to obtain heterogeneous alarm data, including: Based on the preset business attribution knowledge graph, the target heterogeneous alarm request is subjected to alarm anomaly event extraction and business domain matching to obtain heterogeneous alarm indicator information and business domain information. Service tags are generated from the heterogeneous alarm indicator information and the service domain information to obtain alarm service tags; Heterogeneous alarm data is obtained by aggregating the heterogeneous alarm indicator information, the service domain information, and the alarm service tags.
8. A fault-solving device, characterized in that, The fault-solving device includes an acquisition module, a generation module, and an exception handling module, wherein: The acquisition module is used to receive multi-source heterogeneous alarm requests; The generation module is used to perform alarm noise reduction and aggregation on the multi-source heterogeneous alarm requests according to the topological dependency relationship to obtain the target heterogeneous alarm request; The generation module is used to perform abnormal event analysis and business tag matching on the target heterogeneous alarm request based on a preset business attribution knowledge graph to obtain heterogeneous alarm data. The generation module is further configured to perform anomaly root cause identification on the heterogeneous alarm data through a preset anomaly root cause identification model to obtain anomaly root cause information. The preset anomaly root cause identification model is a pre-trained convergent neural network model. The generation module is also used to generate fault solutions from the abnormal root cause information based on a preset business attribution knowledge graph, thereby obtaining fault solutions. The exception handling module is used to handle fault exceptions according to the fault solution.
9. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the steps of the artificial intelligence-based fault solution method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the artificial intelligence-based fault-solving method as described in any one of claims 1 to 7.