Alarm analysis method, device and product applied to distributed system
By classifying and analyzing alarm data in a distributed system, and generating alarm analysis reports using operational knowledge and target models, the problems of misjudgment and high maintenance costs in alarm information processing in distributed systems are solved, achieving efficient and accurate fault location and operational guidance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JINZHUAN INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-10
AI Technical Summary
Existing alarm information processing methods for distributed systems suffer from problems such as misjudgment and high maintenance costs.
By acquiring alarm data from the distributed system, classifying it according to node identifiers, dividing the alarm data into groups and subgroups, using operation and maintenance knowledge sets and target models to determine the identifiers of cause codes, and generating alarm analysis reports.
It enables refined management of alarm data in distributed systems, improves the accuracy and efficiency of fault location, reduces the false judgment rate, provides professional operation and maintenance guidance, and enhances the efficiency and accuracy of system operation and maintenance.
Smart Images

Figure CN121841937A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to an alarm analysis method, apparatus and product applied to distributed systems. Background Technology
[0002] Distributed system software typically involves multiple component types, each with multiple nodes interconnected via a network. With more objects, the alarm scale also increases. A detected fault might trigger an alarm on the node itself or on a connected object; a fault might trigger a single alarm or multiple related alarms. Faced with a large number of alarms, alarm aggregation and grouping are common practices to reduce the number of alarms and ease the burden of on-site maintenance. A common alarm aggregation method is grouping by fault source, where multiple alarms from the same fault source that meet a certain time frame are grouped together. The principle is that multiple alarms occurring sequentially within a sufficiently small time frame may be caused by the same root cause. However, while this method is easy to implement, it can lead to misjudgments, grouping unrelated alarms together, and it doesn't identify the root cause of the fault. Alternatively, based on actual scenario testing, multiple alarm codes generated by the same fault scenario can be treated as a sequence and built into the system as empirical data. When the system generates an alarm, the actual alarm code sequence is compared with the built-in sequence to determine the grouping. Although this method is highly accurate, as the product evolves and features are added or removed, the number of alarms will also increase or decrease accordingly. Before each release, a comprehensive test and alarm sequence reset are required, resulting in high maintenance costs. Therefore, existing alarm information processing methods for distributed systems suffer from problems of misjudgment and high maintenance costs. Summary of the Invention
[0003] This invention provides an alarm analysis method, apparatus, and product for distributed systems, to solve the problems of misjudgment and high maintenance costs in existing alarm information processing methods for distributed systems.
[0004] According to one aspect of the present invention, an alarm analysis method for distributed systems is provided, comprising:
[0005] Obtain at least one alarm data stored in the distributed system, and classify the at least one alarm data according to the node identifier of each execution node in the distributed system to obtain an alarm data group associated with each execution node, wherein the alarm data group includes at least one alarm data and the alarm data includes a cause code sequence.
[0006] For each execution node, the alarm data in the alarm data group is divided into at least one alarm data subgroup based on the reason codes contained in any two alarm data in the alarm data group; wherein the reason codes in the same alarm data subgroup are related.
[0007] The first target model is input into the indicator data corresponding to at least one indicator in the operation and maintenance knowledge set and at least one alarm data in the alarm data group to determine the first identifier corresponding to the cause code in the cause code sequence of at least one alarm data.
[0008] Based on the first identifier in at least one alarm data of the same node identifier, determine the target alarm group, and input the indicator data, operation and maintenance knowledge, repair suggestions and alarm content template prompt words corresponding to the cause code of the first identifier in the target alarm group into the second target model to obtain the alarm analysis report corresponding to the execution node.
[0009] Optionally, at least one alarm data is classified according to the node identifier of each execution node in the distributed system to obtain an alarm data group associated with each execution node, including: dividing alarm data belonging to the same node identifier into a group according to the node identifier in at least one alarm data to obtain an alarm data group associated with the node identifier of each execution node; wherein, the alarm data in the alarm data group includes a cause code sequence, the cause code sequence is composed of at least one cause code, and the cause code is used to characterize the cause of the alarm.
[0010] Optionally, based on the cause codes contained in any two alarm data in the alarm data group, the alarm data in the alarm data group is divided into at least one alarm data subgroup, including: for the alarm data in the alarm data group, obtaining the cause code sequence in any two alarm data; when the cause code sequence includes the same cause identifier, the two alarm data are grouped together, and the divided alarm data are merged according to the lineage lookup method to obtain the data subgroup to be processed associated with the alarm data group; wherein, the data subgroup to be processed includes alarm data with a lineage relationship; the frequency of occurrence of each cause code in each data subgroup to be processed is counted, and the data subgroup to be processed is stored according to a preset format to obtain at least one alarm data subgroup associated with the alarm data group; wherein, the preset format includes the cause code and the frequency of occurrence of the cause code.
[0011] Optionally, the first identifier corresponding to the cause code in the cause code sequence of at least one alarm data is determined by inputting the indicator data corresponding to at least one indicator in the operation and maintenance knowledge set and at least one alarm data in the alarm data group into the first target model. This includes: determining at least one indicator and indicator data of at least one indicator from the operation and maintenance knowledge set based on each cause code in the alarm data group; inputting the indicator data, the operation and maintenance knowledge of the cause code, and the cause code into the first target model to obtain the first identifier corresponding to each cause code in each alarm data; wherein, the first identifier includes an identifier for correct cause code or an identifier for incorrect cause code.
[0012] Optionally, determining a target alarm group based on a first identifier in at least one alarm data with the same node identifier includes: for at least one alarm data with the same node identifier, aggregating at least one alarm data based on the first identifier being a correct reason code to obtain a target alarm group; wherein the target alarm group includes alarm data corresponding to at least one common reason code.
[0013] Optionally, the indicator data, operation and maintenance knowledge, repair suggestions, and alarm content template prompts corresponding to the cause code of the first identifier in the target alarm group are input into the second target model to obtain an alarm analysis report corresponding to the execution node. This includes: inputting the alarm data, cause code, indicator data corresponding to the cause code, operation and maintenance knowledge corresponding to the cause code, repair suggestions, and alarm content template prompts from the target alarm group into the second target model so that the second target model outputs the corresponding alarm analysis report based on the alarm content template prompts.
[0014] According to another aspect of the present invention, an alarm analysis device for a distributed system is provided, comprising:
[0015] The alarm data group determination module is used to obtain at least one alarm data stored in the distributed system, and classify the at least one alarm data according to the node identifier of each execution node in the distributed system to obtain the alarm data group associated with each execution node. The alarm data group includes at least one alarm data, and the alarm data includes a cause code sequence.
[0016] The alarm data subgroup determination module is used to divide the alarm data in the alarm data group corresponding to each execution node into at least one alarm data subgroup based on the reason codes contained in any two alarm data in the alarm data group; wherein the reason codes in the same alarm data subgroup are related.
[0017] The reason code identifier determination module is used to input the indicator data corresponding to at least one indicator in the operation and maintenance knowledge set and at least one alarm data in the alarm data group into the first target model to determine the first identifier corresponding to the reason code in the reason code sequence of at least one alarm data.
[0018] The alarm analysis report determination module is used to determine the target alarm group based on the first identifier in at least one alarm data of the same node identifier, and input the indicator data corresponding to the cause code of the first identifier in the target alarm group, operation and maintenance knowledge, repair suggestions and alarm content template prompt words into the second target model to obtain the alarm analysis report corresponding to the execution node.
[0019] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0020] At least one processor; and
[0021] A memory that is communicatively connected to at least one processor; wherein,
[0022] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor can execute the alarm analysis method for a distributed system according to any embodiment of the present invention.
[0023] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement an alarm analysis method for a distributed system according to any embodiment of the present invention.
[0024] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement an alarm analysis method for a distributed system according to any embodiment of the present invention.
[0025] The technical solution of this invention involves acquiring at least one alarm data stored in a distributed system and classifying the at least one alarm data according to the node identifier of each execution node in the distributed system to obtain an alarm data group associated with each execution node. Each alarm data group includes at least one alarm data, and each alarm data includes a cause code sequence. For each alarm data group corresponding to an execution node, the alarm data in the alarm data group is divided into at least one alarm data subgroup based on the cause codes contained in any two alarm data within the alarm data group. The cause codes within the same alarm data subgroup are associated. The indicator data corresponding to at least one indicator in the operation and maintenance knowledge set and at least one alarm data in the alarm data group are input into a first target model to determine the first identifier corresponding to the cause code in the cause code sequence of at least one alarm data. Based on the first identifier in at least one alarm data with the same node identifier, a target alarm group is determined. The indicator data corresponding to the cause code of the first identifier in the target alarm group, operation and maintenance knowledge, repair suggestions, and alarm content template prompts are input into a second target model to obtain an alarm analysis report corresponding to the execution node. This solution achieves end-to-end processing of distributed alarm data, from global aggregation to node grouping, cause segmentation, precise judgment, and report generation, based on node identifiers, cause code correlation, and primary identifiers. It enables refined and hierarchical management of distributed system alarm data, effectively avoiding analytical interference caused by data clutter. Furthermore, the sequential application of dual models makes cause code judgment more scientific and report generation more standardized and professional, accurately locating the alarm causes and fault chains of each execution node. The automated processing and report generation significantly improve the efficiency of distributed system alarm analysis and fault handling, and provides maintenance personnel with complete analysis reports containing maintenance basis and repair suggestions. This addresses the problems of misjudgment and high maintenance costs in existing alarm information processing methods for distributed systems, helping to effectively shorten the fault diagnosis and resolution cycle and comprehensively improve the accuracy, efficiency, and implementability of distributed system operation and maintenance management.
[0026] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is a flowchart of an alarm analysis method for distributed systems provided in Embodiment 1 of the present invention;
[0029] Figure 2 This is a flowchart of an alarm analysis method applied to a distributed system provided in Embodiment 2 of the present invention;
[0030] Figure 3 This is a schematic diagram of the structure of an alarm analysis device applied to a distributed system provided in Embodiment 3 of the present invention;
[0031] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the alarm analysis method for distributed systems according to embodiments of the present invention. Detailed Implementation
[0032] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0034] Example 1
[0035] Figure 1 This is a flowchart of an alarm analysis method for a distributed system provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of alarm analysis of a distributed system. The method can be executed by an alarm analysis device for a distributed system, which can be implemented in hardware and / or software and can be configured in electronic devices such as computers and servers. Figure 1 As shown, the method includes:
[0036] S110. Obtain at least one alarm data stored in the distributed system, and classify the at least one alarm data according to the node identifier of each execution node in the distributed system to obtain an alarm data group associated with each execution node, wherein the alarm data group includes at least one alarm data and the alarm data includes a cause code sequence.
[0037] Alarm data can be understood as the data generated and stored when anomalies occur at various execution nodes in a distributed system. It includes cause code sequences that reflect the chain and root cause of node anomalies, serving as the core data carrier characterizing the node's operational failure state. It's understood that a distributed system contains multiple different execution nodes. Taking a distributed database system as an example, a distributed database consists of various node types, including Manager (operation and maintenance management node), CN (Coordinator Node), DN (Data Node), and GTM (Global Transaction Manager). Each node type has multiple nodes, interconnected to form a mesh network. For example, if a DN node experiences an SQL (Structured Query Language) execution anomaly, it can generate an SQL execution anomaly alarm. A corresponding cause code can be set for this alarm, and the node identifier, corresponding alarm information, and cause code can be stored in the node's alarm storage space or transmitted to a centralized alarm information management center for subsequent alarm analysis and processing. The node identifier of an execution node can be understood as a unique identity pre-assigned to each execution node within the distributed system. It serves as a specific basis for distinguishing different execution nodes, enabling precise location and identification of each node. For example, the node identifier can consist of the node name and the node address. An alarm data group can be understood as a dataset formed by aggregating and integrating all alarm data belonging to the same execution node in the distributed system, based on the node identifier. Each alarm data group is associated with a corresponding single execution node and contains at least one alarm data point generated by that node.
[0038] Specifically, based on the unified data acquisition interface of the distributed system, the full alarm data can be retrieved in batches from the storage modules of each execution node through distributed data retrieval or active reporting by each execution node. Then, the execution node identifier corresponding to each alarm data is extracted, which is used as the core classification basis. With the help of hash mapping and group aggregation processing logic, all alarm data are classified and sorted, and alarm data belonging to the same node identifier are integrated into independent alarm data groups. This ensures that each execution node can be associated with its own dedicated alarm data group, and that each group contains at least one alarm data generated by that node. At the same time, the core information of the cause code sequence recorded in each alarm data is completely preserved. For example, an alarm data group includes alarms A, B, C, D, E, and F. The cause code sequence for each alarm is as follows: A cause code sequence {1, 7, 8, 9}; B cause code sequence {2, 7, 10}; C cause code sequence {3, 10, 11}; D cause code sequence {4, 12}; E cause code sequence {5, 12}; F cause code sequence {6, 13}. It should be noted that the numbers in parentheses represent cause codes, and the representation of cause codes can be set as needed, which is not limited here.
[0039] In this embodiment, at least one alarm data is classified according to the node identifier of each execution node in the distributed system. This effectively integrates the scattered alarm data in the distributed system, achieving precise binding between alarm data and execution nodes. This not only allows for rapid location of abnormal alarms on each node, significantly improving the efficiency of anomaly investigation in the distributed system, but also provides complete data support for root cause tracing and fault link analysis of node anomalies by leveraging the cause code sequences retained in the alarm data. Furthermore, it adapts to the dynamic changes of distributed nodes, exhibiting good adaptability and scalability. It can also reduce the overall processing pressure of massive alarm data by grouping data by node, making alarm data management in the distributed system more targeted and efficient.
[0040] Optionally, at least one alarm data is classified according to the node identifier of each execution node in the distributed system to obtain an alarm data group associated with each execution node, including: dividing alarm data belonging to the same node identifier into a group according to the node identifier in at least one alarm data to obtain an alarm data group associated with the node identifier of each execution node.
[0041] The alarm data in the alarm data group includes a cause code sequence, which consists of at least one cause code. The cause code is used to identify the cause of the alarm.
[0042] Specifically, based on the node identifier of each execution node in the distributed system, the node identifier information corresponding to at least one alarm data is extracted. According to the principle of grouping the same node identifier into a group, all alarm data are classified and divided to directly form alarm data groups that correspond one-to-one with each execution node. Moreover, the alarm data in each alarm data group carries a cause code sequence consisting of at least one alarm cause identifier, which can represent the alarm cause corresponding to the alarm.
[0043] In this embodiment, the alarm data is categorized and divided according to the node identifiers, which can quickly complete the node-dimensional aggregation of scattered alarm data in the distributed system. This enables precise binding of alarm data with the corresponding execution nodes, which can not only intuitively present the full picture of alarms of each node, greatly improving the pertinence and efficiency of anomaly investigation in the distributed system, but also clearly trace the specific causes of alarms of each node based on the complete cause code sequence in the alarm data. This provides accurate data support for fault root cause location and problem analysis and review, making the alarm management of the distributed system more systematic.
[0044] S120. For each execution node, the alarm data in the alarm data group is divided into at least one alarm data subgroup based on the reason codes contained in any two alarm data in the alarm data group; wherein the reason codes in the same alarm data subgroup are related.
[0045] It's important to note that alarms are notifications triggered by various system anomalies, and the alarm content provides a clear description of the anomaly. An alarm's fault symptom can be caused by one or more reasons. For example, an alarm indicating that the standby machine's replication latency exceeds a threshold describes an anomaly where the standby machine's replication latency exceeds the threshold. The corresponding reasons could include one or more of the following: large transactions, tables without primary key keys, executing DDL statements, and standby machine lock waits. In other words, there is a one-to-many relationship between the alarm's symptom and its cause; different alarm symptoms may originate from the same fault cause. Based on this, various fault causes can be globally encoded, called cause codes. An alarm may have multiple cause codes. Alarms with the same cause code are likely caused by the same root cause, while alarms without the same cause code are unlikely to be caused by the same root cause. Therefore, alarm data can be further classified based on the correlation of cause codes to determine the root cause of the corresponding alarm.
[0046] Specifically, an alarm data subgroup can be understood as a dataset formed by further subdividing the alarm data group corresponding to the execution node based on the correlation of the cause codes contained in the alarm data within the group. All alarm data gathered in the same alarm data subgroup have correlated cause codes, which can reflect a series of alarm situations on the execution node caused by similar related causes.
[0047] Specifically, for each alarm data group corresponding to each execution node in the distributed system, the cause code contained in each alarm data in the group is extracted one by one. By analyzing the correlation of the cause codes of any two alarm data in the group, the alarm data group of that node is further subdivided. Alarm data with correlation of cause codes are integrated and collected. Finally, at least one alarm data subgroup is divided into each alarm data group, and it is ensured that the cause codes of all alarm data in the same alarm data subgroup are correlated. For example, it can be set as follows: if any two alarm data in the alarm data group have at least one identical cause code, then it is determined that the two alarm data are correlated and can be divided into an alarm data subgroup. For example, for an alarm data group, the alarm data group includes: A cause code sequence {1, 7, 8, 9}; B cause code sequence {2, 7, 10}; C cause code sequence {3, 10, 11}; D cause code sequence {4, 12}; E cause code sequence {5, 12}; F cause code sequence {6, 13}. The cause code sequences of any two alarm data are compared. Alarm data with the same cause code are divided into one alarm data subgroup, and alarm data without the same cause code are divided into a separate group. That is, the division result includes 3 alarm data subgroups g1{A, B, C}, g2{D, E}, and g3{F}.
[0048] In this embodiment, based on the aggregation of alarm data at the node level, the alarm data groups are further divided to achieve a refined breakdown of the fault cause dimension. This allows for the rapid identification of a series of alarms caused by similar related causes under the same execution node, significantly improving the accuracy and efficiency of fault tracing and problem localization within a single node. Simultaneously, the subgroup sorting based on the cause code correlation can clearly identify the associated causes and fault propagation links of node anomalies, providing a clear data framework for systematically investigating deep-seated node problems. It also makes the management hierarchy of alarm data clearer and the dimensions richer, facilitating differentiated handling of alarms of different cause types, and further improving the refinement level of alarm data processing and fault management in distributed systems.
[0049] Optionally, based on the cause codes contained in any two alarm data in the alarm data group, the alarm data in the alarm data group is divided into at least one alarm data subgroup, including: for the alarm data in the alarm data group, obtaining the cause code sequence in any two alarm data; when the cause code sequence includes the same cause identifier, the two alarm data are grouped together, and the divided alarm data are merged according to the lineage lookup method to obtain the data subgroup to be processed associated with the alarm data group; wherein, the data subgroup to be processed includes alarm data with a lineage relationship; the frequency of occurrence of each cause code in each data subgroup to be processed is counted, and the data subgroup to be processed is stored according to a preset format to obtain at least one alarm data subgroup associated with the alarm data group; wherein, the preset format includes the cause code and the frequency of occurrence of the cause code.
[0050] Specifically, for each execution node's corresponding alarm data group, the cause code sequence is first extracted from any two alarm data within the alarm data group. When the cause code sequences of two alarm data contain the same cause identifier, they are grouped into the same group. Then, the lineage lookup method is used to merge the divided groups, integrating alarm data with lineage relationships to form a data subgroup to be processed corresponding to the alarm data group. Subsequently, the frequency of cause code occurrences in each data subgroup to be processed is counted, and the data subgroup to be processed is stored according to a preset format containing the cause code and its corresponding frequency. Finally, at least one alarm data subgroup associated with the alarm data group is generated. For example, when the alarm data group includes: A cause code sequence {1, 7, 8, 9}; B cause code sequence {2, 7, 10}; C cause code sequence {3, 10, 11}; D cause code sequence {4, 12}; E cause code sequence {5, 12}; F cause code sequence {6, 13}, and three alarm data subgroups g1{A, B, C}, g2{D, E}, and g3{F} are obtained, frequency statistics are performed on each cause code in each alarm data subgroup to obtain the division results g1{7:2, 10:2, 1:1, 2:1, 3:1, 8:1, 9:1, 11:1}, g2{12:2, 4:1, 5:1}, and g3{6:1, 13:1}, where the values are represented as follows: the part before the colon is the cause code, and the part after the colon is the frequency of occurrence of the cause code.
[0051] In this embodiment, by matching identical identifiers and merging lineage relationships in the cause code sequence, alarm data with causal and link associations under the execution node can be accurately aggregated. At the same time, by combining cause code frequency statistics and standardized storage, the occurrence patterns and impact of various fault causes within the node can be clearly presented, greatly improving the accuracy and efficiency of fault root cause location and abnormal link tracing. It also makes the organization of alarm data more standardized and systematic, which is conducive to subsequent data mining and in-depth analysis. Furthermore, through refined subgroup decomposition and quantitative statistics, it provides detailed and intuitive data support for fault classification and targeted optimization of execution nodes in distributed systems, further improving the scientific nature of system alarm governance and operation and maintenance management.
[0052] S130. Input the indicator data corresponding to at least one indicator in the operation and maintenance knowledge set and at least one alarm data in the alarm data group into the first target model to determine the first identifier corresponding to the cause code in the cause code sequence of at least one alarm data.
[0053] Specifically, the operation and maintenance knowledge set can be understood as a structured knowledge base that integrates professional information related to the entire process of distributed system operation and maintenance. It covers the core indicator definitions of each execution node of the system, the normal threshold range of the indicators, the specific fault causes corresponding to different reason codes, the fault propagation links, the association mapping rules between indicators and reason codes, as well as professional content such as operation and maintenance handling plans and repair strategies for various faults. It can provide comprehensive and authoritative theoretical and practical basis for the analysis of early warning data and alarm data, the determination of the validity of reason codes, the matching of indicator data, the location of the root cause of the fault, and the generation of analysis reports, so as to ensure the professionalism and accuracy of distributed system operation and maintenance data analysis. The first target model can be understood as a large model specifically designed for integrating and analyzing indicator data from the operation and maintenance knowledge set with alarm data from the alarm data group. It can rely on built-in analysis logic to complete feature matching, correlation calculation, and result determination of the two types of data, achieving accurate output of the identifier corresponding to each cause code in the cause code sequence of the alarm data. Specifically, it is used to determine whether the cause code in the alarm data is the cause code that caused the alarm. The initial large model can be pre-trained based on a preset number of sample alarm data to obtain a model that can accurately determine the identifier of the cause code. The identifier of the cause code can be data used to characterize whether the cause code is the root cause of the alarm data to which it belongs. In this embodiment, a first identifier is used to represent the identifier of the cause code. After calculation by the first target model, a corresponding exclusive identifier, namely the first identifier, can be determined for each cause code in the cause code sequence of the alarm data. This first identifier can be used to accurately define and classify the alarm cause represented by the cause code, clearly reflecting the specific alarm connotation corresponding to the cause code.
[0054] Specifically, the system first integrates the indicator data corresponding to at least one metric from the operation and maintenance knowledge set, and at the same time retrieves at least one alarm data from the alarm data group. The two types of data are then input into the preset first target model. Based on the algorithm logic of the model, the data fusion analysis and feature matching are completed, and finally the first identifier corresponding to each reason code in the reason code sequence of the alarm data is accurately determined. For example, when the alarm data group includes: A reason code sequence {1, 7, 8, 9}; B reason code sequence {2, 7, 10}; C reason code sequence {3, 10, 11}; D reason code sequence {4, 12}; E reason code sequence {5, 12}; F reason code sequence {6, 13}, and three alarm data subgroups g1{A, B, C}, g2{D, E}, and g3{F} are obtained, the first identifier corresponding to each reason code is bound to the corresponding reason code, resulting in the following: A reason code sequence {1: No, 7: Yes, 8: No, 9: No}; B reason code sequence {2: No, 7: Yes, 10: No}; C reason code sequence {3: Yes, 10: No, 11: No}; D reason code sequence {4: No, 12: Yes}; E reason code sequence {5: No, 12: Yes}; F reason code sequence {6: Yes, 13: No}.
[0055] In this embodiment, by integrating indicator data from the operation and maintenance knowledge system with actual alarm data to conduct model analysis, the identification of the cause code can be more scientific and accurate with the professional support of operation and maintenance knowledge. This effectively avoids the judgment bias caused by single data dimension analysis. At the same time, the model-based analysis method greatly improves the determination efficiency of the first identifier, adapts to the processing needs of massive alarm data in distributed systems, and the determined first identifier can accurately interpret and classify the cause code, providing a clear judgment basis for subsequent fault root cause location and alarm handling strategy generation, further enhancing the professionalism and practicality of distributed system alarm analysis.
[0056] S140. Based on the first identifier in at least one alarm data of the same node identifier, determine the target alarm group, and input the indicator data, operation and maintenance knowledge, repair suggestions and alarm content template prompt words corresponding to the cause code of the first identifier in the target alarm group into the second target model to obtain the alarm analysis report corresponding to the execution node.
[0057] The alarm content template prompts can be understood as standardized guiding text that defines the output format of alarm analysis reports and clarifies the core presentation dimensions of the reports. It specifies the core content direction and layout requirements that the report should cover, such as alarm causes, abnormal indicators, and remediation suggestions, providing a unified content framework for report generation. Alarm content template prompts can be pre-set for generating alarm analysis reports and directly invoked later. The second target model can be understood as an algorithmic model that integrates indicator data, operational knowledge, remediation suggestions, and alarm content template prompts corresponding to the first identifier within the target alarm group, completing multi-dimensional information integration, analysis, and content generation. It can automatically output complete text-based analysis results that conform to the specifications based on the input information. The alarm analysis report can be understood as a dedicated analysis document generated by the second target model and matched to the corresponding execution node of the distributed system. It fully integrates the core causes of node alarms, abnormal indicator characteristics, professional operational basis, and specific fault remediation suggestions, presenting a clear and systematic picture of node alarm issues and providing direct and professional reference for operational personnel to handle faults.
[0058] Specifically, based on the first identifier identified in all alarm data under the same node identifier, clustering and integration are performed to define the corresponding target alarm group. Then, the indicator data corresponding to the cause code associated with the first identifier in the target alarm group, the relevant operation and maintenance knowledge in the operation and maintenance knowledge set, the corresponding fault repair suggestions, and the alarm content template prompt words are extracted. All of the above data and information are uniformly input into the second target model, which performs multi-dimensional data fusion analysis, content integration, and report generation, and finally produces an alarm analysis report that accurately corresponds to the execution node.
[0059] For example, the alarm code content, cause code, corresponding indicator data, corresponding operation and maintenance knowledge, corresponding repair suggestions, and alarm content template prompts of the target alarm group are input into the large model. The large model then outputs an alarm analysis report based on the prompts. The input information format can be set as follows: Fault object: {'Fault object':'?'}, Abnormal monitoring indicator: [{'Monitoring indicator':'?','Indicator description':'?','Judgment result':'?','Diagnostic formula':'?','Indicator value':'?'}], Cause code information: ['Cause code':'?','Monitoring indicator':'?', The input data is structured as follows: 'Diagnostic Formula': '?', 'Result Description': '?', 'Detailed Cause Description': '?', 'Repair Suggestions': '?'}], Alarm Group 1: [{'Cause Code': '?', [{'Alarm Level': '?', 'Alarm Code': '?', 'Alarm Name': '?', 'Alarm Time': '?', 'Alarm Content': '?'}]}]. Note that question marks in the input information can be used as placeholders. The required information is obtained from the target alarm group and the operations and maintenance knowledge base. The number of inputs corresponds to the number of target alarm groups. Additionally, the prompt can be set to: "Analyze the above input content and output an analysis report according to the report template below." Alarm Analysis Report Template: Alarm root cause diagnosis mainly analyzes abnormal monitoring indicators. Alarms with the same cause code are grouped together, and the severity of each group is identified by the most severe alarm level in that group. Alarm Level Severity: Urgent > Important > Minor > Warning. Groups are sorted by severity. 1. For the pre-assigned alarm groups, the title displays: Fault Object: 'Fault Object' Alarm Group 'Sequence Number'. A table listing the alarms for this group is displayed, showing: Sequence Number: Alarm Level, Alarm Code, Alarm Name, Alarm Time, Alarm Content. 2. Next, the cause name and cause code are displayed. 3. The evidence chain is then displayed, showing: Diagnostic Formula, Indicator Name: Indicator Value. 4. A detailed description of the cause code corresponding to the abnormal indicator is displayed. 5. Repair suggestions for the cause code corresponding to the abnormal indicator are displayed. Further, the input data and prompts obtained above are input into the second target model. After processing by the second target model, an alarm analysis report is output.
[0060] In this embodiment, the node alarm data is accurately grouped based on the first identifier. Combined with multi-dimensional core data and preset prompts, a model-based report is generated. This ensures that the alarm analysis report closely matches the actual alarm and fault conditions of the execution node, accurately presenting the core causes, abnormal indicator characteristics, and handling directions of the node alarm, significantly improving the professionalism and completeness of the alarm analysis. Furthermore, by automatically generating reports using the model, the tedious process of manual analysis is eliminated, greatly improving the output efficiency of the alarm analysis report and adapting to the batch analysis needs of multiple nodes in a distributed system. At the same time, the operation and maintenance knowledge and repair suggestions integrated in the report can directly provide standardized and targeted fault handling guidance for operation and maintenance personnel, effectively shortening the troubleshooting and resolution cycle of node faults, and further enhancing the efficiency and practicality of distributed system operation and maintenance management.
[0061] Optionally, determining a target alarm group based on a first identifier in at least one alarm data with the same node identifier includes: for at least one alarm data with the same node identifier, aggregating at least one alarm data based on the first identifier being a correct reason code to obtain a target alarm group; wherein the target alarm group includes alarm data corresponding to at least one common reason code.
[0062] Specifically, for all alarm data belonging to the same node identifier, the cause code determined by the first identifier is selected. Based on this, the alarm data under the node is classified and integrated. Alarm data containing the same common cause code are grouped into the same group, and finally the target alarm group of the corresponding node is formed. Each group contains at least one alarm data corresponding to a common cause code. For example, by binding the first identifier corresponding to each reason code to the corresponding reason code, the following results are obtained: Reason code sequence A {1: No, 7: Yes, 8: No, 9: No}; Reason code sequence B {2: No, 7: Yes, 10: No}; Reason code sequence C {3: Yes, 10: No, 11: No}; Reason code sequence D {4: No, 12: Yes}; Reason code sequence E {5: No, 12: Yes}; Reason code sequence F {6: Yes, 13: No}. Aggregating according to the first identifier being the Yes identifier, the target alarm group is obtained as: Newg1{A,B}{7: Yes}; Newg2{C}{3: Yes}; Newg3{D,E}{12: Yes}; Newg4{F}{6: Yes}.
[0063] In this embodiment, at least one alarm data is aggregated based on the first correct cause code to obtain a target alarm group. This can accurately filter and aggregate valid and correctly judged alarm data under a node. The target grouping is completed based on the common cause code, which can quickly focus on the core alarm causes of the node, eliminate interference from invalid judgment data, and make the grouping results of node alarms more accurate and targeted. At the same time, it can clearly present the full picture of alarms caused by the same core causes under the same node, which can significantly improve the efficiency of root cause location of node failures. It can also make subsequent alarm analysis based on grouping more focused on core issues, provide an accurate and high-quality data foundation for alarm analysis report generation, and further ensure the scientific nature of node alarm handling and operation and maintenance decisions in distributed systems.
[0064] Optionally, the indicator data, operation and maintenance knowledge, repair suggestions, and alarm content template prompts corresponding to the cause code of the first identifier in the target alarm group are input into the second target model to obtain an alarm analysis report corresponding to the execution node. This includes: inputting the alarm data, cause code, indicator data corresponding to the cause code, operation and maintenance knowledge corresponding to the cause code, repair suggestions, and alarm content template prompts from the target alarm group into the second target model so that the second target model outputs the corresponding alarm analysis report based on the alarm content template prompts.
[0065] Specifically, the alarm data within the target alarm group, the cause codes corresponding to each alarm, the indicator data matching the cause codes, the professional operation and maintenance knowledge and fault repair suggestions adapted to the cause codes, and the standardized alarm content template prompts are all integrated and input into the second target model. The second target model then integrates, sorts, analyzes, and generates content based on the framework and presentation requirements specified by the alarm content template prompts, and finally automatically outputs an alarm analysis report that is accurately matched with the corresponding execution node.
[0066] In this embodiment, the implementation method, through the integrated model input of multi-dimensional core data and the standardized guidance of template prompts, ensures that the alarm analysis report fully covers the core content of node alarms, such as the cause of the fault, abnormal indicators, operation and maintenance basis, and repair plan. This makes the report content more professional, complete, and standardized. At the same time, it can rely on the model to automatically generate reports, saving the tedious process of manual sorting and analysis, greatly improving the efficiency of report output, and adapting to the actual needs of batch analysis of multiple nodes in distributed systems. Meanwhile, the report is accurately bound to the execution node, and the content is closely related to the actual alarm fault situation of the node. It can directly provide clear and specific fault handling guidance for operation and maintenance personnel, effectively shorten the node fault investigation and resolution cycle, and further improve the efficiency and implementation of distributed system operation and maintenance management.
[0067] The technical solution of this embodiment involves acquiring at least one alarm data stored in a distributed system and classifying the at least one alarm data according to the node identifier of each execution node in the distributed system to obtain an alarm data group associated with each execution node. Each alarm data group includes at least one alarm data, and each alarm data includes a cause code sequence. For each alarm data group corresponding to an execution node, the alarm data in the alarm data group is divided into at least one alarm data subgroup based on the cause codes contained in any two alarm data within the alarm data group. The cause codes within the same alarm data subgroup are associated. The indicator data corresponding to at least one indicator in the operation and maintenance knowledge set and at least one alarm data in the alarm data group are input into a first target model to determine the first identifier corresponding to the cause code in the cause code sequence of at least one alarm data. Based on the first identifier in at least one alarm data with the same node identifier, a target alarm group is determined. The indicator data corresponding to the cause code of the first identifier in the target alarm group, operation and maintenance knowledge, repair suggestions, and alarm content template prompts are input into a second target model to obtain an alarm analysis report corresponding to the execution node. This solution achieves end-to-end processing of distributed alarm data, from global aggregation to node grouping, cause segmentation, precise judgment, and report generation, based on node identifiers, cause code correlation, and primary identifiers. It enables refined and hierarchical management of distributed system alarm data, effectively avoiding analytical interference caused by data clutter. Furthermore, the sequential application of dual models makes cause code judgment more scientific and report generation more standardized and professional, accurately locating the alarm causes and fault chains of each execution node. The automated processing and report generation significantly improve the efficiency of distributed system alarm analysis and fault handling, and provides maintenance personnel with complete analysis reports containing maintenance basis and repair suggestions. This addresses the problems of misjudgment and high maintenance costs in existing alarm information processing methods for distributed systems, helping to effectively shorten the fault diagnosis and resolution cycle and comprehensively improve the accuracy, efficiency, and implementability of distributed system operation and maintenance management.
[0068] Example 2
[0069] Figure 2 This is a flowchart of an alarm analysis method applied to a distributed system provided in Embodiment 2 of the present invention. The method in this embodiment is a further optimization of the method in the above embodiments. Optionally, at least one indicator and indicator data of at least one indicator are determined from the operation and maintenance knowledge set based on each reason code in the alarm data group; the indicator data, the operation and maintenance knowledge of the reason code, and the reason code are jointly input into a first target model to obtain a first identifier corresponding to each reason code in each alarm data; wherein, the first identifier includes an identifier indicating that the reason code is correct or an identifier indicating that the reason code is incorrect. Figure 2 As shown, the method includes:
[0070] S210. Obtain at least one alarm data stored in the distributed system, and classify the at least one alarm data according to the node identifier of each execution node in the distributed system to obtain an alarm data group associated with each execution node, wherein the alarm data group includes at least one alarm data and the alarm data includes a cause code sequence.
[0071] S220. For each execution node, the alarm data in the alarm data group is divided into at least one alarm data subgroup based on the reason codes contained in any two alarm data in the alarm data group; wherein the reason codes in the same alarm data subgroup are related.
[0072] S230. Determine at least one metric from the operation and maintenance knowledge set based on each reason code in the alarm data group, as well as the metric data of at least one metric.
[0073] It should be noted that the reason code structure definition in the operations and maintenance knowledge set includes the following information: reason code, reason name, detailed reason description, remediation suggestion, and monitoring metric. A reason code can be associated with at least one related monitoring metric, which can be obtained by matching from the operations and maintenance knowledge base based on the reason code. For example, if the alarm is a standby machine replication latency exceeding the threshold alarm, the corresponding reason code sequence is {[10004, large transaction],[10005, no primary key table],[10006, DDL],[10007, standby machine lock wait]}. The metric corresponding to reason code 1004 could be big_trans_count, the metric corresponding to reason code 1005 could be replaying_trx_without_unique_key, the metric corresponding to reason code 1006 could be replaying_ddl_statements, and the metric corresponding to reason code 1007 could be replaying_lock_wait. After determining the metric corresponding to each reason code, matching can be performed from the operations and maintenance knowledge base or a preset storage space to obtain the corresponding metric data.
[0074] Specifically, for each alarm data group corresponding to each execution node, all reason codes contained in each alarm data in the group are first extracted completely. Then, based on each extracted reason code, a precise matching and retrieval is performed in a preset operation and maintenance knowledge set. According to the association mapping relationship between reason codes and indicators in the operation and maintenance knowledge set, at least one associated indicator corresponding to each reason code is determined. Real-time or historical indicator data corresponding to these indicators are synchronously retrieved through the preset API interface corresponding to the reason code. The retrieved real-time or historical data is the data closest to the timestamp information of the corresponding alarm, or real-time or historical data whose time difference with the timestamp information of the corresponding alarm meets a preset threshold.
[0075] In this embodiment, a precise association between cause codes and indicators / indicator data is established based on the operation and maintenance knowledge set. This provides corresponding quantitative indicators to support the analysis of alarm data, enabling the determination of corresponding indicator data based on abstract cause codes. This effectively strengthens the data foundation for subsequent alarm cause determination and fault root cause location. Furthermore, the retrieval and retrieval process adapts to the operating characteristics of distributed systems, allowing for rapid batch matching and data acquisition of multiple cause codes and indicators, significantly improving indicator association efficiency. The mapping relationship of the operation and maintenance knowledge set can be flexibly updated to adapt to the dynamic adjustment of system indicators and cause codes, exhibiting good scalability. It also allows for precise locking of associated indicators through cause codes, avoiding redundant retrieval of invalid indicator data, making alarm analysis data processing more targeted, and further improving the accuracy and efficiency of distributed system alarm judgment.
[0076] S240. Input the indicator data, the cause code operation and maintenance knowledge, and the cause code into the first target model to obtain the first identifier corresponding to each cause code in each alarm data; wherein, the first identifier includes an identifier for correct cause code or an identifier for incorrect cause code.
[0077] Specifically, the extracted indicator data, the operation and maintenance knowledge matching the cause codes, and the cause codes are first integrated. The three types of core information are then uniformly input into the first target model. The model relies on its built-in algorithm to complete the fusion verification, feature matching, and logical judgment of multi-dimensional data. Finally, it outputs the first identifier corresponding to each cause code in each alarm data. This identifier clearly marks whether the corresponding cause code is correct or incorrect.
[0078] In this embodiment, the model-based verification of cause code validity is completed through multi-dimensional data linkage. This relies on the dual support of operational knowledge and indicator data, making the determination of cause code correctness more objective and accurate. It effectively eliminates the interference of erroneous cause codes on subsequent analysis, ensuring the accuracy of alarm data analysis. Simultaneously, the model automates the rapid determination of batch cause codes, significantly improving determination efficiency and adapting to the processing needs of massive alarm data in distributed systems. Furthermore, the standardized first identifier output after determination can directly serve as the core basis for subsequent target alarm grouping, making the hierarchical aggregation of alarm data more clearly directional. This further solidifies the foundation for alarm analysis and fault location in distributed systems, enhancing the professionalism and efficiency of overall operational analysis work.
[0079] Optionally, the cause code structure definition in the operation and maintenance knowledge set also includes associated indicator diagnostic formulas. These formulas are diagnostic rules that link indicators, operators, thresholds, and results. They can be combined with actually collected indicator data to determine whether an indicator is correct. For example, the indicator corresponding to cause code 1004 could be `big_trans_count`, where `big_trans_count_inc` represents the increment of large transaction statistics within a certain sampling period. The diagnostic formula is: `big_trans_count_inc>0`. If the result is True, it indicates that there is a large transaction alarm, meaning the cause code is correctly identified. If the result is False, it indicates that there is no large transaction alarm, meaning the cause code is incorrectly identified. When the cause codes and corresponding data to be determined exceed a preset data volume, the associated diagnostic formulas can also be used as input data for the first target model. The first target model is then used to determine the first identifier corresponding to each cause code. When the cause codes and corresponding data to be determined are less than a preset data volume, the first identifier corresponding to each cause code can be determined directly based on the diagnostic formulas.
[0080] S250. Based on the first identifier in at least one alarm data of the same node identifier, determine the target alarm group, and input the indicator data, operation and maintenance knowledge, repair suggestions and alarm content template prompt words corresponding to the cause code of the first identifier in the target alarm group into the second target model to obtain the alarm analysis report corresponding to the execution node.
[0081] The technical solution of this embodiment extracts at least one alarm data from the distributed system, classifies the alarm data according to the node identifier of each execution node, and generates alarm data groups corresponding to each execution node. Each group contains at least one alarm data with a cause code sequence. Then, for each alarm data group, further subdivision is performed based on the correlation of cause codes of any two alarm data within the group, dividing at least one alarm data subgroup with related cause codes. Next, based on each cause code in the alarm data group, at least one corresponding indicator is matched and determined in the operation and maintenance knowledge set, and the indicator data of these indicators are obtained synchronously. Then, the indicator data, the operation and maintenance knowledge corresponding to the cause code, and the cause code itself are input into the first target model to complete the validity judgment of each cause code in each alarm data, and output a first identifier indicating whether the cause code is correct or incorrect. Then, the first identifier of the alarm data under the same node identifier is combined to complete data aggregation, delineate the corresponding target alarm group, and then input the indicator data, operation and maintenance knowledge, repair suggestions, and alarm content template prompt words matching the first identifier in the group into the second target model, finally generating an alarm analysis report exclusive to each execution node. This solution establishes a fully automated processing system for distributed alarm data, encompassing node aggregation, cause segmentation, indicator matching, cause code verification, and report generation. This system enables hierarchical and refined analysis of dispersed alarm data. Leveraging the collaborative application of operational knowledge sets and a dual-objective model, it ensures the objectivity and accuracy of cause code determination, effectively eliminating erroneous data interference. Furthermore, it makes alarm analysis reports more standardized and professional, accurately presenting the causes, indicator characteristics, and handling suggestions for alarm faults at each node. This significantly improves the efficiency of alarm analysis and troubleshooting in distributed systems. Adaptable to the dynamic changes in distributed nodes and the massive dispersion of data, it possesses excellent scalability and adaptability. It also provides operational personnel with detailed and practical fault handling guidelines, effectively shortening the fault resolution cycle and comprehensively improving the accuracy and practicality of distributed system operation and maintenance management.
[0082] Example 3
[0083] Figure 3 This is a schematic diagram of an alarm analysis device applied to a distributed system, provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes:
[0084] The alarm data group determination module 310 is used to obtain at least one alarm data stored in the distributed system, and classify the at least one alarm data according to the node identifier of each execution node in the distributed system to obtain the alarm data group associated with each execution node. The alarm data group includes at least one alarm data, and the alarm data includes a cause code sequence.
[0085] The alarm data subgroup determination module 320 is used to divide the alarm data in the alarm data group into at least one alarm data subgroup based on the reason codes contained in any two alarm data in the alarm data group corresponding to each execution node; wherein the reason codes in the same alarm data subgroup are related.
[0086] The reason code identifier determination module 330 is used to input the indicator data corresponding to at least one indicator in the operation and maintenance knowledge set and at least one alarm data in the alarm data group into the first target model to determine the first identifier corresponding to the reason code in the reason code sequence of at least one alarm data.
[0087] The alarm analysis report determination module 340 is used to determine the target alarm group based on the first identifier in at least one alarm data of the same node identifier, and input the indicator data corresponding to the cause code of the first identifier in the target alarm group, operation and maintenance knowledge, repair suggestions and alarm content template prompt words into the second target model to obtain the alarm analysis report corresponding to the execution node.
[0088] The technical solution of this embodiment involves obtaining at least one alarm data stored in the distributed system through an alarm data group determination module, and classifying the at least one alarm data according to the node identifier of each execution node in the distributed system to obtain an alarm data group associated with each execution node. Each alarm data group includes at least one alarm data, and each alarm data includes a cause code sequence. For each execution node's corresponding alarm data group, the alarm data subgroup determination module divides the alarm data in the alarm data group into at least one alarm data subgroup based on the cause codes contained in any two alarm data within the alarm data group. The same alarm data subgroup... The reason codes in the data are related; the reason code identification module inputs the indicator data corresponding to at least one indicator in the operation and maintenance knowledge set and at least one alarm data in the alarm data group into the first target model to determine the first identifier corresponding to the reason code in the reason code sequence of at least one alarm data; the alarm analysis report determination module determines the target alarm group based on the first identifier in at least one alarm data with the same node identifier, and inputs the indicator data, operation and maintenance knowledge, repair suggestions and alarm content template prompt words corresponding to the reason code of the first identifier in the target alarm group into the second target model to obtain the alarm analysis report corresponding to the execution node. This solution enables end-to-end processing of distributed alarm data, from global aggregation to node grouping, cause segmentation, precise judgment, and report generation, based on node identifiers, cause code correlation, and primary identifiers. It achieves refined and hierarchical management of distributed system alarm data, effectively avoiding analytical interference caused by data clutter. Furthermore, the sequential application of dual models makes cause code judgment more scientific and report generation more standardized and professional, accurately locating the alarm causes and fault chains of each execution node. The automated processing and report generation significantly improve the efficiency of distributed system alarm analysis and fault handling, and provides maintenance personnel with complete analysis reports containing maintenance basis and repair suggestions, effectively shortening the fault investigation and resolution cycle and comprehensively improving the accuracy, efficiency, and implementability of distributed system operation and maintenance management.
[0089] Based on the above embodiments, optionally, the alarm data group determination module 310 is specifically used to divide alarm data belonging to the same node identifier into a group according to the node identifier in at least one alarm data, to obtain an alarm data group associated with the node identifier of each execution node; wherein, the alarm data in the alarm data group includes a cause code sequence, the cause code sequence is composed of at least one cause code, and the cause code is used to characterize the cause identifier of the alarm.
[0090] Optionally, the alarm data subgroup determination module 320 is specifically used to obtain the cause code sequence from any two alarm data in the alarm data group; when the cause code sequence includes the same cause identifier, the two alarm data are divided into a group, and the divided alarm data are merged according to the lineage lookup method to obtain the data subgroup to be processed associated with the alarm data group; wherein, the data subgroup to be processed includes alarm data with a lineage relationship; the frequency of each cause code in each data subgroup to be processed is counted, and the data subgroup to be processed is stored according to a preset format to obtain at least one alarm data subgroup associated with the alarm data group; wherein, the preset format includes the cause code and the frequency of the cause code.
[0091] Optionally, the reason code identifier determination module 330 is specifically used to determine at least one indicator from the operation and maintenance knowledge set based on each reason code in the alarm data group, as well as the indicator data of at least one indicator; input the indicator data, the reason code operation and maintenance knowledge, and the reason code together into the first target model to obtain the first identifier corresponding to each reason code in each alarm data; wherein, the first identifier includes an identifier for correct reason code or an identifier for incorrect reason code.
[0092] Optionally, the alarm analysis report determination module 340 is specifically used to aggregate at least one alarm data with the same node identifier according to the cause code that the first identifier is correct, to obtain a target alarm group; wherein, the target alarm group includes alarm data corresponding to at least one common cause code.
[0093] Optionally, the alarm analysis report determination module 340 further inputs the alarm data, cause code, indicator data corresponding to the cause code, operation and maintenance knowledge corresponding to the cause code, repair suggestions, and alarm content template prompts from the target alarm group into the second target model, so that the second target model outputs the corresponding alarm analysis report based on the alarm content template prompts.
[0094] The alarm analysis device for distributed systems provided in the embodiments of the present invention can execute the alarm analysis method for distributed systems provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0095] Example 4
[0096] Figure 4This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0097] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0098] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0099] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as alarm analysis methods applied to distributed systems.
[0100] In some embodiments, the alarm analysis method for a distributed system can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the alarm analysis method for a distributed system described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the alarm analysis method for a distributed system by any other suitable means (e.g., by means of firmware).
[0101] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0102] Computer programs for implementing the alarm analysis method for distributed systems of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0103] Example 5
[0104] Embodiment 5 of the present invention also provides a computer-readable storage medium storing computer instructions for causing a processor to execute an alarm analysis method applied to a distributed system, the method comprising:
[0105] Obtain at least one alarm data stored in the distributed system, and classify the at least one alarm data according to the node identifier of each execution node in the distributed system to obtain an alarm data group associated with each execution node, wherein the alarm data group includes at least one alarm data and the alarm data includes a cause code sequence.
[0106] For each execution node, the alarm data in the alarm data group is divided into at least one alarm data subgroup based on the reason codes contained in any two alarm data in the alarm data group; wherein the reason codes in the same alarm data subgroup are related.
[0107] The first target model is input into the indicator data corresponding to at least one indicator in the operation and maintenance knowledge set and at least one alarm data in the alarm data group to determine the first identifier corresponding to the cause code in the cause code sequence of at least one alarm data.
[0108] Based on the first identifier in at least one alarm data of the same node identifier, determine the target alarm group, and input the indicator data, operation and maintenance knowledge, repair suggestions and alarm content template prompt words corresponding to the cause code of the first identifier in the target alarm group into the second target model to obtain the alarm analysis report corresponding to the execution node.
[0109] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0110] To provide interaction with an object, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the object; and a keyboard and pointing device (e.g., a mouse or trackball) through which the object provides input to the electronic device. Other types of devices can also be used to provide interaction with the object; for example, feedback provided to the object can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the object can be received in any form (including sound input, voice input, or tactile input).
[0111] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., a computer with a graphical user interface or web browser through which an item can interact with the implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0112] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0113] Example 6
[0114] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the alarm analysis method for distributed systems according to any embodiment of this invention.
[0115] In the implementation of a computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages as well as conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0116] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0117] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. An alarm analysis method applied to a distributed system, characterized by, The method comprises: obtaining at least one alarm data stored in a distributed system, and classifying the at least one alarm data according to a node identifier of each execution node in the distributed system to obtain an alarm data group associated with each execution node, wherein the alarm data group comprises at least one alarm data, and the alarm data comprises a reason code sequence; for the alarm data group corresponding to each execution node, dividing the alarm data in the alarm data group into at least one alarm data subgroup according to the reason codes contained in any two alarm data in the alarm data group; wherein the reason codes in the same alarm data subgroup are associated; inputting the index data corresponding to at least one index in a set of operation and maintenance knowledge and at least one alarm data in the alarm data group into a first target model to determine a first identifier corresponding to a reason code in the reason code sequence of the at least one alarm data; determining a target alarm group according to the first identifiers in at least one alarm data of the same node identifier, and inputting the reason code corresponding to the first identifier in the target alarm group, the index data, the operation and maintenance knowledge, the repair suggestion and the alarm content template prompt word into a second target model to obtain an alarm analysis report corresponding to the execution node.
2. The method of claim 1, wherein, The method comprises: dividing the alarm data belonging to the same node identifier into a group according to the node identifier in the at least one alarm data to obtain the alarm data group associated with the node identifier of each execution node; wherein the alarm data in the alarm data group comprises a reason code sequence, the reason code sequence is composed of at least one reason code, and the reason code is used to represent the reason identifier of the alarm reason.
3. The method of claim 1, wherein, The method comprises: for the alarm data in the alarm data group, obtaining the reason code sequence in any two alarm data; when the same reason identifier is included in the reason code sequence, dividing the two alarm data into a group, and merging the divided alarm data according to a blood relationship searching method to obtain a to-be-processed data subgroup associated with the alarm data group; wherein the to-be-processed data subgroup comprises alarm data having a blood relationship; counting the frequency of occurrence of each reason code in each to-be-processed data subgroup, and storing the to-be-processed data subgroup according to a preset format to obtain at least one alarm data subgroup associated with the alarm data group; wherein the preset format comprises the reason code and the frequency of occurrence of the reason code.
4. The method of claim 1, wherein, The method comprises: determine at least one indicator from the operation and maintenance knowledge set according to each reason code in the alarm data set, and indicator data of the at least one indicator; input the indicator data, operation and maintenance knowledge of the reason code, and the reason code into the first target model to obtain a first identifier corresponding to each reason code in each alarm data; wherein the first identifier includes an identifier of a correct reason code or an identifier of an incorrect reason code.
5. The method of claim 1, wherein, The first identifier in at least one alarm data of the same node identifier is used to determine a target alarm group, including: For at least one alarm data of the same node identifier, the at least one alarm data is aggregated according to the first identifier of the correct reason code to obtain a target alarm group; wherein the target alarm group includes alarm data corresponding to at least one common reason code.
6. The method of claim 1, wherein, The reason code corresponding to the first identifier in the target alarm group is input into the second target model to obtain an alarm analysis report corresponding to the execution node, including: The alarm data, reason code, indicator data corresponding to the reason code, operation and maintenance knowledge corresponding to the reason code, repair suggestion, and alarm content template prompt word in the target alarm group are input into the second target model, so that the second target model outputs a corresponding alarm analysis report according to the alarm content template prompt word.
7. An alarm analysis device for distributed systems, characterized in that, including: An alarm data set determination module is configured to obtain at least one alarm data stored in a distributed system, and classify the at least one alarm data according to a node identifier of each execution node in the distributed system to obtain an alarm data set associated with each execution node, wherein the alarm data set includes at least one alarm data, and the alarm data includes a reason code sequence; An alarm data sub-set determination module is configured to divide the alarm data in the alarm data set into at least one alarm data sub-set according to the reason codes contained in any two alarm data in the alarm data set for each alarm data set corresponding to each execution node; wherein the reason codes in the same alarm data sub-set are associated; A reason code identifier determination module is configured to input at least one alarm data in the alarm data set and indicator data corresponding to at least one indicator in an operation and maintenance knowledge set into a first target model to determine a first identifier corresponding to a reason code in the reason code sequence of the at least one alarm data; An alarm analysis report determination module is configured to determine a target alarm group according to the first identifier in at least one alarm data of the same node identifier, and input the reason code corresponding to the first identifier in the target alarm group, the indicator data, the operation and maintenance knowledge, the repair suggestion, and the alarm content template prompt word into a second target model to obtain an alarm analysis report corresponding to the execution node.
8. An electronic device, comprising: The electronic device includes: at least one processor; and a memory connected in communication with the at least one processor; wherein The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the alarm analysis method applied to a distributed system according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the alarm analysis method applied to a distributed system according to any one of claims 1-6 when executed.
10. A computer program product, characterised in that, The computer program product comprises a computer program which, when executed by a processor, implements the alarm analysis method applied to a distributed system according to any one of claims 1-6.