A data intelligent analysis method and system
By establishing a cross-domain, real-time unified operation and maintenance data model and machine learning rule engine, combined with artificial intelligence algorithms, the problem of low efficiency in fault location and analysis in complex IT environments has been solved, achieving fast and accurate fault location and analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUA XIA BANK
- Filing Date
- 2023-12-25
- Publication Date
- 2026-08-04
AI Technical Summary
In complex and cross-domain IT environments, fault location and analysis rely on human experience, which is time-consuming and inefficient, making it difficult to quickly and accurately locate and analyze faults.
Establish a unified operation and maintenance data model that spans multiple domains and operates in real time. Configure a rule engine through data analysis and machine learning, and combine it with artificial intelligence algorithms to locate faults and analyze their impact, thereby creating fault profiles.
It enables rapid and accurate fault location and analysis in complex cross-domain IT environments, reducing manual intervention time and improving fault location and analysis efficiency.
Smart Images

Figure CN117785530B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data analysis, and in particular to a data intelligent analysis method and system. Background Technology
[0002] With the widespread application of big data and cloud computing technologies in bank data centers, bank IT systems have become extremely complex. They typically deploy a wide variety of basic hardware and software products from various brands and use custom-developed application systems. Simultaneously, with the deepening of internet-based services, bank IT systems are required to provide 24 / 7 uninterrupted service to users. Therefore, in the current environment, IT operations and maintenance in banking enterprises are characterized by extremely high timeliness and complexity. Furthermore, during IT operations and maintenance, anomalies and alarms often occur simultaneously in business operations, applications, systems, networks, storage, and data center environments. This necessitates collaborative analysis by technical personnel from various disciplines to locate the fault. In other words, the entire fault location process relies on human intervention and the expertise of those involved, which often consumes significant time and effort.
[0003] Therefore, how to quickly locate and analyze faults in complex and cross-domain IT environments has become a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0004] To address the aforementioned issues, this application provides a data intelligence analysis method and system that can quickly locate and analyze faults in complex and cross-domain IT environments through data analysis.
[0005] The embodiments of this application disclose the following technical solutions:
[0006] In a first aspect, embodiments of this application provide a data intelligent analysis method, the method comprising:
[0007] Establish a unified operation and maintenance data model that is real-time and cross-domain;
[0008] Using the various types of data and the relationships between them in the unified operation and maintenance data model as parameters, the rules in the fault location rule base are executed to obtain the fault location result; the rules in the fault location rule base are rules obtained by configuring expert experience into the rule engine through data analysis and machine learning.
[0009] Based on the unified operation and maintenance data model, the model describes the architecture of each business, and the impact parameters and rules are configured for the business impact model. Combined with the health score calculated by artificial intelligence algorithms, an impact analysis is formed.
[0010] The fault profile is obtained by using the fault location results, the impact analysis, and fault-related information.
[0011] Optionally, establishing a unified operation and maintenance data model across domains in real time includes:
[0012] Through the operation platform and data acquisition engine, data is monitored and captured from the data source;
[0013] The forwarding function of the data integration engine is used to forward the monitored and captured data to the data storage engine in the form of the source data layer;
[0014] The data is extracted from the source data layer by the data integration engine according to the data extraction rules in the data rule base;
[0015] The extracted data is standardized and transformed based on the data standard library;
[0016] The converted data is then validated and cleaned.
[0017] Based on the data tag library, business tags and technical tags are added to the verified and cleaned data;
[0018] Based on data with business and technical tags, data dimensions are supplemented to the factual details to form a wide data table;
[0019] Perform multi-level calculations on the wide data table to form a data theme;
[0020] The data themes are then fused and correlated across multiple themes to obtain data topics.
[0021] The operation and maintenance objects, attributes, and relationships in the configuration information management database of the operation and maintenance domain are combined and associated with the content of the real-time data warehouse and the offline data warehouse to form an operation and maintenance data model;
[0022] Based on the aforementioned data topics, relationship mining, data mining, and OLAP analysis are performed on the data to supplement and enhance the aforementioned operation and maintenance data model.
[0023] Optionally, the step of using various types of data in the unified operation and maintenance data model and the relationships between these types of data as parameters to execute rules in the fault location rule base to obtain fault location results includes:
[0024] After a fault alarm is triggered, the percentage and matching degree ranking of different historical root causes are obtained through fault feature matching.
[0025] Use historical fault handling data to match the current fault location;
[0026] Based on a fault knowledge base, intelligent classification is performed through manual definition or natural language parsing, serving as the basis for fault analysis and root cause recommendations, resulting in handling suggestions, automated toolkits, and solution recommendations; the fault knowledge base is formed by collecting fault handling process data.
[0027] Optionally, based on the unified operation and maintenance data model, the step involves using the model's description of each architecture within the business, configuring impact parameters and rules for the business impact model, and combining this with the health score calculated using artificial intelligence algorithms to form an impact analysis, including:
[0028] Construct a health impact model using the object-level impact model in the business impact model;
[0029] Calculate the relationships between maintenance indicators using expert experience or artificial intelligence algorithms;
[0030] By analyzing the mean and variance of the current distribution of related indicators, the health score is calculated using the Gaussian formula.
[0031] Based on the health score, the impact is propagated according to the impact relationships and strategies in the business impact model to form an impact analysis.
[0032] Optionally, the method further includes:
[0033] The fault profile is pushed to the user;
[0034] When the user provides feedback, NLP word segmentation technology is used to extract the feedback.
[0035] Secondly, embodiments of this application provide a data intelligent analysis system, the system comprising:
[0036] The unified operation and maintenance data model building module is used to build a unified operation and maintenance data model across domains in real time.
[0037] The fault location result acquisition module is used to execute the rules in the fault location rule base with various types of data and the relationships between them in the unified operation and maintenance data model as parameters to obtain the fault location result; the rules in the fault location rule base are rules obtained by configuring expert experience into the rule engine through data analysis and machine learning.
[0038] The impact analysis module is used to form an impact analysis based on the unified operation and maintenance data model, by using the model to describe the architecture of each business, configuring impact parameters and impact rules for the business impact model, and combining the health score calculated by artificial intelligence algorithms.
[0039] The fault profile acquisition module is used to profile the fault using the fault location results, the impact analysis, and fault-related information to obtain a fault profile.
[0040] Optionally, the unified operation and maintenance data model module includes:
[0041] The data acquisition submodule is used to listen to and capture data from the data source through the job platform and data acquisition engine;
[0042] The data storage submodule is used to forward the monitored and captured data to the data storage engine's form-painted source data layer using the forwarding function of the data integration engine;
[0043] The data extraction submodule is used to extract data from the source data layer using the data integration engine according to the data extraction rules in the data rule base;
[0044] The data standardization and transformation submodule is used to perform data standardization and transformation on the extracted data based on the data standard library.
[0045] The data cleaning submodule is used to verify and clean the transformed data;
[0046] The tag supplementation submodule is used to supplement the verified and cleaned data with business tags and technical tags based on the data tag library;
[0047] The data dimension supplementation submodule is used to supplement the factual detail data with data that has business tags and technical tags, forming a wide data table.
[0048] The multi-level data calculation submodule is used to perform multi-level calculations on the wide data table to form a data theme.
[0049] The topic fusion submodule is used to fuse and associate the data topics to obtain data themes;
[0050] The Operation and Maintenance Data Model Acquisition Submodule is used to combine and associate the operation and maintenance objects, attributes, and relationships in the configuration information management database of the operation and maintenance domain with the content of the real-time data warehouse and the offline data warehouse to form an operation and maintenance data model.
[0051] The operation and maintenance data model update submodule is used to supplement and enhance the operation and maintenance data model by performing relationship mining, data mining and OLAP analysis on the data topic.
[0052] Optionally, the fault location result acquisition module includes:
[0053] The root cause percentage acquisition submodule is used to obtain the percentage and matching degree ranking of different historical root causes by matching fault features after a fault alarm is triggered.
[0054] The fault location matching submodule is used to match the current fault location with historical fault handling data;
[0055] The solution recommendation submodule is used to intelligently classify faults based on a fault knowledge base, either through manual definition or natural language parsing, as a basis for fault analysis and root cause recommendations, resulting in handling suggestions, an automated toolbox, and solution recommendations; the fault knowledge base is formed by collecting fault handling process data.
[0056] Optionally, the impact analysis module includes:
[0057] The Health Impact Module Construction Submodule is used to construct a health impact model using the object-level impact model in the business impact model.
[0058] The maintenance indicator relationship acquisition submodule is used to calculate the relationships between maintenance indicators through expert experience or artificial intelligence algorithms;
[0059] The health score calculation submodule is used to calculate the health score by analyzing the mean and variance of the current distribution of related indicators and applying the Gaussian formula.
[0060] The impact analysis submodule is used to propagate the impact based on the health score value and the impact relationships and strategies in the business impact model to form an impact analysis.
[0061] Optionally, the system further includes:
[0062] The fault profile push module is used to push the fault profile to the user;
[0063] The feedback module is used to extract feedback from users using NLP word segmentation technology.
[0064] Compared with the prior art, this application has the following beneficial effects:
[0065] This application provides a data intelligence analysis method that establishes a cross-domain, real-time unified operation and maintenance data model. Using various types of data and the relationships between them as parameters, it executes rules from a fault location rule base to obtain fault location results. The rules in the fault location rule base are obtained by configuring expert experience into a rule engine through data analysis and machine learning. Based on the unified operation and maintenance data model, it utilizes the model's description of various architectures in the business, configures impact parameters and rules for the business impact model, and combines this with health scores calculated using artificial intelligence algorithms to form an impact analysis. The method then uses the fault location results, the impact analysis, and fault-related information to create a fault profile. When a fault occurs, it can locate the fault using the unified operation and maintenance data model and the fault location rule base, obtain fault location results, and then perform impact analysis. Based on the fault location results, impact analysis, and fault-related information, a fault profile is created. This fault profile is equivalent to an analysis of the fault from its root causes, handling suggestions, and solution recommendations, thereby enabling rapid fault location and analysis in complex and cross-domain IT environments.
[0066] It should be noted that the data intelligent analysis system provided in this application can implement the steps of the above-mentioned data intelligent analysis method, and therefore also has the above-mentioned beneficial effects. Attached Figure Description
[0067] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0068] Figure 1 This is a schematic diagram of a data intelligent analysis method provided in an embodiment of this application;
[0069] Figure 2 This application provides a schematic diagram of a method for constructing a unified operation and maintenance data model.
[0070] Figure 3 This is a schematic diagram of a unified operation and maintenance data model output provided in an embodiment of this application;
[0071] Figure 4 This is a schematic diagram of a fault location processing flow provided in an embodiment of this application;
[0072] Figure 5 A schematic diagram of an object-level health impact model provided in an embodiment of this application;
[0073] Figure 6 This application provides a schematic diagram illustrating the influence relationship between indicators in an embodiment.
[0074] Figure 7 This is a schematic diagram of a set of health-related indicators provided in an embodiment of this application;
[0075] Figure 8 This application provides a schematic diagram of the distribution of various indicator data in an embodiment.
[0076] Figure 9 This is a schematic diagram illustrating an effect propagation method provided in an embodiment of this application;
[0077] Figure 10 A schematic diagram illustrating the principle of a data intelligent analysis methodology provided in this application embodiment;
[0078] Figure 11 A schematic diagram of the technical architecture of a data intelligent analysis method provided in this application embodiment;
[0079] Figure 12 This is a schematic diagram of the structure of a data intelligent analysis system provided in an embodiment of this application. Detailed Implementation
[0080] As described above, during IT operations and maintenance, there are often simultaneous abnormal alarms from business, applications, systems, networks, storage, and data center environments. In such cases, technical personnel from various disciplines need to collaborate to analyze the data and locate the fault. In other words, the entire process of locating the fault relies on human intervention and the knowledge and experience of their respective professional fields. However, manually locating and analyzing faults often consumes a lot of time and energy.
[0081] Through research, the inventors have developed a data intelligence analysis method and system that can quickly locate and analyze faults in complex and cross-domain IT environments.
[0082] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0083] Method Implementation Examples
[0084] See Figure 1 The figure is a schematic diagram of a data intelligent analysis method provided in an embodiment of this application, including the following steps:
[0085] S101, establish a unified operation and maintenance data model that is real-time across domains.
[0086] It should be noted that the unified operation and maintenance data model can take the configuration model in the configuration information management database (CMDB) of the operation and maintenance domain as the core, organize multi-source operation and maintenance data such as alarms, indicators, logs, traffic, IT service management work order data, automated operations, and manual operations of the audit platform, and form relationships between various types of data, such as the relationship between configuration and alarms, the relationship between alarms and indicators, the relationship between indicators, and the relationship between indicators and automated and manual operations, etc., and finally form a cross-domain real-time unified operation and maintenance data model.
[0087] Specifically, data can be monitored and captured from data sources through an operational platform and a data acquisition engine. The data integration engine's forwarding function then forwards the monitored and captured data to the data storage engine in the form of a source data layer. The data integration engine extracts data from this source data layer according to data extraction rules in the data rule base. Based on a data standard library, the extracted data undergoes data standardization transformation. The transformed data is then validated and cleaned. Based on a data tag library, business and technical tags are added to the validated and cleaned data. Based on the data with business and technical tags, data dimensioning is performed on the factual details to form a wide data table. Multi-level calculations are performed on the wide data table to form data themes. These data themes are then fused and correlated to obtain data topics. The operation and maintenance objects, attributes, and relationships in the operation and maintenance domain configuration information management database are combined and correlated with the content of the real-time and offline data warehouses to form an operation and maintenance data model. Based on these data topics, relationship mining, data mining, and OLAP analysis are performed on the data to supplement and enhance the operation and maintenance data model.
[0088] The process for building a unified operation and maintenance data model can be found in [link to documentation]. Figure 2 This figure is a schematic diagram of a unified operation and maintenance data model construction method provided in an embodiment of this application, mainly including three parts: model input, model construction, and model output. The model input mainly includes five dimensions: configuration management data, key business indicators, basic monitoring data, ITIL flow data, and knowledge data, where ITIL stands for IT operation and maintenance management. The model construction mainly includes nine steps: data collection, data extraction, data standardization, data cleaning, data tagging, data thematication, data specialization, operation and maintenance data modeling, and data analysis. The model output can be categorized into four parts: entities and entity relationships forming a description of the real world; continuous entity activities forming big data; logically linked activities forming business processes; and the temporality of the data.
[0089] Specifically, regarding the input to the model, the data can first be configured using a configuration management model. This model covers the business service layer, IT service layer, logical resource layer, physical resource layer, and infrastructure layer. It represents the logical relationship diagram between applications and technical architecture, clearly mapping and combining the entity model of the IT system with the logical model of the business layers. This automatically and quickly maps each IT component failure or performance event to the status of different layers of business services, enabling rapid and accurate determination of the scope and extent of business impact. Next, key metrics including monitoring alarms, monitoring indicators, and logs are obtained. Monitoring alarms include application alarms, system alarms, network alarms, and hardware alarms; monitoring indicators include business applications, capacity indicators, status indicators, and performance indicators; and logs include transaction logs, application logs, security audit logs, middleware logs, system logs, device logs, and network traffic. Then, IT service ticket data containing data on events, problems, changes, and routine operations is obtained. Knowledge data includes expert experience, knowledge base data, and fault rules.
[0090] Specifically, for model construction, the system uses an operational platform and data acquisition engine to monitor and capture data from data sources. This data is then forwarded to the data storage engine's source data layer via a data integration engine, providing raw data for data analysis and artificial intelligence. The data integration engine extracts data according to the data extraction rules in the data rule base. Data standardization is achieved using data standard libraries, such as metadata standards and data item standards. Data is then validated, removing incorrect attributes, missing data, and duplicate data. Finally, business and technical tags are added to the cleaned data using a data tag library, ensuring each piece of data has a corresponding business and technical tag. This process is repeated to achieve the desired data structure. After data is labeled, it undergoes dimensionality extraction, classification, summarization, and degradation. Detailed factual data is supplemented with additional dimensions to form a wide data table. Multi-level calculations are then performed to create data themes, achieving data thematization. Data themes are then fused and correlated to form richer, more multi-dimensional data topics, achieving data thematization. It should be noted that data topics form the foundation for subsequent data analysis steps. Combining the operational objects, attributes, and relationships in the CMDB with the content of real-time and offline data warehouses, correlations are formed to create an operational data model, achieving operational data modeling. Based on the thematized data, relationship mining, data mining, and OLAP analysis are performed. Artificial intelligence algorithms are used for intelligent analysis to supplement and enhance the operational data model, achieving data analysis. It should be noted that OLAP (Online Analytical Processing) is a data processing technology used to support complex analytical operations.
[0091] For the model's output, please refer to Figure 3 This diagram illustrates the output of a unified operation and maintenance data model provided in this application embodiment. It exemplarily shows the relationships between various types of data. It should be noted that, from the perspective of entities and entity relationships constituting a description of the real world, the model's data can be understood as all physical and logical entities within an enterprise being abstracted as entities, i.e., the subjects of enterprise activities. For example, entities such as people, equipment, computers, cloud services, and business systems have relationships with each other, such as connection relationships, dependency relationships, installation relationships, and colleague relationships. These relationships form a network of relationships between entities. From the perspective of the continuous activities of entities forming big data, entities are constantly engaged in activities, which may manifest as server CPU utilization, etc. While entities are finite, activities are infinite. Big data is caused by a large number of entity activities and can be stored in versioned form according to time. From the perspective of logically linked activities forming business processes, there are various logics between activities, such as AND, OR, and NOT. Entity activities are logically linked into processes, which are driven by "events" generated by the activities, drawing inspiration from EPC (Event-Driven Process). Chain theory; from the perspective of data temporality, since information is time-sensitive, knowledge and data that change at a point in time also represent meaningful information. By taking "time" as an important variable, we can locate the time of fact generation, the effective time period, the state of knowledge at a specific historical moment, and predict and analyze development trends.
[0092] S102, using the various types of data in the unified operation and maintenance data model and the relationships between them as parameters, execute the rules in the fault location rule base to obtain the fault location result.
[0093] It should be noted that the rules in the fault location rule base are obtained by configuring expert experience into the rule engine through data analysis and machine learning.
[0094] Specifically, after a fault alarm is triggered, fault feature matching can be used to obtain the proportion and matching degree ranking of different historical root causes. Historical fault handling data can then be used to match the current fault location. Based on a fault knowledge base, intelligent classification is performed through manual definition or natural language parsing, serving as the basis for fault analysis and root cause recommendations. This yields handling suggestions, an automated toolkit, and solution recommendations. It should be noted that the fault knowledge base is formed by collecting fault handling process data.
[0095] See Figure 4 The figure is a schematic diagram of a fault location processing flow provided in an embodiment of this application, which mainly includes five parts: fault discovery, fault analysis, fault handling, fault knowledge base, and auxiliary location and handling recommendations.
[0096] First, fault location is triggered. When a new event occurs, such as a fault alarm, the information is filtered through scenario activation rules before the root cause rule is executed. Filtering using scenario activation rules filters out fault alarm information. The information used by the root cause rule is obtained in real-time from the operations and maintenance data model using a lazy loading model, which then performs rule calculations to arrive at the final result.
[0097] Then, historical root cause analysis and recommendations for handling are provided. This can be achieved by collecting fault handling process data, such as event tickets, automated toolkits, and solutions, through manual and automated methods, forming a fault knowledge base. Then, by matching fault characteristics, the proportion and matching degree ranking of different historical root causes are displayed for user reference. It should be noted that the intelligent matching of historical fault handling solutions uses historical fault handling data to match the current fault location, using those with high similarity as handling suggestions. Based on the fault knowledge base, intelligent classification is performed through manual definition or natural language parsing to provide users with historical root cause rankings as a basis for fault analysis and root cause recommendations, thereby providing handling suggestions, automated toolkits, and solution recommendations.
[0098] S103, based on the unified operation and maintenance data model, the model describes the architecture of each business, and the impact parameters and rules are configured for the business impact model. Combined with the health score calculated by artificial intelligence algorithms, an impact analysis is formed.
[0099] In the embodiments provided in this application, a health impact model can be constructed using the object-level impact model in the business impact model. Through expert experience or artificial intelligence algorithms, the relationship between maintenance indicators can be calculated. By analyzing the mean and variance of the current distribution of the related indicators, the health score is calculated according to the Gaussian formula. Based on the health score, the impact is propagated according to the impact relationship and impact strategy in the business impact model to form an impact analysis.
[0100] Specifically, business impact analysis can be explained from three aspects: business impact model, health status calculation, and model operation.
[0101] Firstly, regarding business impact models, the current business service tree model can describe objects at various levels of business, technology, and resources, as well as the relationships between them, laying a solid foundation for business impact analysis and root cause analysis. However, in actual production, the business impact model is not entirely the same as the business service tree model. For example, in the relationship between service A and service B, the access to the call relationship is positive in the business service tree model, but the direction of impact is reversed. Some relationships exist in the business service tree model but have no impact, while others do not have direct or regular relationships in the business service tree model but still have an impact. In these cases, an impact relationship can be configured. Therefore, a business impact model needs to be derived from the business service tree model. Specific attributes of impact relationships can be found in Table 1 below:
[0102] Table 1
[0103]
[0104] It should be noted that business impact models can be divided into object-level impact models and metric-level impact models. Object-level impact models, also known as object-level health response models, can be configured based on expert experience. For example, in the context of a call relationship between two services, "Is there an impact?" can be configured as "Yes", "Impact direction" as "Reverse impact", and "Impact propagation mode" as "Normal mode," i.e., modeling is based on expert experience. Alternatively, artificial intelligence algorithms can be used to analyze logs, traffic, metrics, and other data to derive the impact direction and relationships. As an example, specific instantiation of impact relationships can be found in Table 2 below:
[0105] Table 2
[0106]
[0107] The final object-level health impact model can be found in [reference needed]. Figure 5 The figure is a schematic diagram of an object-level health impact model provided in an embodiment of this application.
[0108] For indicator-level impact models, relationships between indicators can be calculated and maintained through expert experience or artificial intelligence algorithms. For example, AI algorithms can determine the correlation between indicators by identifying those with similar or opposite fluctuation trends within the same time window. The influence relationships between indicators can be found in [reference needed]. Figure 6 The figure is a schematic diagram of the influence relationship between indicators provided in an embodiment of this application, showing the influence relationship between the database session count indicator and the local disk busyness indicator.
[0109] For health score calculation, the health score is determined by related indicators. These indicators can be analyzed by examining their current distribution's mean and variance, and then the score is calculated using the Gaussian formula based on these values. It should be noted that the weighting of multiple related indicators can be the average of their total weights.
[0110] See Figure 7 The figure is a schematic diagram of a set of health-related indicators provided in an embodiment of this application. For the selection of the set of health-related indicators, the relevant two-dimensional matrix of indicators can be selected by using the Pearson correlation coefficient without changing the health object.
[0111] It should be noted that, regarding the data feature extraction for health score calculation, since the data for each indicator is continuous from a data distribution perspective, it theoretically follows a normal distribution. (See [link to relevant documentation]). Figure 8 The figure is a schematic diagram of the distribution of various indicator data provided in the embodiment of this application. During extraction, the feature values and scores of each indicator can be extracted at 20-minute intervals, and the feature values and scores are used to construct a training set.
[0112] Then, the health score can be calculated by combining offline and online models. For the offline model, the main consideration is the features extracted from each indicator, and a normal distribution statistical model is used for model construction and training. The features extracted in the above steps are used as the model input, and the score is used as the result, calculated using a formula. The model is trained to obtain the mean μ and standard deviation σ corresponding to the indicator. Then, using the indicator as the primary key, the mean μ and standard deviation σ corresponding to each indicator of the object are stored in MySQL. In the above formula, Index is the health value of the indicator, x is the feature value extracted by the indicator, mean μ is the mean obtained by the indicator, and standard deviation σ is the standard deviation obtained by the indicator.
[0113] Then, online models can be used for prediction. Online model prediction mainly involves finding the mean μ and standard deviation σ of the server object and its index input by the user from MySQL, and then substituting these values into the formula. The corresponding health value is calculated from the formula, and then the health values of all indicators are calculated using the formula. The health value of the server object is calculated, where HMI_SCORE is the health value of the server object, and index is the health value calculated for each indicator. After softmax normalization, the values are summed to obtain the health value of the server object. Finally, the calculated health value is converted into availability status.
[0114] For model computation, the influence model algorithm rules represent the impact of business-oriented events or indicators on the status of resource nodes and business nodes. Based on the node's health status attribute (STATUS), it can be categorized into five types: unavailable, slightly damaged, damaged, slightly damaged, and normal. The health of resource nodes can be calculated through the analysis of key indicators. Based on the influence relationships and influence strategies in the influence model, influence propagation is performed. It should be noted that the influence relationships in the influence model propagate influence based on the direction and propagation mode. There are three propagation modes: direct propagation mode, health-based propagation mode, and influence rule mode. The direct propagation mode directly influences the node without considering its health status. The health-based propagation mode checks the health of each node during propagation; if the health is normal, propagation stops; if the health is abnormal, propagation continues. The influence rule mode requires reading the influence rules when an influence reaches a node, performing rule calculations, and deciding whether to propagate the influence. For details on influence propagation methods, please refer to [link to relevant documentation]. Figure 9 This figure is a schematic diagram of an influence propagation method provided in an embodiment of this application.
[0115] S104, using the fault location results, the impact analysis, and fault-related information, a fault profile is created to obtain a fault profile.
[0116] Specifically, fault profiles can be created by combining fault location, impact analysis, fault tagging, fault tracing, fault handling knowledge base + fault handling recommendations, related alarms, related work orders, related audit logs, related automated operation logs, and various information of nodes. This means analyzing faults from the perspectives of their root causes, handling suggestions, and solution recommendations.
[0117] This application provides a data intelligence analysis method that establishes a cross-domain, real-time unified operation and maintenance data model. Using various types of data and their relationships as parameters, it executes rules from a fault location rule base to obtain fault location results. Based on the unified operation and maintenance data model, it utilizes the model's description of various architectures within the business, configures impact parameters and rules for the business impact model, and combines this with health scores calculated using artificial intelligence algorithms to form an impact analysis. Finally, it uses the fault location results, impact analysis, and fault-related information to create a fault profile. When a fault occurs, the fault profile is created based on the fault location results, impact analysis, and fault-related information. This fault profile is equivalent to an analysis of the fault from its root causes, handling suggestions, and solution recommendations, thereby enabling rapid fault location and analysis in complex and cross-domain IT environments.
[0118] As an example, after obtaining the fault profile, the fault profile can be pushed to the user; when the user provides feedback, NLP word segmentation technology can be used to extract the feedback.
[0119] It should be noted that NLP (Natural Language Processing) technology is a natural language processing technology that pushes fault profile data analysis reports to users in real time, which can achieve the effect of immediate user engagement. It can also receive user feedback, extract feedback through NLP word segmentation technology, form a structured fault analysis and judgment, and form a closed-loop management of data verification.
[0120] The methodological principle of the data intelligent analysis method provided in this application embodiment can be found in [reference needed]. Figure 10 The figure is a schematic diagram of the data intelligent analysis methodology provided in the embodiment of this application. It mainly observes the external system environment and external information and intelligence input into the business system, and then makes judgments and takes actions based on the interaction between key business objectives, logs and service trees, analysis and diagnosis, experience and knowledge base and basic monitoring data.
[0121] The overall technical architecture of the data intelligence method provided in this application embodiment can be found in [reference needed]. Figure 11 This diagram illustrates the technical architecture of a data intelligence analysis method provided in this application embodiment. First, data is acquired from a data source containing performance metrics, status metrics, log information, alarm information, and inspection data. This data is then sent to the Flink engine and data platform via the Kafka platform. It should be noted that the Flink engine is a framework and distributed processing engine used for stateful computation on both unbounded and bounded data streams. The Flink engine can perform steps such as scenario initiation judgment, rule parameter collection, fault location rule execution, and impact analysis rule execution. The data is then transmitted to various databases for storage via the Kafka platform, and also sent to platforms such as the operation and maintenance management platform and workbench to assist operation and maintenance personnel in handling faults. After being sent to the data platform, the data undergoes standardization, dimension completion, and metric calculation. After being stored in relational databases, graph databases, and time-series databases, the data is sent to the operation and maintenance data model and rule management platform. The rule management platform includes a rule base, KIE Server, KIE Workbench, and artificial intelligence. It should be noted that the KIE Server is a container containing an application programming interface (REST API) and an execution engine. Workbench is a rules engine that includes rule configuration and testing modules, and its rule base contains rules and models. Through... Figure 11The architecture shown enables functions such as fault profiling, impact rule management, fault location rule management, and rule testing.
[0122] This application provides a data intelligence analysis method that establishes a unified operation and maintenance data model, a fault location rule base, and business impact analysis. Based on these, a fault profile is formed. Using a big data platform as its foundation, it can enhance powerful data capabilities such as data system visualization construction capabilities, integrated data flow batch computing capabilities, multi-type data storage and query capabilities, unified resource scheduling capabilities, and data task scheduling capabilities. Furthermore, it can combine artificial intelligence algorithms to achieve functions such as anomaly detection, data correlation analysis, health analysis, fault correlation analysis, and fault location, enabling rapid fault location and analysis in complex and cross-domain IT environments.
[0123] System Implementation Examples
[0124] See Figure 12 The figure is a schematic diagram of the structure of a data intelligent analysis system provided in an embodiment of this application, including: a same operation and maintenance data model establishment module 1201, a fault location result acquisition module 1202, an impact analysis module 1203, and a fault profile acquisition model 1204.
[0125] Among them, the unified operation and maintenance data model establishment module 1201 is used to establish a cross-domain real-time unified operation and maintenance data model;
[0126] The fault location result acquisition module 1202 is used to execute the rules in the fault location rule base with various types of data and the relationships between various types of data in the unified operation and maintenance data model as parameters to obtain the fault location result; the rules in the fault location rule base are rules obtained by configuring expert experience into the rule engine through data analysis and machine learning.
[0127] The impact analysis module 1203 is used to form an impact analysis based on the unified operation and maintenance data model, by using the model to describe the architecture of each business, configuring impact parameters and impact rules for the business impact model, and combining the health score calculated by artificial intelligence algorithms.
[0128] The fault profile acquisition module 1204 is used to profile the fault using the fault location results, the impact analysis, and fault-related information to obtain a fault profile.
[0129] Optionally, the unified operation and maintenance data model module 1201 includes:
[0130] The data acquisition submodule is used to listen to and capture data from the data source through the job platform and data acquisition engine;
[0131] The data storage submodule is used to forward the monitored and captured data to the data storage engine's form-painted source data layer using the forwarding function of the data integration engine;
[0132] The data extraction submodule is used to extract data from the source data layer using the data integration engine according to the data extraction rules in the data rule base;
[0133] The data standardization and transformation submodule is used to perform data standardization and transformation on the extracted data based on the data standard library.
[0134] The data cleaning submodule is used to verify and clean the transformed data;
[0135] The tag supplementation submodule is used to supplement the verified and cleaned data with business tags and technical tags based on the data tag library;
[0136] The data dimension supplementation submodule is used to supplement the factual detail data with data that has business tags and technical tags, forming a wide data table.
[0137] The multi-level data calculation submodule is used to perform multi-level calculations on the wide data table to form a data theme.
[0138] The topic fusion submodule is used to fuse and associate the data topics to obtain data themes;
[0139] The Operation and Maintenance Data Model Acquisition Submodule is used to combine and associate the operation and maintenance objects, attributes, and relationships in the configuration information management database of the operation and maintenance domain with the content of the real-time data warehouse and the offline data warehouse to form an operation and maintenance data model.
[0140] The operation and maintenance data model update submodule is used to supplement and enhance the operation and maintenance data model by performing relationship mining, data mining and OLAP analysis on the data topic.
[0141] Optionally, the fault location result acquisition module 1202 includes:
[0142] The root cause percentage acquisition submodule is used to obtain the percentage and matching degree ranking of different historical root causes by matching fault features after a fault alarm is triggered.
[0143] The fault location matching submodule is used to match the current fault location with historical fault handling data;
[0144] The solution recommendation submodule is used to intelligently classify faults based on a fault knowledge base, either through manual definition or natural language parsing, as a basis for fault analysis and root cause recommendations, resulting in handling suggestions, an automated toolbox, and solution recommendations; the fault knowledge base is formed by collecting fault handling process data.
[0145] Optionally, the impact analysis module 1203 includes:
[0146] The Health Impact Module Construction Submodule is used to construct a health impact model using the object-level impact model in the business impact model.
[0147] The maintenance indicator relationship acquisition submodule is used to calculate the relationships between maintenance indicators through expert experience or artificial intelligence algorithms;
[0148] The health score calculation submodule is used to calculate the health score by analyzing the mean and variance of the current distribution of related indicators and applying the Gaussian formula.
[0149] The impact analysis submodule is used to propagate the impact based on the health score value and the impact relationships and strategies in the business impact model to form an impact analysis.
[0150] Optionally, the system further includes:
[0151] The fault profile push module is used to push the fault profile to the user;
[0152] The feedback module is used to extract feedback from users using NLP word segmentation technology.
[0153] This application provides a data intelligence analysis system that utilizes a unified operation and maintenance data model establishment module, a fault location result acquisition module, an impact analysis module, and a fault profile acquisition module. It establishes a cross-domain, real-time unified operation and maintenance data model. Using various types of data and the relationships between them as parameters, it executes rules from a fault location rule base to obtain fault location results. The rules in the fault location rule base are obtained by configuring expert experience into a rule engine through data analysis and machine learning. Based on the unified operation and maintenance data model, it uses the model's description of various architectures in the business, configures impact parameters and rules for the business impact model, and combines this with health scores calculated using artificial intelligence algorithms to form an impact analysis. Finally, it uses the fault location results, the impact analysis, and fault-related information to create a fault profile. When a fault occurs, it can be located through a unified operation and maintenance data model and fault location rule base to obtain the fault location result. Then, an impact analysis is performed. Based on the fault location result, impact analysis, and fault-related information, a fault profile is created. The fault profile is equivalent to an analysis of the fault from the aspects of its root cause, handling suggestions, and solution recommendations. This enables rapid fault location and analysis in complex and cross-domain IT environments.
[0154] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The system embodiments described above are merely illustrative, and the modules described as separate components may or may not be physically separate. The components indicated as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the solution in this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0155] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A data intelligent analysis method, characterized in that, The method includes: Establish a unified operation and maintenance data model that is real-time across domains, including data monitoring and capture from data sources through an operation platform and data acquisition engine; The forwarding function of the data integration engine is used to forward the monitored and captured data to the data storage engine in the form of the source data layer; The data is extracted from the source data layer by the data integration engine according to the data extraction rules in the data rule base; The extracted data is standardized and transformed based on the data standard library; The converted data is then validated and cleaned. Based on the data tag library, business tags and technical tags are added to the verified and cleaned data; Based on data with business and technical tags, data dimensions are supplemented to the factual details to form a wide data table; Perform multi-level calculations on the wide data table to form a data theme; The data themes are then fused and correlated across multiple themes to obtain data topics. The operation and maintenance objects, attributes, and relationships in the configuration information management database of the operation and maintenance domain are combined and associated with the content of the real-time data warehouse and the offline data warehouse to form an operation and maintenance data model; Based on the aforementioned data topics, relationship mining, data mining, and OLAP analysis are performed on the data to supplement and enhance the aforementioned operation and maintenance data model; Using the various types of data and the relationships between them in the unified operation and maintenance data model as parameters, the rules in the fault location rule base are executed to obtain the fault location result; the rules in the fault location rule base are rules obtained by configuring expert experience into the rule engine through data analysis and machine learning; the various types of data include at least one or more of the following: organizational alarms, indicators, logs, traffic, IT service management work order data, automated operations, and manual operations from the audit platform; Based on the unified operation and maintenance data model, the model describes the architecture of each business, and the impact parameters and rules are configured for the business impact model. Combined with the health score calculated by artificial intelligence algorithms, an impact analysis is formed, including the construction of a health score impact model using the object-level impact model in the business impact model. Calculate the relationships between maintenance indicators using expert experience or artificial intelligence algorithms; By analyzing the mean and variance of the current distribution of related indicators, the health score is calculated using the Gaussian formula. Based on the health score, the impact is propagated according to the impact relationships and strategies in the business impact model to form an impact analysis. The fault profile is obtained by using the fault location results, the impact analysis, and fault-related information.
2. The method according to claim 1, characterized in that, The step involves using various types of data and the relationships between these types of data in the unified operation and maintenance data model as parameters to execute rules in the fault location rule base, thereby obtaining fault location results, including: After a fault alarm is triggered, the percentage and matching degree ranking of different historical root causes are obtained through fault feature matching. Use historical fault handling data to match the current fault location; Based on a fault knowledge base, intelligent classification is performed through manual definition or natural language parsing, serving as the basis for fault analysis and root cause recommendations, resulting in handling suggestions, automated toolkits, and solution recommendations; the fault knowledge base is formed by collecting fault handling process data.
3. The method according to claim 1, characterized in that, The method further includes: The fault profile is pushed to the user; When the user provides feedback, NLP word segmentation technology is used to extract the feedback.
4. A data intelligent analysis system, characterized in that, The system includes: The unified operations and maintenance data model building module is used to build a cross-domain, real-time unified operations and maintenance data model, including: The data acquisition submodule is used to listen to and capture data from the data source through the job platform and data acquisition engine; The data storage submodule is used to forward the monitored and captured data to the data storage engine's form-painted source data layer using the forwarding function of the data integration engine; The data extraction submodule is used to extract data from the source data layer using the data integration engine according to the data extraction rules in the data rule base; The data standardization and transformation submodule is used to perform data standardization and transformation on the extracted data based on the data standard library. The data cleaning submodule is used to verify and clean the transformed data; The tag supplementation submodule is used to supplement the verified and cleaned data with business tags and technical tags based on the data tag library; The data dimension supplementation submodule is used to supplement the factual detail data with data that has business tags and technical tags, forming a wide data table. The multi-level data calculation submodule is used to perform multi-level calculations on the wide data table to form a data theme. The topic fusion submodule is used to fuse and associate the data topics to obtain data themes; The Operation and Maintenance Data Model Acquisition Submodule is used to combine and associate the operation and maintenance objects, attributes, and relationships in the configuration information management database of the operation and maintenance domain with the content of the real-time data warehouse and the offline data warehouse to form an operation and maintenance data model. The operation and maintenance data model update submodule is used to perform relationship mining, data mining and OLAP analysis on the data topic to supplement and enhance the operation and maintenance data model. The fault location result acquisition module is used to execute the rules in the fault location rule base with various types of data and the relationships between them in the unified operation and maintenance data model as parameters to obtain the fault location result. The rules in the fault location rule base are rules obtained by configuring expert experience into the rule engine through data analysis and machine learning. The various types of data include at least one or more of the following: organizational alarms, indicators, logs, traffic, IT service management work order data, automated operations, and manual operations from the audit platform. The impact analysis module, based on the unified operation and maintenance data model, utilizes the model's description of various architectures within the business, configures impact parameters and rules for the business impact model, and combines this with health scores calculated using artificial intelligence algorithms to generate impact analysis. The impact analysis module includes: The Health Impact Module Construction Submodule is used to construct a health impact model using the object-level impact model in the business impact model. The maintenance indicator relationship acquisition submodule is used to calculate the relationships between maintenance indicators through expert experience or artificial intelligence algorithms; The health score calculation submodule is used to calculate the health score by analyzing the mean and variance of the current distribution of related indicators and applying the Gaussian formula. The impact analysis submodule is used to propagate the impact based on the health score value and the impact relationships and strategies in the business impact model to form an impact analysis. The fault profile acquisition module is used to profile the fault using the fault location results, the impact analysis, and fault-related information to obtain a fault profile.
5. The system according to claim 4, characterized in that, The fault location result acquisition module includes: The root cause percentage acquisition submodule is used to obtain the percentage and matching degree ranking of different historical root causes by matching fault features after a fault alarm is triggered. The fault location matching submodule is used to match the current fault location with historical fault handling data; The solution recommendation submodule is used to intelligently classify faults based on a fault knowledge base, either through manual definition or natural language parsing, as a basis for fault analysis and root cause recommendations, resulting in handling suggestions, an automated toolbox, and solution recommendations; the fault knowledge base is formed by collecting fault handling process data.
6. The system according to claim 4, characterized in that, The system also includes: The fault profile push module is used to push the fault profile to the user; The feedback module is used to extract feedback from users using NLP word segmentation technology.