FAULT DETECTION AND ANALYSIS SYSTEM

TR202607189A2Pending Publication Date: 2026-06-22TURKIYE GARANTI BANKASI ANONIM SIRKETI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
TR · TR
Patent Type
Applications
Current Assignee / Owner
TURKIYE GARANTI BANKASI ANONIM SIRKETI
Filing Date
2026-05-07
Publication Date
2026-06-22

Smart Images

  • Figure 00000019_0000
    Figure 00000019_0000
Patent Text Reader

Abstract

This invention relates to a system (1) that enables monitoring the operational status of services by analyzing the data from scattered logs, metrics, traces and software update processes generated in the information processing infrastructures of institutions, determining the root causes and impact areas of events such as errors, slowdowns or interruptions affecting user operations, identifying anomalies and situations that pose a risk of performance degradation and offering solutions and / or improvement suggestions.
Need to check novelty before this filing date? Find Prior Art

Description

1 TARIFF FAULT DETECTION AND ANALYSIS SYSTEM Technical Area This invention analyzes the distributed logs, metrics, and data generated within the IT infrastructures of organizations. By analyzing data related to monitoring and software update processes, the services' operation is evaluated. monitoring their status, errors, slowdowns or other issues affecting user operations the root causes and areas of impact of events that occur in the form of disruptions Identifying anomalies and risks that could lead to decreased performance. 10 identifying the situations that cause this and suggesting solutions and / or improvements It is related to a system that enables its presentation. Previous Technique Today, operational monitoring processes are generated in software and service infrastructures. Collection and indexing of log data, metric data, and monitoring outputs. and is ensured through questioning. Splunk, used in current systems, Dynatrace, Instanta, and Kibana-like solutions offer basic monitoring, performance tracking, and It provides event notifications and information about the real-time status of target systems. It provides visibility. However, data obtained from different sources... integration, removal of service dependencies and cause-and-effect analysis of events The representation of domains in a graph-based manner is limited in current systems. In contrast, warning generation based on fixed thresholds and manual rules is variable load. In these conditions, it leads to noise and stimulus fatigue. 25 Therefore, considering the studies and shortcomings in the current technique... When considered, log records generated in IT infrastructures improve performance. Real-time output of measurements, transaction traces, and software update processes. monitoring, determining the root cause of events, identifying the scope of impact, and similar 30 2 Developing proactive testing recommendations to prevent the recurrence of problems. It appears that a system providing this is needed. Chinese patent CN114139747A, which is included in the known state of the art. in the document, failures, performance degradation or 5 experienced in large IT systems a system that enables the early detection and resolution of capacity problems using artificial intelligence The system in question is discussed. This invention concerns the operation of large IT infrastructures. and an intelligent AIOps that automates maintenance processes with the support of artificial intelligence. This describes the Artificial Intelligence for IT Operations system. The system is a The center operates on a four-layer middleware architecture, and the AIOps center provides data and artificial intelligence. It consists of intelligence, technology, and business middle layers, as well as a business application system. This The system brings together log, performance, hardware, and software data from different sources into a single structure. It analyzes the system's overall status by gathering data at the central location. The AIOps center analyzes the overall status of the system. a health score measurement module, and a root cause analysis module that identifies the reasons for errors, Fault prediction module that predicts future malfunctions, system 15 capacity assessment module that monitors and plans resources and unnecessary It includes a smart alarm management module that reduces alarms. Artificial intelligence is central to the system. Data preprocessing, natural language processing, association, clustering, located in the layer, Time series and tree-based forecasting algorithms analyze the collected data. It detects anomalies regardless of threshold. The system's past performance is 20. It offers a self-improving analysis infrastructure that learns from its records. Thanks to these analyses, system performance drops are detected early, and the root cause of the problem is identified. quickly identifies the causes, optimizes resource utilization, and increases capacity. It is involved in planning. It also handles alarm management, recurring or insignificant alarms. By filtering notifications using AI algorithms, only critical events are reported. It enables the display of the front-end, middleware, and back-end interfaces. through integration, all operational and maintenance activities can be managed through a single portal. It enables monitoring and management. 3 Brief Description of the Invention The purpose of this invention is to enable computing performed in distributed or hybrid computing infrastructures. showing user actions, target system calls, and data exchanges that occurred. 5 errors and problems that manifest as slowdowns, pauses, or inaccessibility Log data recording performance status indicates the speed and error rate of target systems. metrics showing rates and software update processes By analyzing the records with the relevant transaction number information, the source of the error can be determined. automated suggestions to identify and prevent error occurrence The goal is to create a system that enables its presentation. 10 Detailed Description of the Invention The "Error Detection and Analysis System" was developed to achieve the purpose of this invention. shown in the attached figure; this figure is 15 Figure 1. Schematic view of the system described in the invention. The parts shown in the figure are individually numbered, and the corresponding numbers correspond to these numbers. given below: 20 1. System 2. Data collection module 3. Normalization module 4. Analysis module 25 5. Warning module 6. Visualization module 7. Suggestion module 8. Risk assessment module 9. Feedback module 30 4 Dispersed logs, metrics, traces, and software generated within organizations' IT infrastructures. By analyzing data related to update processes, the operational status of services is determined. monitoring of errors, slowdowns, or interruptions affecting user operations. Identifying the root causes and areas of impact of the emerging events, 5 abnormalities and risk situations that could lead to decreased performance to identify and provide solutions and / or improvement suggestions The developed invention subject system (1); - logs are transaction records generated within an organization's IT infrastructure. data expressing numerical performance values ​​in the form of speed and error rate Metric measurements are traces that show the progress of processes within the target system. 10 gathering information and records generated during software update processes at least one data collection system structured to ensure that data is collected by bringing it in. module (2), - indicating the time when the data collected by the data collection module (2) was generated. 15 timestamps showing the time, the source of the data, and the service to which the data belongs. Service information and records belonging to the same transaction are used with the transaction ID information. at least one normalization module configured to enable its regulation (3), - Operation of services with data organized by the normalization module (3) extracting the relationship pattern during the process and visualizing the connections between services. at least one analysis module (4) configured to enable monitoring as such, 20 - the processes generated by the services during operation by the analysis module (4) Speed ​​data showing the completion time, failures that occurred during the process. error rates and whether the service is accessible Calculating the overall operational status of each service using accessibility information. at least one warning module configured to provide (5), 25 - errors occurring while the services are running, or the processing time being extended, or By analyzing the warnings generated in cases where service access is interrupted to enable repeated warnings to be combined into a single notification at least one visualization module configured (6), - Inter-service relationship information revealed by the analysis module (4) and 30 with alerts combined and prioritized by the visualization module (6) Identifying an error or performance issue that affects user operations and at least one structured to enable the creation of improvement suggestions. suggestion module (7), - initiating and completing software update processes. Records generated during commissioning, as well as before and after the update (update 5). subsequently, error logs and service processing appear on the target system. By analyzing the changes in their durations, the update is applied to the target system. at least one risk assessment structured to determine whether it is working or not module (8), - Errors occurring in services by the target system, processing times unusually short. with warnings generated in cases where the duration of the service extension or interruption of service access According to these warnings, the solutions and improvement suggestions offered should be evaluated by users. To enable tracking of information related to acceptance or rejection processes. It includes at least one feedback module (9) structured accordingly. The data collection module (2) in the system (1) which is the subject of the invention, any communication to communicate with the normalization module (3) using the protocol and data It is configured to carry out the shopping. Data collection module (2), log data (transaction record) generated in the IT infrastructure of organizations, speed and error Metric measurements and tracking information related to numerical performance values ​​in the form of ratios 20 and it asynchronously retrieves logs generated during software update processes; Each record obtained includes a timestamp indicating the time the event occurred, and the data... The source type indicates the origin of the data, and the service indicates the service to which the data belongs. the tag and records belonging to the same transaction together via the transaction ID. It is structured to enable evaluation. Data collection module 25 (2), normalization module according to the order of occurrence of the collected records in time (3) by transferring each software service between a node and services Call and / or error relationships are the connections (edges) between these nodes. the creation of a topology that shows the relationship pattern between services as expressed. It is structured to provide. 30 6 The normalization module (3) in the system (1) of the invention, any communication data collection module (2), analysis module (4) and recommendation using the protocol to communicate with module (7) and exchange data It is structured. Normalization module (3), data collection module (2) The event logs transmitted by 5 will include common domain names and data types. JSON (JavaScript Object Notation) data format that enables representation It is structured to enable conversion. Normalization module (3), Event logs include time information, log characteristics reflecting error conditions, and process data. through identity and dependency information that describes inter-service call relationships by matching the temporal differences between records belonging to the same process or the same service chain. Time series to determine ranking and inter-service interaction patterns. By processing event data sequentially, LSTM (Long Short-Term) Using an artificial intelligence model in the form of (memory) to understand the relationships between events a multidimensional system including time interval, relationship type, and event intensity dimensions It is structured to convert it into a correlation output. 15 The analysis module (4) in the system (1) which is the subject of the invention, any communication Communication with the normalization module (3) and the warning module (5) using the protocol It is structured to establish and exchange data. Analysis module (4) indicates whether each software service is available. Availability information indicates that operations on the service have been completed. latency information indicating the duration of transactions A metric in the form of error-rate information that shows the rate of failures. by evaluating the measurements, these metrics represent the service operational status. based on their levels using a weighted scoring model Calculating a quantitative health score (HS) for each service, the calculated 25 historical scores show the trends of increase and decrease in health scores over time. by examining the overall operational status of the service in comparison with its values. a health trend that reflects a tendency towards continuity, deterioration, or improvement It is structured to produce output. 7 The warning module (5) in the system (1) which is the subject of the invention, any communication Communication with the analysis module (4) and visualization module (6) using the protocol It is configured to establish and exchange data. Alert module (5), by the analysis module (4) regarding the operational status of the services produced and errors occurring, processing time extended or service access interruption 5 Warning logs describing the situations; the time the warning occurred, and the situation it represents. Similarity between alert logs based on type and service information We calculate using the cosine similarity method and the resulting similarity values. Collecting the frequency of warnings over time into a single warning log It is structured accordingly. The warning module (5) is closed by users or 10 By monitoring disregarded warning records, recurring and / or insignificant issues Noise-reduced and prioritized warnings to suppress other warnings. By creating records, notifications are sent to the user. It is configured to transfer to the visualization module (6). The visualization module (6) in the system (1) which is the subject of the invention, analysis and warning 15 Using the event propagation information obtained as a result, a bug, slowdown, or The chain of effects created by an outage event between services can be plotted on a topology graph. It is structured to present to the user. Visualization module (6), service between event logs organized by normalization module (3) Graph navigation performed according to the temporal sequence of calls (graph 20) traversal) and AI inference of data obtained from this traversal The resulting root cause identification output, the depth of the calls, and the occurrence The radius of impact is calculated based on the frequency and the level of propagation of the delay effect. Color on the graphic structure representing the relationship structure between services using radius information. and is structured to visually express differences in location. 25 The proposal module (7) and the normalization module (3) in the system (1) are the subject of the invention. historical log records organized by and event data derived from these records Obtaining sequences, recurring error patterns and unusual occurrences on log records It is structured to identify anomalies expressing behaviors. Proposal module (7) displays the detected failure patterns in the failure topology 30 8 classifying each failure generation topology as a corresponding chaos matching it with the test scenario and performing this matching process as a time series log An LSTM-based artificial intelligence learning system that transforms patterns into a representation space. It is structured to be implemented using the model. The suggestion module (7), For each chaos test proposal created, the proposal must be tested 5 times on the target system. Calculating the uncertainty of the potential impact using the Monte Carlo Dropout method, This Chaos Impact Score (CIS) depends on the level of uncertainty. To generate value, CIS prioritizes testing chaos test scenarios with high CIS values. to define as a scenario and prioritize chaos test proposals by presenting the feedback to the user with the visualization module (6) and monitoring the feedback 10 It is configured to share with the feedback module (9). The risk assessment module (8) in the system (1) which is the subject of the invention, software CI / CD (Continuous Integration / Continuous) related to update processes Deployment – ​​Continuous Integration / Continuous Distribution) line records, software 15 code change record (commit) messages, which are executed during the CI / CD process. test failure records, which refer to records of unsuccessful test results, and records of errors that occurred during the transfer of the software to the target system To collect deployment error logs and compile them into a dataset. It is structured. The risk assessment module (8) sets the code complexity criterion to 20 (code complexity), existing bug logs from previous versions representing the repetition and similarity relationship between the bug logs that appeared in the release the failure similarity metric and the results run for the relevant version. the scope of the tests, the percentage of code sections tested, and the release of test results. 25 to produce test coverage, a measure of how widespread the virus is. It is structured. The risk assessment module (8) contains the code covered by the tests. a test that expresses the proportion of its parts and the spread of test results across the entire version criteria, the status of the relevant software version running on the target system to consider it as a representative indicator, release based on these criteria RSI (Release 30 Stability Indicator), also called an indicator of stability, Calculating the Stability Index (RSI) and visualizing the calculated RSI value. 9 to transfer to the instrument panel to be provided by module (6) It is being structured. The feedback module (9) in the system (1) which is the subject of the invention, any communication Communicate and exchange data with the suggestion module (7) using the protocol. It is configured to perform. Feedback module (9), user 5 information regarding acceptance or rejection processes carried out by this the role and responsibilities of the user performing the operations within the target system to obtain role information that describes the field and to make recommendations based on this feedback data. It is configured to share with the feedback module (7). The feedback module (9), Depending on the role information received, different user groups will evaluate it. 10 In order to distinguish the criteria, Site Reliability Engineering (SRE) Service latency and accessibility for the Reliability Engineering role. (availability) values, error density and sources of errors for the developer role. stack traces showing the location of the code, security role access violation 15 refers to unauthorized access attempts and unusual access behavior. role-based representation vectors (embedding) of features representing anomalies It is configured to convert into vectors. Feedback module (9), the role created for the acceptance or rejection actions performed by users a feedback loop based on reinforcement learning with representation vectors by including and updating the weights of these representation vectors, and these 20 As a result of the update, the chaos test suggestions generated by the suggestion module (7), redefining based on attribute weights corresponding to the user's role It is structured to provide this. Industrial Application of the Invention Thanks to the system (1) which is the subject of the invention, banking, e-commerce, telecommunications, manufacturing 25 and data from institutions operating in different sectors such as healthcare. log and record data produced in their centers and / or cloud-based infrastructures By analyzing and continuously monitoring operational status, user actions Early detection of errors affecting service can prevent service interruptions. Operational processes are ensured to be carried out quickly and safely. 30 Around these basic concepts, the subject of the invention is the "Error Detection and Analysis System (1)". It is possible to develop a wide variety of related applications, and the invention described herein It cannot be limited to examples; it is essentially as stated in the requests.

Claims

11 REQUESTS 1. Dispersed logs, metrics, traces, and data generated in the IT infrastructures of organizations. By analyzing data related to software update processes, the services... Monitoring operational status, errors affecting user operations, 5 The root causes of events that manifest as slowdowns or interruptions, and Identifying areas of impact, abnormalities, and factors leading to performance degradation. Identifying and resolving / / or improving situations that pose a risk enabling the submission of proposals; - log 10 refers to transaction records generated within the IT infrastructure of organizations. expressing the numerical performance values ​​of the data in the form of speed and error rate. metric measurements that track the progress of processes within the target system tracking information and data generated during software update processes structured to enable the collection and aggregation of records at least one data collection module (2), 15 - the time when the data collected by the data collection module (2) was generated time information, data source information, and information to which the data belongs. Service information showing the service and transaction ID of the records belonging to the same transaction at least one structured to enable its regulation using information normalization module (3), 20 - data and services organized by the normalization module (3) extracting the relationship pattern during the operation and between the services the most structured to enable visual tracking of connections a small analysis module (4), - 25 generated by the services during operation by the analysis module (4) Speed ​​data, which shows the time it takes for transactions to be completed, during the transactions error rates, which indicate the failures that occurred, and whether the service is available. general operation of each service, along with accessibility information indicating that it is not available. at least one alert configured to enable the calculation of its status module (5), 30 12 - Errors occurring during service operation, increased processing time. or analysis of alerts generated in case of service interruption to ensure that repeated warnings are combined into a single notification. at least one visualization module structured to (6), - Inter-service relationship information revealed by the analysis module (4) 5 combined and prioritized by the visualization module (6) Warnings indicate an error or performance issue affecting user operations. Identifying the problem and developing improvement suggestions at least one recommendation module configured to provide (7), - Initiating the update process for software updates, 10 Records and updates generated during completion and commissioning. Error that occurred on the target system before and after the update. by analyzing changes in records and service processing times to determine whether the update works on the target system at least one structured risk assessment module (8), 15 - Errors occurring in services by the target system, processing times being unusual. generated in cases of extended or interrupted service access. warnings and the solutions and improvement suggestions offered in response to these warnings, acceptance or rejection actions performed by users At least one feedback structured to enable monitoring of information 20 a system characterized by module (9) (1).

2. Using any communication protocol with the normalization module (3) data structured for communication and data exchange A system like the one in Claim 1 characterized by the collection module (2) (1). 25 3. Log data generated in the IT infrastructure of organizations, speed and error rate. metric measurements related to numerical performance values ​​in the form of tracking information and records generated during software update processes It receives data asynchronously; each received record is set to the time the event occurred. timestamp indicating the source, source type indicating the origin of the data, 13 a service tag indicating the service to which the data belongs and records belonging to the same process to enable joint evaluation via transaction ID the above is characterized by the structured normalization module (3) a system like any of the requests (1).

4. Normalization of the collected records according to their chronological order. by transferring each software service to module (3) a node and services call and / or error relationships between these nodes a diagram showing the relationship structure between services, expressed as links Normalization structured to enable the creation of the topology 10 in any of the above requests characterized by module (3) such a system (1).

5. Data collection module (2) using any communication protocol, Communicating and exchanging data with the analysis module (4) and the recommendation module (7) 15 normalization module (3) configured to perform a system like any of the above characterized claims (1).

6. Event logs transmitted by the data collection module (2) are in common area 20 JSON data allows data to be represented in a way that includes names and data types. normalization structured to enable conversion to the specified format in any of the above requests characterized by module (3) such a system (1).

7. Event logs should include time information and log characteristics reflecting error conditions. Transaction ID and dependency refers to inter-service call relationships. by matching the information belonging to the same process or the same service chain the temporal ordering of records and interaction between services To identify patterns, time series event data were analyzed over 30 consecutive weeks. by processing in this way, using an artificial intelligence model in the form of an LSTM. 14 relationships between events: time interval, type of relationship, and event intensity. transforming it into a multidimensional correlation output that includes its dimensions characterized by the normalization module (3) structured on a system like any of the above requests (1).

8. Using any communication protocol, normalization module (3) and to communicate with the alert module (5) and exchange data The above is characterized by the structured analysis module (4) a system like any of the requests (1).

9. Indicates whether each software service is available. Accessibility information refers to the time it takes for operations on the service to complete. information about the delay and failures that occur during the processes metric measurements in the form of error rate information showing the ratio by evaluating how these metrics represent the service operational status. 15 for each service based on a weighted scoring model according to their levels calculating a quantitative health score, the calculated health scores over time increasing and decreasing trends within it compared to past score values by examining the overall operational status of the service comparatively a health trend reflecting a tendency towards continuity, deterioration, or improvement 20 characterized by the analysis module (4) which is structured to produce the output a system like any of the above requests (1).

10. Using any communication protocol, the analysis module (4) and Communicating with the visualization module (6) and exchanging data 25 characterized by the warning module (5) configured to perform a system like any of the above requests (1).

11. Analysis module (4) regarding the operational status of the services produced and errors occurring, processing time extended or service access 30 Warning logs indicating interruption situations; the time the warning occurred, A warning based on the type of situation it represents and the service information it belongs to. calculating the similarity between the records using the cosine similarity method and the similarity values ​​obtained and the occurrence of warnings over time alerts configured to group their frequencies under a single alert log any of the above requests characterized by module (5) such a system (1).

12. Alert logs that have been closed or ignored by users. by monitoring and suppressing repetitive and / or insignificant stimuli 10 noise-reduced and prioritized warning logs Creating notification visualizations to deliver notifications to the user. Characterized by the warning module (5) configured to transfer to module (6) a system like any of the above-mentioned requests (1).

13. Using the event propagation information obtained as a result of analysis and warning, 15 the impact a fault, slowdown, or interruption event has on services structured to present the chain to the user on a topology graph from the above requests characterized by the visualization module (6) a system like any other (1).

14. Normalization module (3) between event logs organized by graph generated according to the temporal sequence of service calls circulation and artificial intelligence extraction of data obtained from this circulation The resulting root cause identification output, the depth of the calls, 25 calculated according to the frequency of occurrence and the level of spread of the delay effect A graphical structure representing the relationship pattern between services using radius of influence information. to express visually through differences in color and position the above is characterized by the structured visualization module (6) a system like any of the requests (1). 16 15. Historical log records organized by the normalization module (3) and extracting event sequences derived from these records, on log records referring to recurring error patterns and unusual behaviors with the suggestion module (7) structured to detect anomalies A system like any of the above characterized claims 5 (1).

16. Classifying the identified error patterns as error occurrence topologies. Each failure occurrence topology corresponds to a chaos test scenario. matching and performing this matching process in time series log 10 an LSTM-based artificial intelligence that transforms patterns into a representation space a proposal structured to be implemented using a learning model any of the above requests characterized by module (7) such a system (1).

17. For each chaos test proposal created, the proposal's impact on the target system... The uncertainty of the impact it may create can be determined using the Monte Carlo Dropout method. to calculate the Chaos Impact Score based on this uncertainty rate To generate and prioritize testing chaos test scenarios with high CIS values. to define as a scenario and prioritize chaos test proposals 20 by presenting the feedback to the user with the visualization module (6) configured to share with the feedback module (9) for monitoring any of the above requests characterized by the suggestion module (7) a system like one of them (1).

18. CI / CD line records for software update processes, software Explanation of source code changes made during the development process. Code change log messages, which represent the records, in the CI / CD process. tests that record the results of failed tests failure logs and 30 errors that occurred during the transfer of the software to the target system The dataset retrieves distribution error records, which represent records of errors. 17 risk assessment module (8) structured to make it into a system like any of the above characterized claims (1).

19. Code complexity is measured by the number of bugs found in previous versions. Repetition and similarity between the bug reports that appeared in the current version. error similarity metric representing the relationship and run for the relevant version the scope of the tests, the proportion of code sections tested, and the test results. to generate a test coverage metric that shows the rollout across the entire release 10 characterized by the structured risk assessment module (8) a system like any of the above requests (1).

20. Percentage of code sections covered by the tests and version of test results. test criteria that express its widespread adoption, of the relevant software version 15 as an indicator representing the working status on the target system. to evaluate, based on these criteria, as an indicator of version stability calculating the so-called RSI value, the calculated RSI value to the dashboard to be presented by the visualization module (6) characterized by the risk assessment module (8) structured to transfer a system like any of the above requests (1). 20 21. Communicate with the suggestion module (7) using any communication protocol. feedback structured for establishing and exchanging data any of the above requests characterized by module (9) such a system (1). 25 22. Regarding acceptance or rejection actions performed by the user. information and the user performing these operations within the target system to receive role information that expresses the area of ​​duty and responsibility and to receive this feedback 30 configured to share notification data with the suggestion module (7) 18 any of the above requests characterized by the notification module (9) a system like one of them (1).

23. Evaluation of different user groups based on the role information received. To be able to distinguish the criteria, 5 for the Site Reliability Engineering role. Service latency and availability values, error for developer role. code that shows the density and location of errors in the source code stack traces, unauthorized access attempts for the security role, and unusual features representing access anomalies that describe access behaviors back 10 structured to convert into role-based representation vectors any of the above requests characterized by the notification module (9) a system like one of them (1).

24. The acceptance or rejection actions performed by users are created. Role-based representation vectors and reinforcement-based feedback 15 by including the weights of these representation vectors in the loop to ensure its updating and as a result of this update the suggestion module (7) Chaos test suggestions generated by the program, corresponding to the user's role. to enable redefinition based on feature weights The above 20 is characterized by the structured feedback module (9). a system like any of the requests (1). 30