Failure node identification method and device of financial clearing system, and electronic equipment
By pre-processing and abnormal detection of the operation and maintenance data of the financial clearing system, combining topological relationship diagrams and fault identification models, fault nodes are identified and located, the problem of low identification efficiency in the existing technology is solved, and efficient fault handling and business continuity is achieved.
Patent Information
- Application Number
- CN202411873642.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, the fault identification method of financial clearing system based on human experience and single-point monitoring is low in recognition efficiency, making it difficult to meet the business continuity and operation and maintenance efficiency requirements of distributed financial clearing systems.
By collecting operation and maintenance data of the financial clearing system, pre-processing is performed to obtain the operation and maintenance feature vector, and input it into the abnormality detection model for preliminary detection. If an exception is detected, the topology diagram is called for updates, a real-time operation and maintenance view is built, and the fault nodes are identified based on the graph recognition algorithm through the fault identification model.
It realizes rapid and accurate positioning of fault nodes, effectively shortens fault processing time, ensures business continuity of financial clearing system, and improves fault identification efficiency.
Smart Images

Figure CN119961820A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence or other related technical fields, and in particular to a method for identifying a faulty node in a financial clearing system, a device thereof, and an electronic device. Background Art
[0002] With the rapid development of financial technology and the continuous advancement of digital transformation, the clearing business system of financial institutions is undergoing a transformation from a traditional centralized architecture to a distributed architecture and a microservice architecture. In the traditional centralized architecture, financial clearing business mainly relies on the host system for processing, the system structure is relatively simple, and the operation and maintenance work is relatively centralized. However, with the surge in business volume and the increase in system complexity, this architecture has gradually shown performance bottlenecks and insufficient scalability. In order to meet these challenges, a distributed technology platform was adopted and core applications were gradually migrated to a microservice architecture. The microservice architecture splits a single application into multiple small, independent services, each of which can be deployed and expanded independently, thereby significantly improving the flexibility and response speed of the system. However, this architecture also leads to a multiple increase in the total amount of server resources, and the complexity and difficulty of system operation have increased significantly. This transformation not only greatly improves the system's processing power and scalability, but also brings new challenges, especially in system operation and maintenance and fault handling.
[0003] In the related technologies, the fault identification method of performing anomaly detection on the financial clearing system based on the work experience of operation and maintenance personnel or based on single-point monitoring has low processing efficiency and is difficult to meet the digital transformation of the financial clearing system. For large-scale distributed financial clearing systems, it is difficult to ensure the business continuity and operation and maintenance efficiency of the financial clearing system.
[0004] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0005] The embodiment of the present invention provides a method and device for identifying faulty nodes in a financial clearing system, and an electronic device thereof, so as to at least solve the technical problem of low identification efficiency in the related art of the fault identification method of the financial clearing system based on human experience and single-point monitoring.
[0006] According to one aspect of an embodiment of the present invention, a method for identifying faulty nodes of a financial clearing system is provided, comprising: collecting operation and maintenance data of the financial clearing system, and preprocessing the operation and maintenance data to obtain an operation and maintenance feature vector, wherein the financial clearing system adopts a distributed architecture; inputting the operation and maintenance feature vector into an anomaly detection model, and outputting an anomaly detection result, wherein the anomaly detection model is a model pre-constructed based on a generative adversarial network for detecting operation and maintenance anomalies of the financial clearing system; when the anomaly detection result indicates that the financial clearing system has an anomaly, calling a topological relationship graph of the financial clearing system, and updating the topological relationship graph based on the operation and maintenance data to obtain a real-time operation and maintenance view of the financial clearing system, wherein the topological relationship graph is used to record the topological relationship between applications and network resources in the financial clearing system; inputting the real-time operation and maintenance view into a fault identification model, and outputting a faulty node, wherein the fault identification model is a model pre-constructed based on a graph recognition algorithm for determining a faulty node.
[0007] Optionally, before collecting the operation and maintenance data of the financial clearing system, it also includes: obtaining all applications and network resources in the financial clearing system, and constructing an initial topological relationship graph using the applications and the network resources as nodes; obtaining the mutual access relationships between the applications in the financial clearing system, and constructing a horizontal topological relationship graph based on the mutual access relationships between the applications; obtaining the mutual dependence relationships between the network resources in the financial clearing system, and constructing a vertical topological relationship graph based on the mutual dependence relationships between the network resources; and merging the horizontal topological relationship graph and the vertical topological relationship graph on the basis of the initial topological relationship graph to obtain a topological relationship graph of the financial clearing system.
[0008] Optionally, after outputting the faulty node, the method further includes: traversing the nodes in the topological relationship graph based on the faulty node to obtain associated nodes that have a topological relationship with the faulty node; and constructing a fault propagation link based on the faulty node and the associated nodes.
[0009] Optionally, after constructing a fault propagation link based on the fault node and the associated node, it also includes: calling a scenario recognition rule base to identify the fault scenario of the financial clearing system based on the fault propagation link and the scenario recognition rule base; querying a decision database based on the fault scenario to obtain a solution strategy corresponding to the fault scenario; generating a fault identification report based on the fault node, the fault propagation link, the fault scenario and the solution strategy, and sending the fault identification report to the operation and maintenance terminal.
[0010] Optionally, the fault identification model is pre-constructed, and the steps of constructing the fault identification model include: obtaining historical operation and maintenance data of the financial clearing system within a historical time period, and constructing a historical operation and maintenance view for the financial clearing system based on the historical operation and maintenance data; configuring fault labels for the historical operation and maintenance view, wherein the fault labels are used to mark fault nodes in the historical operation and maintenance view; constructing sample data based on the historical operation and maintenance view and the fault labels of the historical operation and maintenance view, and dividing the sample data into a training set and a test set; constructing an initial fault identification model based on a graph recognition algorithm, and using the training set to train the initial fault identification model to obtain the trained fault identification model; using the test set to test the trained fault identification model to obtain a test result, and when the test result indicates that the recognition accuracy of the fault identification model is greater than or equal to a preset accuracy threshold, it is determined that the model training is completed to obtain the final fault identification model.
[0011] Optionally, the step of preprocessing the operation and maintenance data to obtain an operation and maintenance feature vector includes: cleaning the operation and maintenance data to obtain cleaned operation and maintenance data; performing feature extraction on the cleaned operation and maintenance data to obtain operation and maintenance features, wherein feature extraction on the cleaned operation and maintenance data includes at least one of the following: calculating statistical features of the operation and maintenance data, extracting time series features from the operation and maintenance data, and identifying event features from the operation and maintenance data; encoding the operation and maintenance features to obtain an operation and maintenance feature vector.
[0012] Optionally, the operation and maintenance data includes at least one of the following: business tracking link data, performance monitoring data of the financial clearing system, log data of the application and the network resources, real-time transaction flow data, and system alarm information.
[0013] According to another aspect of an embodiment of the present invention, a fault node identification device for a financial clearing system is further provided, comprising: a collection unit, configured to collect operation and maintenance data of the financial clearing system, and preprocess the operation and maintenance data to obtain an operation and maintenance feature vector, wherein the financial clearing system adopts a distributed architecture; a first output unit, configured to input the operation and maintenance feature vector into an anomaly detection model, and output an anomaly detection result, wherein the anomaly detection model is a model pre-constructed based on a generative adversarial network for detecting operation and maintenance anomalies of the financial clearing system; a calling unit, configured to call a topological relationship graph of the financial clearing system when the anomaly detection result indicates that the financial clearing system has an anomaly, and update the topological relationship graph based on the operation and maintenance data to obtain a real-time operation and maintenance view of the financial clearing system, wherein the topological relationship graph is used to record the topological relationship between applications and network resources in the financial clearing system; a second output unit, configured to input the real-time operation and maintenance view into a fault identification model, and output a faulty node, wherein the fault identification model is a model pre-constructed based on a graph recognition algorithm for determining a faulty node.
[0014] Optionally, the fault node identification device of the financial clearing system also includes: a first construction module, used to obtain all applications and network resources in the financial clearing system, and use the applications and the network resources as nodes to construct an initial topological relationship graph; a second construction module, used to obtain the mutual access relationship between the applications in the financial clearing system, and construct a horizontal topological relationship graph based on the mutual access relationship between the applications; a third construction module, used to obtain the mutual dependence relationship between the network resources in the financial clearing system, and construct a vertical topological relationship graph based on the mutual dependence relationship between the network resources; a first fusion module, used to fuse the horizontal topological relationship graph and the vertical topological relationship graph on the basis of the initial topological relationship graph to obtain the topological relationship graph of the financial clearing system.
[0015] Optionally, the fault node identification device of the financial clearing system also includes: a first traversal module, used to traverse the nodes in the topological relationship graph based on the faulty node to obtain associated nodes that have a topological relationship with the faulty node; and a fourth construction module, used to construct a fault propagation link based on the faulty node and the associated nodes.
[0016] Optionally, the fault node identification device of the financial clearing system also includes: a first identification module, used to call a scenario identification rule base, and identify the fault scenario of the financial clearing system based on the fault propagation link and the scenario identification rule base; a first acquisition module, used to query a decision database based on the fault scenario, and obtain a solution strategy corresponding to the fault scenario; a first generation module, used to generate a fault identification report based on the fault node, fault propagation link, fault scenario and the solution strategy, and send the fault identification report to the operation and maintenance terminal.
[0017] Optionally, the fault node identification device of the financial clearing system also includes: a second acquisition module, used to acquire historical operation and maintenance data of the financial clearing system within a historical time period, and construct a historical operation and maintenance view for the financial clearing system based on the historical operation and maintenance data; a first configuration module, used to configure fault labels for the historical operation and maintenance view, wherein the fault labels are used to mark fault nodes in the historical operation and maintenance view; a fifth construction module, used to construct sample data based on the historical operation and maintenance view and the fault labels of the historical operation and maintenance view, and divide the sample data into a training set and a test set; a first training module, used to construct an initial fault identification model based on a graph recognition algorithm, and use the training set to train the initial fault identification model to obtain the trained fault identification model; a first testing module, used to use the test set to test the trained fault identification model to obtain a test result, and when the test result indicates that the recognition accuracy of the fault identification model is greater than or equal to a preset accuracy threshold, it is determined that the model training is completed to obtain the final fault identification model.
[0018] Optionally, the acquisition unit includes: a first cleaning module, used to clean the operation and maintenance data to obtain cleaned operation and maintenance data; a first extraction module, used to extract features from the cleaned operation and maintenance data to obtain operation and maintenance features, wherein feature extraction from the cleaned operation and maintenance data includes at least one of the following: calculating statistical features of the operation and maintenance data, extracting time series features from the operation and maintenance data, and identifying event features from the operation and maintenance data; a first encoding module, used to encode the operation and maintenance features to obtain an operation and maintenance feature vector.
[0019] Optionally, the operation and maintenance data includes at least one of the following: business tracking link data, performance monitoring data of the financial clearing system, log data of the application and the network resources, real-time transaction flow data, and system alarm information.
[0020] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is also provided, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned methods for identifying faulty nodes in a financial clearing system.
[0021] According to another aspect of an embodiment of the present invention, there is also provided an electronic device, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement any one of the above-mentioned methods for identifying faulty nodes in a financial clearing system.
[0022] In the present application, the following steps are performed: collecting operation and maintenance data of a financial clearing system, and preprocessing the operation and maintenance data to obtain an operation and maintenance feature vector, wherein the financial clearing system adopts a distributed architecture, and inputs the operation and maintenance feature vector into an anomaly detection model, and outputs an anomaly detection result, wherein the anomaly detection model is a model pre-built based on a generative adversarial network for detecting operation and maintenance anomalies of the financial clearing system, and then, when the anomaly detection result indicates that there is an anomaly in the financial clearing system, calling the topological relationship graph of the financial clearing system, and updating the topological relationship graph based on the operation and maintenance data to obtain a real-time operation and maintenance view of the financial clearing system, wherein the topological relationship graph is used to record the topological relationship between applications and network resources in the financial clearing system, and finally inputting the real-time operation and maintenance view into a fault identification model, and outputting a fault node, wherein the fault identification model is a model pre-built based on a graph recognition algorithm for determining a fault node.
[0023] In the present application, for a distributed financial clearing system, a topological relationship diagram is pre-constructed, and the complex topological relationships in the distributed system are recorded through the topological relationship diagram. In the process of continuous operation and maintenance of the financial clearing system, the operation and maintenance data of the system is obtained in real time, and preliminary anomaly detection is performed through an anomaly detection model. When an anomaly exists in the system, a real-time operation and maintenance view is constructed based on the operation and maintenance data and the topological relationship diagram, and the nodes in the operation and maintenance view are analyzed through a graph recognition algorithm model to identify faulty nodes, thereby achieving rapid and accurate positioning of faulty nodes, effectively shortening fault handling time, and ensuring the business continuity of the financial clearing system, achieving the technical effect of improving fault identification efficiency, and thereby solving the technical problem of low identification efficiency in the related technology of financial clearing system fault identification method based on human experience and single-point monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0025] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for identifying a faulty node in a financial clearing system is shown;
[0026] Figure 2 is a flow chart of an optional method for identifying a faulty node in a financial settlement system according to an embodiment of the present invention;
[0027] Figure 3 is an architecture diagram of a fault node identification system of an optional financial settlement system according to an embodiment of the present invention;
[0028] Figure 4 is a schematic diagram of an optional fault node identification device of a financial settlement system according to an embodiment of the present invention;
[0029] Figure 5 It is a hardware structure block diagram of an optional electronic device (or mobile device) for executing a faulty node identification method of a financial settlement system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0030] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0032] It should be noted that the faulty node identification method and device of the financial clearing system in the present application can be used in the field of artificial intelligence when anomaly detection and fault identification are performed on distributed systems based on artificial intelligence, and can also be used in any field other than the field of artificial intelligence when anomaly detection and fault identification are performed on distributed systems based on artificial intelligence. The present application does not limit the application field of the faulty node identification method and device of the financial clearing system.
[0033] It should be noted that the collected information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data are in compliance with relevant laws, regulations and standards, necessary confidentiality measures are taken, and public order and good customs are not violated, and corresponding operation entrances are provided for users to choose to authorize or refuse. For example, an interface is set up between this system and relevant users or institutions to provide users with corresponding operation entrances for users to choose to agree or refuse the results of automated decision-making; if the user chooses to refuse, the expert decision-making process will be entered.
[0034] The following embodiments of the present invention can be applied to various fault node identification systems / applications / devices of financial clearing systems. On the basis of ensuring the security and stability of the financial clearing system, the present invention realizes the full process automation and standardized processing of the clearing system exception processing based on the generative adversarial network and graph recognition algorithm, improves the efficiency of exception detection and fault identification, and based on the fault view and graph recognition algorithm, can not only effectively handle a single fault, but also, when multiple systems have related faults, can accurately match based on the operation and maintenance view, generate a fault link, realize fault tracing, improve fault handling efficiency, and ensure the continuity of clearing business.
[0035] The present invention is described in detail below in conjunction with various embodiments.
[0036] Embodiment 1
[0037] According to an embodiment of the present invention, an embodiment of a method for identifying a faulty node in a financial clearing system is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0038] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1The hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for identifying a faulty node in a financial clearing system is shown. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.
[0039] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0040] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the fault node identification method of the financial settlement system in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, the fault node identification method of the financial settlement system described above is realized. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0041] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0042] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0043] Under the above operating environment, this application provides Figure 2 The fault node identification method of the financial clearing system shown is implemented by the fault node identification system of the financial clearing system.
[0044] Figure 2 is a flow chart of an optional method for identifying a faulty node in a financial clearing system according to an embodiment of the present invention. Figure 2 As shown, the method comprises the following steps:
[0045] Step S201, collecting operation and maintenance data of the financial clearing system, and preprocessing the operation and maintenance data to obtain an operation and maintenance feature vector.
[0046] It should be noted that the financial clearing system refers to the software system used by financial institutions to process transactions such as transfers, payments, and settlements. With the exponential growth and complexity of the business volume of financial institutions, the traditional centralized architecture has been unable to meet the needs of high concurrency, low latency, and high reliability. Therefore, the financial clearing system has gradually transformed to a distributed architecture. By distributing tasks and data across multiple network nodes, the distributed system can improve the overall processing power, scalability, and fault tolerance of the system without increasing the burden on a single component. For financial clearing systems, the distributed architecture can better support the processing of massive transactions while ensuring the accuracy and security of transactions. Even if some components fail, the system can continue to operate.
[0047] It should be noted that while the distributed financial clearing system improves business processing capabilities and system scalability, it also brings unprecedented challenges to operation and maintenance. The distributed architecture decomposes a single system into multiple components or services that collaborate with each other. These components may be distributed on different physical servers or geographical locations, which increases the complexity of the system. Operation and maintenance personnel need to monitor and manage more service instances and network connections at the same time, making it difficult to intuitively understand the operating status of the entire system. At the same time, in a distributed system, a fault may occur in a certain service or component, but due to the dependencies between components, the fault may quickly spread to the entire system. This fault propagation path is often nonlinear and difficult to predict, which increases the difficulty of fault location and recovery.
[0048] Optionally, the operation and maintenance data includes at least one of the following: business tracking link data, performance monitoring data of the financial clearing system, log data of applications and network resources, real-time transaction flow data, and system alarm information.
[0049] It should be noted that when conducting real-time operation and maintenance of the financial clearing system, the monitoring center conducts real-time monitoring, health checks, and continuous monitoring of the operation of the main business processes of the financial clearing system, collects business tracking link data corresponding to the financial business, performance monitoring data of the financial clearing system, log data of applications and network resources, real-time transaction flow data, system alarm information, etc., and collects operation and maintenance data of the financial clearing system from multiple dimensions to provide a solid data foundation for subsequent fault identification.
[0050] It should be noted that preprocessing is the first step in operation and maintenance data analysis, and its purpose is to ensure data quality and improve the accuracy and efficiency of subsequent analysis. The preprocessing stage mainly includes three key links: data cleaning, feature extraction, and feature encoding.
[0051] Optionally, before collecting the operation and maintenance data of the financial clearing system, it also includes: obtaining all applications and network resources in the financial clearing system, and constructing an initial topological relationship graph using the applications and network resources as nodes; obtaining the mutual access relationships between the applications in the financial clearing system, and constructing a horizontal topological relationship graph based on the mutual access relationships between the applications; obtaining the mutual dependence relationships between the network resources in the financial clearing system, and constructing a vertical topological relationship graph based on the mutual dependence relationships between the network resources; and merging the horizontal topological relationship graph and the vertical topological relationship graph based on the initial topological relationship graph to obtain a topological relationship graph of the financial clearing system.
[0052] It should be noted that the topology diagram, as a visualization tool, can describe the logical relationship and physical connection between various applications and network resources in the financial clearing system, which helps to assist the fault identification model in understanding the system architecture and dependencies between components, so as to quickly locate faulty nodes and fault propagation links.
[0053] Specifically, when constructing the topology diagram, based on the distributed architecture of the financial clearing system, it is divided into horizontal topology and vertical topology. Applications and services mainly focus on the implementation of business logic. Applications and microservices interact through API calls, message passing, etc., forming a horizontal logical association. Network resources, such as servers, storage devices, network devices, etc., constitute the physical infrastructure of the distributed system. These resources are the basis for the operation of applications and services, supporting the execution of business logic from the physical level. The topology network constructed based on horizontal topology and vertical topology can quickly locate the faulty node and trace the fault when the system fails.
[0054] The specific steps of building a topology diagram include: first, conduct a comprehensive system inventory of the financial clearing system, identify all applications and services in the financial clearing system, and the network resources (such as servers, storage, network devices, etc.) that support the operation of these applications. These applications and services and network resources are used as nodes in the diagram. Then, based on the above nodes, an initial topology diagram is created, which reflects the distribution of applications and services on network resources. Build a horizontal topology diagram. The horizontal topology diagram focuses on the mutual access relationship between applications and services. By monitoring network traffic and request logs, analyzing the call paths and dependencies between applications, and building a horizontal topology diagram at the application level, it helps to identify the key paths and possible fault propagation modes in the business flow. Build a vertical topology diagram. The vertical topology diagram focuses on the mutual dependencies between network resources. This includes the server's dependence on storage, the connection relationship between network devices, etc. The vertical topology diagram can reveal potential bottlenecks and failure points at the resource level. Finally, the horizontal topology diagram and the vertical topology diagram are integrated to create a comprehensive topology diagram of the financial clearing system. The diagram includes not only the relationship between applications and services, but also the network resources they depend on and their connections with each other. This comprehensive view can provide more accurate context information for fault identification.
[0055] Optionally, the step of preprocessing the operation and maintenance data to obtain an operation and maintenance feature vector includes: cleaning the operation and maintenance data to obtain cleaned operation and maintenance data; performing feature extraction on the cleaned operation and maintenance data to obtain operation and maintenance features, wherein the feature extraction on the cleaned operation and maintenance data includes at least one of the following: calculating statistical features of the operation and maintenance data, extracting time series features from the operation and maintenance data, and identifying event features from the operation and maintenance data; and encoding the operation and maintenance features to obtain an operation and maintenance feature vector.
[0056] It should be noted that when preprocessing the data, the data is first cleaned to remove noise and irrelevant data in the financial data, fill in missing values, and correct outliers. Then the financial data is feature extracted, including: extraction of statistical features, such as calculating mean, standard deviation, maximum value, minimum value, etc., which reflect static change features of system resource usage and performance indicators; it also includes extraction of time series features, such as trends, cyclical changes, seasonal fluctuations, etc., which reflect dynamic change trends of the system. In addition, feature extraction can also include extraction of event features, which can be based on specific events in the system log. Event features, such as system startup, shutdown, error messages, etc. Feature extraction can also include extraction of relationship features, based on the horizontal topological relationship between applications and the vertical topological relationship between resources, extract dependency and association information between nodes to help identify fault propagation paths.
[0057] It should be noted that after the financial features are extracted based on the financial data, it is necessary to vector encode the financial features to convert the feature data into coded data recognizable by the model to obtain the operation and maintenance feature vector.
[0058] Step S202: input the operation and maintenance feature vector into an anomaly detection model and output an anomaly detection result.
[0059] It should be noted that the preprocessed operation and maintenance feature vector is first input into the anomaly detection model for preliminary anomaly detection. If the detection determines that there is no anomaly in the financial clearing system, there is no need for subsequent fault definition, saving computing resources. On the contrary, if the initial detection shows that there is an anomaly in the financial clearing system, it is necessary to further define the fault, locate the fault node, trace the fault propagation path, and avoid clearing business anomalies caused by system failures.
[0060] It should be noted that the anomaly detection model is constructed using a generative adversarial network, which includes a generator and a discriminator. When training the model, historical operation and maintenance data under normal conditions is used as a training data set, and the generator and the discriminator are trained alternately. The generator learns to simulate the normal data distribution, while the discriminator learns to distinguish between generated data and real data. The network parameters are optimized through multiple rounds of iterations until the model reaches the preset convergence criteria. When performing real-time anomaly detection, the real-time operation and maintenance data is input into the anomaly detection model. The generative adversarial network will analyze the input real-time operation and maintenance data, and use the discriminator to distinguish the difference between the real-time operation and maintenance data and the normal state operation and maintenance data, so as to determine whether there is an anomaly in the financial clearing system and obtain the anomaly detection result.
[0061] Step S203, when the anomaly detection result indicates that the financial clearing system has an anomaly, the topology diagram of the financial clearing system is called, and the topology diagram is updated based on the operation and maintenance data to obtain a real-time operation and maintenance view of the financial clearing system.
[0062] It should be noted that when the anomaly detection results indicate that there are anomalies in the financial clearing system, the operation and maintenance view is generated using real-time operation and maintenance data and a pre-built topological relationship diagram. The operation and maintenance view can be used for subsequent fault node identification based on the fault identification model, and can also provide a visual display of the fault propagation path and impact range, thereby supporting more accurate fault identification.
[0063] It should be noted that when building a real-time operation and maintenance view, the pre-built topology relationship diagram is called, and the collected real-time operation and maintenance data is integrated with the topology relationship diagram to obtain an operation and maintenance view, so that the operation and maintenance view not only records the topological relationship of each system node, but also records the real-time operation and maintenance status and detailed status information of each system node.
[0064] Step S204: input the real-time operation and maintenance view into the fault identification model and output the fault node.
[0065] It should be noted that after obtaining the real-time operation and maintenance view, the real-time operation and maintenance view is input into the pre-built fault identification model. The fault identification model traverses and analyzes the nodes in the real-time operation and maintenance view based on the graph recognition algorithm, quickly locates the faulty nodes, and shortens the fault detection and response time.
[0066] Optionally, the fault identification model is pre-constructed, and the steps of constructing the fault identification model include: obtaining historical operation and maintenance data of the financial clearing system within a historical time period, and constructing a historical operation and maintenance view for the financial clearing system based on the historical operation and maintenance data; configuring fault labels for the historical operation and maintenance views, wherein the fault labels are used to mark fault nodes in the historical operation and maintenance views; constructing sample data based on the historical operation and maintenance views and the fault labels of the historical operation and maintenance views, and dividing the sample data into a training set and a test set; constructing an initial fault identification model based on a graph recognition algorithm, and using the training set to train the initial fault identification model to obtain a trained fault identification model; using the test set to test the trained fault identification model to obtain test results, and when the test results indicate that the recognition accuracy of the fault identification model is greater than or equal to a preset accuracy threshold, determining that the model training is complete to obtain a final fault identification model.
[0067] It should be noted that the core of the fault identification model is the graph recognition algorithm, which can extract patterns and features from complex graph structures, process the features of nodes and edges in the graph structure, and thus identify system nodes with faults. The fault identification model is pre-built. During the model training phase, historical operation and maintenance data and fault events need to be used as training sets. Through supervised learning, the model learns the pattern changes from normal state to fault state. The training data should contain the structural information of the topology graph, node status information, and fault event markers to ensure that the model can accurately identify faulty nodes. Specifically, first, the historical operation and maintenance data of the financial clearing system within the historical time period can be read from the database, and a historical operation and maintenance view is constructed based on the historical operation and maintenance data and a pre-generated topological relationship diagram. Then, fault labels are configured for the historical operation and maintenance view, and the fault nodes in the historical operation and maintenance view are marked. Then, sample data is constructed based on the historical operation and maintenance view and its fault labels, and the data is divided into a training set and a test set according to a pre-set division ratio. The training set is used to iteratively train the fault identification model to seek the optimal model parameters, and the test set is used to test the trained fault identification model to determine the fault identification model's recognition accuracy for the fault nodes, until the accuracy of the fault identification model is greater than the accuracy threshold, and it is determined that the fault identification model training is complete.
[0068] Optionally, after outputting the faulty node, the method further includes: traversing nodes in the topology relationship graph based on the faulty node to obtain associated nodes that have a topological relationship with the faulty node; and constructing a fault propagation link based on the faulty node and the associated nodes.
[0069] It should be noted that after determining the specific faulty node, the associated nodes that have a topological relationship with the faulty node are obtained through the topological relationship diagram. Since each system node in the distributed system executes financial services by interacting with other nodes, a failure of a system node may cause other system nodes that have an interactive relationship or dependency relationship with the faulty node to also fail, thereby affecting the execution of financial services. The faulty node is traced to obtain the fault propagation link. The fault propagation link may involve a service call chain, a data flow path, or a network connection path. The fault propagation link can clearly show the path of the fault propagating from one node to another, as well as the impact of each node. By constructing a fault propagation link, operation and maintenance personnel can quickly identify the impact range of the fault, accurately locate the source of the fault and determine the affected components, providing a basis for rapid response to fault recovery, while also helping to reduce the risk of further spread of the fault.
[0070] Optionally, after constructing a fault propagation link based on the fault node and associated nodes, it also includes: calling a scenario recognition rule base to identify the fault scenario of the financial clearing system based on the fault propagation link and the scenario recognition rule base; querying a decision database based on the fault scenario to obtain a solution strategy corresponding to the fault scenario; generating a fault identification report based on the fault node, fault propagation link, fault scenario and solution strategy, and sending the fault identification report to the operation and maintenance terminal.
[0071] It should be noted that after determining the fault node and obtaining the fault propagation link, the corresponding fault resolution strategy can be automatically matched based on the fault node and the fault propagation link to assist the operation and maintenance personnel in fault recovery. Specifically, the pre-created scenario identification rule library can be called, and the scenario identification rule library can be queried according to the fault propagation link to identify the current fault scenario of the financial clearing system. The decision database can be queried according to the fault scenario to obtain the resolution strategy corresponding to the fault scenario, thereby generating a fault identification report based on the fault node, fault propagation path, fault scenario and resolution strategy, and sending it to the operation and maintenance terminal, where the operation and maintenance personnel execute the corresponding resolution strategy to enable the financial clearing system to quickly recover from the fault state and improve the operation and maintenance efficiency and business continuity of the financial clearing system.
[0072] Through the above steps, the operation and maintenance data of the financial clearing system is collected, and the operation and maintenance data is preprocessed to obtain the operation and maintenance feature vector, wherein the financial clearing system adopts a distributed architecture, and the operation and maintenance feature vector is input into the anomaly detection model, and the anomaly detection result is output, wherein the anomaly detection model is a model pre-built based on a generative adversarial network for detecting operation and maintenance anomalies of the financial clearing system, and then when the anomaly detection result indicates that there is an anomaly in the financial clearing system, the topological relationship graph of the financial clearing system is called, and the topological relationship graph is updated based on the operation and maintenance data to obtain a real-time operation and maintenance view of the financial clearing system, wherein the topological relationship graph is used to record the topological relationship between applications and network resources in the financial clearing system, and finally the real-time operation and maintenance view is input into the fault identification model, and the fault node is output, wherein the fault identification model is a model pre-built based on a graph recognition algorithm for determining the fault node.
[0073] In this embodiment, for a distributed financial clearing system, a topological relationship diagram is pre-constructed, and the complex topological relationships in the distributed system are recorded through the topological relationship diagram. In the process of continuous operation and maintenance of the financial clearing system, the operation and maintenance data of the system is obtained in real time, and preliminary anomaly detection is performed through an anomaly detection model. When an anomaly exists in the system, a real-time operation and maintenance view is constructed based on the operation and maintenance data and the topological relationship diagram, and the nodes in the operation and maintenance view are analyzed through a graph recognition algorithm model to identify faulty nodes, thereby achieving rapid and accurate positioning of faulty nodes, effectively shortening fault handling time, and ensuring the business continuity of the financial clearing system, achieving the technical effect of improving fault identification efficiency, and further solving the technical problem of low identification efficiency in the related technology of financial clearing system fault identification method based on human experience and single-point monitoring.
[0074] Another optional specific implementation is described in detail below.
[0075] Figure 3 is an architecture diagram of a fault node identification system of an optional financial settlement system according to an embodiment of the present invention, such as Figure 3 As shown, the fault node identification system of the financial clearing system mainly includes: anomaly detection module, real-time operation and maintenance view construction module and fault definition module. The anomaly detection module and the real-time operation and maintenance view construction module interact with the financial clearing system, and can obtain node construction topology relationship diagrams and collect life and death indicators from the financial clearing system. The anomaly detection module transmits the anomaly detection results to the real-time operation and maintenance view construction module. The real-time operation and maintenance view construction module constructs a real-time operation and maintenance view based on real-time operation and maintenance data and relationship topology diagrams. The fault definition module analyzes the real-time operation and maintenance view and finally outputs a fault identification report generated based on the faulty node and fault propagation path. Specifically,
[0076] The financial clearing system refers to the application system used by financial institutions to handle clearing business, which is a collection of hardware and software including IT infrastructure, network equipment, and tool software.
[0077] The fault node identification system of the financial clearing system also includes a monitoring center, which is responsible for real-time monitoring, health checks, and continuous monitoring of the operation of the main business processes of the financial clearing system, and collecting full-link data, performance data, log data, transaction flow, warning information and other operation and maintenance data corresponding to the financial business.
[0078] During the real-time operation and maintenance process, the anomaly detection module uses an anomaly detection model based on a generative adversarial network to perform preliminary detection of abnormal conditions in the financial clearing system based on operation and maintenance data, and extract life and death indicators in the operation and maintenance data. The life and death indicators are used to determine whether there are abnormalities in the financial clearing system. If there are abnormalities, the real-time operation and maintenance view construction module and the fault definition module will perform subsequent fault definition. If there are no abnormalities, there is no need to perform fault definition, saving computing resources.
[0079] The fault node identification system of the financial clearing system also includes a scenario customization center. The scenario customization center can call various operation and maintenance services provided by the operation and maintenance center according to the emergency manual or abnormal handling plan, and pre-formulate fault scenarios and fault resolution strategies. At the same time, the scenario customization center builds a relationship topology diagram adapted to the financial clearing system, that is, the relationship between each sub-module in the financial clearing system is displayed in the form of a view. Constructing two topological relationships of the clearing system, horizontal topology (inter-application access relationship) and vertical topology (interdependence between network resources), lays the foundation for subsequent fault definition.
[0080] The real-time operation and maintenance view construction module constructs a real-time operation and maintenance view according to the anomaly detection results output by the anomaly detection module. The real-time operation and maintenance view is generated based on a pre-built relationship topology diagram and real-time operation and maintenance data of the financial clearing system.
[0081] The real-time operation and maintenance view construction module transmits the real-time operation and maintenance view to the fault definition module. The fault definition module uses a fault identification model built based on a graph recognition algorithm to identify the real-time operation and maintenance view, outputs the fault node, and then combines the relationship topology graph to trace the fault based on the fault node to obtain the fault propagation path.
[0082] The fault node identification system of the financial clearing system also includes a scenario verification center. The scenario verification center identifies the fault scenario through the fault propagation path, matches the pre-configured emergency plan elements (corresponding to the above-mentioned solution strategy) based on the fault scenario, and finally generates a fault identification report based on the fault node, fault propagation path, fault scenario and emergency plan elements.
[0083] On the basis of ensuring the security and stability of the financial clearing system, the embodiments of the present invention realize the full-process automation and standardization of the clearing system exception handling based on the generative adversarial network and graph recognition algorithm, improve the efficiency of anomaly detection and fault identification, and based on the fault view and graph recognition algorithm, not only can a single fault be effectively handled, but also when multiple systems have related faults, accurate matching based on the operation and maintenance view can be achieved to generate a fault link, realize fault tracing, improve fault handling efficiency, and ensure the continuity of clearing business.
[0084] The following is a detailed description in conjunction with another embodiment.
[0085] Embodiment 2
[0086] A fault node identification device of a financial clearing system provided in this embodiment includes multiple implementation units, each implementation unit corresponds to each implementation step in the above-mentioned embodiment 1. Its specific implementation method and beneficial effects can refer to the above-mentioned method embodiment and will not be repeated here.
[0087] Figure 4 is a schematic diagram of an optional fault node identification device of a financial settlement system according to an embodiment of the present invention, such as Figure 4 As shown, the fault node identification device of the financial settlement system may include: a collection unit 41, a first output unit 42, a calling unit 43, and a second output unit 44, wherein:
[0088] A collection unit 41 is used to collect operation and maintenance data of the financial clearing system and pre-process the operation and maintenance data to obtain an operation and maintenance feature vector, wherein the financial clearing system adopts a distributed architecture;
[0089] A first output unit 42 is used to input the operation and maintenance feature vector into an anomaly detection model and output an anomaly detection result, wherein the anomaly detection model is a model pre-built based on a generative adversarial network for detecting operation and maintenance anomalies of a financial clearing system;
[0090] The calling unit 43 is used to call the topological relationship diagram of the financial clearing system when the abnormality detection result indicates that the financial clearing system has an abnormality, and update the topological relationship diagram based on the operation and maintenance data to obtain a real-time operation and maintenance view of the financial clearing system, wherein the topological relationship diagram is used to record the topological relationship between the application programs and network resources in the financial clearing system;
[0091] The second output unit 44 is used to input the real-time operation and maintenance view into the fault identification model and output the fault node, wherein the fault identification model is a model pre-built based on the graph recognition algorithm for determining the fault node.
[0092] The above-mentioned fault node identification device of the financial clearing system collects operation and maintenance data of the financial clearing system through the collection unit 41, and pre-processes the operation and maintenance data to obtain an operation and maintenance feature vector, wherein the financial clearing system adopts a distributed architecture; the operation and maintenance feature vector is input into the anomaly detection model through the first output unit 42, and the anomaly detection result is output, wherein the anomaly detection model is a model pre-built based on the generative adversarial network for detecting operation and maintenance anomalies of the financial clearing system; when the anomaly detection result indicates that the financial clearing system has an anomaly, the topological relationship graph of the financial clearing system is called through the calling unit 43, and the topological relationship graph is updated based on the operation and maintenance data to obtain a real-time operation and maintenance view of the financial clearing system, wherein the topological relationship graph is used to record the topological relationship between applications and network resources in the financial clearing system; the real-time operation and maintenance view is input into the fault identification model through the second output unit 43, and the fault node is output, wherein the fault identification model is a model pre-built based on the graph recognition algorithm for determining the fault node.
[0093] In this embodiment, for a distributed financial clearing system, a topological relationship diagram is pre-constructed, and the complex topological relationships in the distributed system are recorded through the topological relationship diagram. In the process of continuous operation and maintenance of the financial clearing system, the operation and maintenance data of the system is obtained in real time, and preliminary anomaly detection is performed through an anomaly detection model. When an anomaly exists in the system, a real-time operation and maintenance view is constructed based on the operation and maintenance data and the topological relationship diagram, and the nodes in the operation and maintenance view are analyzed through a graph recognition algorithm model to identify faulty nodes, thereby achieving rapid and accurate positioning of faulty nodes, effectively shortening fault handling time, and ensuring the business continuity of the financial clearing system, achieving the technical effect of improving fault identification efficiency, and further solving the technical problem of low identification efficiency in the related technology of financial clearing system fault identification method based on human experience and single-point monitoring.
[0094] Optionally, the fault node identification device of the financial clearing system also includes: a first construction module, used to obtain all applications and network resources in the financial clearing system, and use the applications and network resources as nodes to construct an initial topological relationship graph; a second construction module, used to obtain the mutual access relationship between each application in the financial clearing system, and construct a horizontal topological relationship graph based on the mutual access relationship between the applications; a third construction module, used to obtain the mutual dependence relationship between each network resource in the financial clearing system, and construct a vertical topological relationship graph based on the mutual dependence relationship between the network resources; a first fusion module, used to fuse the horizontal topological relationship graph and the vertical topological relationship graph on the basis of the initial topological relationship graph to obtain a topological relationship graph of the financial clearing system.
[0095] Optionally, the fault node identification device of the financial clearing system also includes: a first traversal module, used to traverse the nodes in the topological relationship graph based on the fault node, and obtain associated nodes that have a topological relationship with the fault node; and a fourth construction module, used to construct a fault propagation link based on the fault node and the associated nodes.
[0096] Optionally, the fault node identification device of the financial clearing system also includes: a first identification module, used to call the scenario identification rule base, and identify the fault scenario of the financial clearing system based on the fault propagation link and the scenario identification rule base; a first acquisition module, used to query the decision database based on the fault scenario, and obtain the solution strategy corresponding to the fault scenario; a first generation module, used to generate a fault identification report based on the fault node, fault propagation link, fault scenario and solution strategy, and send the fault identification report to the operation and maintenance terminal.
[0097] Optionally, the fault node identification device of the financial clearing system also includes: a second acquisition module, used to acquire historical operation and maintenance data of the financial clearing system within a historical time period, and construct a historical operation and maintenance view for the financial clearing system based on the historical operation and maintenance data; a first configuration module, used to configure fault labels for the historical operation and maintenance view, wherein the fault labels are used to mark fault nodes in the historical operation and maintenance view; a fifth construction module, used to construct sample data based on the historical operation and maintenance view and the fault labels of the historical operation and maintenance view, and divide the sample data into a training set and a test set; a first training module, used to construct an initial fault identification model based on a graph recognition algorithm, and use the training set to train the initial fault identification model to obtain a trained fault identification model; a first testing module, used to use the test set to test the trained fault identification model to obtain a test result, and when the test result indicates that the recognition accuracy of the fault identification model is greater than or equal to a preset accuracy threshold, it is determined that the model training is completed to obtain a final fault identification model.
[0098] Optionally, the acquisition unit 41 includes: a first cleaning module, used to clean the operation and maintenance data to obtain cleaned operation and maintenance data; a first extraction module, used to extract features from the cleaned operation and maintenance data to obtain operation and maintenance features, wherein feature extraction of the cleaned operation and maintenance data includes at least one of the following: calculating statistical features of the operation and maintenance data, extracting time series features from the operation and maintenance data, and identifying event features from the operation and maintenance data; a first encoding module, used to encode the operation and maintenance features to obtain an operation and maintenance feature vector.
[0099] Optionally, the operation and maintenance data includes at least one of the following: business tracking link data, performance monitoring data of the financial clearing system, log data of applications and network resources, real-time transaction flow data, and system alarm information.
[0100] It should be noted that the above-mentioned acquisition unit 41, first output unit 42, calling unit 43, and second output unit 44 correspond to steps S201 to S204 in Embodiment 1, and the above-mentioned units and corresponding steps implement the same instances and application scenarios, but are not limited to the contents disclosed in Embodiment 1. It should be noted that the above-mentioned modules or units may be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n), and the above-mentioned modules or units may also be part of a device and may be run in the computer terminal 10 provided in Embodiment 1.
[0101] The present invention is described below in conjunction with another optional embodiment.
[0102] Embodiment 3
[0103] An embodiment of the present invention may also provide an electronic device, Figure 5 is a hardware structure block diagram of an optional electronic device (or mobile device) for executing a fault node identification method of a financial settlement system according to an embodiment of the present invention, such as Figure 5 As shown, the electronic device may include: one or more ( Figure 5 (only one is shown) processor 502, memory 504, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0104] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application, and the processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0105] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: collect operation and maintenance data of the financial clearing system, and pre-process the operation and maintenance data to obtain an operation and maintenance feature vector, wherein the financial clearing system adopts a distributed architecture; input the operation and maintenance feature vector into an anomaly detection model, and output an anomaly detection result, wherein the anomaly detection model is a model pre-built based on a generative adversarial network for detecting operation and maintenance anomalies of the financial clearing system; when the anomaly detection result indicates that there is an anomaly in the financial clearing system, call the topological relationship graph of the financial clearing system, and update the topological relationship graph based on the operation and maintenance data to obtain a real-time operation and maintenance view of the financial clearing system, wherein the topological relationship graph is used to record the topological relationship between applications and network resources in the financial clearing system; input the real-time operation and maintenance view into a fault identification model, and output a fault node, wherein the fault identification model is a model pre-built based on a graph recognition algorithm for determining a fault node.
[0106] The processor can call the information and applications stored in the memory through the transmission device to perform the following steps: before collecting the operation and maintenance data of the financial clearing system, it also includes: obtaining all applications and network resources in the financial clearing system, and using the applications and network resources as nodes to construct an initial topological relationship graph; obtaining the mutual access relationship between each application in the financial clearing system, and constructing a horizontal topological relationship graph based on the mutual access relationship between the applications; obtaining the mutual dependence relationship between each network resource in the financial clearing system, and constructing a vertical topological relationship graph based on the mutual dependence relationship between the network resources; based on the initial topological relationship graph, the horizontal topological relationship graph and the vertical topological relationship graph are merged to obtain a topological relationship graph of the financial clearing system.
[0107] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: after outputting the fault node, it also includes: traversing the nodes in the topological relationship graph based on the fault node to obtain associated nodes that have a topological relationship with the fault node; and building a fault propagation link based on the fault node and the associated nodes.
[0108] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: after building a fault propagation link based on the fault node and the associated nodes, it also includes: calling the scenario recognition rule library to identify the fault scenario of the financial clearing system based on the fault propagation link and the scenario recognition rule library; querying the decision database based on the fault scenario to obtain the solution strategy corresponding to the fault scenario; generating a fault identification report based on the fault node, fault propagation link, fault scenario and solution strategy, and sending the fault identification report to the operation and maintenance terminal.
[0109] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: the fault identification model is pre-constructed, and the steps of constructing the fault identification model include: obtaining historical operation and maintenance data of the financial clearing system within a historical time period, and constructing a historical operation and maintenance view for the financial clearing system based on the historical operation and maintenance data; configuring fault labels for the historical operation and maintenance views, wherein the fault labels are used to mark fault nodes in the historical operation and maintenance views; constructing sample data based on the historical operation and maintenance views and the fault labels of the historical operation and maintenance views, and dividing the sample data into a training set and a test set; constructing an initial fault identification model based on a graph recognition algorithm, and using the training set to train the initial fault identification model to obtain a trained fault identification model; using the test set to test the trained fault identification model to obtain a test result, and when the test result indicates that the recognition accuracy of the fault identification model is greater than or equal to a preset accuracy threshold, determining that the model training is completed to obtain a final fault identification model.
[0110] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: pre-processing the operation and maintenance data, and the step of obtaining the operation and maintenance feature vector includes: cleaning the operation and maintenance data to obtain the cleaned operation and maintenance data; extracting features from the cleaned operation and maintenance data to obtain operation and maintenance features, wherein the feature extraction from the cleaned operation and maintenance data includes at least one of the following: calculating the statistical features of the operation and maintenance data, extracting the time series features from the operation and maintenance data, and identifying the event features in the operation and maintenance data; encoding the operation and maintenance features to obtain the operation and maintenance feature vector.
[0111] The processor can call the information and applications stored in the memory through the transmission device to perform the following steps: the operation and maintenance data includes at least one of the following: business tracking link data, performance monitoring data of the financial clearing system, log data of applications and network resources, real-time transaction flow data, and system alarm information.
[0112] By adopting the embodiment of the present invention, a solution for identifying faulty nodes in a financial clearing system is provided. The complex topological relationships in a distributed system are recorded through a topological relationship graph. In the process of continuous operation and maintenance of the financial clearing system, the operation and maintenance data of the system is obtained in real time, and preliminary anomaly detection is performed through an anomaly detection model. In the case of anomalies in the system, a real-time operation and maintenance view is constructed based on the operation and maintenance data and the topological relationship graph, and the nodes in the operation and maintenance view are analyzed through a graph recognition algorithm model to identify faulty nodes, thereby achieving rapid and accurate positioning of faulty nodes, effectively shortening the fault handling time, and ensuring the business continuity of the financial clearing system, achieving the technical effect of improving fault identification efficiency, and thus solving the technical problem of low identification efficiency in the related technology of the financial clearing system fault identification method based on human experience and single-point monitoring.
[0113] It can be understood by those skilled in the art that Figure 5 The structure shown is for illustration only, and the electronic device may also be a smart phone, a tablet computer, a PDA, a mobile Internet device (MID), a PAD or other terminal device. Figure 5 The structure of the electronic device is not limited. Figure 5 More or fewer components (such as network interfaces, display devices, etc.) shown in, or having Figure 5 Different configurations shown.
[0114] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0115] The present invention is described below in conjunction with another optional embodiment.
[0116] Embodiment 4
[0117] The embodiment of the present invention further provides a computer-readable storage medium. Optionally, in the embodiment of the present invention, the computer-readable storage medium can be used to store the program code executed by the fault node identification method of the financial settlement system provided in the first embodiment.
[0118] Optionally, in an embodiment of the present invention, the above-mentioned storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0119] An embodiment of the present invention also provides a computer program product, which, when executed on a data processing device, is suitable for executing the steps of a method for identifying faulty nodes in a financial clearing system: collecting operation and maintenance data of the financial clearing system, and preprocessing the operation and maintenance data to obtain an operation and maintenance feature vector, wherein the financial clearing system adopts a distributed architecture; inputting the operation and maintenance feature vector into an anomaly detection model, and outputting an anomaly detection result, wherein the anomaly detection model is a model pre-constructed based on a generative adversarial network for detecting operation and maintenance anomalies of the financial clearing system; when the anomaly detection result indicates that there is an anomaly in the financial clearing system, calling a topological relationship graph of the financial clearing system, and updating the topological relationship graph based on the operation and maintenance data to obtain a real-time operation and maintenance view of the financial clearing system, wherein the topological relationship graph is used to record the topological relationship between applications and network resources in the financial clearing system; inputting the real-time operation and maintenance view into a fault identification model, and outputting a faulty node, wherein the fault identification model is a model pre-constructed based on a graph recognition algorithm for determining a faulty node.
[0120] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0121] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0122] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0123] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0124] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0125] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.
[0126] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for identifying faulty nodes in a financial clearing system, characterized in that: include: Collecting operation and maintenance data of a financial clearing system and preprocessing the operation and maintenance data to obtain an operation and maintenance feature vector, wherein the financial clearing system adopts a distributed architecture; Inputting the operation and maintenance feature vector into an anomaly detection model and outputting an anomaly detection result, wherein the anomaly detection model is a model pre-built based on a generative adversarial network for detecting operation and maintenance anomalies of the financial clearing system; When the anomaly detection result indicates that the financial clearing system has an anomaly, calling a topological relationship diagram of the financial clearing system, and updating the topological relationship diagram based on the operation and maintenance data to obtain a real-time operation and maintenance view of the financial clearing system, wherein the topological relationship diagram is used to record the topological relationship between application programs and network resources in the financial clearing system; The real-time operation and maintenance view is input into a fault identification model, and a fault node is output, wherein the fault identification model is a model pre-built based on a graph recognition algorithm for determining a fault node.
2. The identification method according to claim 1, characterized in that: Before collecting the operation and maintenance data of the financial clearing system, it also includes: Acquire all applications and network resources in the financial clearing system, and construct an initial topology relationship graph using the applications and network resources as nodes; Acquire the mutual access relationship between the applications in the financial clearing system, and construct a horizontal topological relationship graph based on the mutual access relationship between the applications; Obtaining the interdependence relationship between network resources in the financial clearing system, and constructing a vertical topology relationship diagram based on the interdependence relationship between the network resources; The horizontal topology relationship diagram and the vertical topology relationship diagram are merged on the basis of the initial topology relationship diagram to obtain the topology relationship diagram of the financial clearing system.
3. The identification method according to claim 1, characterized in that: After outputting the faulty node, it also includes: Traversing the nodes in the topological relationship graph based on the faulty node to obtain associated nodes that have a topological relationship with the faulty node; A fault propagation link is constructed based on the fault node and the associated nodes.
4. The identification method according to claim 3, characterized in that: After the fault propagation link is constructed based on the fault node and the associated node, the method further includes: Calling a scenario recognition rule base to identify a fault scenario of the financial clearing system based on the fault propagation link and the scenario recognition rule base; Querying a decision database based on the fault scenario to obtain a solution strategy corresponding to the fault scenario; A fault identification report is generated based on the fault node, the fault propagation link, the fault scenario and the solution strategy, and the fault identification report is sent to the operation and maintenance terminal.
5. The identification method according to claim 1, characterized in that: The fault identification model is pre-constructed, and the steps of constructing the fault identification model include: Acquire historical operation and maintenance data of the financial clearing system within a historical time period, and construct a historical operation and maintenance view for the financial clearing system based on the historical operation and maintenance data; Configuring a fault label for the historical operation and maintenance view, wherein the fault label is used to mark a faulty node in the historical operation and maintenance view; Constructing sample data based on the historical operation and maintenance view and the fault labels of the historical operation and maintenance view, and dividing the sample data into a training set and a test set; Building an initial fault recognition model based on a graph recognition algorithm, and using the training set to train the initial fault recognition model to obtain the trained fault recognition model; The trained fault identification model is tested using the test set to obtain a test result. When the test result indicates that the recognition accuracy of the fault identification model is greater than or equal to a preset accuracy threshold, it is determined that the model training is completed to obtain the final fault identification model.
6. The identification method according to claim 1, characterized in that: The step of preprocessing the operation and maintenance data to obtain an operation and maintenance feature vector includes: Cleaning the operation and maintenance data to obtain cleaned operation and maintenance data; Performing feature extraction on the cleaned operation and maintenance data to obtain operation and maintenance features, wherein the feature extraction on the cleaned operation and maintenance data comprises at least one of the following: calculating statistical features of the operation and maintenance data, extracting time series features from the operation and maintenance data, and identifying event features from the operation and maintenance data; The operation and maintenance feature is encoded to obtain an operation and maintenance feature vector.
7. The identification method according to claim 1, characterized in that: The operation and maintenance data includes at least one of the following: business tracking link data, performance monitoring data of the financial clearing system, log data of the application and the network resources, real-time transaction flow data, and system alarm information.
8. A fault node identification device for a financial settlement system, characterized in that: include: A collection unit, used for collecting operation and maintenance data of a financial clearing system, and preprocessing the operation and maintenance data to obtain an operation and maintenance feature vector, wherein the financial clearing system adopts a distributed architecture; A first output unit, configured to input the operation and maintenance feature vector into an anomaly detection model and output an anomaly detection result, wherein the anomaly detection model is a model pre-built based on a generative adversarial network for detecting operation and maintenance anomalies of the financial clearing system; a calling unit, configured to call a topological relationship diagram of the financial clearing system when the abnormality detection result indicates that the financial clearing system is abnormal, and update the topological relationship diagram based on the operation and maintenance data to obtain a real-time operation and maintenance view of the financial clearing system, wherein the topological relationship diagram is used to record the topological relationship between application programs and network resources in the financial clearing system; The second output unit is used to input the real-time operation and maintenance view into a fault identification model and output a fault node, wherein the fault identification model is a model pre-built based on a graph recognition algorithm for determining a fault node.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the method for identifying a faulty node in a financial clearing system according to any one of claims 1 to 7.
10. An electronic device, characterized in that: It includes one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the faulty node identification method of the financial clearing system as described in any one of claims 1 to 7.
Citation Information
Cited By
Multi-system real-time monitoring method and system for automobile finance
CN121092401A