Bank IT system fault simulation method based on chaos engineering and related products

By using a fault simulation method based on chaos engineering, the dependency points of the bank's IT system are identified and fault simulations are performed. This solves the problem that traditional testing methods cannot cover faults caused by dependency anomalies, improves system stability and fault tolerance, and reduces fault repair costs.

CN120950384APending Publication Date: 2025-11-14AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511002663.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional testing methods for bank IT systems cannot effectively cover unexpected failure scenarios caused by abnormal dependencies, resulting in insufficient system stability and fault tolerance, and increasing the cost and risk of fault repair.

Method used

Based on the chaos engineering approach, this method obtains the components and fault model types of the bank's IT system, identifies the dependency points using the system dependency graph, designs and executes fault simulation experiments, generates chaos engineering files, quantifies the fault coefficients, and updates the system dependency graph.

Benefits of technology

It enables accurate modeling and efficient experimental design of failure scenarios in bank IT systems, improves the accuracy of system resilience assessment and real-time response capability, and reduces failure operation and maintenance and repair costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950384A_ABST
    Figure CN120950384A_ABST
Patent Text Reader

Abstract

The invention discloses a bank IT system fault simulation method based on chaos engineering and a related product. The method comprises the following steps: acquiring a to-be-simulated system component and a fault model type; determining a relation dependency point list of the to-be-simulated system component based on the system dependency relation graph; the relationship dependency point list comprises a plurality of nodes, and each node has a dependency relationship with the to-be-simulated system component; chaos experiment design is carried out based on the relation dependence point list of the to-be-simulated system component and the fault model type, and a chaos engineering file is obtained; and performing fault simulation on the to-be-simulated system component based on the chaos engineering file to obtain a fault coefficient of the to-be-simulated system component, and updating the system dependency relationship graph based on the fault coefficient of the to-be-simulated system component. The prediction and coping capacity of the bank IT system when facing uncertain faults is effectively enhanced, and the fault operation, maintenance and repair cost of the bank IT system is reduced especially in the aspect of cascade fault simulation caused by the dependency relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of fault simulation technology, and in particular to a fault simulation method and related products for bank IT systems based on chaos engineering. Background Technology

[0002] With the rapid development of fintech, banking IT systems are undergoing a profound transformation from traditional monolithic architectures to microservices, cloud-native, and distributed architectures. This transformation brings unprecedented flexibility and scalability, but it also significantly increases system complexity. A typical banking business process (such as transfers, payments, and loan approvals) may span dozens or even hundreds of internal systems, external services, middleware (message queues, caches, API gateways), and multiple databases.

[0003] These components form a complex web of operational dependencies. A delay, error, or interruption in any dependency can trigger a domino effect, causing some functions to become unavailable or even the entire business to be disrupted, resulting in huge economic losses and reputational risks for the bank.

[0004] Traditional testing methods (such as functional testing and performance testing) often focus on verifying the system's performance under "normal" or "high load" conditions, but they do not adequately cover "unexpected" failure scenarios caused by abnormal dependencies. Summary of the Invention

[0005] To address the aforementioned issues, this application provides a method and related products for simulating faults in bank IT systems based on chaos engineering. The aim is to simulate faults caused by dependencies based on chaos engineering, thereby reducing the cost of fault repair in bank IT systems.

[0006] The embodiments of this application disclose the following technical solutions: The first aspect of this application provides a method for simulating faults in a bank IT system based on chaos engineering, the method comprising: Obtain the system components and fault model types to be simulated; the bank IT system includes multiple of the system components to be simulated. A list of dependency points for the system components to be simulated is determined based on the system dependency graph; the list of dependency points includes multiple nodes, each of which has a dependency relationship with the system components to be simulated; Based on the list of relational dependency points of the system components to be simulated and the fault model type, a chaos experiment design is carried out to obtain a chaos project file; Based on the chaotic engineering file, fault simulation is performed on the components of the system to be simulated to obtain the fault coefficients of the components, and the system dependency graph is updated based on the fault coefficients of the components.

[0007] Optionally, determining the list of dependency points of the system components to be simulated based on the system dependency graph specifically includes: Based on the system dependency graph, multiple nodes that have dependencies on the system components to be simulated are identified, and the dependency attributes between each node and the system components to be simulated are determined. Analyze multiple nodes based on the dependency attributes between each node and the components of the system to be simulated, and obtain the dependency values ​​between each node and the components of the system to be simulated; Multiple nodes are sorted based on the dependency values ​​between each node and the system component to be simulated, and a list of relationship dependency points of the system component to be simulated is constructed according to the sorting results.

[0008] Optionally, the step of analyzing multiple nodes based on the dependency attributes between each node and the component of the system to be simulated, to obtain the dependency value between each node and the component of the system to be simulated, specifically includes: The dependency attributes between each node and the system component to be simulated are input into the trained neural network model to obtain the dependency values ​​between each node and the system component to be simulated.

[0009] Optionally, the method further includes: Obtain the training dataset; the training dataset includes historical dependency attributes and the dependency values ​​corresponding to the historical dependency attributes; Build a neural network model; The historical dependency attributes in the training dataset are used as the input to the neural network model, and the dependency values ​​corresponding to the historical dependency attributes are used as the target output of the neural network model. The neural network model is trained with the goal of minimizing the loss value, and the trained neural network model is obtained.

[0010] Optionally, the method further includes: Retrieve multiple system components and the dependency properties between each system component; Each system component is abstracted as a node, and the dependency properties between each system component are abstracted as edges to construct a system dependency graph.

[0011] Optionally, the chaotic experiment design based on the list of relational dependency points of the components of the system to be simulated and the fault model type, to obtain a chaotic engineering file, specifically includes: The list of relationship dependencies of the system components to be simulated is filtered to determine at least one node as the target node; Get configuration parameters; Based on the configuration parameters, the target node, and the fault model type, a chaos experiment is designed to obtain a chaos project file.

[0012] A second aspect of this application provides a bank IT system fault simulation device based on chaos engineering, the bank IT system fault simulation device based on chaos engineering comprising: The acquisition module is used to acquire the system components to be simulated and the fault model types; the bank IT system includes multiple of the system components to be simulated; The dependency point determination module is used to determine a list of relationship dependency points of the system components to be simulated based on the system dependency graph; the list of relationship dependency points includes multiple nodes, and each node has a dependency relationship with the system components to be simulated; The chaos experiment design module is used to design chaos experiments based on the list of relational dependency points of the components of the system to be simulated and the fault model type, and to obtain chaos project files. The fault simulation module is used to perform fault simulation on the system components to be simulated based on the chaotic engineering file, obtain the fault coefficients of the system components to be simulated, and update the system dependency graph based on the fault coefficients of the system components to be simulated.

[0013] A third aspect of this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the chaos engineering-based bank IT system fault simulation method provided in the first aspect.

[0014] The fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the chaos engineering-based bank IT system fault simulation method provided in the first aspect.

[0015] The fifth aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the chaos engineering-based bank IT system fault simulation method provided in the first aspect.

[0016] Compared with the prior art, this application has the following beneficial effects: This application includes obtaining system components to be simulated and fault model types; a bank IT system includes multiple system components to be simulated; determining a list of dependency points of the system components to be simulated based on a system dependency graph; the list of dependency points includes multiple nodes, each of which has a dependency relationship with the system components to be simulated; designing a chaotic experiment based on the list of dependency points of the system components to be simulated and the fault model types to obtain a chaotic engineering file; performing fault simulation on the system components to be simulated based on the chaotic engineering file to obtain the fault coefficients of the system components to be simulated, and updating the system dependency graph based on the fault coefficients of the system components to be simulated.

[0017] This application acquires the system components to be simulated in a bank's IT system and their corresponding fault model types. Based on a system dependency graph, it conducts in-depth analysis of the dependencies between components, achieving accurate modeling of system fault scenarios and efficient experimental design. By combining chaos engineering methods to generate targeted fault simulation files, it can realistically reproduce the operational behavior of each component in a complex system under abnormal conditions. Based on the simulation results, it quantifies and evaluates the fault coefficients of the components, thereby dynamically updating the system dependency graph and improving the accuracy of system resilience assessment and real-time response capabilities. This effectively enhances the predictive and response capabilities of the bank's IT system to uncertain faults, particularly in simulating cascading faults caused by dependencies, significantly improving system stability and fault tolerance, thereby reducing the fault maintenance and repair costs of the bank's IT system. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart illustrating a method for simulating faults in a bank IT system based on chaos engineering, provided in this application embodiment; Figure 2 This is a structural diagram of a bank IT system fault simulation device based on chaos engineering, provided as an embodiment of this application. Detailed Implementation

[0020] As described earlier, while the industry has begun exploring methods for proactively discovering and verifying system vulnerabilities, chaos engineering has emerged. It involves proactively injecting faults into a system to test its ability to cope with unexpected interruptions. However, effectively implementing the concepts of chaos engineering in the highly regulated and risk-sensitive IT development environment of banks, particularly in systematically designing fault scenarios for complex dependencies, remains a significant technical challenge.

[0021] Existing technologies primarily focus on chaos engineering execution platforms, fault injection tools, and their application in production or near-production environments. There is a lack of design methods for specific dependency-based system fault scenarios. Some solutions utilize monitoring data and alarm information from the production environment to infer potential fault points and dependencies, but this is post-hoc analysis or production environment verification and fails to meet the needs of proactive design and prevention during the R&D phase.

[0022] While some research or patents exist on fault injection targeting specific technology stacks (such as microservices and databases), they often lack a universal methodology for designing fault scenarios for complex dependencies, especially one tailored to the characteristics of banking operations and development processes. Furthermore, while general-purpose chaos engineering tools are powerful, they require users to have a deep understanding of system architecture, potential fault modes, and chaos engineering principles. This often leads to fault scenario design relying on the experience of a few experts, making it difficult to scale and standardize, especially within large banking technology teams. The lack of a systematic design methodology can result in chaotic experiments being designed too randomly or only covering known, obvious fault points, while ignoring more destructive potential faults hidden in complex dependency chains. Chaos engineering often emphasizes production or near-production environments, but in the high-risk banking sector, direct production experimentation is strictly limited. In development environments, due to a lack of effective design guidance, the application of chaos engineering often becomes merely a formality or ineffective.

[0023] In view of the above problems, this application provides a method and related products for simulating and generating faults in a bank IT system based on chaos engineering. The method includes: obtaining system components to be simulated and fault model types; the bank IT system includes multiple system components to be simulated; determining a list of dependency points of the system components to be simulated based on a system dependency graph; the list of dependency points includes multiple nodes, each node having a dependency relationship with the system components to be simulated; designing a chaotic experiment based on the list of dependency points of the system components to be simulated and the fault model types to obtain a chaotic engineering file; simulating faults in the system components to be simulated based on the chaotic engineering file to obtain fault coefficients of the system components to be simulated, and updating the system dependency graph based on the fault coefficients of the system components to be simulated.

[0024] This application overcomes the aforementioned shortcomings by providing a structured method for designing dependency-based fault scenarios. It lowers the barrier for developers and testers to design high-quality chaos experiments, ensures the systematic nature and coverage of fault scenario design, and promotes the effective implementation of chaos engineering in the banking R&D environment. By proactively and systematically identifying and verifying dependency risks during the R&D phase, the resilience and reliability of banking IT systems can be significantly improved, supporting stable business operations and innovative development. Its value lies in proactive defense rather than passive response. By acquiring the system components to be simulated in the banking IT system and their corresponding fault model types, and conducting in-depth analysis of the dependencies between components based on the system dependency graph, accurate modeling and efficient experimental design of system fault scenarios are achieved. Combined with chaos engineering methods, targeted fault simulation files are generated, which can realistically reproduce the operational behavior of each component in a complex system under abnormal conditions. Based on the simulation results, the fault coefficients of the components are quantitatively evaluated, thereby dynamically updating the system dependency graph and improving the accuracy of system resilience assessment and real-time response capabilities. It effectively enhances the ability of bank IT systems to predict and respond to uncertain failures, especially in simulating cascading failures caused by dependencies, significantly improving system stability and fault tolerance, thereby reducing the failure maintenance and repair costs of bank IT systems.

[0025] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0026] Figure 1 A flowchart of a fault simulation method for a bank IT system based on chaos engineering is provided for an embodiment of this application, as shown below. Figure 1 As shown, a fault simulation method for a bank IT system based on chaos engineering includes: S101: Obtain the system components and fault model types to be simulated.

[0027] Taking a bank IT system as an example, the bank IT system includes multiple system components to be simulated; the system components to be simulated in this application include at least application services, databases, middleware, and external interfaces. For example, the system components to be simulated may be payment gateways, core accounting services, user authentication centers, etc.

[0028] This application does not limit the method for obtaining fault model types. For example, it proposes establishing an extensible Failure Mode Library (FML) tailored to the characteristics of bank IT systems. This library categorizes and parameterizes common dependency problems and recommends relevant fault modes from the FML based on the selected dependency type (e.g., HTTP call, database connection, message queue). Users can select one or more fault modes. The dependency fault mode library includes at least network layer faults, service layer faults, and resource layer faults (indirectly affecting dependencies). Among these, network layer faults include LatencyInjection(Target, Latency, Jitter, Duration): injecting network latency; PacketLoss(Target, Percentage, Duration): injecting network packet loss; Blackhole(Source, Target, Duration): completely disconnecting the network connection; and BandwidthLimit(Target, Limit, Duration): limiting network bandwidth.

[0029] Service layer failures include ErrorResponse(Target, ErrorCode, Percentage, Duration): simulating a dependent service returning a specific error code, ServiceUnavailable(Target, Duration): simulating a dependent service being completely unavailable, and SlowResponse(Target, Delay, Duration): simulating a dependent service responding slowly.

[0030] Resource layer failures (indirectly affecting dependencies) include CPULoad(Target, Percentage, Duration): simulating high CPU load on the node where the dependent service resides, and MemoryLeak(Target, Rate, Duration): simulating memory leaks in the dependent service.

[0031] Each pattern should include a target selection mechanism (how to locate the target in the dependency), parameter definition (fault strength, duration, etc.), and technical implementation mapping (how to implement this fault through specific chaos engineering tools).

[0032] S102: Determine a list of dependency points for the system components to be simulated based on the system dependency graph. The list of dependency points includes multiple nodes, each of which has a dependency relationship with the system components to be simulated.

[0033] The system dependency graph is stored in the form of a graph structure, where nodes represent system components and edges represent dependencies. For each component to be simulated, the system extracts all its dependent nodes from the graph to form a "relationship dependency point list". Dependencies include direct dependencies and indirect dependencies, such as service call chains, database access chains, message queue dependencies, etc.

[0034] For example, when simulating the "core accounting service" component, its dependency list includes: database cluster (data storage dependency), payment gateway (call dependency), and user authentication center (permission dependency). This enables the visualization and structured management of complex dependencies between systems; improves the comprehensiveness and realism of fault simulation; and supports the identification and modeling of multi-level fault propagation paths.

[0035] S103: Design a chaotic experiment based on the list of relational dependency points of the system components to be simulated and the fault model type, and obtain a chaotic project file.

[0036] Based on component dependencies and the selected fault model type, the system automatically generates a chaos engineering experiment configuration file (e.g., YAML format). The file contains parameters such as the target component for fault injection, fault type, injection time, duration, and recovery strategy. This file can be executed by a chaos engineering platform (e.g., Chaos Mesh, Litmus, Chaos Monkey). This achieves standardized and automated generation of chaos experiments; supports seamless integration with mainstream chaos engineering platforms; and improves the repeatability and controllability of fault simulation.

[0037] S104: Based on the chaotic engineering file, perform fault simulation on the system components to be simulated to obtain the fault coefficients of the system components to be simulated, and update the system dependency graph based on the fault coefficients of the system components to be simulated.

[0038] Load the chaos engineering file and perform fault injection; monitor and collect metrics such as response time, error rate, and availability of each component during the fault; calculate the "fault coefficient" of the component based on the collected data (such as the percentage decrease in availability, recovery time, etc.); feed the fault coefficient back to the system dependency graph to dynamically adjust the dependency weights between components. For example, after fault simulation, the fault coefficient of the "core accounting service" in the "network latency" scenario is evaluated as follows: availability decrease: 25%, average recovery time: 30 seconds.

[0039] This application obtains the system components to be simulated in a bank's IT system and their corresponding fault model types. It then identifies the dependencies between components using a system dependency graph. Based on chaos engineering methods, it designs and executes fault simulation experiments, generating executable chaos engineering files. The simulation results are used to quantitatively evaluate the fault coefficients of each component, and the system dependency graph is dynamically updated, enabling continuous assessment and optimization of system resilience. This method not only improves the predictive and coping capabilities of the bank's IT system when facing complex dependencies and uncertain faults, but also provides data support for system architecture optimization, significantly enhancing system stability and fault tolerance, and reducing maintenance and repair costs caused by cascading faults.

[0040] The above describes the main technical solution of this application. Further implementations of the main technical solution are now introduced. Details are as follows: Regarding S102, which determines the list of dependency points of the system components to be simulated based on the system dependency graph, this application provides an optional embodiment: Based on the system dependency graph, multiple nodes that have dependencies on the system components to be simulated are identified, and the dependency attributes between each node and the system components to be simulated are determined.

[0041] Based on the dependency attributes between each node and the system component to be simulated, multiple nodes are analyzed to obtain the dependency value between each node and the system component to be simulated.

[0042] Multiple nodes are sorted based on the dependency values ​​between each node and the system component to be simulated, and a list of relationship dependency points of the system component to be simulated is constructed according to the sorting results.

[0043] This application also provides a specific embodiment: By combining the components of the system to be simulated (such as the core transaction chain), critical dependency paths are identified on the system dependency graph. Based on dependency attributes (such as strong dependency, no degradation / circuit breaker mechanism, single point of failure) and historical problem data of components, potential vulnerable nodes and edges in the SDG are identified. A simple risk scoring model can be introduced (e.g., inputting the dependency attributes between each node and the component to be simulated into a trained neural network model to obtain the dependency value between each node and the component to be simulated), comprehensively considering the business impact and technical vulnerability of the dependencies, and ranking different dependencies to obtain a critical dependency list (CDL) ranked by risk.

[0044] This application also provides an optional embodiment regarding the specific training process of the risk assessment model: Obtain the training dataset; the training dataset includes historical dependency attributes and the dependency values ​​corresponding to the historical dependency attributes.

[0045] Build a neural network model.

[0046] The historical dependency attributes in the training dataset are used as the input to the neural network model, and the dependency values ​​corresponding to the historical dependency attributes are used as the target output of the neural network model. The neural network model is trained with the goal of minimizing the loss value, and the trained neural network model is obtained.

[0047] The loss value is calculated based on the actual output of the neural network model and the target output of the neural network model.

[0048] This application identifies multiple nodes that have dependencies on components of the system to be simulated based on a system dependency graph. It further extracts the dependency attributes between each node and the component, and uses a trained neural network model to quantitatively evaluate these attributes, obtaining numerical dependency values ​​between nodes and components. Based on these dependency values, the nodes are sorted to construct a list of dependency points for the components of the system to be simulated, thereby achieving refined modeling and prioritization of system dependencies. This not only improves the accuracy and relevance of system fault simulation but also enhances the intelligence level of dependency evaluation by introducing machine learning methods, contributing to the optimization of chaotic experiment design and improving the scientific rigor and automation of system resilience assessment.

[0049] The above discussion provides a relatively detailed method for determining the list of dependency points. As can be seen from the above, the determination of the list of dependency points depends on the system dependency graph. Therefore, this application provides an optional embodiment for the method of constructing the system dependency graph: Get multiple system components and the dependency properties between each system component.

[0050] Dependency attributes can be obtained through manual input or maintenance. For example, architects or senior developers can input key applications, middleware, databases, and their call relationships and dependency types (synchronous / asynchronous, strong / weak dependencies) into the system based on design documents and practical knowledge. Alternatively, dependency attributes can be determined by parsing service registries (such as Nacos, Consul), API gateway configurations, application configuration files, etc.; or by analyzing service calls, database connections, message queue usage, etc. in the codebase; or (where permissible) by analyzing tracing data (such as SkyWalking, Zipkin) or application logs in the development environment to identify the actual call relationships.

[0051] Each system component is abstracted as a node, and the dependency properties between each system component are abstracted as edges to construct a system dependency graph.

[0052] The collected information is used to construct a System Dependency Graph (SDG). Nodes in the graph represent system components (application services, databases, middleware, external interfaces, etc.), edges represent the dependencies between them, and dependency attributes (protocol, direction, strength, etc.) are labeled.

[0053] Constructing a relatively complete system dependency graph can provide a guarantee for the subsequent determination of the dependency point list, improve the effectiveness of the dependency point list, effectively complete the design of chaotic experiments based on the dependency point list of the system components to be simulated and the fault model type, and obtain chaotic engineering files.

[0054] Regarding S103: Designing a chaotic experiment based on the list of relational dependency points of the components of the system to be simulated and the fault model type to obtain a chaotic engineering file, this application provides an optional embodiment: The list of relationship dependencies of the system components to be simulated is filtered to determine at least one node as the target node.

[0055] Users can select one or more dependencies to be tested from the list of dependency points of the system components to be simulated, or directly select points on the system dependency graph to determine at least one target node.

[0056] Get configuration parameters.

[0057] Users configure specific parameters for the selected failure mode (e.g., 100ms delay, 5% packet loss, duration 5 minutes). The system can provide recommended parameters based on the risk level.

[0058] Based on the configuration parameters, the target node, and the fault model type, a chaos experiment is designed to obtain a chaos project file.

[0059] Users define the normal business metrics (steady-state assumptions, such as core transaction success rate > 99%) that the system should maintain when a fault is injected, as well as the key performance or business metrics that need to be observed (validation points). Integrate this information into a structured Chaos Experiment Design (CED). This can be a JSON / YAML file or a database record.

[0060] Regarding S104, which performs fault simulation on the components of the system to be simulated based on the chaotic engineering file, obtains the fault coefficients of the components, and updates the system dependency graph based on the fault coefficients of the components, this application provides an optional embodiment: The generated CED is converted into instructions or configuration files that can be recognized and executed by specific chaos engineering tools. This step requires an adapter layer. Integration with the testing platform allows developers / testers to trigger these designed chaos experiments on demand within their testing processes. Experiments are executed in an isolated development environment, and predefined verification points are monitored to collect data and obtain the failure coefficients of the system components to be simulated.

[0061] This application ensures that the design of fault scenarios systematically covers the complex dependencies of the bank's IT systems, avoiding randomness and bias, and making the design process more professional. The guided design process and recommendation mechanism enable developers and testers without in-depth knowledge of chaos engineering or architecture to easily design meaningful and high-value fault scenarios, facilitating the promotion and popularization of chaos engineering within the bank. Deeply integrating fault simulation and reliability verification into the development process allows for the earlier detection and remediation of potential defects caused by dependency issues before system deployment, significantly reducing the risk and cost of production environment failures and ensuring the continuity of banking operations. The dependency modeling and risk assessment process fully considers the characteristics and critical paths of banking operations, making the designed fault scenarios closer to actual risks and improving the effectiveness and relevance of chaos testing. The fault mode library and generated chaos experiment designs can be continuously accumulated and reused, forming an internal reliability knowledge base for the bank, continuously improving the team's fault response capabilities and the overall resilience of the system.

[0062] Figure 2 A structural diagram of a bank IT system fault simulation device based on chaos engineering, provided for an embodiment of this application, is shown below. Figure 2 As shown, based on the chaos engineering-based bank IT system fault simulation method provided in the preceding embodiments, this application also provides a chaos engineering-based bank IT system fault simulation device, comprising: The acquisition module is used to acquire the system components to be simulated and the fault model types; the bank IT system includes multiple of the system components to be simulated.

[0063] The dependency point determination module is used to determine a list of relationship dependency points of the system components to be simulated based on the system dependency graph; the list of relationship dependency points includes multiple nodes, and each node has a dependency relationship with the system components to be simulated.

[0064] The chaos experiment design module is used to design chaos experiments based on the list of relational dependency points of the components of the system to be simulated and the fault model type, and to obtain chaos project files.

[0065] The fault simulation module is used to perform fault simulation on the system components to be simulated based on the chaotic engineering file, obtain the fault coefficients of the system components to be simulated, and update the system dependency graph based on the fault coefficients of the system components to be simulated.

[0066] As an optional embodiment, the dependency point determination module specifically includes: The dependency attribute determination unit is used to determine multiple nodes that have dependencies on the system component to be simulated based on the system dependency graph, and to determine the dependency attribute between each node and the system component to be simulated.

[0067] The dependency value determination unit is used to analyze multiple nodes based on the dependency attributes between each node and the component of the system to be simulated, and to obtain the dependency value between each node and the component of the system to be simulated.

[0068] The list determination unit is used to sort multiple nodes based on the dependency value between each node and the component of the system to be simulated, and to construct a list of relational dependency points of the component of the system to be simulated according to the sorting result.

[0069] As an optional embodiment, the numerical determination unit is specifically used for: The dependency attributes between each node and the system component to be simulated are input into the trained neural network model to obtain the dependency values ​​between each node and the system component to be simulated.

[0070] As an optional embodiment, the device further includes: The training dataset determination module is used to obtain the training dataset; the training dataset includes historical dependency attributes and the dependency values ​​corresponding to the historical dependency attributes.

[0071] The model building module is used to build neural network models.

[0072] The training module is used to take the historical dependency attributes in the training dataset as the input of the neural network model, take the dependency values ​​corresponding to the historical dependency attributes as the target output of the neural network model, and train the neural network model with the goal of minimizing the loss value, so as to obtain the trained neural network model.

[0073] As an optional embodiment, the device further includes: The system dependency graph construction module is used to obtain multiple system components and the dependency attributes between each system component; it abstracts each system component as a node and the dependency attributes between each system component as edges to construct the system dependency graph.

[0074] As an optional embodiment, the chaos experiment design module specifically includes: The target node determination unit is used to filter the list of relational dependency points of the components of the system to be simulated and determine at least one node as a target node.

[0075] The configuration parameter unit is used to obtain configuration parameters.

[0076] The chaos experiment design unit is used to design chaos experiments based on configuration parameters, the target node, and the fault model type, and to obtain chaos project files.

[0077] This application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement a method for simulating faults in a bank IT system based on chaos engineering.

[0078] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a method for simulating faults in a bank IT system based on chaos engineering.

[0079] This application provides a computer program product, including a computer program that, when executed by a processor, implements a method for simulating faults in a bank IT system based on chaos engineering.

[0080] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the device and equipment embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The device and equipment embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components indicated as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0081] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for simulating faults in a bank IT system based on chaos engineering, characterized in that, The method for simulating bank IT system failures based on chaos engineering includes: Obtain the system components and fault model types to be simulated; the bank IT system includes multiple of the system components to be simulated. A list of dependency points for the system components to be simulated is determined based on the system dependency graph; the list of dependency points includes multiple nodes, each of which has a dependency relationship with the system components to be simulated; Based on the list of relational dependency points of the system components to be simulated and the fault model type, a chaos experiment design is carried out to obtain a chaos project file; Based on the chaotic engineering file, fault simulation is performed on the components of the system to be simulated to obtain the fault coefficients of the components, and the system dependency graph is updated based on the fault coefficients of the components.

2. The method for simulating bank IT system faults based on chaos engineering according to claim 1, characterized in that, The process of determining the list of dependency points of the system components to be simulated based on the system dependency graph specifically includes: Based on the system dependency graph, multiple nodes that have dependencies on the system components to be simulated are identified, and the dependency attributes between each node and the system components to be simulated are determined. Analyze multiple nodes based on the dependency attributes between each node and the components of the system to be simulated, and obtain the dependency values ​​between each node and the components of the system to be simulated; Multiple nodes are sorted based on the dependency values ​​between each node and the system component to be simulated, and a list of relationship dependency points of the system component to be simulated is constructed according to the sorting results.

3. The method for simulating bank IT system faults based on chaos engineering according to claim 2, characterized in that, The analysis of multiple nodes based on the dependency attributes between each node and the component of the system to be simulated, to obtain the dependency value between each node and the component of the system to be simulated, specifically includes: The dependency attributes between each node and the system component to be simulated are input into the trained neural network model to obtain the dependency values ​​between each node and the system component to be simulated.

4. The method for simulating bank IT system faults based on chaos engineering according to claim 1, characterized in that, The method further includes: Obtain the training dataset; the training dataset includes historical dependency attributes and the dependency values ​​corresponding to the historical dependency attributes; Build a neural network model; The historical dependency attributes in the training dataset are used as the input to the neural network model, and the dependency values ​​corresponding to the historical dependency attributes are used as the target output of the neural network model. The neural network model is trained with the goal of minimizing the loss value, and the trained neural network model is obtained.

5. The method for simulating bank IT system faults based on chaos engineering according to claim 1, characterized in that, The method further includes: Retrieve multiple system components and the dependency properties between each system component; Each system component is abstracted as a node, and the dependency properties between each system component are abstracted as edges to construct a system dependency graph.

6. The method for simulating bank IT system faults based on chaos engineering according to claim 1, characterized in that, The chaotic experiment design based on the list of relational dependency points of the components of the system to be simulated and the fault model type yields a chaotic project file, which specifically includes: The list of relationship dependencies of the system components to be simulated is filtered to determine at least one node as the target node; Get configuration parameters; Based on the configuration parameters, the target node, and the fault model type, a chaos experiment is designed to obtain a chaos project file.

7. A fault simulation device for a bank IT system based on chaos engineering, characterized in that, The chaos engineering-based bank IT system fault simulation device includes: The acquisition module is used to acquire the system components to be simulated and the fault model types; the bank IT system includes multiple of the system components to be simulated; The dependency point determination module is used to determine a list of relationship dependency points of the system components to be simulated based on the system dependency graph; the list of relationship dependency points includes multiple nodes, and each node has a dependency relationship with the system components to be simulated; The chaos experiment design module is used to design chaos experiments based on the list of relational dependency points of the components of the system to be simulated and the fault model type, and to obtain chaos project files. The fault simulation module is used to perform fault simulation on the system components to be simulated based on the chaotic engineering file, obtain the fault coefficients of the system components to be simulated, and update the system dependency graph based on the fault coefficients of the system components to be simulated.

8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the chaos engineering-based bank IT system fault simulation method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for simulating bank IT system faults based on chaos engineering as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for simulating bank IT system faults based on chaos engineering as described in any one of claims 1-6.