A Chaos Testing Fault Injection Method and System for Multiple Financial Entities
Patent Information
- Application Number
- CN202610975538.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-02
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2046-07-02
AI Technical Summary
[0006]本发明旨在解决现有混沌测试工具在金融多法人架构下需在每个下游节点分别部署故障注入代理所导致的部署成本高、故障协同困难以及无法在ESB/MQ消息传输层面实现统一故障注入的技术问题
[0020]与现有技术相比,本申请的各核心特征具有以下先进性:
Smart Images

Figure CN122507613B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of software testing technology, specifically relating to a method and system for injecting faults in chaotic testing for multiple legal entities in the financial sector. Background Technology
[0002] In a multi-legal entity financial architecture, multiple independent legal entities share the same application instance. The system distinguishes request origination at runtime using fields such as branch_id and syscode, and employs an Enterprise Service Bus (ESB) and Message Queues (MQ) as the core communication links. Under this architecture, chaos engineering testing is used to verify the system's fault tolerance and the effectiveness of fault isolation.
[0003] Chinese patent application CN116225510A discloses a fault injection detection system for financial microservices based on Istio. This system uses the service mesh's reverse proxy, Sidecar, to hijack business traffic and injects faults such as error codes, timeouts, and delays through Lua scripts. This solution achieves fault injection at the inter-service communication level without intruding into business code.
[0004] However, the above solution requires deploying a Sidecar proxy independently alongside each microservice instance. For a multi-legal entity financial system where a transaction needs to pass through multiple nodes such as the gateway, ESB, MQ, wealth management system, and core accounting, deploying a fault injection proxy on each node separately would result in high deployment costs, difficulties in fault coordination, and the inability to achieve unified fault injection at the ESB / MQ message transmission level.
[0005] Therefore, how to achieve unified fault injection for financial multi-legal entity systems at the message transmission level without deploying fault injection agents at multiple points has become a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0006] This invention aims to solve the technical problems of high deployment costs, difficulty in fault coordination, and inability to achieve unified fault injection at the ESB / MQ message transmission level caused by the need to deploy fault injection agents separately on each downstream node in the financial multi-legal entity architecture of existing chaos testing tools.
[0007] To achieve the above objectives, this invention provides a method for chaotic testing and fault injection for multiple financial entities, comprising the following steps: Deploy a unified message interceptor on the ESB bus ingress routing node and the consumer end of the MQ message queue to intercept all passing business messages.
[0008] The message packets are parsed by extracting the strategy chain, which includes a message queue header extraction strategy, a JSON message body extraction strategy, and an Enterprise Service Bus context extraction strategy. Each strategy is executed sequentially according to a preset priority. Business fields are extracted from the message queue header, the JSONPath of the message body, and the ESB routing context. The business fields include at least the legal person identifier, system code, and service identifier. The extracted business fields are then aggregated into a business context object.
[0009] Maintain a fault rule base. Each fault rule in the fault rule base supports any combination of four dimensions: legal person, system, service, and fault type. Each dimension supports three modes: exact match, list match, and wildcard match. Match the business context object of the current message with the fault rules in the fault rule base. If a match is found, the rule is activated. Multiple rules can be combined in two ways: simultaneous triggering and any triggering. The matching process is accelerated by using an inverted index.
[0010] Fault injection is performed at the message transmission level according to the activated fault rules. Fault injection includes one or more of the following: delayed transmission, returning an abnormal response, modifying specified fields in the message body, discarding the message, or retransmitting it. Simultaneously with fault injection, isolation effectiveness verification is initiated, including: forward verification to confirm that requests from the target legal entity are affected by the fault; reverse verification to monitor requests from non-target legal entities and confirm they are not affected; and circuit breaker detection to automatically stop fault injection when the error rate of non-target legal entities exceeds a preset threshold.
[0011] After a fault is injected, a fault injection marker is implanted in the message queue header. The fault injection marker includes at least the fault type, injection time, and rule identifier, which is used by the downstream interceptor to identify and skip the secondary injection. The message carrying the fault injection marker is sent to the downstream business system, and the fault injection marker is cleared after the message is processed. The downstream business system does not need to install any fault injection tools.
[0012] To implement the above method, the present invention also provides a chaos testing fault injection system for multiple financial entities, comprising: The message interception module is deployed on the ESB bus ingress routing node and the consumer end of the MQ message queue to intercept all passing business messages.
[0013] The context extraction module, connected to the message interception module, is used to extract business fields from the message queue header, the JSONPath of the message body, and the ESB routing context by parsing the message messages through the extraction strategy chain, and to aggregate the extracted business fields into a business context object. The business fields include at least the legal person identifier, system code, and service identifier.
[0014] The multi-dimensional matching module, connected to the context extraction module, is used to maintain a fault rule base and match the business context object of the current message with the fault rules in the fault rule base, activating the successfully matched rules. Each fault rule supports any combination of four dimensions: legal person, system, service, and fault type. Each dimension supports three modes: exact matching, list matching, and wildcard matching. Multiple rules support two combination relationships: simultaneous triggering and any triggering. The matching process is accelerated through an inverted index.
[0015] The fault injection module, connected to the multi-dimensional matching module, is used to perform fault injection at the message transmission level according to the activated fault rules.
[0016] The isolation verification module, connected to the fault injection module, is used to initiate isolation effectiveness verification during fault injection, including forward verification, reverse verification, and circuit breaker detection.
[0017] The tag management module, connected to the fault injection module, is used to implant a fault injection tag in the message queue header after fault injection is completed, and to clear the fault injection tag after message processing is completed.
[0018] The message sending module, connected to the tag management module, is used to send a message carrying the fault injection tag to the downstream business system, wherein the downstream business system does not need to install any fault injection tools.
[0019] The configuration management module, connected to the multi-dimensional matching module, provides a fault template library. Users can select a template and fill in parameters through the web console to automatically generate Java fault codes, compile and distribute them. It also supports users uploading custom Java files, which are uniformly distributed after security verification. All fault configurations are centrally managed through a unified web platform, supporting version control, permission management and audit traceability.
[0020] Compared with the prior art, the core features of this application have the following advantages: 1. Unified Injection Architecture for Transport Layer Existing technologies employ a node-level injection architecture, requiring the independent deployment of a Sidecar proxy alongside each microservice instance. This application moves the fault injection point up to the transport layer of the ESB / MQ message bus, deploying a unified interceptor at the message bus entry point. This feature transforms fault injection from "distributed deployment" to "centralized management," solving the technical challenges of multi-point deployment at the architectural level.
[0021] 2. Four-dimensional combination matching ability Existing technologies only support single-dimensional fault injection targeting service interfaces. This application supports any combination of four dimensions: legal entity, system, service, and fault type. Each dimension supports three modes: exact matching, list matching, and wildcard matching. Multiple rules support both simultaneous triggering and any-triggered combinations. This feature expands fault injection from "single point, single fault" to "multiple points, multiple fault combinations."
[0022] 3. Verification of the effectiveness of three-layer isolation Existing technologies lack any isolation verification capabilities and cannot confirm whether a fault affects non-target entities. This application initiates a three-layer verification system—forward verification, reverse verification, and circuit breaker detection—simultaneously with fault injection, quantitatively evaluating the effectiveness of the entity's isolation configuration. This feature fills the technological gap in the field of chaos testing regarding "verifying whether a fault has crossed the boundary."
[0023] 4. Templated configuration and automatic code generation Existing technologies require manually writing Lua scripts or YAML files to define fault behaviors. This application provides a fault template library, allowing users to select a template and fill in parameters via a web console. The system automatically generates Java fault code, compiles it, and distributes it. This feature lowers the barrier to entry for chaos testing from "writing code" to "filling out a form."
[0024] Furthermore, the aforementioned technical features do not operate independently, but rather form an organically synergistic overall technical solution, generating a synergistic effect that goes beyond the simple superposition of existing technical features: Collaboration 1: Collaboration between Injection Architecture and Extraction Strategy The unified injection architecture at the transport layer provides the prerequisite for extracting multi-source business fields, while the multi-source extraction strategy chain provides the precise business context for transport layer injection. Together, they achieve an integrated capability of "full capture and precise parsing": using only the unified injection architecture without multi-source extraction capabilities would be insufficient to parse the diverse message formats in financial systems; conversely, possessing only multi-source extraction capabilities without unified transport layer injection would prevent centralized control across the entire link. This application integrates the two, resulting in a synergistic effect of "capture equals parsing, parsing equals matching."
[0025] Collaboration 2: Collaboration between Combinatorial Matching and Isolation Validation Four-dimensional combination matching provides diverse test scenarios for isolation verification, while isolation verification provides security boundaries and effectiveness evaluation for combination matching. Together, they form a closed-loop testing mechanism of "orchestration, injection, verification, and feedback": combination matching defines "which legal entities and which services to inject which faults into," and isolation verification answers "whether the target legal entity is affected as expected, and whether non-target legal entities are affected beyond their scope." Existing technologies can only achieve the unidirectional action of "injection," while this application achieves a closed-loop effect of "injection equals verification, verification equals feedback" through the synergy of the two.
[0026] Collaboration 3: Collaboration between unified distribution and centralized management Templated configuration solves the threshold problem of fault definition, unified distribution solves the consistency problem of configuration distribution, and centralized management solves the enterprise-level control problems of version control and permission auditing. The collaboration of the three achieves the integrated effect of "one-stop configuration, global effect, and full traceability": the code generated by the template is distributed to all interceptor nodes through a unified platform, the configuration change record is recorded in a complete audit log, and version rollback and permission isolation are supported. In the past, the configuration was scattered across various nodes and acted independently. This application upgrades chaos testing from "a developer's personal tool" to "an enterprise-level standardized platform" through the collaboration of the three.
[0027] The aforementioned collaborative mechanisms work together to create a complete technical closed loop in this application, encompassing "unified interception—multi-source extraction—combination matching—precise injection—isolation verification—marking and pass-through—template configuration—centralized management." Each link in this closed loop provides input to subsequent links, and subsequent links provide feedback to previous links. The synergistic effect between these links enables the overall technical solution to comprehensively address five dimensions of technical problems: high deployment costs, weak combination capabilities, lack of verification capabilities, high usage barriers, and lack of management capabilities. Existing technologies can only solve some of these problems or even cannot solve them at all. Attached Figure Description
[0028] Figure 1 Detailed flowchart of the method in Embodiment 1 of the present invention; Figure 2 : Module architecture diagram of the system in Embodiment 2 of the present invention. Detailed Implementation
[0029] To facilitate understanding of this invention, the following explanations of relevant terms are provided: Fault templates: These are pre-defined, reusable fault modes that support parameterized configuration. Users can generate corresponding fault injection code by filling in the parameters.
[0030] Fault isolation domain: refers to the fault scope defined by the legal entity identifier, system code, and service identifier, used to limit the scope of impact of fault injection.
[0031] Sandbox execution refers to running user-defined scripts in a restricted and secure environment, restricting their file access, network operations, and system call permissions to prevent adverse effects on the host JVM.
[0032] Dynamic code loading: refers to hot updating and executing newly compiled fault-injected code without restarting the JVM.
[0033] Example 1 (Method Example) like Figure 1As shown in the figure, this embodiment provides a chaotic testing fault injection method for multiple financial entities.
[0034] Before performing fault injection testing, the system first distributes fault rules through configuration management, specifically as follows: A fault template library is provided. Users select a template and fill in the parameters through the web console. The system backend automatically generates a complete Java fault source code file using the JavaPoet code generation library, automatically compiles it into a .class file, and distributes it to each interceptor node through the execution engine. The entire process is completed within 3 seconds. When the template cannot meet the testing requirements, users can upload a custom Java file that implements the preset fault behavior interface. The system performs security verification on the uploaded file, including syntax checking, dangerous call scanning, and interface compliance checks. After verification, the file is uniformly distributed to all interceptor nodes through an approval process.
[0035] When a custom Java file is executed, the system establishes a sandbox environment through SecurityManager, restricting its access permissions to the file system, network ports, and system processes. Simultaneously, the system sets a timeout control mechanism for each custom failure behavior. Tasks are executed asynchronously via Future; if the execution time exceeds a preset threshold (e.g., 30 seconds), future.cancel(true) is called to interrupt execution. The system also includes a watchdog thread that forcibly terminates unresponsive tasks after twice the timeout period, preventing infinite loops or prolonged blocking in user-defined code from affecting the normal operation of interceptors.
[0036] All fault configurations are centrally managed through a unified web platform. Configuration changes are pushed to each interceptor node via the Redis publish-subscribe mechanism to achieve second-level synchronization, supporting version control, role-based permission management, and audit traceability.
[0037] Then perform the following steps: Step 1: Deploy a unified message interceptor on the ESB bus ingress routing node and the consumer end of the MQ message queue to intercept all passing business messages.
[0038] Specifically, for ESB, all inbound messages are intercepted through the Processor interface of the Camel framework; for MQ, the message consumption process is intercepted through Channel interceptors or Consumer interceptors for message middleware such as RabbitMQ, Kafka, or RocketMQ. These interceptors are completely transparent to the business code and do not modify any source code of the business system.
[0039] Step 2: Extract and parse the message messages that have passed through the policy chain.
[0040] The extraction strategy chain includes a message queue header extraction strategy, a JSON message body extraction strategy, and an enterprise service bus context extraction strategy, each executed sequentially according to a preset priority. Specifically, the message queue header extraction strategy is executed first to extract business fields such as X-Bank-ID and syscode from the message queue header. If extraction fails or fields are missing, the JSON message body extraction strategy is executed to extract fields such as branch_id and service_id from the message body using JsonPath expressions. If extraction still fails, the enterprise service bus context extraction strategy is executed to extract routing variables from the CamelExchange object. The first successful extraction result is adopted, and the extracted business fields include at least the legal entity identifier, system code, and service identifier. The extracted business fields are then aggregated into a business context object.
[0041] Step 3: Maintain the fault rule base and perform multi-dimensional combination matching.
[0042] Each fault rule in the fault rule base supports any combination of four dimensions: legal entity, system, service, and fault type. Each dimension supports three modes: exact match, list match, and wildcard match. For example, a rule can be defined as an exact match for the legal entity identifier "866", a system identifier matching the list "CORE,FUND", a service identifier wildcard match "", and a fault type of delayed injection. Multiple rules can be combined in two ways: simultaneous triggering and any triggering.
[0043] When matching the current message's business context object with fault rules in the fault rule base, an inverted index is used to accelerate the matching process. The system establishes an inverted index based on the legal entity identifier, system code, and service identifier, reducing the matching time complexity from O(N) to O(1). The matching time is controlled at the microsecond level, with no substantial impact on message transmission performance. Successfully matched rules are activated.
[0044] Step 4: Perform fault injection at the message transport layer according to the activated fault rules.
[0045] The fault injection includes the following five types: (1) Delayed injection: The current interceptor thread is put to sleep for a specified number of milliseconds by calling the Thread.sleep(delayMillis) method. After the sleep ends, the message sending logic continues to be executed, simulating the slow processing of the downstream system.
[0046] (2) Abnormal return: The interceptor does not send the current message to the target downstream system, but directly constructs an abnormal response message that conforms to the caller's expected format, sets the response code to a non-200 status or business exception code, fills the response body with preset exception information, and returns the abnormal response message to the caller to simulate a downstream system failure.
[0047] (3) Return value tampering: the message body is parsed into a JsonNode tree structure, the target field node is located by the JsonPath expression, the field modification method is called to replace the original value with the fake value, and then the modified JsonNode is serialized back to the message body format to simulate data tampering.
[0048] (4) Message discard: directly discard the message and do not send it, simulating message loss.
[0049] (5) Message duplication: After the first successful message transmission, the interceptor calls the send interface again to send an identical copy of the message to the same target queue or topic, simulating message duplication. All fault operations are completed synchronously in the interceptor thread, which is completely transparent to the business system.
[0050] Step 5: Perform fault marking management and message sending, and simultaneously initiate isolation effectiveness verification.
[0051] After a fault is injected, a fault injection flag is inserted into the message queue header. This flag includes at least the fault type, injection time, and rule identifier, and is in the format X-Chaos-Fault-Injected. This flag is used by downstream interceptors to identify the current message as a faulty message and skip secondary injection. The message carrying the fault injection flag is then sent to the downstream business system.
[0052] Simultaneously with fault injection, the system initiates isolation effectiveness verification. Isolation effectiveness verification includes the following three aspects: (1) Positive verification confirms that the target legal entity’s request is indeed affected by the fault, for example, verifying that the response time of the target legal entity’s request has increased to the expected value or returned the expected exception code.
[0053] (2) Reverse verification: monitor requests from non-target legal entities to confirm that their response time and success rate are not affected. If an impact is detected, immediately issue an alarm and record the out-of-bounds log.
[0054] (3) Circuit breaker detection: When the error rate of requests from non-target legal entities exceeds a preset threshold, the circuit breaker mechanism is automatically triggered to stop fault injection and protect the safety of the production environment. The verification results are displayed and recorded in real time.
[0055] After the message reaches the final processing node, the interceptor clears the fault marker in the header before returning a response, ensuring that fault traces are not leaked to external systems. The downstream business systems do not need to install any fault injection tools.
[0056] Example 2 (System Example) like Figure 2As shown, this embodiment provides a chaos testing fault injection system for multiple financial entities. This system is used to implement the method described in Embodiment 1. Since there is a one-to-one correspondence between the system embodiment and the method embodiment, the similarities will not be repeated. The following focuses on describing the module structure and functions of the system.
[0057] The system includes the following modules: The message interception module, deployed at the ESB bus ingress routing node and the consumer end of the MQ message queue, is used to intercept all passing business messages. This module implements ESB message interception through the Processor interface of the Camel framework and MQ message interception through Channel interceptors or consumer interceptors, remaining completely transparent to the business code.
[0058] The context extraction module, connected to the message interception module, is used to parse passing message packets through an extraction strategy chain. This module incorporates message queue header extraction strategies, JSON message body extraction strategies, and Enterprise Service Bus (ESB) context extraction strategies. Each strategy is executed sequentially according to a preset priority, extracting business fields from the message queue header, the JSONPath of the message body, and the ESB routing context. These business fields include at least the legal entity identifier, system code, and service identifier. This module aggregates the extracted business fields into a business context object for use by subsequent modules.
[0059] The multi-dimensional matching module, connected to the context extraction module, maintains a fault rule base and matches the current message's business context object with fault rules in the base, activating successfully matched rules. Each fault rule maintained by this module supports any combination of four dimensions: legal entity, system, service, and fault type. Each dimension supports three modes: exact matching, list matching, and wildcard matching. Multiple rules can be combined for simultaneous triggering or any triggering. This module incorporates an inverted index engine, creating indexes based on legal entity identifier, system code, and service identifier to accelerate the matching process, keeping matching time within microseconds.
[0060] The fault injection module, connected to the multi-dimensional matching module, is used to perform fault injection at the message transmission level according to the activated fault rules. The fault types supported by this module include delayed sending, returning an abnormal response, modifying specified fields in the message body, discarding messages, and retransmitting messages. All fault operations are completed synchronously in the interceptor thread.
[0061] The tag management module, connected to the fault injection module, is used to implant a fault injection tag in the message queue header after fault injection is completed, and to clear the fault injection tag after message processing is completed. The fault injection tag includes at least the fault type, injection time, and rule identifier, in the format X-Chaos-Fault-Injected, which is used by downstream interceptors to identify fault messages and skip secondary injection, while ensuring that fault traces are not leaked to external systems.
[0062] The isolation verification module, connected to the fault injection module and the tagging management module, is used to initiate isolation effectiveness verification simultaneously with fault injection. This module includes a forward verification function to confirm that the target legal entity's request is affected by the fault; a reverse verification function to monitor and confirm that non-target legal entity's requests are unaffected; and a circuit breaker detection function to automatically stop fault injection when the error rate of non-target legal entities exceeds a preset threshold. Verification results are output and recorded in real time.
[0063] The message sending module, connected to the tag management module, is used to send messages carrying the fault injection tag to downstream business systems. The downstream business systems do not need to install any fault injection tools; they already receive messages containing the fault.
[0064] The configuration management module, connected to the multi-dimensional matching module, provides a fault template library. Users can select a template and fill in parameters via a web console to automatically generate Java fault codes, which are then compiled and deployed. This module supports user-uploaded custom Java files, which are then uniformly deployed to all interceptor nodes after security verification and approval. When a custom Java file is executed, this module establishes a sandbox environment through SecurityManager and implements timeout control through Future and watchdog threads. This module also centrally manages all fault configurations through a unified web platform. Configuration changes are pushed to each interceptor node via a Redis publish-subscribe mechanism for second-level synchronization, supporting version control, role-based access control, and audit traceability.
[0065] Example 3 (Application Scenario 1: Verifying the effectiveness of corporate bank fault isolation) This embodiment takes a scenario where a bank's core accounting system simultaneously serves three legal entities (legal entity identifiers are 866, 810, and 812). The test objective is to confirm whether the transactions of legal entities 810 and 812 are affected when all requests from legal entity 866 are delayed by 10 seconds.
[0066] Testers log into the unified chaos testing platform's web console, select the "Interface Delay" fault template, and in the parameter configuration interface, select exact match "866" for the legal entity dimension, select "CORE" for the system code, select "2000" for the service identifier, enter 10000 milliseconds for the delay time, and click the "Deploy" button. The system backend automatically generates the Java source code file for the delay fault using the JavaPoet code generation library, automatically compiles it, and pushes it to the interceptor node deployed at the ESB entry point through the execution engine. The entire process is completed within 3 seconds without restarting any services.
[0067] After receiving the configuration, the interceptor begins operation. When the interceptor captures a request message with legal entity identifier 866 that calls the core system service 2000, a match is found. After a 10,000-millisecond delay injection at the message transmission level, the message is sent to the core accounting system. For request messages with legal entity identifiers 810 or 812, the match fails, no delay processing is performed, and the message is directly passed downstream.
[0068] The isolation verification module simultaneously started monitoring. Forward verification results showed that the average response time for requests from legal entity 866 increased from 200 milliseconds to 10200 milliseconds, indicating the forward verification passed. Reverse verification results showed that the average response time for requests from legal entities 810 and 812 remained around 200 milliseconds, with a 100% success rate, unaffected by fault injection, indicating the reverse verification passed. The system automatically generated a chaos test report, outputting the conclusion "Isolation configuration effective".
[0069] Example 4 (Application Scenario 2: Verifying the Fault Tolerance of a Combined Fault System) This embodiment uses an interbank transfer transaction as a scenario. The transaction call chain passes through multiple nodes, including the APP gateway, ESB, wealth management system, MQ, core accounting, and balance notification, involving two legal entities (legal entity identifiers 866 and 810). The test objective is to simulate complex fault combinations and verify the overall fault tolerance capability of the system.
[0070] Testers orchestrated a series of fault scenarios in the web console, named "Interbank Transfer Link Fault Test." This scenario included three rules: Rule 1: Target legal entity precisely matches "866," target system precisely matches "CORE," target service precisely matches "2000," fault type is delayed injection, delay time is 10000 milliseconds, used to simulate core system overload. Rule 2: Target legal entity precisely matches "810," target system precisely matches "FUND," target service precisely matches "3000," fault type is exception return, exception code is "ERR_BALANCE_001," used to simulate insufficient balance. Rule 3: Target legal entity matches the list "866,810," target system precisely matches "NOTIFY," fault type is message discard, used to simulate notification system failure. The three rules were combined using a simultaneous triggering relationship. After configuration, the system was deployed with a single click.
[0071] The system performs fault injection according to the orchestrated scenarios. For the transfer request of legal entity ID 866, calling the core system service 2000 triggers rule one, injecting a 10-second delay; calling the notification system triggers rule three, and the message is discarded. For the transfer request of legal entity ID 810, calling the wealth management system service 3000 triggers rule two, directly returning an insufficient balance exception and not continuing to call subsequent services; calling the notification system also triggers rule three, and the message is discarded.
[0072] The isolation verification module monitored the triggering of three rules. Rule 1 verification showed that the transfer request from bank 866 was successfully completed after a 10-second delay, verifying the system's recovery capability under heavy load in core system scenarios. Rule 2 verification showed that the transfer request from bank 810 was correctly intercepted and returned an insufficient balance exception, verifying the correctness of the business logic after exception injection. Rule 3 verification showed that neither bank received a balance notification message, triggering the system's reconciliation compensation mechanism. The entire test was completed within 5 minutes, and the system automatically generated a complete chaos test report, including the trigger count, success rate, and isolation effectiveness score for each rule.
[0073] Example 5 (Application Scenario 3: Custom Java Fault Simulation of Complex Business Logic) This example uses a scenario that requires simulating complex business logic. Specifically, when the transfer amount exceeds 100,000 yuan and the target account's legal representative is 812, there is a 50% probability of a random delay of 3 to 8 seconds. Since this requirement involves conditional judgment (amount exceeding 100,000 yuan, target legal representative being 812), probability control (50%), and random time generation (3 to 8 seconds), a template-based configuration cannot meet the requirements, necessitating a custom Java extension.
[0074] Testers write Java classes to implement the system's preset fault behavior interface. In the implementation method, the amount and target legal entity identifier fields from the request JSON are retrieved through a message context object. The condition is: the amount is greater than 100,000 and the target legal entity identifier equals "812". If the condition is true, a random number between 0 and 1 is generated. If the random number is less than 0.5, a random delay time between 3000 and 8000 milliseconds is generated, and delayed injection is performed. After completion, the Java source files are uploaded via the web console.
[0075] After receiving the uploaded file, the system performs security checks: It verifies syntax correctness using the Java compiler API, scans for dangerous calls such as `System.exit`, `Runtime.exec`, and `File.delete` using regular expressions, and confirms that it implements the necessary methods of the fault behavior interface. Once the checks pass, the test manager receives an approval notification, logs into the system to view the code content, confirms security, and clicks "Approval." The system compiles the custom Java file and distributes it to all interceptor nodes via the execution engine. After distribution, the system uses the SecurityManager to create a sandbox environment to execute the custom code, setting a 30-second timeout threshold and a 60-second watchdog timer for forced termination mechanism to ensure execution security.
[0076] Testers issued 100 test transactions, 30 of which met the criteria of an amount greater than 100,000 yuan and a target legal entity of 812. Statistics from the isolation verification module showed that among the 30 transactions meeting the criteria, 16 were injected with random delays, an injection probability of approximately 53%, close to the expected 50%; the injection delay time ranged from 3.2 seconds to 7.8 seconds, consistent with the expected range of 3 to 8 seconds. The remaining 70 transactions that did not meet the criteria were not injected with any faults. The test verification passed, and the custom Java fault injection was successful.
[0077] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, or improvements made to the above embodiments within the spirit and principles of the present invention, as well as the application of the technical solutions of the present invention to other similar scenarios (including but not limited to multi-tenant systems in non-financial fields, distributed systems with multiple branches, etc.), should be included within the scope of protection of the present invention. The scope of protection of the present invention is determined by the appended claims. The embodiments and drawings described in the specification are only for interpreting the claims and do not constitute a limitation on the scope of protection of the claims.
[0078] Furthermore, those skilled in the art should understand that the functional division of the modules described in this invention is merely illustrative. In practical applications, they can be merged, split, or reorganized according to the system architecture design. As long as the same function is achieved and the same technical effect is attained, they all fall within the protection scope of this invention. The order of the steps described in this invention can be appropriately adjusted without changing the core technical solution, and the adjusted technical solution still falls within the protection scope of this invention.
Claims
1. A method for chaotic testing and fault injection for multiple legal entities in the financial sector, characterized in that, Includes the following steps: Fault rules are issued through configuration management. Deploy a unified message interceptor at the enterprise service bus entry routing node and message queue consumer end to intercept all passing business messages; The message messages are parsed by extracting the strategy chain. The extraction strategy chain includes multiple extraction strategies that are executed sequentially according to a preset priority. Business fields are extracted from different positions in the message messages, and the extracted business fields are aggregated into a business context object. Maintain a fault rule base, where each fault rule supports multi-dimensional combination matching. Match the current message's business context object with the fault rules in the fault rule base and activate the successfully matched rules. Fault injection is performed at the message transport level based on the activated fault rules; After a fault is injected, a fault injection marker is inserted into the message header, a message carrying the fault injection marker is sent to the downstream business system, and the fault injection marker is cleared after the message processing is completed.
2. The method for chaotic testing and fault injection for multiple financial entities according to claim 1, characterized in that, The extraction strategy chain includes a message queue header extraction strategy, a JSON message body extraction strategy, and an enterprise service bus context extraction strategy; the business fields include at least the legal entity identifier, system code, and service identifier.
3. The method for chaotic testing and fault injection for multiple financial entities according to claim 1, characterized in that, Each fault rule supports multiple dimensions, including four dimensions: legal entity, system, service, and fault type. Each dimension supports three modes: exact matching, list matching, and wildcard matching. Multiple fault rules support two combination relationships: simultaneous triggering and any triggering.
4. The method for chaotic testing and fault injection for multiple financial entities according to claim 1, characterized in that, Simultaneously with fault injection, isolation effectiveness verification is initiated. Isolation effectiveness verification includes: forward verification, confirming that the target legal entity's request is affected by the fault; reverse verification, monitoring non-target legal entity's requests to confirm that they are not affected; and circuit breaker detection, automatically stopping fault injection when the error rate of non-target legal entities exceeds a preset threshold.
5. The method for chaotic testing and fault injection for multiple financial entities according to claim 1, characterized in that, The fault injection marker should include at least the fault type, injection time, and rule identifier; After a message carrying a fault injection flag is sent to a downstream business system, the downstream interceptor identifies that the current message has been injected with a fault based on the fault injection flag and skips the secondary injection; after the message is processed, the interceptor clears the fault injection flag.
6. The method for chaotic testing and fault injection for multiple financial entities according to claim 1, characterized in that, Also includes: Receive user-uploaded custom fault code files and perform security verification on the custom fault code files; After successful verification, custom fault codes are executed in the sandbox environment to perform fault injection, and a timeout control mechanism is set to forcibly terminate execution when the execution time exceeds a preset threshold.
7. The method for chaotic testing and fault injection for multiple financial entities according to claim 1, characterized in that, Also includes: Receive configuration changes for fault rules through a unified web platform; The configuration changes are pushed to all interceptor nodes through a publish-subscribe mechanism to achieve synchronization. Record configuration change information, operation permissions, and audit logs.
8. A chaos testing fault injection system for multiple financial entities, characterized in that, include: The message interception module is deployed at the enterprise service bus ingress routing node and the message queue consumer end to intercept all passing business messages; The context extraction module, connected to the message interception module, is used to parse the passing message messages through the extraction strategy chain, extract business fields from different positions of the message messages and aggregate them into a business context object. The multi-dimensional matching module, connected to the context extraction module, is used to maintain a fault rule base and match the business context object of the current message with the fault rules in the fault rule base, and activate the successfully matched rules. The fault injection module, connected to the multi-dimensional matching module, is used to perform fault injection at the message transmission level according to the activated fault rules. The tag management module, connected to the fault injection module, is used to implant a fault injection tag in the message header after fault injection is completed, and to clear the fault injection tag after message processing is completed. The message sending module, connected to the tag management module, is used to send messages carrying the fault injection tag to downstream business systems.
9. A chaos testing fault injection system for multiple financial entities according to claim 8, characterized in that, It also includes an isolation verification module, which is connected to the fault injection module and is used to start isolation validity verification at the same time as fault injection. The isolation validity verification includes forward verification, reverse verification and circuit breaker detection.
10. A chaos testing fault injection system for multiple financial entities according to claim 8, characterized in that, It also includes a configuration management module, which is connected to the multi-dimensional matching module, to provide a fault template library. Users can select a template and fill in the parameters through the Web console to automatically generate fault codes, compile and distribute them, and support users to upload custom files and distribute them uniformly after security verification.
Citation Information
Patent Citations
Financial micro-service fault injection detection system based on Istio
CN116225510A
Redis chaos fault test method, storage medium and equipment
CN121919081A