Method and system for realizing observability after fault injection, electronic equipment and medium
By setting up acquisition points before Java method injection and collecting fault injection information, the problem of failure injection cannot be directly observed in the existing technology is solved, and direct observation and accurate indicator collection of methods after fault injection are realized, helping to verify the injection effect and locate the root cause.
Patent Information
- Application Number
- CN202311863762.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art cannot directly observe the fault injection information in scenarios such as Java method injection delay and exception, resulting in deviations in the observation results, and it is impossible to accurately count the time and abnormal situations after fault injection.
Set up a collection point before the fault injection, turn on the sampling device to collect status information after the fault injection, and push it to the observation system. A new sampling layer is added to conduct direct observation of the fault injection, including the collection and push of time-consuming and abnormal information.
It realizes the integration of direct observation of the method after fault injection, which can accurately collect and push the time-consuming and abnormal situations caused by fault injection, helps verify the injection effect and quickly locate the root cause.
Smart Images

Figure CN120256277A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer simulation technologies, and particularly to a method for achieving observability after fault injection, a system for achieving observability after fault injection, an electronic device, and a computer-readable storage medium. Background Art
[0002] Currently, most of the fault injection, metric collection, link tracing, etc. in Java-related scenarios are implemented through the Javaagent technology. A Java agent can dynamically modify the compiled Java bytecode during the startup and running of a Java process, and can add custom logic through bytecode modification technology to implement functions such as fault injection and metric collection.
[0003] In scenarios such as injecting delays and exceptions into Java methods, in the existing observation systems, such as link tracing skywalking, opentelemetry, etc., it is impossible to collect the information injected into the method. As Figure 1 shown, because the observation-related probes have modified the bytecode of the method before fault injection, it is actually impossible to count the delays after fault injection and capture the injected exceptions, etc. It can only be observed in the upstream calls; there may be certain deviations in observing the upstream method because the upstream may call many other services at the same time, and it is impossible to locate the time-consuming, exceptions, etc. caused by this method. Summary of the Invention
[0004] In order to at least solve the above technical problems existing in the prior art, the present disclosure provides a method for achieving observability after fault injection, a system for achieving observability after fault injection, an electronic device, and a computer-readable storage medium, which can achieve integration of injection and observation, can directly observe the injected method; can subdivide metrics, and the collected metric data is accurate, and the observation intensity is more detailed than observing the upstream method.
[0005] In a first aspect, the present disclosure provides a method for achieving observability after fault injection, the method comprising:
[0006] Setting a preset collection point for Java method fault injection testing, the collection point being set before fault injection;
[0007] After the start of the fault injection test, turning on the sampling device of the sampling point, and after modifying the bytecode for fault injection, collecting the status information of the fault injection;
[0008] Pushing the collected status information to the observation system.
[0009] Further, the status information includes:
[0010] The time consumption and exception information of the fault injection method;
[0011] The pushing the collected status information to the observation system includes:
[0012] Pushing the time consumption and exception information of the collected fault injection method to Prometheus (an open-source service monitoring system and time series database) to observe the time consumption and exception situations caused by the fault injection.
[0013] Further, the method further includes:
[0014] When the fault injection filters traffic, collecting and parsing the filtered traffic of the fault model by the collection device to determine the traffic status of the fault injection.
[0015] Further, the method further includes:
[0016] Verifying whether the fault injection meets the expectation according to the collected status information;
[0017] If it does not meet the expectation, reporting the status information to locate the failure cause.
[0018] Further, the method further includes:
[0019] Adding a sampling interceptor after the fault injection interceptor and keeping the sampling interceptor during the recovery after the fault injection ends, so that the sampling device continues to work.
[0020] In a second aspect, the present disclosure provides a system enabling observability after fault injection, the system including:
[0021] A setting module configured to set preset collection points for Java method fault injection tests, where the collection points are set before the fault injection;
[0022] A collection module configured to turn on the sampling device at the sampling point after the start of the fault injection test, and collect the status information of the fault injection after the bytecode modification of the fault injection;
[0023] A pushing module configured to push the collected status information to the observation system.
[0024] Further, the status information includes:
[0025] The time consumption and exception information of the fault injection method;
[0026] The pushing module is specifically configured to:
[0027] Push the time consumption and exception information of the collected fault injection method to Prometheus to observe the time consumption and exception situations caused by the fault injection.
[0028] Further, the acquisition module is further configured to:
[0029] When the fault injection filters traffic, the acquisition device acquires and analyzes the filtered traffic of the fault model to determine the traffic status of the fault injection.
[0030] In a third aspect, the present disclosure provides an electronic device, including a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the method for making the fault injection observable as described in any one of the first aspect.
[0031] In a fourth aspect, the present disclosure provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the method for making the fault injection observable as described in any one of the first aspect is implemented.
[0032] Beneficial effects:
[0033] The method for making the fault injection observable, the system for making the fault injection observable, the electronic device and the storage medium provided by the present disclosure can observe the method of fault injection; and through the observation of the injection point, it can help the injection personnel verify whether the injection takes effect, whether the fault reaches the expectation, etc.; it can also help the emergency personnel quickly locate the root cause. Description of the drawings
[0034] Figure 1 It is a schematic flowchart of fault injection in a Java-related scenario in the prior art;
[0035] Figure 2 It is a schematic flowchart of a method for making the fault injection observable provided in Embodiment 1 of the present disclosure;
[0036] Figure 3 It is a schematic diagram of the integrated implementation of injection and observation in a Java-related scenario provided in Embodiment 1 of the present disclosure;
[0037] Figure 4 It is another schematic diagram of the integrated implementation of injection and observation in a Java-related scenario provided in Embodiment 1 of the present disclosure;
[0038] Figure 5 It is an architecture diagram of a system for making the fault injection observable provided in Embodiment 2 of the present disclosure;
[0039] Figure 6 It is an architecture diagram of an electronic device provided in Embodiment 3 of the present disclosure. Detailed implementation manners
[0040] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the present disclosure will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments and drawings described herein are only for explaining the present invention, rather than limiting the present invention.
[0041] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence; and, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other arbitrarily.
[0042] Among them, the terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments, and are not intended to limit the present disclosure. The singular forms of "a", "the", and "said" used in the embodiments of the present disclosure and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0043] In subsequent descriptions, the suffixes such as "module", "component", or "unit" used to represent elements are only for the convenience of explaining the present disclosure, and have no specific meaning in themselves. Therefore, "module", "component", or "unit" can be used interchangeably.
[0044] The problems existing in the existing observations are that the method of fault injection cannot be observed, and even if the observation starts, it is impossible to collect the time-consuming, exceptions, etc. of the method after the fault injection; if the upstream method is observed, there may be a certain deviation, because the upstream calls often involve more than just the call to this method.
[0045] The following will specifically describe the technical solutions of the present disclosure and how the technical solutions of the present disclosure solve the above-mentioned technical problems existing in the prior art with specific embodiments. It can be understood that in the embodiments of the present application, the execution subject can execute some or all of the steps in the embodiments of the present application. These steps or operations are only examples, and the embodiments of the present application can also execute other operations or various deformations of the operations. In addition, each step can be executed in a different order presented in the embodiments of the present application, and it is possible not to execute all the operations in the embodiments of the present application. And, these several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0046] Figure 2 A flowchart of a method for realizing observability after fault injection provided for Embodiment 1 of the present disclosure is shown as Figure 2 shown, and the method includes:
[0047] Step S101: Set a preset collection point for the Java method fault injection test, and the collection point is set before the fault injection;
[0048] Step S102: After the start of the fault injection test, turn on the sampling device at the sampling point. After the bytecode modification for fault injection, collect the status information of the fault injection.
[0049] Step S103: Push the collected status information to the observation system.
[0050] In the current observation method, bytecode enhancement is statically mounted after the application starts, and operations such as sampling and recording have been performed on a certain method; fault injection has been performed before sampling and observation. As Figure 1 shown, the fault injection is outside the sampling and observation. The fault injection is outside the dotted box, and the current sampling method cannot collect the external fault injection.
[0051] In order to achieve integrated injection and observation in the embodiments of the present disclosure, a collection point is set before fault injection. After the start of the fault injection test, the sampling device at the sampling point is turned on to collect the information of the fault injection, and the collected information is pushed to the observation system to directly observe the injection method.
[0052] As Figure 3 shown, by adding a new layer of sampling, an additional layer of sampling is performed outside the fault injection to collect the fault scenario and metrics, so as to directly observe the method of fault injection.
[0053] Further, the status information includes:
[0054] The time consumption and exception information of the fault injection method;
[0055] The pushing the collected status information to the observation system includes:
[0056] Push the collected time consumption and exception information of the fault injection method to Prometheus to observe the time consumption and exceptions caused by the fault injection.
[0057] By setting the sampling layer outside the fault injection, it is possible to perform statistics on the method delay and exception rate for some fault scenarios, collect the delay after the fault injection, capture the injected exceptions, and add observation tags, and push the data collected during the whole process to Prometheus to determine the time consumption, exceptions, etc. of the method after the fault injection.
[0058] Further, the method further includes:
[0059] When the fault injection filters traffic, the collection device collects and analyzes the filtered traffic of the fault model to determine the traffic status of the fault injection.
[0060] When the fault injection filters traffic, the acquisition device can also parse the fault model to filter traffic. For example, when 30% of the MySQL traffic is called abnormally, the read operation (select operation) is specified. At the same time, it is also possible to observe the MySQL read operation, such as Figure 4 As shown, when a fault is injected, sampling is enabled. According to the fault model, sampling tags for MySQL read requests are added, etc., to achieve refined metrics. For example, among 100 operations when a user request comes in, 30 operations return abnormally, and the other 70% are processed as normal traffic. 70% will not reach the abnormal interceptor, but can be collected by the acquisition device at the sampling point. The curve graph sampled can calculate how much of the currently injected traffic is abnormal.
[0061] Further, the method further includes:
[0062] Verify whether the fault injection meets the expectation according to the collected status information;
[0063] If it does not meet the expectation, report the status information to locate the cause of failure.
[0064] By observing the injection point, it can help the injection personnel verify whether the injection takes effect, whether the fault meets the expectation, etc.; it can also help the emergency personnel quickly locate the root cause.
[0065] Further, the method further includes:
[0066] Add a sampling interceptor after the fault injection interceptor, and keep the sampling interceptor during the recovery after the fault injection ends, so that the sampling device continues to work.
[0067] The switching of the fault injection is implemented based on the interceptor of the process. A cut point is made before a certain method. By adding an interceptor manager, a fault interceptor is inserted into it for delay, and then a sampling interceptor is inserted. A sampling interceptor is added after the fault injection interceptor, and the sampling interceptor is kept during the recovery after the fault injection ends, so that the sampling device continues to work; the sampling interceptor is not removed during the recovery. Sampling has nothing to do with fault injection and recovery. It is installed when the fault is first executed and is used for a long time later. The entire process of sampling is realized on the outer layer of the fault injection through the sampling point and two layers of interceptors.
[0068] The embodiments of the present disclosure can observe the method of fault injection; achieve integration of injection and observation, and can refine metrics, and can dynamically add observation tags according to the fault model; and the directly collected index data is more accurate, and the observation intensity is more detailed than that of the upstream method. By observing the injection point, it can help the injection personnel verify whether the injection takes effect, whether the fault meets the expectation, etc.; it can also help the emergency personnel quickly locate the root cause.
[0069] Figure 5 The architecture diagram of a system that can be observed after fault injection provided in the second embodiment of the present disclosure is as follows Figure 5 shown. The system includes:
[0070] A setting module 11, which is set to set preset collection points for Java method fault injection testing, and the collection points are set before fault injection;
[0071] A collection module 12, which is set to turn on the sampling device of the sampling point after the start of the fault injection test, and collect the condition information of the fault injection after the bytecode modification of the fault injection;
[0072] A push module 13, which is set to push the collected condition information to the observation system.
[0073] Furthermore, the condition information includes:
[0074] The time consumption and exception information of the fault injection method;
[0075] The push module 13 is specifically set as:
[0076] Push the time consumption and exception information of the collected fault injection method to prometheus to observe the time consumption and exception conditions caused by the fault injection.
[0077] Furthermore, the collection module 12 is also set as:
[0078] When the fault injection filters traffic, collect and analyze the filtered traffic of the fault model through the collection device to determine the traffic condition of the fault injection.
[0079] Furthermore, the collection module 12 is also set as:
[0080] Add a sampling interceptor after the fault injection interceptor, and keep the sampling interceptor when recovering after the end of the fault injection, so that the sampling device continues to work.
[0081] The system that can be observed after fault injection in the embodiments of the present disclosure is used to implement the method that can be observed after fault injection in the first method embodiment, so the description is relatively simple. For specific details, please refer to the relevant description in the first method embodiment above, and will not be elaborated here.
[0082] In addition, as Figure 6 shown, the third embodiment of the present disclosure also provides an electronic device, including a memory 100 and a processor 200. A computer program is stored in the memory 100. When the processor 200 runs the computer program stored in the memory 100, the processor 200 executes the above various possible methods.
[0083] Among them, the memory 100 is connected to the processor 200. The memory 100 can be a flash memory, a read-only memory, or other memories, and the processor 200 can be a central processing unit or a single-chip microcomputer.
[0084] In addition, an embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by the processor to perform the above various possible methods.
[0085] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, computer program modules, or other data. The computer-readable storage medium includes, but is not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory, or other memory technologies, CD-ROM (Compact Disc Read-Only Memory), digital versatile disc (DVD, Digital Video Disc), or other optical disc storage, magnetic cassette, tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.
[0086] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principles of the present disclosure, but the present disclosure is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present disclosure, and these modifications and improvements are also regarded as the protection scope of the present disclosure.
Claims
1. A method for achieving observability after fault injection, characterized in that The method includes: Setting preset collection points for Java method fault injection testing, where the collection points are set before fault injection; After the fault injection test starts, turning on the sampling device at the sampling point, and after modifying the bytecode for fault injection, collecting the status information of the fault injection; Pushing the collected status information to the observation system.
2. The method according to claim 1, wherein The status information includes: The time consumption and exception information of the fault injection method; The pushing the collected status information to the observation system includes: Pushing the collected time consumption and exception information of the fault injection method to Prometheus to observe the time consumption and exception situations caused by the fault injection.
3. The method according to claim 2, wherein The method further includes: When the fault injection filters traffic, collecting and parsing the filtered traffic of the fault model through the collection device to determine the traffic status of the fault injection.
4. The method according to claim 1, characterized in that, The method further includes: Verifying whether the fault injection meets the expectation according to the collected status information; If it does not meet the expectation, reporting the status information to locate the failure cause.
5. The method according to any one of claims 1-4, characterized in that The method further includes: Adding a sampling interceptor after the fault injection interceptor and keeping the sampling interceptor during the recovery after the fault injection ends, so that the sampling device continues to work.
6. A system that enables observability after fault injection, characterized in that, The system includes: A setting module configured to set preset collection points for Java method fault injection testing, where the collection points are set before fault injection; A collection module configured to turn on the sampling device at the sampling point after the fault injection test starts, and after modifying the bytecode for fault injection, collect the status information of the fault injection; A pushing module configured to push the collected status information to the observation system.
7. The system according to claim 6, characterized in that, The status information includes: The time consumption and exception information of the fault injection method; The pushing module is specifically configured to: Push the collected time consumption and exception information of the fault injection method to Prometheus to observe the time consumption and exception situations caused by the fault injection.
8. The system according to claim 7, wherein The collection module is further configured to: When the fault injection filters traffic, collect and parse the filtered traffic of the fault model through the collection device to determine the traffic status of the fault injection.
9. An electronic device, characterized in that, Including a memory and a processor, where a computer program is stored in the memory, and when the processor runs the computer program stored in the memory, the processor executes the observable method after fault injection as described in any one of claims 1-5.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, it implements the observable method after fault injection as described in any one of claims 1-5.