A Fault Warning Method Based on JVM-Sandbox
Through JVM-Sandbox monitoring and dynamic rule filtering, the problems of long fault processing cycles and complicated alarms in the existing technology are solved, and accurate fault events are collected and selective alarms are realized, and the efficiency and accuracy of fault processing are improved.
Patent Information
- Application Number
- CN202510082125.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-01-20
AI Technical Summary
The existing technology is difficult to accurately capture the specific scenarios of failures, resulting in a long fault processing cycle, low efficiency, and a large number of invalid or non-serious fault alarm events, which cannot identify the same fault problems, resulting in duplicate alarms and complicated error logs.
Through JVM-Sandbox, listen for events in business applications, intercept BeforeEvent, ReturnEvent and ThrowEvent events, configure dynamic rules for matching filtering, upload them to the cloud alarm platform, and count the times within the time window to achieve selective alarms.
Accurate collection of fault events is realized, invalid or unimportant fault events are reduced, the number of alarms from developers is reduced, and the efficiency and accuracy of fault handling is improved.
Smart Images

Figure CN119537161B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fault warning, and particularly relates to a fault warning method based on JVM-Sandbox. Background Art
[0002] Nowadays, with the wide application of large-scale complex systems, the stability and reliability of the systems have become extremely important concerns. However, due to the complexity and diversity of the systems, fault troubleshooting has become extremely difficult. Conventional monitoring means often fail to accurately capture the specific scenarios of fault occurrence, resulting in a long fault handling cycle and low efficiency.
[0003] Existing methods usually include analyzing alarm notifications based on collected faults or alarm notifications based on error logs. The above methods have several technical problems: 1. Alarm events are not converged and denoised, resulting in too many fault event alarms; 2. There are a large number of fault events with low weights, increasing the number of fault alarms; 3. The same fault problems cannot be accurately identified, resulting in repeated alarms for the same fault; 4. Error log alarms are too redundant, mixed with a large number of invalid error logs.
[0004] To solve the above problems, the present invention proposes a scenario-based accurate fault warning method based on JVM-Sandbox. This method mainly lists common fault scenarios and uses the Agent method of JVM-Sandbox to achieve real-time collection and alarm notification of fault information. Summary of the Invention
[0005] In view of the above technical problems, the present invention provides a fault warning method based on JVM-Sandbox, which can help the development team more accurately discover fault problems and solve the problem of being interfered by a large number of invalid or not serious fault alarm events.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] A fault warning method based on JVM-Sandbox includes the following steps:
[0008] S1. Listen for events that need to be intercepted by the business application based on JVM-Sandbox;
[0009] S2. Intercept the BeforeEvent event, ReturnEvent event, and ThrowEvent event of each request, and obtain the data information of the events that need to be intercepted;
[0010] S3. Configure dynamic rules for fault warning, match and filter the data information based on the dynamic rules, obtain alarm events, and upload them to the cloud alarm platform;
[0011] S4. After receiving the alarm event, the cloud alarm platform counts the number of different alarm events within the time window and configures alarm rules for selective alarm.
[0012] In some embodiments of the present invention, in S1, the ModuleEventWatcher mechanism based on JVM-Sandbox listens for events that need to be intercepted by the business application.
[0013] In some embodiments of the present invention, in S1, the business application is loaded in the form of an Agent.
[0014] In some embodiments of the present invention, in S2: for the Dubbo call method, intercept the BeforeEvent event, ReturnEvent event, and ThrowEvent event of FutureFilter and GenericFilter in the Dubbo framework; for the Controller call method, intercept the BeforeEvent event, ReturnEvent event, and ThrowEvent event of DispatcherServlet in the Spring framework; for the MyBatis call method, intercept the BeforeEvent event, ReturnEvent event, and ThrowEvent event of MapperMethod in the MyBatis framework; for the MQ call method, intercept the BeforeEvent event, ReturnEvent event, and ThrowEvent event of MessageListenerConcurrently and MessageListenerOrderly; for the Job call method, intercept the BeforeEvent event, ReturnEvent event, and ThrowEvent event of IJobHandler.
[0015] In some embodiments of the present invention, in S2, the data information includes metadata information, the business result at the time of call, and the fault event information at the time of call. The metadata information is obtained through the BeforeEvent event, the business result of the call is parsed through the ReturnEvent event, and the fault event information at the time of call is obtained through the ThrowEvent event.
[0016] In some embodiments of the present invention, step S3 includes: S31, configuring dynamic rules for fault alarms, where the dynamic rules include custom first keywords and second keywords, and there are multiple first keywords and second keywords; S32, using the dynamic rules to perform matching and filtering on metadata information, service results, and fault event information. If at least one first keyword is included in the string of the exception stack in the metadata information or fault event information, or at least one second keyword is included in the service result, then the event corresponding to the current metadata information, fault event information, or service result that needs to be intercepted is marked as an alarm event and the alarm event is asynchronously uploaded to the cloud alarm platform.
[0017] In some embodiments of the present invention, the first keywords include strings representing null pointer exception, database deadlock exception, no service provider for Dubbo call, SQL execution exception, and internal server error respectively.
[0018] In some embodiments of the present invention, the second keywords include strings representing insufficient API call times and please recharge first.
[0019] In some embodiments of the present invention, in step S4, the number of alarm events of different fault types is counted within a time window.
[0020] In some embodiments of the present invention, configuring alarm rules for selective alarm includes: within a specified time window, if an alarm event of a certain fault type has been alarmed, then after counting the current alarm event, directly ignore the alarm event; if not, directly alarm; within a specified time window, call a specified interface of a specified business application, and based on the service result, if there are multiple call failures within the specified time window, directly alarm.
[0021] Compared with the prior art, the present invention has the following beneficial effects:
[0022] 1. The present invention listens for events to be intercepted through JVM-Sandbox, and then loads the business application, realizing asynchronous listening, collection, and reporting of fault events to the cloud alarm platform.
[0023] 2. The present invention realizes the collection of accurate fault events through dynamic rules for fault alarms, thereby reducing a large number of invalid or less important fault events.
[0024] 3. The present invention realizes the noise reduction and convergence ability of faults within a time window on the cloud fault alarm platform according to data statistics within the time window, thereby greatly reducing the fault events received by developers.
[0025] 4. The present invention filters out the alarm events that require manual intervention through the matching and filtering of the dynamic rules of fault alarms, and alarms the fault problems that developers really need to care about. The final achieved effect is that an alarm is a problem. Description of the Drawings
[0026] Figure 1 It is the flowchart of the method of the present invention. Detailed Implementation Modes
[0027] Term Explanation:
[0028] JVM-Sandbox is a non-invasive runtime aspect-oriented programming (AOP) solution for the Java Virtual Machine (JVM) platform.
[0029] The Agent method is a method that uses an agent workflow and artificial intelligence agents to complete tasks.
[0030] The ModuleEventWatcher mechanism is a technology used to monitor and respond to module-level events. It allows developers or system administrators to set specific observers to monitor the activities of modules and trigger corresponding operations or reactions when specific events occur.
[0031] FutureFilter is a filter in the Dubbo framework for handling event notifications.
[0032] GenericFilter is a custom filter in the Dubbo framework.
[0033] DispatcherServlet is a front controller in the Spring framework.
[0034] MapperMethod is an inner class in the MyBatis framework.
[0035] MessageListenerConcurrently is an interface in RocketMQ for concurrent message consumption.
[0036] MessageListenerOrderly is a consumer interface in RocketMQ.
[0037] RocketMQ is a distributed message middleware.
[0038] IJobHandler is an interface in the XXL-JOB framework.
[0039] XXL-JOB is a distributed task scheduling platform.
[0040] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts fall within the scope of protection of the present invention. In the description of the present invention, it should be noted that the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0041] As Figure 1 shown, a fault warning method based on JVM-Sandbox provided by the present invention includes the following steps:
[0042] S1. Listen for events that need to be intercepted by the business application based on JVM-Sandbox;
[0043] S2. Intercept the BeforeEvent event, ReturnEvent event and ThrowEvent event of each request, and obtain the data information of the events that need to be intercepted;
[0044] S3. Configure dynamic rules for fault warning, match and filter the data information based on the dynamic rules, obtain warning events and upload them to the cloud warning platform;
[0045] S4. After the cloud warning platform receives the warning events, count the number of different warning events within the time window, and configure warning rules for selective warning.
[0046] The present invention listens for events that need to be intercepted through JVM-Sandbox, then loads the business application, and then matches and filters out warning events that require human intervention through dynamic rules for fault warning, reducing a large number of ineffective or less important fault events. Finally, on the cloud fault warning platform, according to the data statistics within the time window, the noise reduction and convergence ability of faults within the time window is realized, thus greatly reducing the number of fault events received by developers.
[0047] In one of the embodiments, the S1 listens for events that need to be intercepted by the business application based on the ModuleEventWatcher mechanism of JVM-Sandbox. The ModuleEventWatcher mechanism is an event listening mechanism based on the observer pattern, which allows corresponding observers to be notified when the target object changes.
[0048] In one of the embodiments, the S1 loads business applications in the Agent mode. The Agent mode is a method that uses an agent workflow and artificial intelligence agents to complete tasks.
[0049] In one of the embodiments, events that need to be intercepted under different call methods are intercepted. By intercepting the BeforeEvent event, ReturnEvent event, and ThrowEvent event of different filters or interfaces, the event data information that needs to be intercepted is obtained. The data information includes metadata information, business results during the call, and fault event information during the call. Among them, metadata information is obtained through the BeforeEvent event; the business results of the call are parsed through the ReturnEvent event, which facilitates the realization that in some scenarios, it is possible to selectively determine whether an alarm is required based on the business results; the fault event information during the call is obtained through the ThrowEvent event, and then selective fault alarms can be realized according to information such as the fault event type.
[0050] For the Dubbo call method, the BeforeEvent event, ReturnEvent event, and ThrowEvent event of FutureFilter and GenericFilter in the Dubbo framework are intercepted to obtain the data information of the events that need to be intercepted (RPC events). Among the data information, the metadata information includes request time, request URL, request parameters, response parameters, the application to which it belongs, client IP, exception stack, current user ID, and current user name.
[0051] For the Controller call method, the BeforeEvent event, ReturnEvent event, and ThrowEvent event of DispatcherServlet in the Spring framework are intercepted to obtain the data information of the events that need to be intercepted (HTTP events). Among the data information, the metadata information includes request time, request URL, request parameters, response parameters, the application to which it belongs, client IP, exception stack, current user ID, and current user name.
[0052] For the MyBatis call method, the BeforeEvent event, ReturnEvent event, and ThrowEvent event of MapperMethod in the MyBatis framework are intercepted to obtain the data information of the events that need to be intercepted (SQL execution events). Among the data information, the metadata information includes request time, request Mapper address, request parameters, response parameters, the application to which it belongs, executed SQL, exception stack, current user ID, and current user name.
[0053] For the MQ call method (taking RocketMQ as an example), intercept the BeforeEvent event, ReturnEvent event, and ThrowEvent event of MessageListenerConcurrently and MessageListenerOrderly to obtain the data information of the events to be intercepted (MQ consumption execution events). The metadata information in the data information includes the request time, Topic (the first-level classification of messages), Group (a type of producer or consumer), Tags (the second-level classification of messages), request parameters, response parameters, the application to which it belongs, the exception stack, the current user ID, and the current user name.
[0054] For the Job call method (taking XXL-JOB as an example), intercept the BeforeEvent event, ReturnEvent event, and ThrowEvent event of IJobHandler to obtain the data information of the events to be intercepted (scheduled task execution events). The metadata information in the data information includes the request time, task address, task name, request parameters, response parameters, the application to which it belongs, and the exception stack.
[0055] Further, step S3 includes: S31, configuring dynamic rules for fault alarms, where the dynamic rules include custom first keywords and second keywords, and there are multiple first keywords and second keywords; S32, using the dynamic rules to match and filter metadata information, service results, and fault event information. If at least one first keyword is included in the string of the exception stack in the metadata information or / and fault event information, the event to be intercepted corresponding to the current metadata information or fault event information is marked as an alarm event and the alarm event is asynchronously uploaded to the cloud alarm platform. The exception stack represents the scenario when a call throws an exception. By setting keywords indicating that a call throws an exception, if the string of the exception stack contains the set keywords, it indicates that the event to be intercepted is an exception and an alarm is required. The first keywords include the string representing a null pointer exception (NullPointerException), the string representing a database deadlock exception (Deadlock), the string representing no service provider available for Dubbo calls (No provider available), the string representing an SQL execution exception (SQLException), the string representing an internal server error (a custom string according to the specific error), and other first keywords can also be set or customized according to the actual situation. The above mainly matches the scenario when a call throws an exception, but in actual situations, there may be no exception thrown, but the service result does not meet the expectation, and an alarm is also required. So here, the alarm is based on the service result. If at least one second keyword is included in the service result, the second keywords include the string representing insufficient API call times and please recharge first. That is, when calling a third-party API interface, if the number of times is exhausted, instead of throwing an exception, it is returned through the parameter status or description. When the keyword is found in the service result, the event to be intercepted corresponding to the current service result can also be marked as an alarm event. Other second keywords can also be set or customized according to the actual situation.
[0056] In one of the embodiments, after receiving the alarm event, the cloud alarm platform counts the number of alarm events of different fault types within the time window. If within the specified time window, if an alarm event of a certain fault type has been alarmed, the current alarm event is counted and then directly ignored. If it has not been alarmed (for example, only counted once within a 5-minute time window), it is directly alarmed, so as to prevent a large number of repeated alarms of the same fault type; within the specified time window, call the specified interface of the specified business application. Based on the service result, if there are multiple call failures within the specified time window, directly alarm (for example, if there are 20 consecutive call failures within a 5-minute time window).
[0057] The present invention realizes data information collection based on JVM-Sandbox, and then by configuring dynamic rules for fault warning, where the dynamic rules contain keywords that require human intervention in the scenario, so as to achieve the effect of accurate warning. Only by configuring the keywords can the warning events be filtered out, greatly improving the flexibility of the solution, and thus more abnormal scenarios can be dealt with.
[0058] Finally, it should be noted that the above embodiments are only preferred embodiments of the present invention to illustrate the technical solutions of the present invention, rather than limiting it, and certainly not limiting the patent scope of the present invention; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention; that is to say, any meaningless changes or polishings made in the main design concept and spirit of the present invention, as long as the technical problems solved are still the same as those of the present invention, should be included in the protection scope of the present invention; in addition, directly or indirectly applying the technical solutions of the present invention to other related technical fields shall also be included in the patent protection scope of the present invention by the same token.
Claims
1. A fault warning method based on JVM-Sandbox, characterized in that, It includes the following steps: S1. Listen for events that need to be intercepted in the business application based on JVM-Sandbox; S2. Intercept the BeforeEvent event, ReturnEvent event, and ThrowEvent event of each request, and obtain the data information of the events that need to be intercepted; S3. Configure dynamic rules for fault warning, match and filter the data information based on the dynamic rules, obtain warning events and upload them to the cloud warning platform; S4. After the cloud warning platform receives the warning events, count the number of different warning events within the time window, and configure warning rules for selective warning; In S2: For the Dubbo call method, intercept the BeforeEvent event, ReturnEvent event, and ThrowEvent event of FutureFilter and GenericFilter in the Dubbo framework; for the Controller call method, intercept the BeforeEvent event, ReturnEvent event, and ThrowEvent event of DispatcherServlet in the Spring framework; for the MyBatis call method, intercept the BeforeEvent event, ReturnEvent event, and ThrowEvent event of MapperMethod in the MyBatis framework; for the MQ call method, intercept the BeforeEvent event, ReturnEvent event, and ThrowEvent event of MessageListenerConcurrently and MessageListenerOrderly; for the Job call method, intercept the BeforeEvent event, ReturnEvent event, and ThrowEvent event of IJobHandler; In S2, the data information includes metadata information, business results at the time of call, and fault event information at the time of call. Obtain the metadata information through the BeforeEvent event, parse the business results of the call through the ReturnEvent event, and obtain the fault event information at the time of call through the ThrowEvent event; In S3, it includes: S31. Configure dynamic rules for fault warning. The dynamic rules include custom first keywords and second keywords, and there are multiple first keywords and second keywords; S32. Use the dynamic rules to match and filter the metadata information, business results, and fault event information. If at least one first keyword is included in the string of the exception stack in the metadata information or fault event information, or at least one second keyword is included in the business results, then mark the event that needs to be intercepted corresponding to the current metadata information, fault event information, or business results as a warning event and asynchronously upload the warning event to the cloud warning platform; The first keyword includes strings representing null pointer exception, database deadlock exception, no service provider for Dubbo call, SQL execution exception, and internal server error respectively; The second keyword includes strings representing insufficient API calls and please recharge first.
2. The fault warning method based on JVM-Sandbox according to claim 1, wherein In S1, the ModuleEventWatcher mechanism based on JVM-Sandbox listens for events that need to be intercepted by the business application.
3. The fault warning method based on JVM-Sandbox according to claim 1, wherein In S1, the business application is loaded through the Agent method.
4. The fault warning method based on JVM-Sandbox according to claim 1, characterized in that In S4, the number of alarm events of different fault types is counted within the time window.
5. The fault warning method based on JVM-Sandbox according to claim 4, characterized in that, Configuring alarm rules for selective alarm includes: within the specified time window, if an alarm event of a certain fault type has been alarmed, the current alarm event is counted and then directly ignored; if not, it is directly alarmed; within the specified time window, the specified interface of the specified business application is called, and if there are multiple call failures within the specified time window based on the business result, it is directly alarmed.
Citation Information
Patent Citations
Alarm monitoring method in cloud environment
CN112532456A
Service room remote scheduling method and device, equipment and storage medium
CN118585352A