An abnormal acquisition method, device, electronic device and storage medium
By collecting stack trace data and internal error data of virtual machines, combined with Bloom filters and message queues, efficient collection and accurate positioning of Java virtual machine exceptions is achieved, solving the problem of impossible to accurately locate exceptions in the existing technology, and improving the accuracy of exception analysis and data processing efficiency.
Patent Information
- Application Number
- CN202411531122.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-10-30
AI Technical Summary
In the prior art, Java virtual machine exception analysis cannot accurately locate the location of the exception, and only the error code can be obtained, and the cause of the exception cannot be effectively located.
By collecting stack trace data and internal error data of virtual machine, using the Bloom filter to count the occurrence of exceptions and error data, sampling according to the sampling frequency, and obtaining exception data by filling the stack trace information, combining message queues and uploading arrays for batch upload.
It improves the acquisition integrity and analysis accuracy of abnormal data, reduces data processing volume, and reduces memory consumption and network transmission pressure.
Smart Images

Figure CN119493684B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to an abnormal data collection method, device, electronic device and storage medium. Background Art
[0002] All kinds of network services are generally implemented through underlying codes. When an abnormality occurs in a network service, it is necessary to analyze the underlying code of the network service to obtain the cause of the abnormality of the network service, so as to improve the underlying code of the network service according to the cause of the abnormality.
[0003] The underlying code of a network service can be implemented in the Java language. The running environment of the Java language is usually implemented through a Java Virtual Machine (JVM). The Java Virtual Machine (JVM) is an abstract computer, which is implemented by simulating various computer functions on an actual computer. It is a hypothetical computer that can run Java code. As long as the interpreter is transplanted to a specific computer according to the JVM specification description, it can be ensured that any compiled Java code can run on this system.
[0004] In related technologies, when obtaining an exception and an error, usually only an error code can be obtained, and this error code can only identify a general exception type, but cannot locate the position where the exception occurs. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide an abnormal data collection method, device, electronic device and storage medium to improve the completeness of abnormal data collection and reduce the amount of data processing.
[0006] According to one aspect of the present invention, there is provided an abnormal data collection method, the method including:
[0007] Collecting basic data through a preset method, where the basic data includes stack trace data of abnormal data and internal error data of the virtual machine, and the stack trace data is used to identify the call chain of the abnormal data;
[0008] Based on the occurrence times of various types of abnormal data and error data in the basic data, sampling the abnormal data and error data corresponding to each type according to a sampling frequency preset for the occurrence times to obtain data to be uploaded; where the occurrence times are obtained through a Bloom filter;
[0009] Adding the data to be uploaded to an upload array to perform batch upload of the data to be uploaded in the upload array.
[0010] In a possible embodiment, the method further includes:
[0011] When the upload array is full, encapsulate the data to be uploaded into the data block corresponding to the service identifier of the data to be uploaded based on the service identifier, and add the data block to the message queue;
[0012] When a preset callback condition is met, call a preset callback function to process the data blocks in the message queue. Among them, the logic of the preset callback function at least includes sending data blocks or storing data blocks, and the preset callback condition is that the number of data blocks in the message queue reaches a preset number or the waiting time of the data blocks in the message queue exceeds a preset duration;
[0013] In a possible embodiment, the collecting of the basic data in the preset manner includes:
[0014] Embed code in the fillInStackTrace() method to obtain the exception data captured by throwable, and obtain the stack trace data corresponding to the exception data through the getStackTrace() method of throwable; and,
[0015] Add an uncaught error data operation function to each thread of the JVM virtual machine. The uncaught error data operation function is used to capture error events when the thread runs abnormally, and obtain the stack trace data of the error events through the getStackTrace() method of throwable.
[0016] In a possible embodiment, the method further includes:
[0017] For each piece of basic data, use a preset hash function to map the basic data to obtain each mapping position of the basic data in the Bloom filter, and increment the counter count result of the mapping position by 1;
[0018] Based on the statistical value of the counter count results corresponding to the basic data in the Bloom filter as the occurrence times of the basic data, where the statistical value is the minimum value or the average value.
[0019] In a possible embodiment, the sampling of the exception data and error data of each type corresponding to the basic data according to a preset sampling frequency for the occurrence times to obtain the data to be uploaded includes:
[0020] For each type of abnormal data and error data in the basic data, sample the abnormal data and error data of the type according to the sampling order corresponding to the frequency interval in which the number of occurrences of the type is located, where the sampling order is generated based on the frequency interval and the sampling frequency, and is used to identify the order of the data to be uploaded that needs to be sampled in the abnormal data and error data of the type.
[0021] According to another aspect of the present invention, there is provided an abnormal data collection device, the device includes:
[0022] A collection module, configured to collect basic data in a preset manner, where the basic data includes stack trace data of abnormal data and internal virtual machine error data, and the stack trace data is used to identify the call link of the abnormal data;
[0023] A sampling module, configured to sample the abnormal data and error data corresponding to each type according to a preset sampling frequency based on the number of occurrences of each type of abnormal data and error data in the basic data, to obtain data to be uploaded; where the number of occurrences is obtained through a Bloom filter;
[0024] An upload module, configured to add the data to be uploaded to an upload array, so as to batch upload the data to be uploaded in the upload array.
[0025] In a possible embodiment, the device further includes:
[0026] A waiting module, configured to, when the upload array is full, encapsulate the data to be uploaded into a data block corresponding to the service identifier of the data to be uploaded based on the service identifier of the data to be uploaded, and add the data block to a message queue; when a preset callback condition is reached, call a preset callback function to process the data block in the message queue, where the logic of the preset callback function at least includes sending the data block or storing the data block, and the preset callback condition is that the number of data blocks in the message queue reaches a preset number or the waiting time of the data block in the message queue exceeds a preset duration;
[0027] In a possible embodiment, the collecting the basic data in a preset manner includes:
[0028] By embedding code in the fillInStackTrace() method to obtain abnormal data captured by throwable, and obtaining the stack trace data corresponding to the abnormal data through the getStackTrace() method of throwable; and,
[0029] Add an uncaught error data operation function to each thread of the JVM virtual machine. The uncaught error data operation function is used to capture error events when the thread runs abnormally, and obtain the stack trace data of the error event through the getStackTrace() method of throwable;
[0030] The sampling module is used to map each piece of basic data using a preset hash function to obtain each mapping position of the basic data in the Bloom filter, and increment the counter count result of the mapping position by 1;
[0031] Use the statistical value of the counter count result corresponding to the basic data in the Bloom filter as the occurrence times of the basic data, where the statistical value is the minimum value or the average value;
[0032] Sampling the abnormal data and error data corresponding to each type according to the sampling frequency preset for the occurrence times based on the occurrence times of each type of abnormal data and error data in the basic data to obtain the data to be uploaded, including:
[0033] For each type of abnormal data and error data in the basic data, sample the abnormal data and error data of the type according to the sampling order corresponding to the range of the occurrence times of the type. The sampling order is generated based on the range of times and the sampling frequency, and is used to identify the order of the data to be uploaded that needs to be sampled in the abnormal data and error data of the type.
[0034] According to another aspect of the present invention, there is provided an electronic device, including:
[0035] A processor; and
[0036] A memory storing a program,
[0037] wherein the program includes instructions that, when executed by the processor, cause the processor to execute any one of the above-mentioned abnormal data collection methods.
[0038] According to another aspect of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute any one of the above-mentioned abnormal data collection methods.
[0039] One or more technical solutions provided in the embodiments of the present invention can collect stack data of abnormal data and internal error data of the virtual machine by setting different collection methods for different exceptions and error types, improving the integrity of the collected basic data, facilitating the tracing of exceptions, and thus improving the accuracy of subsequent exception analysis. Furthermore, by sampling the corresponding basic data based on the occurrence times of abnormal data and error data, while retaining the abnormal and error information, the amount of abnormal data and error data that needs to be uploaded is greatly reduced, and thus the data processing amount for subsequent exception analysis is reduced. Description of the Drawings
[0040] In the following description of exemplary embodiments with reference to the accompanying drawings, more details, features, and advantages of the present invention are disclosed. In the drawings:
[0041] Figure 1 It is a schematic flowchart of an abnormal data collection method provided by an embodiment of the present invention;
[0042] Figure 2 It is a schematic flowchart of collecting general abnormal data in an embodiment of the present invention;
[0043] Figure 3 It is a schematic flowchart of collecting internal error data of the virtual machine in an embodiment of the present invention;
[0044] Figure 4 It is a schematic diagram of counting basic data in an embodiment of the present invention;
[0045] Figure 5 It is a schematic flowchart of obtaining the occurrence times of abnormal and error data in an embodiment of the present invention;
[0046] Figure 6 It is a schematic flowchart of collecting abnormal and error data in an embodiment of the present invention
[0047] Figure 7 It is a schematic flowchart of uploading data to be uploaded in an embodiment of the present invention;
[0048] Figure 8 It is a schematic structural diagram of an abnormal data collection device provided by an embodiment of the present invention;
[0049] Figure 9 It shows a structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present invention. Detailed Embodiments
[0050] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.
[0051] It should be understood that the various steps recited in the method embodiments of the present invention can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.
[0052] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependent relationships.
[0053] It should be noted that the modifications of "one" and "a plurality" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0054] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0055] In order to improve the integrity of abnormal data collection and reduce the amount of abnormal analysis data, embodiments of the present invention provide an abnormal collection method, device, electronic device and storage medium. The abnormal collection method provided by the embodiments of the present invention can be applied to any electronic device with abnormal collection function. The electronic device can be a server, a computer or a mobile terminal, etc. The solution of the present invention will be described below with reference to the accompanying drawings:
[0056] Figure 1 It is a schematic flowchart of an abnormal collection method provided by an embodiment of the present invention, and may include the following steps:
[0057] S101. Collect basic data through a preset method. Among them, the basic data includes stack trace data of abnormal data and internal error data of the virtual machine. The stack trace data is used to identify the call chain of the abnormal data;
[0058] S102. Based on the occurrence times of various types of abnormal data and error data in the basic data, sample the abnormal data and error data corresponding to each type according to the sampling frequency preset for the occurrence times to obtain data to be uploaded. Among them, the occurrence times are obtained through a Bloom filter;
[0059] S103. Add the data to be uploaded to the upload array to batch upload the data to be uploaded in the upload array.
[0060] Applying the embodiment of the present invention, by setting different collection methods for different types of exceptions and errors, stack data of abnormal data and internal error data of the virtual machine can be collected, improving the integrity of the collected basic data, facilitating the tracing of exceptions, and thus improving the accuracy of subsequent exception analysis. Moreover, by sampling the corresponding basic data based on the occurrence times of abnormal data and error data, while retaining the exception and error information, the amount of abnormal data and error data to be uploaded is greatly reduced, and further the data processing volume of subsequent exception analysis is reduced.
[0061] The following is an exemplary description of the above S101 - S103:
[0062] The abnormal data collection method provided by the present invention can be applied to scenarios such as operating systems and application APPs to collect abnormal data and error data in these scenarios. In a possible embodiment, the above preset method may include a collection method for general abnormal and error data and a collection method for internal error data of the JVM virtual machine. Among them, an exception is a mechanism used to represent errors or unexpected situations that occur during the execution of a program. Exceptions provide a way to separate error handling code from the normal code execution path, enabling the program to better manage and handle error conditions. The exception mechanism in Java is implemented through try - catch blocks and exception classes (such as Exception and Error). An error refers to serious problems that occur during the execution of a program. These problems usually go beyond the control scope of the program itself and cannot be recovered through the normal exception handling mechanism. Errors usually represent system - level problems, such as virtual machine errors and memory shortages. Such errors are represented by the Error class and its subclasses.
[0063] As a possible implementation, the above basic data can be collected through the following steps:
[0064] By instrumenting the `fillInStackTrace()` method to obtain the exception data captured by the `throwable`, and obtaining the stack trace data corresponding to the exception data through the `getStackTrace()` method of the `throwable`; and,
[0065] Add an uncaught error data operation function to each thread of the JVM virtual machine. The uncaught error data operation function is used to capture error events when the thread runs abnormally, and obtain the stack trace data of the error event through the `getStackTrace()` method of the `throwable`.
[0066] In a possible embodiment, for ordinary exception data and error data, the `public synchronized Throwable fillInStackTrace()` method in `java.lang.Throwable` can be instrumented. Specifically, when calling `fillInStackTrace()`, the data of the `Throwable` can be collected through the return value, and the corresponding stack trace data can be obtained through the `getStackTrace()` method of the `Throwable`. Among them, `Throwable` is the superclass of all errors and exceptions in Java. It is a throwable object representing an error or exception condition that occurs during the execution of a program. The `Throwable` class has two main subclasses: `Error` (error) and `Exception` (exception), and ordinary exception data and error data are exceptions that can be captured by the `throwable`. The stack trace data includes the location where the exception occurred and the dependency parameter path of the exception. The above location where the exception occurred is the specific line of the specific code where the exception is located, and the dependency parameter path of the exception is the parameter on which the exception depends and the code information that generates the parameter.
[0067] The `fillInStackTrace()` method is a synchronized method of the `java.lang.Throwable` class in Java, which is used to fill in the stack trace information of an exception. This method is usually automatically called when an exception is caught to obtain the stack trace at the time of the exception.
[0068] Such as Figure 2 , Figure 2It is a schematic flowchart for collecting abnormal and error data in an embodiment of the present invention. Specifically, in the case of an abnormality or error, the constructor of Throwable is called to call the fillInStackTrace() method to collect the Throwable object, that is, the data where the abnormality or error occurs, and the getStackTrace() method is called to obtain the stack trace data of the object. The stack trace data includes the specific location of the data where the abnormality or error occurs in the code and the call chain of the data, that is, the code location where the parameters on which the data depends are generated.
[0069] In the related art, generally, through the pre-buried point method, only the exceptions and errors thrown by the buried point can be collected. If the exceptions and errors are not thrown, they cannot be collected. In the embodiment of the present invention, the Throwable object is actively obtained through the fillInStackTrace() method to achieve the acquisition of all exception and error objects. In addition, in the related art, generally only the exception code is returned for an exception, resulting in the inability to locate the root cause of the problem even if the exception is seen. However, in the embodiment of the present invention, the exception information and the stack trace data of the exception are returned to achieve the traceability of the data causing the exception, facilitating exception location.
[0070] In the related art, generally, the error data inside the JVM virtual machine cannot be collected, such as the errors that cause the thread to exit or crash cannot be collected. In the present invention, the constructor method of java.lang.Thread can be embedded with code, and the custom UncaughtExceptionHandler is set to the thread to collect the uncaught exceptions in the custom UncaughtExceptionHandler. The above uncaught exceptions generally refer to the exceptions not caught by Throwable. Specifically, as Figure 3 shown Figure 3 It is a schematic flowchart for collecting the error data inside the virtual machine provided by an embodiment of the present invention. Specifically, after the program runs and the thread starts, the thread constructor method is called to set the custom UncaughtExceptionHandler. If an uncaught exception occurs during the thread running, the uncaught exception can be captured and recorded through the UncaughtExceptionHandler.
[0071] Each piece of the above-collected basic data can include its exception information and stack trace data. The above exception information can include the type of the exception, the occurrence time, the occurrence location, and so on.
[0072] Through the above steps, all abnormal and error data can be collected. However, there may be the same abnormal or error data in this abnormal and error data. For example, the same error is collected multiple times, which will greatly affect the efficiency of data processing. Moreover, if a map in the key-value form is used and there are too many abnormal types within a time period, the occupied memory space will be uncontrollable, which is a great waste of memory space. Therefore, the above-collected basic data can be filtered.
[0073] In a possible embodiment, the filtering of duplicate basic data can be achieved through a Bloom filter. A Bloom filter is a combination of a very long binary vector and a series of random mapping functions, which can be used to retrieve whether an element is in a set. And the Bloom filter can record the occurrence times of data by counting, without storing the specific content of the data, which greatly saves memory space.
[0074] As a possible implementation, for each piece of basic data, a preset hash function can be used to map the basic data to obtain each mapping position of the basic data in the Bloom filter, and the counter count result at the mapping position is incremented by 1;
[0075] The statistical value of the count result of the counter corresponding to the basic data in the Bloom filter is used as the occurrence times of the basic data, where the statistical value is the minimum value or the average value.
[0076] In a possible embodiment, the mapping of different types of basic data can be performed separately. The type of the basic data can be distinguished by an error code identifier or a code identifier for an error or an exception occurring. The above error code identifier can illustrate the specific abnormal type. For example, 404 indicates that the page is not found, 401 indicates unauthorized access, etc. The above code identifier can be the code name of the exception occurring and the identifier of the line where the exception code is located. The basic data with the same error code identifier or code identifier is the same type of data.
[0077] There can be multiple above-mentioned preset hash functions. The specific number and the specific function content of the hash functions can be flexibly set according to the actual application scenario. By using the multiple preset hash functions to map the basic data, the hash value corresponding to the basic data can be obtained, and this hash value can identify the position of the basic data in the Bloom filter. The same type of basic data can obtain the same hash value after being mapped by the hash function.
[0078] In a possible embodiment, in order to avoid the hash value overflowing the storage range of the Bloom filter, the length of the Bloom filter can be set to a relatively large value, which can be specifically set according to the actual business scenario.
[0079] Each position in the Bloom filter is a counter. After obtaining the hash value of the basic data, the counter at the position corresponding to the hash value can be incremented by 1. As Figure 4 shown, for the basic data, it is mapped through hash function 1, hash function 2, and hash function 3 to obtain the corresponding counter positions 1, 2, and 3. Therefore, the counts at counter positions 1, 2, and 3 can be incremented by 1.
[0080] Subsequently, the occurrence times of the basic data of the same type can be queried based on this Bloom filter. Exemplarily, the statistical value of the count results at the counter positions corresponding to the basic data of the type can be used as the occurrence times of the basic data of this type. This statistical value can be the average value, the median value, or the minimum value, etc. As a possible implementation manner, since each basic data may be mapped to the same position after being mapped by the multiple preset hash functions, the minimum value of the count values in the counter corresponding to the basic data may be the position where other basic data is mapped the least, that is, the value closer to the occurrence times of the basic data. Therefore, the minimum count value in each counter can be used as the occurrence times of the basic data to improve the accuracy of the judgment of the occurrence times.
[0081] In a possible embodiment, the above Bloom filter can store each position counter and its count result through a counter array. As Figure 5 shown, Figure 5 is a schematic flowchart for taking the minimum value of the count results of each counter as the occurrence times of the basic data. Specifically, the basic data to be queried is determined, and it is mapped based on the preset hash function to determine the count value in the counter position recording this basic data, and the minimum value among the count values is determined as the approximate value of the occurrence times of this basic data. Exemplarily, when recording the occurrence times of the basic data, it is mapped through hash function 1, hash function 2, and hash function 3, and the count is incremented once in the corresponding counter positions 1, counter position 2, and counter position 3. When querying an element, it can be mapped through hash function 1, hash function 2, and hash function 3 to the basic data to be queried, and the count values of counter positions 1, counter position 2, and counter position 3 are obtained from the counter array. The counter array includes counter positions 1, counter position 2, counter position 3, counter position 4, and counter position 5. Based on the count values of counter positions 1, counter position 2, and counter position 3, the minimum value calculation is performed to obtain the number of elements, that is, the occurrence times, of the basic data to be queried.
[0082] In a possible embodiment, a new counting Bloom filter can be initialized per unit time. During this unit time, for each exception handling, it is necessary to first obtain the number of current type records through the counting Bloom filter. If the number of records is 0, then insert directly and collect this exception directly. If the number of records is not 0, then the basic data can be sampled and collected. The above unit time can be set according to the actual application scenario, such as 1 hour, 6 hours, etc.
[0083] In a possible embodiment, for each type of exception data and error data in the basic data, the exception data and error data of the type can be sampled based on the sampling order corresponding to the number range in which the occurrence times of the type are located, where the sampling order is generated based on the number range and the sampling frequency and is used to identify the order of the data to be uploaded that needs to be sampled among the exception data and error data of the type.
[0084] Both the above number range and the sampling frequency can be set according to the actual application scenario, as long as the rule that the higher the number of exception occurrences, the lower the sampling frequency is followed. Exemplarily, it can be set that in the 1 - 10 stage, the sampling rate is 10%, in the 10 - 100 stage, the sampling rate is 2%, and in the 100 - 1000 stage, the sampling rate is 0.5%. In a possible embodiment, the sampling order can be generated based on the number range and the sampling rate. Exemplarily, the sampling rate for 10 - 100 is 2%, so two random numbers in the range of 10 - 100 can be generated as the sampling order.
[0085] When the number of occurrences of the basic data reaches the number range, the basic data that conforms to the sampling order corresponding to the number range can be collected. Exemplarily, based on the above example, the sampling order generated for the 10 - 100 range is 45 and 76. Therefore, the 45th and 76th occurrences of the basic data of this type can be collected.
[0086] As Figure 6 shown, Figure 6 This is a schematic flowchart of a process for collecting basic data in an embodiment of the present invention. Specifically, a new counting Bloom filter is initialized per unit time, and a sampling rule is preset. Specifically, in the 1 - 10 stage, the sampling rate is 10%, in the 10 - 100 stage, the sampling rate is 2%, and in the 100 - 1000 stage, the sampling rate is 0.5%. After obtaining the exception data, the value of the type record of the exception data is obtained through the counting Bloom filter. If the number of records is 0, then this exception is directly collected. If the number of records is not 0, then it is calculated according to the sampling rule whether to collect. Specifically, it is determined whether the occurrence order of the exception is the preset sampling order. If so, then this exception data can be collected. If not, then this exception data is not collected.
[0087] For the convenience of description, in the present invention, the collected basic data is referred to as data to be uploaded. The data to be uploaded refers to abnormal or error data that needs to be uploaded to the server for further anomaly analysis. The data to be uploaded may include anomaly information and corresponding stack trace data.
[0088] Since the collected stack trace data is a relatively large amount of data, if too much data is sent at one time, it will cause the data packet to be too large, resulting in a long network transmission time and container timeout failure. If the sent packet is small but the frequency is too high, it is easy to cause waste of resources, resulting in frequent network and high CPU consumption. Based on this, an upload array transferDataArray can be preset, and the maximum amount of data that can be uploaded at one time can be specified in the upload array. As a possible implementation, the data to be uploaded can be added to the transferDataArray by calling the add method.
[0089] In a possible embodiment, the above method may further include, when the upload array is full, encapsulating the data to be uploaded into a data block corresponding to the service identifier of the data to be uploaded, and adding the data block to a message queue.
[0090] When a preset callback condition is met, a preset callback function is called to process the data block in the message queue. The logic of the preset callback function at least includes sending the data block or storing the data block. The preset callback condition is that the number of data blocks in the message queue reaches a preset number or the waiting time of the data block in the message queue exceeds a preset duration.
[0091] When the upload array is full or the preset upload time is reached, the data in the upload array can be uploaded. In a possible embodiment, if the upload array is full or batchadd is false, that is, data cannot be uploaded, the data to be uploaded can be encapsulated into a data block (in actual code, it can be encapsulated into a TransferDataBlock object), and the data block is added to the message queue TransferDataBlockQueue, which is a blocking queue. As a possible implementation, the data to be uploaded belonging to the same service can be encapsulated into the same data block, so that it is convenient to store the abnormal data of the cloud storage service.
[0092] In a possible embodiment, a consuming thread can be set up. This consuming thread can retrieve data blocks from the TransferDataBlockQueue queue and process these data blocks. The processing method can be to send the data in the data block to the server by calling the process method of the TransferDataCallBack interface.
[0093] In a possible embodiment, a callback function can also be registered in the message queue, and the callback condition of the callback function can be set. Both the callback function and the callback condition can be set according to the actual application scenario. Exemplarily, the callback function can be set to send the data in the data block to the server by calling the process method of the TransferDataCallBack interface, or to call the batchToQueue method to add the temporarily stored data in the message queue to the upload array. The above callback condition can be that the number of data blocks in the message queue reaches a preset number or the waiting time of the data blocks in the message queue exceeds a preset duration.
[0094] In a possible embodiment, the callback condition can be set through the poll method of the TransferDataBlockQueue class. The poll method in the TransferDataBlockQueue class is usually used to retrieve and remove the head element from the queue. If the queue is empty and the time since the last callback exceeds the pollCallbackTimeOut (callback timeout), a callback will be triggered, and the poll methods of all registered NotificationCallbacks will be called. NotificationCallback is an interface that defines a set of callback methods. In the poll method of the TransferDataBlockProducer class, an attempt will be made to call the batchToQueue method to add the temporarily stored data to the queue.
[0095] As Figure 7 shown, Figure 7 This is a schematic diagram of a process for uploading data in an embodiment of the present invention. Specifically, after generating the data to be uploaded in the TransferDataBlockProducer class, when the add method is called, the data is added to the transferDataArray array. If the array is full or batchAdd is false, the data is immediately encapsulated into a TransferDataBlock object and added to the queue. When there is no data in the queue, the consumer thread will be blocked until a new data block is added to the queue.
[0096] Callback functions are pre-registered in the TransferDataBlockQueue. When the data in the queue meets the preset callback conditions, these callback functions will be triggered. The TransferDataTask is the data-consuming thread, responsible for taking out data blocks from the TransferDataBlockQueue and processing this data. The processing method can be to send the data in the data block to the server by calling the process method of the TransferDataCallBack interface.
[0097] Applying the embodiments of the present invention, by adding a control function in the JVM virtual machine to capture the internal error data of the virtual machine, the integrity of obtaining basic exception data is improved. Furthermore, by using a Bloom filter to count the basic data, the memory consumption is reduced. By sampling the duplicate exception data, the amount of exception data to be sent is greatly reduced. In addition, by setting up an array and a queue to upload the collected data, the balance of system performance and transmission efficiency is achieved, thus realizing efficient, rapid, and reasonable collection of exception data.
[0098] Based on the same inventive concept, the embodiments of the present invention also provide an exception collection device, as Figure 8 shown. The device 800 may include:
[0099] A collection module 801, configured to collect basic data in a preset manner, where the basic data includes stack trace data of exception data and internal error data of the virtual machine, and the stack trace data is used to identify the call chain of the exception data;
[0100] A sampling module 802, configured to sample the exception data and error data corresponding to each type according to a preset sampling frequency for the occurrence times of each type of exception data and error data in the basic data to obtain data to be uploaded; where the occurrence times are obtained through a Bloom filter;
[0101] An upload module 803, configured to add the data to be uploaded to an upload array to batch upload the data to be uploaded in the upload array.
[0102] In a possible embodiment, the device further includes:
[0103] A waiting module, configured to, when the upload array is full, encapsulate the data to be uploaded into the data block corresponding to the service identifier based on the service identifier of the data to be uploaded, and add the data block to a message queue; and when a preset callback condition is met, call a preset callback function to process the data blocks in the message queue, where the logic of the preset callback function at least includes sending data blocks or storing data blocks, and the preset callback condition is that the number of data blocks in the message queue reaches a preset number or the waiting time of the data blocks in the message queue exceeds a preset duration;
[0104] In a possible embodiment, the collecting basic data by a preset method includes:
[0105] By embedding codes in the fillInStackTrace() method to obtain the exception data captured by throwable, and obtaining the stack trace data corresponding to the exception data through the getStackTrace() method of throwable; and,
[0106] Adding an uncaught error data operation function to each thread of the JVM virtual machine, where the uncaught error data operation function is used to capture error events when the thread runs abnormally, and obtain the stack trace data of the error events through the getStackTrace() method of throwable;
[0107] The sampling module is configured to, for each piece of basic data, map the basic data by using a preset hash function to obtain each mapping position of the basic data in the Bloom filter, and increment the counter count result of the mapping position by 1;
[0108] Based on the statistical value of the counter count results corresponding to the basic data in the Bloom filter as the occurrence times of the basic data, where the statistical value is the minimum value or the average value;
[0109] The sampling of the exception data and error data of each type in the basic data according to a preset sampling frequency for the occurrence times to obtain the data to be uploaded includes:
[0110] For the exception data and error data of each type in the basic data, sample the exception data and error data of the type according to the sampling order corresponding to the number range where the occurrence times of the type are located, where the sampling order is generated based on the number range and the sampling frequency and is used to identify the order of the data to be uploaded that needs to be sampled in the exception data and error data of the type.
[0111] Among them, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved in the present invention all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0112] An exemplary embodiment of the present invention further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method according to the embodiment of the present invention.
[0113] An exemplary embodiment of the present invention further provides a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiment of the present invention.
[0114] An exemplary embodiment of the present invention further provides a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiment of the present invention.
[0115] Reference Figure 9 , the structural block diagram of the electronic device 900 that can be used as the server or client of the present invention will now be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0116] As Figure 9 shown, the electronic device 900 includes a computing unit 901, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 902 or the computer program loaded from the storage unit 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.
[0117] Multiple components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, an output unit 907, a storage unit 908, and a communication unit 909. The input unit 906 can be any type of device capable of inputting information into the electronic device 900. The input unit 906 can receive input digital or character information and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 907 can be any type of device capable of presenting information and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 908 can include, but is not limited to, a magnetic disk and an optical disc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0118] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above. For example, in some embodiments, the above-mentioned anomaly collection method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 900 via the ROM 902 and / or the communication unit 909. In some embodiments, the computing unit 901 can be configured to execute the above-mentioned anomaly collection method by any other suitable means (e.g., by means of firmware).
[0119] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0120] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0121] As used in the present invention, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a disk, optical disk, memory, programmable logic device (PLD)) that can be used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal that can be used to provide machine instructions and / or data to a programmable processor.
[0122] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0123] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0124] A computer system can include clients and servers. The clients and servers are generally far apart from each other and typically interact through a communication network. The client - server relationship is generated by computer programs that run on the respective computers and have a client - server relationship with each other.
Claims
1. An abnormal acquisition method, characterized in that, The method includes: Collecting basic data through a preset method, where the basic data includes stack trace data of abnormal data and internal error data of the virtual machine, and the stack trace data is used to identify the call chain of the abnormal data; Based on the occurrence times of various types of abnormal data and error data in the basic data, sampling the abnormal data and error data corresponding to each type according to the sampling frequency preset for the occurrence times to obtain data to be uploaded; where the occurrence times are obtained through a Bloom filter; Adding the data to be uploaded to an upload array to batch upload the data to be uploaded in the upload array; When the upload array is full, based on the service identifier of the data to be uploaded, encapsulating the data to be uploaded into a data block corresponding to the service identifier, and adding the data block to a message queue; When a preset callback condition is met, calling a preset callback function to process the data blocks in the message queue, where the logic of the preset callback function at least includes sending or storing the data blocks, and the preset callback condition is that the number of data blocks in the message queue reaches a preset number or the waiting time of the data blocks in the message queue exceeds a preset duration; The step of based on the occurrence times of various types of abnormal data and error data in the basic data, sampling the abnormal data and error data corresponding to each type according to the sampling frequency preset for the occurrence times to obtain data to be uploaded includes: For each type of abnormal data and error data in the basic data, sampling the abnormal data and error data of the type according to the sampling order corresponding to the range of the occurrence times of the type, where the sampling order is generated based on the range and the sampling frequency, and is used to identify the order of the data to be uploaded that needs to be sampled in the abnormal data and error data of the type; the sampling frequency is preset for different ranges, and the higher the upper limit value of the range, the lower the sampling frequency.
2. The method according to claim 1, wherein The step of collecting basic data through a preset method includes: Embedding code in the fillInStackTrace() method to obtain abnormal data captured by throwable, and obtaining the stack trace data corresponding to the abnormal data through the getStackTrace() method of throwable; and Adding an uncaught error data operation function to each thread of the JVM virtual machine, where the uncaught error data operation function is used to capture error events when the thread runs abnormally, and obtaining the stack trace data of the error events through the getStackTrace() method of throwable.
3. The method according to claim 1, characterized in that, The method further includes: For each basic data, using a preset hash function to map the basic data to obtain each mapping position of the basic data in the Bloom filter, and incrementing the counter count result of the mapping position by 1; Using the statistical value of the count result of the counter corresponding to the basic data in the Bloom filter as the occurrence times of the basic data, where the statistical value is the minimum value or the average value.
4. An abnormal acquisition device, characterized in that, The device includes: An acquisition module, configured to acquire basic data through a preset method, where the basic data includes stack trace data of abnormal data and internal error data of the virtual machine, and the stack trace data is used to identify the call chain of the abnormal data; A sampling module, configured to sample the abnormal data and error data corresponding to each type according to a sampling frequency preset for the occurrence times based on the occurrence times of various types of abnormal data and error data in the basic data to obtain data to be uploaded; where the occurrence times are obtained through a Bloom filter; the higher the occurrence times, the lower the sampling frequency; the sampling of the abnormal data and error data corresponding to each type according to the sampling frequency preset for the occurrence times based on the occurrence times of various types of abnormal data and error data in the basic data to obtain data to be uploaded includes: for each type of abnormal data and error data in the basic data, sampling the abnormal data and error data of the type according to the sampling order corresponding to the occurrence time range of the type, where the sampling order is generated based on the occurrence time range and the sampling frequency and is used to identify the order of the data to be uploaded that needs to be sampled in the abnormal data and error data of the type; and the higher the upper limit value of the occurrence time range, the lower the sampling frequency; An upload module, configured to add the data to be uploaded to an upload array to perform batch upload of the data to be uploaded in the upload array; A waiting module, configured to, when the upload array is full, encapsulate the data to be uploaded into a data block corresponding to the service identifier based on the service identifier of the data to be uploaded, and add the data block to a message queue; and when a preset callback condition is met, call a preset callback function to process the data block in the message queue, where the logic of the preset callback function at least includes sending the data block or storing the data block, and the preset callback condition is that the number of data blocks in the message queue reaches a preset number or the waiting time of the data block in the message queue exceeds a preset duration.
5. The device according to claim 4, characterized in that, The acquisition of the basic data through the preset method includes: By embedding code in the fillInStackTrace() method to obtain the abnormal data captured by throwable, and obtaining the stack trace data corresponding to the abnormal data through the getStackTrace() method of throwable; and, Adding an uncaught error data operation function to each thread of the JVM virtual machine, where the uncaught error data operation function is used to capture error events when the thread runs abnormally and obtain the stack trace data of the error events through the getStackTrace() method of throwable; The sampling module is configured to map each piece of basic data by using a preset hash function to obtain respective mapping positions of the basic data in the Bloom filter, and increment the count result of the counter at the mapping position by 1; Based on the statistical value of the count result of the counter corresponding to the basic data in the Bloom filter as the occurrence times of the basic data, wherein the statistical value is the minimum value or the average value.
6. An electronic device, comprising: A processor; And A memory storing a program, Wherein the program includes instructions that, when executed by the processor, cause the processor to execute the method according to any one of claims 1-3.
7. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-3.
Citation Information
Patent Citations
Method and device for monitoring and analyzing application fault reason, and storage medium
CN113849330A
Monitoring system for sampling exception data with a controlled data rate
US20220342790A1