Load Exception Handling Method, Electronic Device, Storage Medium, and Program Product
By collecting multiple process information of the equipment to be tested during the load detection cycle and automatically determining the target process, the problems of long load abnormal positioning time and low efficiency in the prior art are solved, and fast and efficient load abnormal root cause positioning is achieved.
Patent Information
- Application Number
- CN202411687164.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-11-25
AI Technical Summary
When the prior art detects that the system load is too high, it usually relies on post-event reproduction and manual analysis, resulting in a long time and inefficient positioning of the root cause.
By collecting multiple process information of the device to be tested during the load detection cycle, and when the load is in an abnormal state, the target process that causes the load abnormality is automatically determined based on this information.
It realizes automated positioning of the root causes of load abnormalities, shortens the positioning time, improves positioning efficiency, and reduces the need for storage space.
Smart Images

Figure CN119179582B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a method for processing abnormal load, an electronic device, a computer-readable storage medium, and a computer program product, which can be applied to the field of load management technologies. Background Art
[0002] With the rapid development of the Internet, more and more service scenarios require high-performance system support. System load is one of the important indicators for measuring system performance. When the system load is too high, phenomena such as long service response time and unstable service are likely to occur. Therefore, in order to ensure the effective operation of services, system load is often detected. However, in related technologies, when it is detected that the system load is too high, it is often through postmortem reproduction methods such as redeploying the site and combined with manual analysis to locate the root cause of the high load, with a long positioning time and low efficiency. Summary of the Invention
[0003] Embodiments of this application provide a method for processing abnormal load, an electronic device, a computer-readable storage medium, and a computer program product to alleviate or solve one or more technical problems existing in the prior art.
[0004] In a first aspect, embodiments of this application provide a method for processing abnormal load, including: obtaining load information of a device under test during a load detection period, where at least one process is running in the device under test; if the load information indicates that the load of the device under test is in an abnormal state, determining a target process in the at least one process that causes the abnormal state according to at least one process information of the device under test, where the at least one process information is collected during the load detection period, and the process information includes performance data of the at least one process.
[0005] In a second aspect, embodiments of this application provide a method for processing abnormal load, including: obtaining load information of an operating system during a load detection period, where at least one process is running in the operating system; if the load information indicates that the load of the operating system is in an abnormal state, determining a target process in the at least one process that causes the abnormal state according to at least one process information of the operating system, where the at least one process information is collected during the load detection period, and the process information includes performance data of the at least one process.
[0006] In a third aspect, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored on the memory, where the processor implements the method according to any one of the embodiments of this application when executing the computer program.
[0007] Fourthly, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method of any one of the embodiments of the present application is implemented.
[0008] Fifthly, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method of any one of the embodiments of the present application is implemented.
[0009] According to the technical solution of the embodiment of the present application, by collecting multiple process information of the device under test during the load detection period, and when the load of the device under test is in an abnormal state, determining the target process that causes the load abnormality according to the multiple process information collected during the load detection period. It not only realizes the automatic positioning of the root cause of the load abnormality, but also shortens the positioning time of the root cause of the load abnormality and improves the positioning efficiency of the root cause of the load abnormality because multiple process information of the device under test is collected during the load detection period without the need for post-event reproduction.
[0010] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In the drawings, unless otherwise specified, the same reference numerals throughout the drawings denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present application and should not be regarded as limiting the scope of the present application.
[0012] Figure 1 Shows the first scenario flowchart of the load abnormality handling method provided by the embodiment of the present application;
[0013] Figure 2 Shows the flowchart of the load abnormality handling method 200 provided by the embodiment of the present application;
[0014] Figure 3 Shows the second scenario flowchart of the load abnormality handling method provided by the embodiment of the present application;
[0015] Figure 4A 、 Figure 4B and Figure 4C Shows the schematic diagram of the timing diagram provided by the embodiment of the present application;
[0016] Figure 5 Shows the schematic diagram of the flame graph provided by the embodiment of the present application;
[0017] Figure 6 The flowchart of the load exception handling method 600 provided by the embodiments of the present application is shown;
[0018] Figure 7 The block diagram of the electronic device provided by the embodiments of the present application is shown. Detailed implementation manners
[0019] In the following, only some exemplary embodiments are briefly described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the concept or scope of the present application. Therefore, the drawings and descriptions are considered to be exemplary in nature and not restrictive.
[0020] To facilitate the understanding of the technical solutions of the embodiments of the present application, the related technologies of the embodiments of the present application are described below. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present application as optional solutions, and they all fall within the protection scope of the embodiments of the present application.
[0021] The following terms will be used hereinafter:
[0022] Operating System (OS): It is the core software in a computer system responsible for managing computer hardware and software resources. It provides an interface between users and computer hardware, and reasonably organizes and schedules the working process of the computer, thereby controlling the execution of other programs. According to the running environment, operating systems can be divided into desktop operating systems, mobile operating systems, server operating systems, embedded operating systems, etc.
[0023] Load: It refers to the measure of the number of tasks waiting to obtain processor resources within a certain time interval by the operating system, reflecting the busy degree of the operating system.
[0024] High load: It means that the operating system has performance bottlenecks under high concurrent access, the response time becomes slower, and the service is unstable.
[0025] Performance bottleneck: It means that the operating system has poor performance in a certain aspect, resulting in a decline in overall performance.
[0026] Response time: It refers to the time from when the operating system receives a request to when it returns a response.
[0027] The embodiments of the present application aim to provide a load exception handling method to shorten the positioning duration of the root cause of load exceptions and improve the positioning efficiency. Figure 1 For the application scenario schematic diagram of a load exception handling method provided by the embodiments of the present application, as Figure 1 shown, this scenario includes: a device under test 110 and a load exception handling device 120.
[0028] Among them, when the load anomaly detection device 120 determines that it reaches the end time of the collection period, it collects the process information of the device under test 110 and saves it in the data warehouse. And obtains the load information of the device under test 110 during the load detection period. And when the load information indicates that the load of the device under test is in an abnormal state, it obtains at least one piece of process information collected during the corresponding load detection period from the data warehouse. And determines, according to the at least one piece of process information obtained, at least one process running in the device under test 110 that causes the load to be in an abnormal state, that is, the target process. Among them, the load detection period is N times the collection period, and N is an integer greater than 1. To avoid occupying too much storage space due to storing too much process information collected during historical load detection periods in the data warehouse, the load anomaly detection device 120 can also delete the process information collected during a preset historical duration from the data warehouse when it determines that it reaches the data cleaning time. Exemplarily, the data cleaning time is 8:00 am every day, and the preset historical duration is the previous day. For example, at 8:00 am on November 10, 2024, the process information collected on November 9, 2024 is deleted from the data warehouse.
[0029] There are M processes running in the device under test 110, where M is a positive integer, and any process includes at least one thread. Exemplarily, the device under test 110 can be a terminal device or a server. Among them, the terminal device can be a mobile phone, a desktop computer, a portable notebook, a tablet computer, etc., and the server can be a physical server, a cloud server, etc.
[0030] In one application example, the load anomaly detection device 120 is set in the device under test 110. In another application example, the load anomaly detection device 120 is a device independent of the device under test 110, or the load anomaly detection device 120 is set in other devices, and the other devices are devices different from the device under test 110; at this time, the load anomaly detection device 120 can perform data interaction with the device under test 110 through a network. The specific application architecture can be built according to requirements in actual applications.
[0031] It should be noted that the above application scenarios or application examples provided in the embodiments of the present application are for easy understanding, and the embodiments of the present application do not specifically limit the application of the technical solutions. In addition, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0032] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the foregoing technical problems. The several specific embodiments listed may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0033] Figure 2 The flowchart of the load anomaly handling method 200 according to an embodiment of this application is shown. As Figure 2 shown, the load anomaly handling method 200 may include step S201 and step S202.
[0034] Step S201: Obtain the load information of the device under test during the load detection period, where at least one process is running in the device under test.
[0035] Among them, the load information may be obtained at the middle time of the load detection period or at the end time of the load detection period. In the following, it is described that the load information is obtained at the end time of the load detection period. In one implementation, the load anomaly handling device is disposed in the device under test. Correspondingly, in response to reaching the end time of the load detection period, the load anomaly handling device accesses the target file of the device under test and obtains the load information of the device under test from the target file. Alternatively, in response to reaching the end time of the load detection period, the load anomaly handling device runs a preset command to obtain the load information of the device under test. Exemplarily, the target file is the / proc / loadavg file; the preset command is the uptime command, and the uptime command is used to display the running time of the operating system, the current time, the number of logged-in users, and the average load of the operating system in the past 1 minute, 5 minutes, and 15 minutes.
[0036] In another implementation, the load anomaly handling device is independent of the device under test. Correspondingly, in response to reaching the end time of the load detection period, the load anomaly handling device sends a load information acquisition request to the device under test and receives the load information sent by the device under test.
[0037] Among them, the duration of the load detection period can be set as needed in practical applications. Exemplarily, the load detection period is 60 seconds. The load information may be a decimal between 0 and 1, and the larger the value of the load information, the higher the load of the device under test.
[0038] Step S202: If the load information indicates that the load of the device under test is in an abnormal state, determine the target process that causes the abnormal state among at least one process according to at least one process information of the device under test, where at least one process information is collected during the load detection period, and the process information includes the performance data of at least one process.
[0039] Specifically, the load anomaly handling device can determine whether the value of the load information is greater than the threshold. If it is greater than the threshold, it is determined that the load of the device under test is in an abnormal state; if it is not greater than the threshold, it is determined that the load of the device under test is in a normal state. When it is determined that the load of the device under test is in an abnormal state, at least one process information of the device under test collected during the load detection period is obtained from the data warehouse according to the start time and end time of the load detection period. And based on the at least one obtained process information, the target process that causes the abnormal state is determined among at least one process running in the current device under test. Any process information includes the performance data of at least one process running in the device under test, and the performance data is used to indicate the running state of the process, and the running state of the process is closely related to the load of the device under test. Therefore, the target process that causes the load anomaly can be determined according to the performance data of the process.
[0040] According to the technical solution of the embodiment of the present application, by collecting multiple process information of the device under test during the load detection period, and when the load of the device under test is in an abnormal state, determining the target process that causes the load anomaly according to the multiple process information collected during the load detection period. It not only realizes the automatic positioning of the root cause of the load anomaly, but also shortens the positioning time of the root cause of the load anomaly and improves the positioning efficiency of the root cause of the load anomaly because multiple process information of the device under test is collected during the load detection period without the need for post hoc reproduction.
[0041] In order to collect the process information of the device under test in a timely manner, thereby improving the positioning efficiency of the root cause of the load anomaly. In some embodiments, a collection period can also be preset. Correspondingly, the method can further include the following steps A1 to A3:
[0042] Step A1, in response to reaching the end time of the collection period, collect the subprocess information of each process currently running in the device under test.
[0043] Wherein, the load detection period is N times the collection period, and N is an integer greater than 1. Both the load detection period and the collection period can be set according to requirements in actual applications. Exemplarily, the load detection period is 60 seconds and the collection period is 10 seconds.
[0044] In some embodiments, collecting the subprocess information of each currently running process in the device under test may include: for any currently running process in the device under test: calling a first function interface to obtain the sub-performance data of the any process in multiple dimensions; calling a second function interface to obtain the sub-call relationship of the any process; and determining the sub-performance data and the sub-call relationship as the subprocess information of the any process. Exemplarily, the first function interface is task_thread_info, and the second function interface is getppid(). Among them, the task_thread_info function is used to obtain the process structure, which includes information such as the total number of threads included in the process, the number of threads in the process that are in the running state, and the resource occupation duration of the process. The getppid() function is used to return the upper-layer call relationship of the process, that is, to return the process identifier of the parent process (calling party) of the process (callee). It should be noted that the first function interface and the second function interface are not limited to the above examples, and can be set according to needs in actual applications.
[0045] Among them, the multiple dimensions are selected from the following dimensions: the resource occupation duration dimension, the first quantity dimension, and the second quantity dimension. The resource occupation duration dimension is used to characterize the occupation duration of at least one process for computing resources during the corresponding collection period; the first quantity dimension is used to characterize the total number of threads included in at least one process during the corresponding collection period; the second quantity dimension is used to characterize the number of threads in the running state among the threads included in at least one process during the corresponding collection period. The threads in the running state include the threads that are in the running state throughout the corresponding collection period, and also include the threads that are in the running state in the middle of the corresponding collection period.
[0046] That is to say, the sub-performance data of any process includes any combination of the occupation duration of the any process for computing resources, the total number of threads included in the any process, and the number of threads in the running state among the threads included in the any process. Among them, the computing resources may include processor resources, memory resources, etc.
[0047] It should be noted that when the total number of threads included in any process is multiple, the sub-performance data obtained by calling the first function interface may include the occupation durations of the computing resources by the threads included in the any process, and the load exception handling device may add up the occupation durations of the computing resources by the threads to obtain the occupation duration of the any process for computing resources.
[0048] Exemplarily, the subprocess information includes the duration of the above process's occupation of computing resources, the total number of threads included in the process, and the number of threads in the running state among the threads included in the process. The computing resources are processor resources. Process 1 includes 3 threads at the end time of collection period 1, namely Thread 1, Thread 2, and Thread 3. Thread 1 is in the running state during collection period 1, and the duration of its occupation of computing resources is 5 seconds. Thread 2 is in the running state during collection period 1, and the duration of its occupation of computing resources is 3 seconds. Process 3 is in the waiting state during collection period 1, that is, the duration of its occupation of computing resources is 0 seconds. It can be determined that the sub-performance data of Process 1 during collection period 1 includes: 8 seconds in the dimension of resource occupation duration (i.e., the total duration of occupation of computing resources is 5 + 3 + 0 = 8 seconds), 3 in the first quantity dimension (i.e., the total number of threads included in Process 1 is 3), and 2 in the second quantity dimension (i.e., the number of threads in the running state is 2).
[0049] Furthermore, the sub-call relationship in the subprocess information of any process may include: the upper call relationship of the any process. In the upper call relationship, the any process is the callee, that is, the upper call relationship is used to represent the parent process that calls the any process.
[0050] Step A2: Aggregate the subprocess information of each process to obtain the initial process information.
[0051] Specifically, for any dimension among multiple dimensions, respectively obtain the sub-performance data of the any dimension from the subprocess information of each process, accumulate the obtained sub-performance data of the any dimension to obtain the performance data corresponding to the any dimension; integrate the sub-call relationships included in the subprocess information of each process to obtain the call relationships among the processes; determine the performance data corresponding to each dimension, the call relationships among the processes, and the collection time as the process information of the device under test.
[0052] Exemplarily, at the end time of collection cycle 5, there are 3 processes running in the device under test, namely process 4, process 6, and process 7. The subprocess information 4 of process 4 includes 7 seconds in the dimension of resource occupation duration, 3 in the first quantity dimension, and 1 in the second quantity dimension, and the upper-layer call relationship is process 2. The subprocess information 6 of process 6 includes 8 seconds in the dimension of resource occupation duration, 4 in the first quantity dimension, and 2 in the second quantity dimension, and the upper-layer call relationship is process 4. The subprocess information 7 of process 7 includes 5 seconds in the dimension of resource occupation duration, 2 in the first quantity dimension, and 2 in the second quantity dimension, and the upper-layer call relationship is process 6. Then, for the dimension of resource occupation duration, the sub-performance data of the dimension of resource occupation duration obtained from the subprocess information 4 is 7 seconds, the sub-performance data of the dimension of resource occupation duration obtained from the subprocess information 6 is 8 seconds, and the sub-performance data of the dimension of resource occupation duration obtained from the subprocess information 7 is 5 seconds. Then, the performance data of the dimension of resource occupation duration is determined to be 7 + 8 + 5 = 20 seconds. For the first quantity dimension, the sub-performance data of the first quantity dimension obtained from the subprocess information 4 is 3, the sub-performance data of the first quantity dimension obtained from the subprocess information 6 is 4, and the sub-performance data of the first quantity dimension obtained from the subprocess information 7 is 2. Then, the performance data of the first quantity dimension is determined to be 3 + 4 + 2 = 9. For the second quantity dimension, the sub-performance data of the second quantity dimension obtained from the subprocess information 4 is 1, the sub-performance data of the second quantity dimension obtained from the subprocess information 6 is 2, and the sub-performance data of the second quantity dimension obtained from the subprocess information 7 is 2. Then, the performance data of the second quantity dimension is determined to be 1 + 2 + 2 = 5. Integrate the sub-call relationships in the subprocess information 4, subprocess information 6, and subprocess information 7, and the obtained call relationship is "process 2 - process 4 - process 6 - process 7". That is, the initial process information of the device under test in collection cycle 5 includes 20 seconds in the dimension of resource occupation duration, 9 in the first quantity dimension, 5 in the second quantity dimension, and the call relationship "process 2 - process 4 - process 6 - process 7". Among them, 20 seconds in the dimension of resource occupation duration, 9 in the first quantity dimension, and 5 in the second quantity dimension are performance data.
[0053] Step A3, if it is determined that the performance data in the initial process information is different from the performance data in the target process information, then determine the initial process information as the process information of the device under test, where the target process information is the process information determined according to the collection order and is the currently last collected process information.
[0054] In one embodiment, a data warehouse is used to save the process information of the device under test. Accordingly, the load anomaly handling device can obtain the target process information from the data warehouse. For any dimension, determine whether the performance data of any dimension in the initial process information is the same as the performance data of any dimension in the target process information. If the determination results of each dimension are the same, it is determined that the performance data in the initial process information is the same as the performance data in the target process information, and the initial process information is deleted. If the determination result of at least one dimension is different, it is determined that the performance data in the initial process information is different from the performance data in the target process information, the initial process information is determined as the process information of the device under test, and the process information is saved in the data warehouse.
[0055] Continuing with the above example, according to the collection order, it is determined that the last process information collected is process information 4 collected in collection cycle 4. For example, the performance data in process information 4 is 20 seconds in the resource occupation time dimension, 6 in the first quantity dimension, and 4 in the second quantity dimension, which is different from the 20 seconds in the source occupation time dimension, 9 in the first quantity dimension, and 5 in the second quantity dimension in the aforementioned initial process information, and the performance data in both the first quantity dimension and the second quantity dimension. Therefore, the aforementioned initial process information is determined as the process information of the device under test.
[0056] Therefore, during the load detection cycle, the process information of the device under test is collected according to the collection cycle. When it is detected that the load of the device under test is in an abnormal state, the root cause of the load abnormality can be determined according to the various process information collected during the load detection cycle, without the need to redeploy the site afterwards or perform manual analysis. Therefore, the time for locating the root cause of the load abnormality is greatly shortened, and the efficiency of locating the root cause of the load abnormality is improved. Furthermore, since data screening is performed before saving the performance data, that is, the performance data in the initial process information is compared with the performance data in the target process information, and when the two are different, the initial process information is determined as the process information of the device under test and saved, instead of saving each collected initial process information as process information, the demand for storage space is greatly reduced, the waste of storage resources caused by storing too much identical process information is avoided, and the amount of process information that needs to be maintained is reduced, thereby reducing the difficulty of maintaining the process information.
[0057] As mentioned above, any process information includes the calling relationship between processes and the performance data of the processes in multiple dimensions. In order to improve the determination rate and accuracy of the target process, Figure 3As shown, when the load anomaly handling device determines that the load of the device under test is in an abnormal state based on the obtained load information, it obtains multiple process information collected during the corresponding load detection period from the data warehouse, generates a visual graph for characterizing the call relationship and performance data based on the obtained multiple process information, and determines, based on the visual graph, the target process among the multiple processes running in the device under test that causes the load anomaly.
[0058] To avoid generating too many visual graphs and taking up too much time, in one implementation, the target dimension that causes the abnormal state can be determined based on the timing diagram first, and then the target process can be determined based on the target dimension and the call relationship. Specifically, the step of determining the target process among at least one process that causes the abnormal state according to at least one process information of the device under test in the foregoing step S202 may include the following steps S2021 and S2022:
[0059] Step S2021: Determine the target dimension that causes the abnormal state according to the performance data in the multiple process information.
[0060] In one implementation, for any dimension, a timing diagram corresponding to the any dimension can be generated according to the collection time included in the multiple process information and the performance data of the any dimension, where the timing diagram characterizes the change state of the performance data of the any dimension over time; the dimension corresponding to the target timing diagram is determined as the target dimension, where the target timing diagram includes a timing diagram with an abnormal change state.
[0061] In one implementation, for any dimension, the generated timing diagram can be saved as a first image, and a pre-trained first analysis model is called to analyze and process the first image to obtain a first analysis result. And whether the any dimension is the target dimension is determined according to the first analysis result. Wherein, the first analysis result is used to characterize whether the change state in the first image is an abnormal change state. The structure and training process of the first analysis model can be referred to the related art, and will not be elaborated in this application.
[0062] Exemplarily, the multiple dimensions include the foregoing resource occupation duration dimension, the first quantity dimension, and the second quantity dimension. The timing diagram corresponding to the resource occupation duration dimension is as Figure 4A shown, in Figure 4A the horizontal axis represents the collection time, and the vertical axis represents the unit of the resource occupation duration, 1 unit = 25 milliseconds. The timing diagram corresponding to the first quantity dimension is as Figure 4B shown, in Figure 4B the horizontal axis represents the collection time, and the vertical axis represents the total number of threads included in the process. The timing diagram corresponding to the second quantity dimension is as Figure 4C shown, in Figure 4C the horizontal axis represents the collection time, and the vertical axis represents the number of threads in the running state. It can be seen that Figure 4Aand Figure 4B The overall fluctuation of the curve in [specific context] is relatively uniform, without any particularly prominent increase or decrease, so it belongs to the normal change state. And Figure 4C the curve in [specific context] has an obvious upward fluctuation, which belongs to the abnormal change state. Therefore, it is determined that the second quantity dimension is the target dimension leading to the abnormal state.
[0063] Since the time series diagram can intuitively show the change state of the performance data, therefore, by generating the time series diagram corresponding to any dimension, it is possible to quickly and accurately determine the target dimension leading to the load anomaly based on the time series diagram, thereby improving the positioning speed of the root cause of the load anomaly.
[0064] Step S2022, determine the target process according to the performance data of the target dimension and the call relationship in the multiple process information.
[0065] In some embodiments, it is possible to generate a flame graph according to the performance data of the target dimension and the call relationship in the multiple process information, and determine the target process from the multiple processes according to the flame graph. Among them, the flame graph includes multiple call stacks, any call stack includes multiple grids, the multiple grids correspond to the multiple processes one by one, the arrangement order of the multiple grids represents the call relationship of the multiple processes, and the width of any grid is related to the performance data of the corresponding process in the target dimension, and is used to represent the probability that the corresponding process causes the abnormal state. The larger the width of the grid, the greater the probability that the corresponding process causes the abnormal state, or the greater the possibility that the corresponding process causes the abnormal state.
[0066] In some embodiments, determining the target process from the multiple processes according to the flame graph may include: for any call stack in the flame graph, in the order from top to bottom, if it is determined that the width of the last grid of the call stack is greater than the preset width, then determine the process corresponding to the last grid as the target process. Specifically, the generated time series diagram can be saved as a second image, and the pre-trained second analysis model can be called. For any call stack in the flame graph, in the order from top to bottom, if it is determined that the width of the last grid of the call stack is greater than the preset width, then determine the process corresponding to the last grid as the target process, and output the process name and / or process identifier of the target process. Among them, the structure and training process of the second analysis model can refer to the related technology, and will not be elaborated in this application.
[0067] In a flame graph, in the order from top to bottom, the grid shown in the first row can represent the total amount of performance data corresponding to the target dimension. Starting from the second row, the data in each grid is in the form of "process name: process identifier (total quantity)", where the process name represents the process name of the grid-corresponding process, the process identifier represents the process identifier of the grid-corresponding process, and (total quantity) represents the total amount of performance data of each process (i.e., the callee) called by the grid-corresponding process (i.e., the caller) in the target dimension. In particular, when the grid is the last grid, (total quantity) represents the performance data of the grid-corresponding process in the target dimension. In addition, it should be noted that in a flame graph, when the width of the grid is small, only part of the data in the grid can be displayed, such as only the process name, or only the process name and the process identifier, etc.
[0068] Exemplarily, the target dimension is the second quantity dimension. According to the performance data of the target dimension and the call relationship in the multiple process information, the generated flame graph is as Figure 5 shown. In Figure 5 , in the order from top to bottom, the first row shows a grid, denoted as grid 1. The total quantity (3196) in grid 1 represents that the total number of threads in the running state is 3196. The second row shows a grid, denoted as grid 2. Grid 2 corresponds to process 1. The (3196) shown in grid 2 represents that among the threads included in each process (the callee) called by process 1 (the caller), the total number of threads in the running state is 3196. The meaning of the data shown in other grids can be referred to here and will not be elaborated one by one. The third row shows three grids. In the order from left to right, the first grid is denoted as grid 3, and grid 3 corresponds to process 2; the second grid is denoted as grid 4, and grid 4 corresponds to process 3; the third grid is denoted as grid 5, and grid 5 corresponds to process 4. The fourth row shows three grids. In the order from left to right, the first grid is denoted as grid 6, and grid 6 corresponds to process 5; the second grid is denoted as grid 7, and grid 7 corresponds to process 6; the third grid is denoted as grid 8, and grid 8 corresponds to process 7. The fifth row shows 2 grids. In the order from left to right, the first grid is denoted as grid 9, and grid 9 corresponds to process 8; the second grid is denoted as grid 10, and grid 10 corresponds to process 9. In Figure 5 , the widths of grids 3, 4, 6 to 10 are small, so only the process name, or the process name and the process identifier are displayed.
[0069] Among them, grid 2 and grid 3 form call stack 1. In the order from top to bottom, the arrangement order of the grids in call stack 1 is grid 2 and grid 3 in sequence, indicating that the corresponding process 1 calls process 2. Grid 2, grid 4, and grid 6 form call stack 2. In the order from top to bottom, the arrangement order of the grids in call stack 2 is grid 2, grid 4, and grid 6 in sequence, indicating that the corresponding process 1 calls process 3, and process 3 calls process 5. Grid 2, grid 5, grid 7, and grid 9 form call stack 3. In the order from top to bottom, the arrangement order of the grids in call stack 3 is grid 2, grid 5, grid 7, and grid 9 in sequence, indicating that the corresponding process 1 calls process 4, process 4 calls process 6, and process 6 calls process 8. Grid 2, grid 5, grid 8, and grid 10 form call stack 4. In the order from top to bottom, the arrangement order of the grids in call stack 4 is grid 2, grid 5, grid 8, and grid 10 in sequence, indicating that the corresponding process 1 calls process 4, process 4 calls process 7, and process 7 calls process 9.
[0070] According to the above call stacks, it can be determined that the width of grid 3 (i.e., the last grid) of call stack 1 is greater than the preset width, and the width of grid 10 of call stack 4 is greater than the preset width (i.e., the last grid). Therefore, process 2 corresponding to grid 3 and process 9 corresponding to grid 10 are determined as the target processes, that is, process 2 and process 9 are the direct sources causing the abnormal load of the device under test.
[0071] In addition, according to the flame graph, the calling process (i.e., the caller) of the target process can also be determined, and this calling process is determined as the indirect source causing the abnormal load of the device under test. That is, in the above example, process 1 and process 7 are determined as the indirect sources causing the abnormal load of the device under test.
[0072] It should be noted that in some embodiments, based on the multiple process information of the device under test, the time series diagram and flame graph corresponding to each dimension can be generated simultaneously. After determining the target dimension causing the abnormal load according to the time series diagram, according to the flame graph corresponding to the target dimension, the target process causing the abnormal load among the multiple processes running on the device under test can be determined.
[0073] Since the flame graph can intuitively display the call relationship between processes and the possibility of each process causing abnormal load, therefore, it can quickly and accurately determine the direct source causing the abnormal load of the device under test, that is, the target process, improving the positioning speed of the source of abnormal load. Furthermore, the indirect source causing the abnormal load of the device under test, that is, the calling process of the target process, can also be determined based on the flame graph, so the cause of abnormal load can be analyzed more comprehensively.
[0074] Figure 6 The flowchart of the load abnormal processing method 600 according to the embodiment of the present application is shown, asFigure 6 As shown in Figure 6 , the load anomaly handling method 600 may include step S601 and step S602:
[0075] Step S601: Obtain the load information of the operating system during a load detection period, where at least one process is running in the operating system;
[0076] Step S602: If the load information indicates that the load of the operating system is in an abnormal state, then determine, according to the at least one process information of the operating system, a target process in the at least one process that causes the abnormal state, where the at least one process information is collected during the load detection period, and the process information includes performance data of the at least one process.
[0077] The operating system in step S601 and step S602 may be the operating system of the aforementioned device under test. The specific implementation processes of step S601 and step S602 may refer to the relevant descriptions above, and the repeated parts will not be elaborated here.
[0078] That is to say, the load anomaly handling method 600 of the embodiment of the present application and the load anomaly handling method 400 are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by them.
[0079] Corresponding to the load anomaly handling method 200 provided in the embodiment of the present application, the embodiment of the present application further provides a load anomaly handling device, including: a load information acquisition module, configured to acquire the load information of the device under test during a load detection period, where at least one process is running in the device under test. A target process determination module, configured to, if the load information indicates that the load of the device under test is in an abnormal state, then determine, according to the at least one process information of the device under test, a target process in the at least one process that causes the abnormal state, where the at least one process information is collected during the load detection period, and the process information includes performance data of the at least one process.
[0080] In one implementation, the process information includes the call relationship between processes and the performance data of the processes in multiple dimensions. Correspondingly, the load information acquisition module is specifically configured to: determine a target dimension that causes the abnormal state according to the performance data in the multiple process information; determine the target process according to the performance data of the target dimension and the call relationship in the multiple process information.
[0081] In one embodiment, the process information further includes the collection time of the process information. Correspondingly, the load information acquisition module is further specifically configured to: for any dimension, generate a time series graph corresponding to the any dimension according to the collection time included in multiple pieces of the process information and the performance data of the any dimension, where the time series graph represents the change state of the performance data of the any dimension over time; determine the dimension corresponding to the target time series graph as the target dimension, where the target time series graph includes the time series graph with an abnormal change state.
[0082] In one embodiment, the load information acquisition module is further specifically configured to: generate a flame graph according to the performance data of the target dimension and the call relationship in multiple pieces of the process information, where the flame graph includes multiple call stacks, any call stack includes multiple grids, the multiple grids correspond to multiple processes one by one, the arrangement order of the multiple grids represents the call relationship of the multiple processes, and the width of any grid is related to the performance data of the corresponding process in the target dimension and is used to represent the probability that the corresponding process causes the abnormal state; determine the target process from the multiple processes according to the flame graph.
[0083] In one embodiment, the apparatus further includes: a process information collection module, an aggregation module, and a process information determination module. The process information collection module is configured to collect the sub-process information of each currently running process in the device under test in response to reaching the end time of the collection period, where the load detection period is N times the collection period, and N is an integer greater than 1; the aggregation module is configured to aggregate the sub-process information of each process to obtain initial process information; the process information determination module is configured to, if it is determined that the performance data in the initial process information is different from the performance data in the target process information, determine the initial process information as the process information, where the target process information is the process information that is determined according to the collection order and is the last collected process information currently.
[0084] In one embodiment, the process information collection module is specifically configured to: for any currently running process in the device under test: call a first function interface to obtain the sub-performance data of the any process in multiple dimensions; call a second function interface to obtain the sub-call relationship of the any process; determine the sub-performance data and the sub-call relationship as the sub-process information of the any process.
[0085] In one embodiment, the multiple dimensions are selected from the following dimensions:
[0086] The resource occupation duration dimension, which is used to represent the occupation duration of a process for computing resources;
[0087] The first quantity dimension is used to represent the total number of threads included in a process;
[0088] The second quantity dimension is used to represent the number of threads in a running state among the threads included in a process.
[0089] For the functions of each module in each device of the embodiments of the present application, reference may be made to the corresponding descriptions in the above methods, and the corresponding beneficial effects are achieved, which will not be elaborated herein.
[0090] Corresponding to the load anomaly handling method 600 provided by the embodiments of the present application, the embodiments of the present application further provide a load anomaly handling device, including: an acquisition module, configured to acquire load information of an operating system within a load detection period, where at least one process is running in the operating system. A determination module, configured to, if the load information indicates that the load of the operating system is in an abnormal state, determine a target process that causes the abnormal state among the at least one process according to at least one process information of the operating system, where the at least one process information is collected within the load detection period, and the process information includes performance data of the at least one process.
[0091] Figure 7 It is a block diagram of an electronic device for implementing the embodiments of the present application. As Figure 7 shown, the electronic device includes: a memory 701 and a processor 702, and a computer program that can run on the processor 702 is stored in the memory 701. When the processor 702 executes the computer program, the method in the above embodiments is implemented. The number of the memory 701 and the processor 702 may be one or more. In a specific implementation, the electronic device may further include a communication interface 703 for communicating with external devices and performing data interaction and transmission.
[0092] In a specific implementation, if the memory 701, the processor 702, and the communication interface 703 are independently implemented, the memory 701, the processor 702, and the communication interface 703 may be connected to each other through a bus and complete communication therebetween. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0093] Optionally, in a specific implementation, if the memory 701, the processor 702, and the communication interface 703 are integrated on a single chip, the memory 701, the processor 702, and the communication interface 703 can communicate with each other through an internal interface.
[0094] An embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the method provided in the embodiment of the present application.
[0095] An embodiment of the present application provides a computer program product including a computer program, which when executed by a processor implements the method provided in the embodiment of the present application.
[0096] An embodiment of the present application further provides a chip, which includes a processor for calling and running instructions stored in a memory from the memory, so that a communication device installed with the chip executes the method provided in the embodiment of the present application.
[0097] An embodiment of the present application further provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is configured to execute code in the memory, and when the code is executed, the processor is configured to execute the method provided in the embodiment of the application.
[0098] It should be understood that the above-mentioned processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor supporting the advanced risc machines (ARM) architecture.
[0099] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0100] In the above embodiments, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.
[0101] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0102] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of these features. In the description of the present application, "a plurality of" means two or more unless otherwise specifically defined.
[0103] Any process or method described in the flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed.
[0104] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in connection with these instruction execution systems, apparatus, or devices.
[0105] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the method in the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0106] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. This storage medium may be a read-only memory, a magnetic disk or an optical disc, etc.
[0107] As mentioned above, it is only an exemplary embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope recorded in the present application can easily think of various changes or substitutions thereof, and these should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for handling load anomalies, comprising: Obtaining load information of a device under test during a load detection period, wherein at least one process is running in the device under test; If the load information indicates that the load of the device under test is in an abnormal state, a target dimension causing the abnormal state is determined according to multiple process information of the device under test, wherein the multiple process information is collected within the load detection cycle, the process information includes a calling relationship between processes and performance data of the at least one process in multiple dimensions, the multiple dimensions include a resource occupancy time dimension, a first number dimension of threads included in a process, and a second number dimension of threads in a running state, and the target dimension is at least one dimension among the multiple dimensions; Generate a flame graph according to the performance data of the target dimension and the call relationship, wherein the flame graph includes multiple call stacks, each call stack includes multiple grids, each of the multiple grids corresponds to multiple processes one by one, an arrangement order of the multiple grids represents the call relationship of the multiple processes, and a width of each grid is related to the performance data of the corresponding process in the target dimension, and is used to represent the probability that the corresponding process causes the abnormal state; According to the flame graph, a target process causing the abnormal state is determined from the multiple processes.
2. The method according to claim 1, wherein the process information further includes a collection time of the process information, and the determining the target dimension causing the abnormal state according to the multiple process information of the device under test comprises: For any dimension, a timing diagram corresponding to the any dimension is generated according to the collection time included in the plurality of process information and the performance data of the any dimension, wherein the timing diagram represents the change state of the performance data of the any dimension over time; The dimension corresponding to the target timing diagram is determined as the target dimension, wherein the target timing diagram includes a timing diagram in which the change state is an abnormal change state.
3. The method according to claim 1 or 2, further comprising: In response to reaching the end time of the collection cycle, collecting sub-process information of each process currently running in the device under test, wherein the load detection cycle is N times the collection cycle, N is an integer greater than 1, and any sub-process information includes sub-performance data of the corresponding process in multiple dimensions; Aggregating the sub-process information of each process to obtain initial process information, wherein the initial process information includes performance data obtained by accumulating each sub-performance data; If it is determined that the performance data in the initial process information is different from the performance data in the target process information, the initial process information is determined as the process information, wherein the target process information is the last collected process information determined according to the collection order.
4. The method according to claim 3, wherein collecting sub-process information of each process currently running in the device under test comprises: For any process currently running in the device under test: calling a first function interface to obtain sub-performance data of the process in multiple dimensions; Call the second function interface to obtain the sub-call relationship of any one of the processes; determine the sub-performance data and the sub-call relationship as the sub-process information of any one of the processes.
5. A method for handling load anomalies, comprising: Obtaining load information of an operating system within a load detection period, wherein at least one process is running in the operating system; If the load information indicates that the load of the operating system is in an abnormal state, a target dimension causing the abnormal state is determined according to multiple process information of the operating system, wherein the multiple process information is collected within the load detection cycle, the process information includes a call relationship between processes and performance data of the at least one process in multiple dimensions, the multiple dimensions include a resource occupancy time dimension, a first number dimension of threads included in a process, and a second number dimension of threads in a running state, and the target dimension is at least one dimension among the multiple dimensions; Generate a flame graph according to the performance data of the target dimension and the call relationship, wherein the flame graph includes multiple call stacks, each call stack includes multiple grids, each of the multiple grids corresponds to multiple processes one by one, an arrangement order of the multiple grids represents the call relationship of the multiple processes, and a width of each grid is related to the performance data of the corresponding process in the target dimension, and is used to represent the probability that the corresponding process causes the abnormal state; According to the flame graph, a target process causing the abnormal state is determined from the multiple processes.
6. An electronic device comprising a memory, a processor and a computer program stored in the memory, wherein the processor implements the method according to any one of claims 1 to 5 when executing the computer program.
7. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
8. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Process control method and apparatus
CN107368326A
Service process control method, service process control device and terminal
CN108733465A
Service execution method and device, terminal equipment and storage medium
CN115951954A
System load adjusting method and device, electronic equipment and readable storage medium
CN117827439A