Storage performance monitoring method, electronic equipment, readable storage medium and program product
By recording operation information with processing time exceeding the threshold in the storage device and generating segmented delay statistics tables, the accuracy and cost problems of existing storage performance monitoring methods are solved, and timely capture and efficient monitoring of storage failures are achieved.
Patent Information
- Application Number
- CN202510905968.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-08-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing storage performance monitoring methods are difficult to accurately capture the problem of instantaneous high latency under high load conditions, and require manual intervention, resulting in high monitoring costs and inability to record and store fault information in a timely manner.
By obtaining the input/output operation processing time of the storage device, the operation information whose processing time exceeds the threshold is recorded to the exception log pool, and a segmented delay statistics table is generated, and the exception log and device information are automatically transferred to the target file in combination with the trigger judgment condition.
It realizes timely capture and record storage failures, reduces monitoring costs, and improves monitoring accuracy and efficiency of storage performance.
Smart Images

Figure CN120406859A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a storage performance monitoring method, a storage performance monitoring device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] A system for managing storage devices is a storage system. The storage system needs to be able to monitor the performance of storage devices at all times to ensure that when the performance shows an abnormality, the abnormal information can be captured and recorded in a timely manner. This enables maintenance personnel to notice performance changes in a timely manner and is also beneficial to solving the problem of restoring the normal performance of storage failures.
[0003] However, in the process of monitoring storage performance in related technologies, maintenance personnel need to manually analyze the specific location where a failure occurs from a large amount of logs, resulting in a high cost of performance monitoring. Summary of the Invention
[0004] In view of the above problems, the present invention provides a storage performance monitoring method, a storage performance monitoring device, an electronic device, a computer-readable storage medium, and a computer program product.
[0005] According to a first aspect of the present invention, there is provided a storage performance monitoring method, including: obtaining a processing time for completing an input / output operation when the operating system performs an input / output operation on a storage device; for any input / output operation, when the processing time of the input / output operation is greater than a first processing duration, storing operation-related information of the input / output operation in an exception log in an exception log pool; for input / output operations performed within a first time period, obtaining a segmented delay statistical table of the input / output operations within the first time period based on their respective processing times; associating device-related information of the storage device for performing the input / output operation within the first time period with the segmented delay statistical table and storing them in a target file in the storage device; and according to multiple processing times within the first time period and a trigger judgment condition, transferring a dump file to a target address, where the dump file includes at least one of the exception logs in the exception log pool and the target file.
[0006] The second aspect of the present invention provides a storage performance monitoring device, including: an acquisition module, configured to acquire the processing time for completing an input / output operation when the operating system performs an input / output operation on a storage device; a first storage module, configured to, for any input / output operation, store the operation-related information of the input / output operation in an exception log in an exception log pool when the processing time of the input / output operation is greater than a first processing duration; a obtaining module, configured to, for the input / output operations performed within a first time period, obtain a segmented delay statistical table of the input / output operations within the first time period based on their respective processing times; a second storage module, configured to associatively store the device-related information of the storage device performing the input / output operation within the first time period and the segmented delay statistical table in a target file in the storage device; a transfer storage module, configured to transfer a dump file to a target address according to multiple processing times within the first time period and a trigger judgment condition, where the dump file includes at least one of the exception logs in the exception log pool and the target file.
[0007] The third aspect of the present invention provides an electronic device, including: one or more processors; a memory, configured to store one or more computer programs, where the above-mentioned one or more processors execute the above-mentioned one or more computer programs to implement the steps of the above-mentioned method.
[0008] The fourth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above-mentioned method are implemented.
[0009] The fifth aspect of the present invention further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above-mentioned method are implemented.
[0010] According to the embodiments of the present invention, by storing operation-related information in an exception log in an exception log pool according to the processing time of the storage device performing an input / output operation, and at the same time generating a segmented delay statistical table according to multiple processing times within a first time period, and storing the device-related information and the segmented delay statistical table within the first time period in a target file, at least one of the exception log and the target file is timely transferred to a target address when a trigger judgment condition is met. Since the exception log and the segmented delay statistical table are used to save the data during a storage failure and are timely transferred according to the trigger judgment condition, the information of the storage failure is effectively guaranteed, which helps to improve the performance of the operating system and at the same time reduces the monitoring cost of performance monitoring. Description of the Drawings
[0011] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above content and other objects, features, and advantages of the present invention will become clearer.
[0012] Figure 1 Shows an application scenario diagram of the storage performance monitoring method according to an embodiment of the present invention.
[0013] Figure 2 Shows a flowchart of the storage performance monitoring method according to an embodiment of the present invention.
[0014] Figure 3 Shows a block diagram of a storage device according to an embodiment of the present invention.
[0015] Figure 4 Shows a schematic block diagram of the circular overwrite of the exception log pool according to an embodiment of the present invention.
[0016] Figure 5 Shows a flowchart of the generation method of the segmented delay statistical table according to an embodiment of the present invention.
[0017] Figure 6 Shows a block diagram of the structure of the storage performance monitoring device according to an embodiment of the present invention.
[0018] Figure 7 Shows a block diagram of an electronic device suitable for implementing the above method according to an embodiment of the present invention. Detailed implementation manners
[0019] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a thorough understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.
[0020] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0021] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0022] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0023] In the current storage performance monitoring method, since the input / output volume of the storage device per unit time can reach hundreds of thousands or even millions, the data flushing speed is very fast. Even if a large amount of space is allocated to record the logs, the massive data logs are easily overwritten, resulting in a very short recordable time. Secondly, the performance monitoring is not precise enough to effectively capture the instantaneous high latency problem within a short period of time. Moreover, after a high latency problem occurs in the current storage performance monitoring method, maintenance personnel need to actively collect the system logs, which leads to the start of log collection being much later than a long time after the problem occurs. At this time, starting to collect logs may not record the information at the time of the problem occurrence and may have been overwritten by subsequent logs.
[0024] In view of this, embodiments of the present invention provide a storage performance monitoring method, an electronic device, a readable storage medium, and a program product, which can be applied to the field of information technology. The method includes: obtaining the processing time for completing an input / output operation when the operating system performs an input / output operation on a storage device; in the case where the processing time of the input / output operation is greater than a first processing duration, storing the operation-related information of the input / output operation in an abnormal log in an abnormal log pool; obtaining a segmented latency statistical table of input / output operations within a first time period based on the processing time of the input / output operation; associating and storing the device-related information of the storage device performing the input / output operation within the first time period with the segmented latency statistical table in a target file; and transferring a dump file to a target address according to multiple processing times within the first time period and a trigger judgment condition, where the dump file includes at least one of the abnormal logs in the abnormal log pool and the target file.
[0025] In the technical solution of the present invention, the involved user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all complies with relevant laws, regulations, and standards, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or refuse.
[0026] Figure 1The figure shows an application scenario diagram of the storage performance monitoring method according to an embodiment of the present invention.
[0027] As Figure 1 shown, the application scenario 100 according to this embodiment may include a storage service providing scenario. The network 104 is a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0028] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).
[0029] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0030] The server 105 may be a server providing various data storage and reading services, such as a background management server (only for example) that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process data such as received user requests, etc., and feedback the processing results (such as web pages, information, or data, etc. obtained or generated according to user requests) to the terminal device.
[0031] It should be noted that the storage performance monitoring method provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the storage performance monitoring device provided by the embodiments of the present invention can generally be set in the server 105. The storage performance monitoring method provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the storage performance monitoring device provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0032] It should be understood, Figure 1The numbers of the first terminal device, the second terminal device, the third terminal device, the network, and the server in [it] are only [as described]. According to the implementation requirements, there can be any number of the first terminal device, the second terminal device, the third terminal device, the network, and the server.
[0033] Based on the Figure 1 scenario described below, through Figures 2 to 5 a detailed description of the storage performance monitoring method of the embodiment will be given.
[0034] Figure 2 FIG. shows a flowchart of the storage performance monitoring method according to an embodiment of the present invention. Figure 3 FIG. shows a block diagram of a storage device according to an embodiment of the present invention. Figure 4 FIG. shows a schematic block diagram of circular overwrite of an exception log pool according to an embodiment of the present invention.
[0035] As Figure 2 shown, the storage performance monitoring method of this embodiment includes operations S210 to S250, and this storage performance monitoring method can be executed by a server.
[0036] In operation S210, obtain the processing time for completing an input / output operation when the operating system performs an input / output operation on the storage device.
[0037] In operation S220, for any input / output operation, when the processing time of this input / output operation is greater than the first processing duration, store the operation-related information of this input / output operation in the exception log of the exception log pool.
[0038] In operation S230, for the input / output operations performed within the first time period, based on their respective processing times, obtain a segmented delay statistical table of the input / output operations within the first time period.
[0039] In operation S240, associate and store the device-related information of the storage device performing the input / output operation within the first time period with the segmented delay statistical table in a target file in the storage device.
[0040] In operation S250, according to multiple processing times within the first time period and a trigger judgment condition, transfer the dump file to a target address, where the dump file includes at least one of the exception logs in the exception log pool and the target file.
[0041] As Figure 3As shown, the storage device can be a centralized storage device, which mainly consists of a processor, memory, motherboard, and disks connected through a hard disk interface. The disks for input / output operations in the centralized storage device can be hard disks, a hard disk array composed of multiple hard disks, or different hard disk volumes divided from the capacities of multiple hard disks. The input / output operations can be to write processed data to the storage device or read processed data from the storage device.
[0042] The operating system can refer to a storage management system, which can store or read data from the storage device. The specific time length of the first time period can be set according to actual needs. For example, it can be ten minutes.
[0043] When the storage device executes an input / output operation for a certain processed data, after completing the input / output operation, the processing time of the input / output operation can be obtained. For example, when storing processed data, the time consumed by the storage operation can be known after storage, such as 8 ms.
[0044] An exception log buffer, that is, an exception log pool, is divided in the memory area of the operating system. If the processing time of a certain input / output operation is greater than the first processing duration, the operation-related information of the input / output operation can be stored in the exception log of the exception log pool. Among them, the first processing duration can be set according to actual needs. For example, it can be 1 ms.
[0045] It should be noted that the memory size occupied by the exception log pool is fixed, so the number of exception logs that can be recorded is also fixed. The exception logs adopt a log cyclic overwrite method. For example, first store exception log 1, and then record exception log 2, exception log 3... exception log N in sequence based on time. When the exception logs are full, overwrite starts from the earliest exception log, as Figure 4 shown. Therefore, the time period that the exception log pool can record = the maximum number of logs recorded / the number of logs that need to be printed per second on average. For example, if the exception log pool can record 100,000 logs and there are 100 input / output operations with a latency greater than 1 ms per second, then a total of 1000 seconds of input / output operations can be recorded.
[0046] If in a certain first time period, such as the first time period from 8:00 to 8:10, multiple input / output operations are executed within this first time period, and the processing times of the multiple input / output operations are statistically analyzed, then a segmented latency statistical table for this first time period can be obtained. The segmented latency statistical table statistically counts the number of input / output operations in different processing duration intervals. For example, the processing duration interval is (1 ms, 10 ms], and the corresponding number of input / output operations is 25.
[0047] While counting the segmented delay statistical table, the device-related information of the storage device during the input / output operations in the first time period can also be recorded, and the device-related information is associated with the segmented delay statistical table of the first time period, and then the two are stored in the target file of the storage device. Among them, the device-related information may refer to the information of the resources used for the input / output operations, such as the resource occupancy information of the processor, etc.
[0048] In a specific embodiment, the target file can be a file in any format. For example, it can be a text-format file. In the target file, the device-related information and the segmented delay statistical table for different time periods can be arranged according to the time periods.
[0049] According to the multiple processing times and the trigger judgment condition in the first time period, at least one of the abnormal logs in the dump file and the target file is transferred and stored at the target address. Among them, the trigger judgment condition can be a judgment condition based on the processing duration threshold. For example, if the processing time of a certain input / output operation is greater than the processing duration threshold (such as 100 ms), the target file can be transferred and stored at the target address. The target address can be the system disk of the maintenance system of the maintenance personnel. The maintenance personnel can check the files at the target address in time for system maintenance to avoid the decline of the input / output performance caused by system failures.
[0050] According to the embodiment of the present invention, by storing the operation-related information in the abnormal log in the abnormal log pool according to the processing time of the input / output operation of the storage device, generating a segmented delay statistical table according to the multiple processing times in the first time period, and storing the device-related information and the segmented delay statistical table in the first time period in the target file, at least one of the abnormal log and the target file is transferred and stored at the target address in time when the trigger judgment condition is satisfied. Since the abnormal log and the segmented delay statistical table are used to save the data during the storage failure and are transferred and stored in time through the trigger judgment condition, the information of the storage failure is effectively guaranteed, which helps to improve the performance of the operating system and reduce the monitoring cost of performance monitoring at the same time.
[0051] According to the embodiment of the present invention, the operation-related information includes the operation parameters of the input / output operation and the data parameters corresponding to the input / output operation. The operation parameters include the operation time parameters for executing the input / output operation, and the data parameters include the data information of the processed data corresponding to the input / output operation.
[0052] According to an embodiment of the present invention, when the processing time of the input / output operation is greater than the first processing duration, the operation-related information of the input / output operation is stored in the exception log of the exception log pool, including: when the processing time is greater than the first processing duration, obtaining an operation time parameter from a storage device; performing an association operation on the operation time parameter and data information in terms of time scale to obtain storage association information; and writing the storage association information into the exception log.
[0053] The operation time parameter may refer to the time when the input / output operation is executed, such as the start time, end time, etc. of the execution of the input / output operation. The data information may refer to the basic information of the processed data, such as the data name, data size of the processed data, and the data file of the processed data.
[0054] If the processing time of a certain input / output operation is greater than the first processing duration, it indicates that there is an exception in the input / output operation, such as the time spent on the input / output exceeding the normal processing duration. At this time, the operation time parameter corresponding to the input / output operation can be obtained from the storage device, and at the same time, an association operation on the operation time parameter and data information is performed in terms of time scale, such as associating the operation time parameter and data information according to the start time, so as to obtain storage association information, and write the storage association information into the exception log.
[0055] If there are multiple input / output operations with processing times greater than the first processing duration, at this time, the multiple storage association information can be sorted in time according to the start time of different input / output operations, and then the sorted multiple storage association information is written into the exception log.
[0056] According to an embodiment of the present invention, since the exception log in the exception log pool is independent of the regular log pool, and the exception log pool only stores data with operation exceptions, it is convenient to know the specific information of the data with operation exceptions from the exception log in the dump file after the forwarding of the dump file, effectively improving the efficiency of daily maintenance.
[0057] According to an embodiment of the present invention, when the processing time of the input / output operation is greater than the first processing duration, the operation-related information of the input / output operation is stored in the exception log of the exception log pool, including: when the input / output operation corresponding to the processed data fails to execute, obtaining an operation failure identifier corresponding to the input / output operation; and associatively storing the operation failure identifier and the data information of the processed data in the exception log.
[0058] If, when a storage device performs an input / output operation for processing data, the input / output operation fails to execute successfully, an operation failure identifier corresponding to the input / output operation can be obtained from the storage device at this time. The operation failure identifier is associated with the data information of the corresponding processed data, and the associated information is stored in an exception log.
[0059] According to an embodiment of the present invention, by storing the relevant information of the processed data with operation failures in the input / output operation in an exception log, after the dump file is transferred, it is convenient for maintenance personnel to promptly know the specific processed data with problems, thereby improving the maintenance efficiency.
[0060] Figure 5 The flowchart of the method for generating a segmented delay statistics table according to an embodiment of the present invention is shown.
[0061] According to an embodiment of the present invention, the segmented delay statistics table includes a plurality of different processing time intervals and the data quantity parameter values corresponding to the processing time intervals.
[0062] According to an embodiment of the present invention, based on their respective processing times, a segmented delay statistics table of the input / output operations in the first time period is obtained, including operations S501 to S503.
[0063] In operation S501, for any input / output operation, when the processing time is greater than the first duration threshold, the processing time is determined as the time to be classified.
[0064] In operation S502, for any time to be classified, a target time interval corresponding to the time to be classified is determined from a plurality of processing time intervals.
[0065] In operation S503, according to the input / output operations corresponding to a plurality of times to be classified, the data quantity parameter values of different target time intervals are filled to obtain a segmented delay statistics table.
[0066] The specific time intervals of the plurality of different processing time intervals can be specifically set according to actual needs. For example, a plurality of processing time intervals can be set to (1ms, 10ms] and (10ms, ∞) respectively. The first duration threshold can be specifically set according to specific needs. For example, it can be 1ms.
[0067] If, when a storage device performs an input / output operation for processing data, the processing time of the input / output operation is greater than the first duration threshold, it indicates that the input / output operation is abnormal. At this time, the processing time can be determined as a time to be classified.
[0068] Since the segmented delay statistics table is divided into multiple different processing time intervals, determine which processing time interval the time to be classified belongs to based on the time length of the time to be classified, and thus determine this processing time interval as the target time interval of the time to be classified. Fill and modify the data quantity parameter value of the target time interval based on the input / output operations corresponding to the current time to be classified, thereby obtaining the segmented delay statistics table.
[0069] In a specific embodiment, assume that a total of 100 input / output operations are executed within the first time period, and among them, the processing times of 15 input / output operations are greater than the first processing duration. At this time, determine the target time interval corresponding to each processing time, and fill the data quantity parameter value of this target time interval. For example, the processing times of 8 input / output operations belong to (1ms, 10ms], and the processing times of 7 input / output operations belong to (10ms, ∞). At this time, the data quantity parameter value of the processing time interval (1ms, 10ms] in the obtained segmented delay statistics table is 8, and the data quantity parameter value of the processing time interval (10ms, ∞) is 7.
[0070] According to the embodiment of the present invention, by statistically dividing the processing times greater than the first processing duration into intervals, after transferring and storing the target file including the segmented delay statistics table, maintenance personnel can quickly know the total amount of abnormal data, thereby facilitating the timely positioning of abnormal problems and improving the maintenance efficiency.
[0071] According to the embodiment of the present invention, based on their respective processing times, obtaining the segmented delay statistics table of the input / output operations within the first time period further includes: when the processing time is not greater than the first duration threshold, filling the data quantity parameter value of the processing time interval based on the processing time interval corresponding to the processing time.
[0072] For the case where the input / output operation is normally executed, that is, the processing time of this input / output operation is not greater than the first duration threshold, the number of normally executed input / output operations can also be statistically counted. For the specific filling method of the data quantity parameter value, refer to the filling method for the processing time greater than the first duration threshold.
[0073] When statistically counting the normally executed operations, the corresponding processing time interval can be divided into (0, 100us] and (100us, 1ms], as shown in Table 1. It should be noted that this processing time interval can also be other range intervals.
[0074] Table 1
[0075] Processing time interval Data quantity parameter value (0, 100 μs] 72 (100 μs, 1 ms] 13 (1 ms, 10 ms] 8 (10 ms, ∞) 7
[0076] According to an embodiment of the present invention, by statistically analyzing a segmented delay statistical table for normal input / output operations, after storing the target file including the segmented delay statistical table, maintenance personnel can quickly know the total amount of data processed in the first time period, improving the monitoring performance of storage performance.
[0077] According to an embodiment of the present invention, the first time period includes a plurality of second time periods, and the device-related information includes performance parameters of the storage device.
[0078] According to an embodiment of the present invention, associatively storing the device-related information of the input / output operations performed by the storage device within the first time period and the segmented delay statistical table in a target file in the storage device includes: for any second time period, calculating the average number of operations per unit time and the average bandwidth according to the number of input / output operations corresponding to at least one input / output operation and the amount of bandwidth used for processing at least one piece of data within the second time period; calculating a first average delay according to the number of input / output operations and the total processing duration, where the total processing duration is generated according to at least one processing time of at least one piece of data; and associatively storing the average number of operations, the average bandwidth, the first average delay corresponding to different second time periods, and the segmented delay statistical table in the target file.
[0079] The specific time length of the second time period can be specifically set according to actual requirements. For example, when the first time period is 10 minutes, the second time period can be multiple time periods obtained by dividing the first time period into segments of 5 seconds. The performance parameters of the storage device include the average number of operations, the average bandwidth, the first average delay, etc.
[0080] For each second time period, if a total of m input / output operations are performed within the second time period, m≥1, dividing the number of input / output operations by the duration of the second time period can obtain the average number of operations per unit time. At the same time, according to the amount of bandwidth used for processing the data corresponding to the m input / output operations within the second time period, dividing by the duration of the second time period can also obtain the average bandwidth per unit time.
[0081] Calculating the total processing duration based on the m processing times of the m input / output operations, dividing the total processing duration by the number of input / output operations can obtain the first average delay, which represents the average duration required for a single execution of an input / output operation.
[0082] Arranging the average number of operations, the average bandwidth, and the first average delay corresponding to different second time periods within the first time period in chronological order, and associatively storing them with the segmented delay statistical table in the target file.
[0083] According to an embodiment of the present invention, by dividing the first time period into more detailed second time periods, and counting the average number of operations, average bandwidth, and first average latency within the second time periods, after storing the target file, maintenance personnel can quickly know the storage performance of different second time periods within the first time period, thereby improving the monitoring accuracy and monitoring granularity of the storage performance.
[0084] According to an embodiment of the present invention, the average number of operations, average bandwidth, first average latency, and segmented latency statistical table corresponding to different second time periods are associated and stored in the target file, including: for any second time period, generating the segmented performance parameter of the second time period according to the average number of operations, average bandwidth, and first average latency; sorting the segmented performance parameters of multiple second time periods in chronological order, and associating and storing the sorted multiple segmented performance parameters and the segmented latency statistical table in the target file.
[0085] For each second time period, generate the segmented performance parameter of the second time period according to the average number of operations, average bandwidth, and first average latency of the second time period. The segmented performance parameter can be a description of different parameters in a piece of text. Then sort the different segmented performance parameters in chronological order, so as to associate and store the sorted multiple segmented performance parameters and the segmented latency statistical table in the target file.
[0086] In another specific embodiment, the average number of operations, average bandwidth, and first average latency of different second time periods within the first time period can be plotted into a chart. For example, performance curve graphs and / or pie charts of the average number of operations, average bandwidth, and first average latency can be plotted respectively in the order of the second time periods, and the performance curve graphs and / or pie charts are stored in the target file. The plotting of the chart can enhance the readability of the target file and facilitate maintenance personnel to more conveniently understand the performance fluctuation of the storage device within the first time period.
[0087] According to an embodiment of the present invention, by sorting the segmented performance parameters of each second time period, it is convenient for maintenance personnel to timely understand the performance fluctuation within the first time period from the segmented performance parameters sorted in chronological order.
[0088] According to an embodiment of the present invention, the second time period includes multiple third time periods.
[0089] According to an embodiment of the present invention, the storage performance monitoring method further includes: for each third time period, when the number of input / output operations for processing data within the third time period is greater than a danger threshold, associating the number of input / output operations and the data bandwidth corresponding to the number of input / output operations as the performance fluctuation information of the third time period; sorting at least one piece of performance fluctuation information within the first time period in chronological order and writing it into a target file.
[0090] The time length of the third time period can be specifically set according to actual requirements. For example, when the second time period is 5 seconds, the 5 seconds can be divided into multiple third time periods with a length of 100 ms. The danger threshold can be specifically set according to actual requirements, for example, it can be 15.
[0091] For each third time period, count the number of input / output operations for processing data that perform input / output operations within the third time period. If the number of input / output operations exceeds the danger threshold, at this time, associate the number of input / output operations within the third time period and the data bandwidth corresponding to the number of input / output operations as the performance fluctuation information of the third time period.
[0092] For the performance fluctuation information of different third time periods, it can also be sorted in chronological order to obtain the performance fluctuation information of the second time period. Then, sort the performance fluctuation information of different second time periods in time order to obtain the performance fluctuation information of the first time period, and write the performance fluctuation information of the first time period into the target file.
[0093] In another embodiment, it is possible to draw a chart of the number of input / output operations and / or data bandwidth within different third time periods. For the specific drawing method, refer to the chart drawing method in the second time period.
[0094] According to an embodiment of the present invention, by associating the number of input / output operations and the data bandwidth within the third time period where the number of input / output operations is greater than the danger threshold as the performance fluctuation information, it is convenient for maintenance personnel to promptly know the time period when problems occur and the relevant performance from the target file, thereby improving the maintenance efficiency.
[0095] According to an embodiment of the present invention, the storage performance monitoring method further includes: determining the usage rate of the processor every preset fourth time period; for any fourth time period, when the usage rate meets the processor exception condition, writing the usage rate into the target file.
[0096] When the storage device performs input / output operations, data storage and reading are mainly carried out by the processor. At this time, in each fourth time period, the usage rate of the processor in the fourth time period can be collected.
[0097] If the utilization rate of the processor in a certain fourth time period meets the processor exception condition, write the utilization rate into the target file. For example, when the processor utilization rate is greater than the upper utilization rate limit or less than the lower utilization rate limit, write the processor utilization rate into the target file, where the upper utilization rate limit and the lower utilization rate limit can be specifically set according to the actual situation. For example, the upper utilization rate limit can be set to 85% and the lower utilization rate limit can be set to 1%.
[0098] According to an embodiment of the present invention, by determining whether the utilization rate of the processor in each fourth time period meets the processor exception condition, and thereby writing the utilization rate of the processor in the abnormal situation into the target file, it is convenient for maintenance personnel to know in which time period the processor is in an abnormal state, which helps to improve the maintenance efficiency.
[0099] In a specific embodiment, the utilization rate of the processor is calculated in the following manner: sample the processor file to obtain the total working duration and the idle duration within the fourth time period. Calculate the utilization rate of the processor according to the total working duration and the idle duration.
[0100] The total working duration can be the duration of the fourth time period, such as 5 seconds recorded above. The idle duration can refer to the duration when the processor does not perform effective work. For example, the time when the processor does not transmit and process data.
[0101] After obtaining the total working duration and the idle duration, the processor utilization rate can be calculated based on formula (1).
[0102] Processor utilization rate = (total working duration - idle duration) / total working duration (1)
[0103] If the processor utilization rate of this fourth time period is greater than the upper utilization rate limit or less than the lower utilization rate limit, it indicates that the processor is abnormal at this time. It may be that the processor is overloaded or the processor fails and cannot work. At this time, the processor utilization rate can be written into the target file.
[0104] In another embodiment, for multiple abnormal processor utilization rates within the first time period, the multiple processor utilization rates can be sorted in chronological order and the sorted processor utilization rates can be written into the target file. In addition, the processor utilization rates that are not abnormal can also be written into the target file to facilitate maintenance personnel to understand the resource utilization rates of the processor in each time period.
[0105] According to an embodiment of the present invention, the utilization rate is calculated based on the total working duration and the idle duration, and when the processor utilization rate is greater than the upper limit of the utilization rate or less than the lower limit of the utilization rate, the processor utilization rate is written into a target file, facilitating maintenance personnel to promptly know whether there is an abnormality in the processor and locate performance problems caused by insufficient processor resources, helping maintenance personnel to promptly troubleshoot abnormal problems, thereby improving the maintenance efficiency.
[0106] According to an embodiment of the present invention, the first time period includes a plurality of fifth time periods.
[0107] According to an embodiment of the present invention, based on a plurality of processing times and a trigger judgment condition within the first time period, transferring the dump file to a target address includes: for any fifth time period, when the second average delay within the fifth time period is greater than the first delay threshold, transferring the abnormal logs in the abnormal log pool to the target address, where the second average delay is calculated based on a plurality of processing times within the fifth time period; when any processing time is greater than the second delay threshold, transferring the target file to the target address.
[0108] The time length of the fifth time period can be specifically set according to actual needs. For example, when the first time period is 10 minutes, the fifth time period can be every 1 minute within the first time period. The first delay threshold and the second delay threshold can be specifically set according to the actual situation. For example, they can be set according to the device performance of the storage device, such as setting the first delay threshold to 50 ms and the second delay threshold to 100 ms.
[0109] In the first embodiment, for each fifth time period, the second average delay is calculated based on a plurality of processing times of a plurality of input / output operations executed within the fifth time period, and the second average delay represents the average duration required to execute an input / output operation. If the second average delay of a certain fifth time period is greater than the first delay threshold, it indicates that there is a fault in the operating system at this time, and thus the abnormal logs in the abnormal log pool can be transferred to the target address.
[0110] In the second embodiment, if the processing time of any input / output operation within the first time period is greater than the second delay threshold, the target file can be transferred to the target address at this time.
[0111] According to an embodiment of the present invention, based on a plurality of processing times and a trigger judgment condition within the first time period, transferring the dump file to a target address further includes: when the second average delay is greater than the first delay threshold and the processing time is greater than the second delay threshold, transferring the abnormal logs and the target file to the target address in an associated manner.
[0112] In the third embodiment, if the second average delay within a certain fifth time period is greater than the first delay threshold and the processing time of a certain input / output operation is greater than the second delay threshold, the abnormal log and the target file can be synchronously transferred and stored at the target address. For example, a folder for the first time period can be created at the target address, and the abnormal log and the target file can be placed in this folder.
[0113] It should be noted that the values of the first delay threshold and the second delay threshold cannot be set too small, such as 5 ms. A smaller delay threshold will cause high-frequency transfer and storage of the dump file, and too frequent transfer and storage will also affect the processing time of the input / output operation.
[0114] According to the embodiment of the present invention, the transfer and storage operation of the file is performed through two monitoring dimensions, that is, the instantaneous dimension of whether the processing time of any input / output operation is greater than the second delay threshold, and the statistical dimension of whether the second average delay within the fifth time period is greater than the first delay threshold. This embodiment combines the above two dimensions for monitoring, and the file can be dumped when an abnormality occurs in any dimension, which is convenient for maintenance personnel to analyze and judge the fault location in time through the target address when the storage has an abnormality, thereby improving the maintenance efficiency.
[0115] According to the embodiment of the present invention, according to the multiple processing times within the first time period and the trigger judgment condition, the dump file is transferred and stored at the target address, and it further includes: under the condition of meeting the preset condition, performing backup processing on the data in the memory address space of the operating system to obtain a memory image file, and transferring and storing the memory image file at the target address, where meeting the preset condition includes that any processing time is greater than the third delay threshold, or the number of time periods of the fifth time period in which the second average delay is greater than the first delay threshold is greater than the preset number threshold.
[0116] The third delay threshold can be specifically set according to the actual situation, for example, set to 200 ms. The preset number threshold can be set by the maintenance personnel based on the actual situation, such as 3 or 5, etc.
[0117] If the processing time of a certain input / output operation is greater than the third delay threshold, at this time, all the data in the memory address space of the operating system can be backed up to form a memory image file, and the memory image file is transferred and stored at the target address. The maintenance personnel can perform reverse address resolution on the memory image file at the target address. For example, by parsing the memory image file according to the memory organization structure, the resource situation of each module of the operating system and the input / output operation process of processing data can be seen.
[0118] Among the multiple fifth time periods within the first time period (such as 10 minutes), if the second average delay of n fifth time periods is greater than the first delay threshold, and if n is greater than the preset number threshold, at this time, the memory image file can be transferred and stored at the target address.
[0119] According to an embodiment of the present invention, by dumping the memory image file to a target address for maintenance personnel to locate anomalies, the advantage of this method compared to anomaly logs and target files is that it can provide a more comprehensive view of the operating system's running status. The disadvantages are that the dumping process consumes a portion of the operating system's performance, and the generated memory image file occupies a large amount of memory space. However, in the event of a major anomaly, the anomaly can be accurately located from this memory image file.
[0120] Based on the above storage performance monitoring method, the present invention also provides a storage performance monitoring device. The following will be combined with Figure 6 to describe this device in detail.
[0121] Figure 6 The structural block diagram of the storage performance monitoring device according to an embodiment of the present invention is shown.
[0122] As Figure 6 shown, the storage performance monitoring device 600 of this embodiment includes an acquisition module 610, a first storage module 620, a obtaining module 630, a second storage module 640, and a dumping module 650.
[0123] The acquisition module 610 is used to acquire the processing time for completing an input / output operation when the operating system performs an input / output operation on a storage device.
[0124] The first storage module 620 is used to, for any input / output operation, when the processing time of this input / output operation is greater than a first processing duration, store the operation-related information of this input / output operation in the anomaly log in the anomaly log pool.
[0125] The obtaining module 630 is used to, for the input / output operations performed within a first time period, based on their respective processing times, obtain a segmented delay statistical table of the input / output operations within the first time period.
[0126] The second storage module 640 is used to associate and store the device-related information of the storage device performing input / output operations within the first time period with the segmented delay statistical table in a target file in the storage device.
[0127] The dumping module 650 is used to, according to multiple processing times within the first time period and a trigger judgment condition, dump the dump file to a target address, where the dump file includes at least one of the anomaly logs in the anomaly log pool and the target file.
[0128] According to an embodiment of the present invention, by storing operation-related information in an exception log in an exception log pool according to the processing time for the storage device to perform input / output operations, generating a segmented delay statistical table according to multiple processing times within a first time period, and storing device-related information and the segmented delay statistical table within the first time period in a target file, at least one of the exception log and the target file is transferred and stored at a target address in a timely manner when a trigger judgment condition is satisfied. Since the exception log and the segmented delay statistical table are used to save data during storage failures and are transferred and stored in a timely manner through the trigger judgment condition, the information of storage failures is effectively guaranteed, which helps to improve the performance of the operating system and reduces the monitoring cost of performance monitoring.
[0129] According to an embodiment of the present invention, the operation-related information includes operation parameters of the input / output operation and data parameters corresponding to the input / output operation. The operation parameters include operation time parameters for performing the input / output operation, and the data parameters include data information of the processed data corresponding to the input / output operation.
[0130] According to an embodiment of the present invention, the first storage module 620 includes a first acquisition sub-module, an association sub-module, and a first writing sub-module.
[0131] The first acquisition sub-module is configured to obtain operation time parameters from the storage device when the processing time is greater than a first processing duration.
[0132] The association sub-module is configured to perform an association operation on the operation time parameters and the data information in terms of time scale to obtain storage association information.
[0133] The first writing sub-module is configured to write the storage association information into the exception log.
[0134] According to an embodiment of the present invention, the first storage module 620 further includes a second acquisition sub-module and a first storage sub-module.
[0135] The second acquisition sub-module is configured to obtain an operation failure flag corresponding to the input / output operation when the input / output operation corresponding to the processed data fails to execute.
[0136] The first storage sub-module is configured to store the operation failure flag and the data information of the processed data in an associated manner in the exception log.
[0137] According to an embodiment of the present invention, the segmented delay statistical table includes multiple different processing time intervals and the value of the data quantity parameter corresponding to the processing time interval.
[0138] According to an embodiment of the present invention, the obtaining module 630 includes a first determination sub-module, a second determination sub-module, and a first filling sub-module.
[0139] The first determination sub-module is configured to determine, for any input / output operation, the processing time as the time to be classified when the processing time is greater than the first duration threshold.
[0140] The second determination sub-module is configured to determine, for any time to be classified, the target time interval corresponding to the time to be classified from multiple processing time intervals.
[0141] The first filling sub-module is configured to fill the data quantity parameter values of different target time intervals according to the input / output operations corresponding to multiple times to be classified, so as to obtain a segmented delay statistical table.
[0142] According to an embodiment of the present invention, the obtaining module 630 further includes a second filling sub-module.
[0143] The second filling sub-module is configured to, when the processing time is not greater than the first duration threshold, fill the data quantity parameter values of the processing time interval based on the processing time interval corresponding to the processing time.
[0144] According to an embodiment of the present invention, the first time period includes multiple second time periods, and the device-related information includes the performance parameters of the storage device.
[0145] According to an embodiment of the present invention, the second storage module 640 includes a first calculation sub-module, a second calculation sub-module, and a second storage sub-module.
[0146] The first calculation sub-module is configured to calculate, for each second time period, the average number of operations per unit time and the average bandwidth respectively according to the number of input / output operations of the processed data corresponding to at least one input / output operation within the second time period and the amount of bandwidth used for at least one processed data.
[0147] The second calculation sub-module is configured to calculate the first average delay according to the number of input / output operations and the total processing duration, where the total processing duration is generated according to at least one processing time of at least one processed data.
[0148] The second storage sub-module is configured to store the average number of operations, the average bandwidth, the first average delay, and the segmented delay statistical table corresponding to different second time periods in a target file in an associated manner.
[0149] According to an embodiment of the present invention, the second storage sub-module includes a generation unit and a sorting and storage unit.
[0150] The generation unit is configured to generate the segmented performance parameters of the second time period according to the average number of operations, the average bandwidth, and the first average delay for any second time period.
[0151] The sorting storage unit is used to sort the segmented performance parameters of multiple second time periods in chronological order, and store the sorted multiple segmented performance parameters and the segmented delay statistical table in association in the target file.
[0152] According to an embodiment of the present invention, the second time period includes multiple third time periods.
[0153] According to an embodiment of the present invention, the second storage module 640 further includes an association sub-module and a second writing sub-module.
[0154] The association sub-module is used to, for each third time period, when the number of input / output operations for processing data within the third time period is greater than a danger threshold, associate the number of input / output operations and the data bandwidth corresponding to the number of input / output operations as the performance fluctuation information of the third time period.
[0155] The second writing sub-module is used to sort at least one piece of performance fluctuation information within the first time period in chronological order and write it into the target file.
[0156] According to an embodiment of the present invention, the second storage module 640 further includes a second determination sub-module and a third writing sub-module.
[0157] The third determination sub-module is used to determine the usage rate of the processor every preset fourth time period.
[0158] The third writing sub-module is used to, for any fourth time period, when the usage rate meets the processor exception condition, write the usage rate into the target file.
[0159] According to an embodiment of the present invention, the first time period includes multiple fifth time periods;
[0160] According to an embodiment of the present invention, the transfer module 650 includes a first transfer unit, a second transfer unit, and a third transfer unit.
[0161] The first transfer unit is used to, for any fifth time period, when the second average delay within the fifth time period is greater than the first delay threshold, transfer the exception logs in the exception log pool to the target address, where the second average delay is calculated based on multiple processing times within the fifth time period.
[0162] The second transfer unit is used to transfer the target file to the target address when any processing time is greater than the second delay threshold.
[0163] The third transfer unit is used to transfer the exception logs and the target file in association to the target address when the second average delay is greater than the first delay threshold and the processing time is greater than the second delay threshold.
[0164] According to an embodiment of the present invention, the transfer storage module 650 further includes a fourth transfer storage unit.
[0165] The fourth transfer storage unit is configured to, when a preset condition is satisfied, perform backup processing on data in a memory address space in the operating system to obtain a memory image file, and transfer and store the memory image file at a target address, where the satisfaction of the preset condition includes that any processing time is greater than a third delay threshold, or the number of time periods in a fifth time period where the second average delay is greater than a first delay threshold is greater than a preset number threshold.
[0166] According to an embodiment of the present invention, any multiple of the acquisition module 610, the first storage module 620, the obtaining module 630, the second storage module 640, and the transfer storage module 650 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the acquisition module 610, the first storage module 620, the obtaining module 630, the second storage module 640, and the transfer storage module 650 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the acquisition module 610, the first storage module 620, the obtaining module 630, the second storage module 640, and the transfer storage module 650 may be at least partially implemented as a computer program module, and when the computer program module runs, it can execute corresponding functions.
[0167] Figure 7 The block diagram of an electronic device suitable for implementing the above method according to an embodiment of the present invention is shown.
[0168] As Figure 7As shown, the electronic device 700 according to an embodiment of the present invention includes a processor 701, which can perform various appropriate actions and processes according to a program stored in the read-only memory 702 or a program loaded from the storage section 708 into the random access memory 703. The processor 701 can include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), etc. The processor 701 can also include on-board memory for caching purposes. The processor 701 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0169] In the random access memory 703, various programs and data required for the operation of the electronic device 700 are stored. The processor 701, the read-only memory 702, and the random access memory 703 are connected to each other via a bus 704. The processor 701 performs various operations of the method flow according to an embodiment of the present invention by executing the programs in the read-only memory 702 and / or the random access memory 703. It should be noted that the program can also be stored in one or more memories other than the read-only memory 702 and the random access memory 703. The processor 701 can also perform various operations of the method flow according to an embodiment of the present invention by executing the programs stored in the one or more memories.
[0170] According to an embodiment of the present invention, the electronic device 700 can further include an input / output (I / O) interface 705, and the input / output (I / O) interface 705 is also connected to the bus 704. The electronic device 700 can further include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read from it can be installed into the storage section 708 as needed.
[0171] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist alone without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.
[0172] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the read-only memory 702 and / or the random access memory 703 described above, and / or one or more memories other than the read-only memory 702 and the random access memory 703.
[0173] An embodiment of the present invention further includes a computer program product, which includes a computer program, and the computer program includes program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to cause the computer system to implement the method provided by the embodiments of the present invention.
[0174] When the computer program is executed by the processor 701, the above functions defined in the system / apparatus of the embodiments of the present invention are executed. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0175] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 709, and / or be installed from the removable medium 711. The program code included in the computer program may be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0176] In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, the above-mentioned functions defined in the system of the embodiments of the present invention are executed. According to the embodiments of the present invention, the systems, devices, apparatuses, modules, units, etc. described above can be implemented by computer program modules.
[0177] According to the embodiments of the present invention, the program code for executing the computer program provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).
[0178] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0179] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0180] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A method for monitoring storage performance, characterized in that, The method includes: Obtaining the processing time for completing the input / output operation when the operating system performs an input / output operation on the storage device; For any input / output operation, when the processing time of the input / output operation is greater than the first processing duration, storing the operation-related information of the input / output operation in the exception log of the exception log pool; For the input / output operations performed within the first time period, based on their respective processing times, obtaining a segmented delay statistical table of the input / output operations within the first time period; Associatively storing the device-related information of the storage device when performing the input / output operation within the first time period and the segmented delay statistical table in a target file in the storage device; According to the multiple processing times within the first time period and the trigger judgment condition, transferring the dump file to a target address, where the dump file includes at least one of the exception logs in the exception log pool and the target file.
2. The method according to claim 1, wherein The operation-related information includes operation parameters and data parameters of the input / output operation, the operation parameters include operation time parameters for performing the input / output operation, and the data parameters include data information of the processing data corresponding to the input / output operation; Wherein, when the processing time of the input / output operation is greater than the first processing duration, storing the operation-related information of the input / output operation in the exception log of the exception log pool includes: When the processing time is greater than the first processing duration, obtaining the operation time parameter from the storage device; Performing an association operation on the operation time parameter and the data information in terms of time scale to obtain storage association information; Writing the storage association information into the exception log.
3. The method according to claim 2, wherein The method further includes: When the input / output operation corresponding to the processing data fails, obtaining an operation failure identifier corresponding to the input / output operation; Associatively storing the operation failure identifier and the data information of the processing data in the exception log.
4. The method according to claim 1, wherein The segmented delay statistical table includes multiple different processing time intervals and the data quantity parameter values corresponding to the processing time intervals; Wherein, based on their respective processing times, obtaining the segmented delay statistical table of the input / output operations within the first time period includes: For any input / output operation, when the processing time is greater than the first duration threshold, determining the processing time as the time to be classified; For any time to be classified, determining a target time interval corresponding to the time to be classified from multiple processing time intervals; Filling the data quantity parameter values of different target time intervals according to the input / output operations corresponding to multiple times to be classified to obtain the segmented delay statistical table.
5. The method according to claim 4, wherein The method further includes: When the processing time is not greater than the first duration threshold, filling the data quantity parameter values of the processing time interval based on the processing time interval corresponding to the processing time.
6. The method according to claim 1, characterized in that, The first time period includes multiple second time periods, and the device-related information includes the performance parameters of the storage device; Among them, associating and storing the device-related information of the input / output operations performed by the storage device during the first time period with the segmented delay statistical table in a target file in the storage device includes: For any of the second time periods, calculate the average number of operations per unit time and the average bandwidth respectively according to the number of input / output operations corresponding to at least one input / output operation and the amount of bandwidth used for processing at least one piece of data during the second time period; Calculate a first average delay according to the number of input / output operations and the total processing duration, where the total processing duration is generated according to at least one processing time of at least one piece of data; Associate and store the average number of operations, the average bandwidth, the first average delay corresponding to different second time periods, and the segmented delay statistical table in the target file.
7. The method according to claim 6, wherein Associating and storing the average number of operations, the average bandwidth, the first average delay corresponding to different second time periods, and the segmented delay statistical table in the target file includes: For any of the second time periods, generate the segmented performance parameter of the second time period according to the average number of operations, the average bandwidth, and the first average delay; Sort the segmented performance parameters of multiple second time periods in chronological order, and associate and store the sorted multiple segmented performance parameters and the segmented delay statistical table in the target file.
8. The method according to claim 6 or 7, characterized in that, The second time period includes multiple third time periods; Among them, the method further includes: For each of the third time periods, when the number of input / output operations of at least one piece of data during the third time period is greater than a danger threshold, associate the number of input / output operations and the data bandwidth corresponding to the number of input / output operations as the performance fluctuation information of the third time period; Sort at least one of the performance fluctuation information during the first time period in chronological order and write it into the target file.
9. The method according to claim 6, characterized in that, The method further includes: Determine the usage rate of the processor every preset fourth time period; For any of the fourth time periods, when the usage rate meets the processor exception condition, write the usage rate into the target file.
10. The method according to claim 1, wherein The first time period includes multiple fifth time periods; Among them, according to multiple processing times and a trigger judgment condition during the first time period, transferring the dump file to a target address includes: For any of the fifth time periods, when the second average delay during the fifth time period is greater than a first delay threshold, transfer the exception logs in the exception log pool to the target address, where the second average delay is calculated according to multiple processing times during the fifth time period; When any of the processing times is greater than a second delay threshold, transfer the target file to the target address.
11. The method according to claim 10, wherein The method further includes: When the second average delay is greater than the first delay threshold and the processing time is greater than the second delay threshold, transfer the exception logs and the target file to the target address in an associated manner.
12. The method according to claim 10, wherein The method further includes: When the preset conditions are met, the data in the memory address space of the operating system is backed up to obtain a memory image file, and the memory image file is transferred and stored at the target address, where the meeting of the preset conditions includes that any of the processing times is greater than a third delay threshold, or the number of time periods of the fifth time period that meet the condition that the second average delay is greater than the first delay threshold is greater than a preset number threshold.
13. An electronic device, comprising: One or more processors; A memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 12.
14. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, The computer program or instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 12.
15. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, is used to implement the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Log storage method and device
CN112860195A
Storage device and related instruction time delay statistical method
CN116798480A
Abnormal IO positioning method and system and electronic equipment
CN117193638A
Cited By
Memory disk fault detection method and device and related product
CN121281616A