IO resource monitoring method and device, equipment and storage medium

By recording and counting the time-consuming information of IO requests, the problem of inability to effectively monitor the shared disk IO resources of multiple business processes in the prior art is solved, and more accurate resource usage analysis and scheduling optimization are achieved.

CN120255785APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410014259.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art cannot effectively monitor the use of disk IO resources in multiple business processes sharing systems without modifying the kernel source code, resulting in poor monitoring results.

Method used

By obtaining IO requests of business processes, recording time-consuming information on the IO request on the issuance path, and based on the statistical results within the preset period, a new IO resource usage data structure is added to realize IO resource monitoring of the target process group, avoiding analysis errors caused by monitoring the entire machine or a single process.

Benefits of technology

It improves the accuracy and efficiency of IO resource monitoring, can accurately locate path nodes with high latency problems, and supports resource scheduling optimization in business mixed-department scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255785A_ABST
    Figure CN120255785A_ABST
Patent Text Reader

Abstract

The invention relates to an IO resource monitoring method and device, computer equipment, a storage medium and a computer program product. The method can be applied to various technical fields such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving and the like. The method comprises the following steps: acquiring an IO request from a business process; in the process of issuing the IO request to the target disk, recording time consumption information when the IO request reaches each preset path node on an issuing path; carrying out statistics on IO resource use data of the target process group in each disk based on time consumption information recorded for the plurality of IO requests in a preset time period to obtain a statistical result; according to a statistical result, newly adding an IO resource use data structural body in a resource scheduling control structural body corresponding to the target process group; and monitoring the IO resources used by the plurality of business processes included in the target process group according to the resource scheduling control structure so as to improve the monitoring effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of storage technology, and in particular, to an IO resource monitoring method, device, computer device, storage medium, and computer program product. Background Art

[0002] Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, be flexible and convenient, and various services can be deployed on the same server using cloud technology.

[0003] Therefore, each service needs to share the disk IO resources of the system. Without modifying the kernel source code, users need to use relevant IO tools to understand the IO resource usage of each service, so as to automatically make decisions on the scheduling of mixed services and make adjustments and analyze the reasons when a failure occurs.

[0004] However, the IO tools used in related technologies are all based on the overall machine IO situation or the single-process IO situation of the service for statistics and analysis. Therefore, IO resource monitoring can only be carried out from the single-process or overall machine dimension, and the monitoring effect is not good. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide an IO resource monitoring method, device, computer device, computer-readable storage medium, and computer program product that can improve the monitoring effect of IO resources.

[0006] In a first aspect, the present application provides an IO resource monitoring method. The method includes:

[0007] Obtain an IO request from a service process, where the service process belongs to a target process group, the target process group corresponds to a resource scheduling control structure for scheduling and controlling the resources required by multiple service processes included in the target process group, and the IO request is used to access a target disk;

[0008] During the process of sending the IO request to the target disk, record the time-consuming information when the IO request reaches each preset path node on the sending path;

[0009] Based on the time-consuming information recorded for multiple IO requests within a preset period, statistically analyze the IO resource usage data of the target process group on each disk to obtain a statistical result;

[0010] According to the statistical result, add an IO resource usage data structure to the resource scheduling control structure corresponding to the target process group;

[0011] Monitor the IO resources used by multiple service processes included in the target process group according to the resource scheduling control structure.

[0012] In a second aspect, the present application also provides an IO resource monitoring device. The device includes:

[0013] An acquisition module, configured to acquire an IO request from a service process, where the service process belongs to a target process group, the target process group corresponds to a resource scheduling control structure, and the resource scheduling control structure is used to schedule and control the resources required by multiple service processes included in the target process group, and the IO request is used to access a target disk;

[0014] A recording module, configured to record the time-consuming information when the IO request reaches each preset path node on the issuing path during the process of issuing the IO request to the target disk;

[0015] A statistics module, configured to statistically calculate the IO resource usage data of the target process group on each disk based on the time-consuming information recorded for multiple IO requests within a preset time period to obtain a statistical result;

[0016] An adding module, configured to add an IO resource usage data structure to the resource scheduling control structure corresponding to the target process group according to the statistical result;

[0017] A monitoring module, configured to monitor the IO resources used by multiple service processes included in the target process group according to the resource scheduling control structure.

[0018] In some embodiments, the device further includes a selection module, configured to obtain a set of path nodes of the issuing path during the process of issuing the IO request to the target disk, where the set of path nodes includes multiple preset path nodes; and select a preset number of preset path nodes related to resource scheduling from the set of path nodes.

[0019] In some embodiments, the recording module is configured to, for each selected preset path node, record the start time and arrival time when the IO request reaches the preset path node; and determine the time-consuming information when the IO request reaches the preset path node according to the difference between the arrival time and the start time.

[0020] In some embodiments, the statistics module is configured to filter out the IO requests issued by the service processes in the target process group within a preset time period; determine the disks accessed by each of the filtered IO requests; classify the recorded time-consuming information according to the accessed disks, and statistically calculate the IO resource usage data of the target process group on each disk according to the classification result to obtain a statistical result.

[0021] In some embodiments, the statistical result includes the IO resource usage data of each disk. The addition module is configured to determine the storage object of the IO resource usage data structure corresponding to each disk in the resource scheduling control structure corresponding to the target process group; for each disk, store the corresponding IO resource usage data in the corresponding storage object to obtain the IO resource usage data structure corresponding to the disk.

[0022] In some embodiments, the apparatus further includes a summarization module. The summarization module is configured to determine the process groups with the target process group as the parent process group; summarize the IO resource usage data structure corresponding to the target process group and the IO resource usage data structures corresponding to the process groups with the target process group as the parent process group to obtain the recursive IO resource usage data structure of the target process group.

[0023] In some embodiments, the apparatus further includes a scheduling module. The scheduling module is configured to receive an IO resource query request regarding the target process group; filter and parse the IO resource usage data structures corresponding to each disk of the target process group according to the IO resource query request, and check whether there are IO requests with abnormal time consumption according to the time-consuming records corresponding to each IO request in the parsing result; in the case where there are IO requests with abnormal time consumption, perform scheduling control on the IO resources of the target process group.

[0024] In some embodiments, the apparatus further includes a determination module. The determination module is configured to determine the processing duration of each preset path node corresponding to the IO request and the total duration from the issuance of the IO request to the disk according to the time-consuming information of each preset path node corresponding to the IO request with abnormal time consumption; for each preset path node, calculate the ratio between the corresponding processing duration and the total duration, and determine the preset path nodes with abnormal time consumption on the issuance path based on the ratios of the respective preset path nodes.

[0025] In some embodiments, the determination module is further configured to start a target monitoring tool and determine the temporary storage area corresponding to the target monitoring tool; the recording module is further configured to record the total duration from the issuance of the IO request to the target disk in the temporary storage area.

[0026] In some embodiments, the device further includes a viewing module, configured to receive an IO resource query request regarding a target process group; obtain, from the temporary storage area, the total duration of IO requests for each disk corresponding to the target process group according to the IO resource query request, and view the latency situation of the target process group through a target monitoring tool based on the total duration of each IO request.

[0027] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the IO resource monitoring method are implemented.

[0028] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps of the IO resource monitoring method are implemented.

[0029] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the IO resource monitoring method are implemented.

[0030] The above IO resource monitoring method, device, computer device, storage medium, and computer program product obtain IO requests from a business process. The business process belongs to a target process group, and the target process group corresponds to a resource scheduling control structure. The resource scheduling control structure is used to schedule and control the resources required by multiple business processes included in the target process group. The IO request is used to access a target disk. During the process of sending the IO request to the target disk, by recording in real time the time-consuming information when the IO request reaches each preset path node on the sending path, the time-consuming situation of the IO request at different IO preset path nodes in the sending path is further refined. In this way, once a high-latency problem occurs when the IO request is sent, it can quickly respond based on the time-consuming information of each preset path node to accurately locate the IO preset path node most likely to have problems. Based on the time-consuming information recorded for multiple IO requests within a preset period, the IO resource usage data of the target process group on each disk is statistically analyzed to obtain a statistical result. Then, according to the statistical result, an IO resource usage data structure is newly added to the resource scheduling control structure corresponding to the target process group. That is, before monitoring the IO resources, the IO resource usage data belonging to the target process group is statistically analyzed in the resource scheduling control structure in advance, realizing the statistical analysis of the IO resource usage data with the resource scheduling control structure corresponding to the target process group as the statistical unit. In this way, when scheduling the IO resources, it is possible to directly monitor the IO resources used by multiple business processes included in the target process group according to the resource scheduling control structure, avoiding analysis errors caused by monitoring the IO resources of the entire machine or a single process and improving the effect of IO resource monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1A It is a schematic diagram of IO resource monitoring in the related art;

[0032] Figure 1B It is a schematic diagram of IO resource monitoring provided by the implementation of this application;

[0033] Figure 2 It is an application environment diagram of the IO resource monitoring method in an embodiment;

[0034] Figure 3 It is a flowchart of the IO resource monitoring method in an embodiment;

[0035] Figure 4 It is a schematic diagram of the display of the IO resource monitoring result in an embodiment;

[0036] Figure 5 It is a schematic diagram of the display of the IO resource monitoring result in another embodiment;

[0037] Figure 6Schematic diagram of the display of IO resource monitoring results in another embodiment;

[0038] Figure 7 Schematic diagram of the display of IO resource monitoring results in another embodiment;

[0039] Figure 8 Schematic diagram of the display of IO resource monitoring results in another embodiment;

[0040] Figure 9 Schematic diagram of the IO resource monitoring steps in one embodiment;

[0041] Figure 10 Block diagram of the structure of the IO resource monitoring device in one embodiment;

[0042] Figure 11 Internal structure diagram of a computer device in one embodiment. Detailed implementation manners

[0043] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0044] For the convenience of understanding, relevant concepts will be elaborated first.

[0045] Offline services: Services that are not sensitive to latency and have low real-time requirements. Such as computing services, AI (Artificial Intelligence) training, etc.

[0046] Online services: Services that are sensitive to latency and have high real-time requirements. Such as search and payment services.

[0047] Service co-location: Deploying offline services and online services on the same server.

[0048] Process: An execution activity of a program in a computer with respect to a certain data set, and it is the basic unit for the system to allocate resources.

[0049] cgroup (control group): It is a technology used by the kernel for resource management and isolation. Each cgroup represents a resource management unit that uniformly monitors and limits the resources (such as CPU, memory, disk I / O, network, etc.) of a group of processes. In this application, cgroup is referred to as a resource scheduling control structure. A cgroup represents the association between a group of processes and a group of subsystems with parameters. A subsystem represents a type of resource scheduling controller. For example, the memory subsystem can limit the amount of memory used, and the CPU subsystem can limit the CPU usage time. The subsystem is the basis for actually implementing the limitation of a certain type of resource. For example, if a process uses the CPU subsystem to limit the CPU usage time, the association between this process and the CPU subsystem is called a cgroup.

[0050] IO: Read and write input / output (Input / Output) of the operating system.

[0051] General block layer (Block): A block device abstraction layer located between the file system and the disk driver, used to implement the delivery of IO to the hardware device.

[0052] IO request: The IO read and write requests of the system. When the general block layer processes an IO request, it will control through the data structure of the IO request and perform operations such as merging, sorting, and statistics.

[0053] The IO resource monitoring method provided by the embodiments of this application involves the processing of different services. Exemplarily, the service can be various computer tasks in the field of artificial intelligence, including computer vision tasks, speech processing tasks, natural language processing tasks, and so on. Among them, artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, enabling the machine to have the functions of perception, reasoning, and decision-making.

[0054] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operating / interactive systems, and mechatronics. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. The IO resource monitoring method provided in the embodiments of this application specifically relates to the machine learning technology of artificial intelligence.

[0055] Before introducing this application, first introduce two operating states of the operating system, namely the kernel state and the user state. Among them, the kernel state, also called the kernel space, refers to the area where the kernel process or thread is located, and is mainly responsible for running the system and interacting with the hardware. The user state, also called the user space, refers to the area where the user process or thread is located, and is mainly used to execute user programs. The user state provides a space for applications to run. In order for applications to access the resources managed by the kernel, such as memory resources and IO resources, the user state can use the access interfaces provided by the kernel to call the resources managed by the kernel state.

[0056] In the related technologies, as Figure 1A shown, it is a schematic diagram of IO resource monitoring in the related technologies. When a user or an administrator monitors the entire machine in the user state, it is necessary to monitor the IO resource usage of different services. For example, for Service 1, the corresponding cgroup1 manages the IO resources of Service 1. Among them, the processes corresponding to Service 1 are Process 11, Process 12, and Process 13. For Service 2, the corresponding cgroup2 manages the IO resources of Service 2. Among them, the processes corresponding to Service 2 are Process 21, Process 22, and Process 23. For Service 3, the corresponding cgroup3 manages the IO resources of Service 3. Among them, the processes corresponding to Service 3 are Process 31, Process 32, and Process 33. Taking Service 1 as an example, when the kernel state collects the IO resource usage data of each of Process 11, Process 12, and Process 13, although Process 11, Process 12, and Process 13 are all managed by cgroup1, the IO resource usage data of each of Process 11, Process 12, and Process 13 is directly transmitted through cgroup1 to the user state. That is to say, the related technologies monitor the IO resource usage data by individual processes, and thus control and adjust the IO resources related to the service. However, in the actual scenario of multi-service hybrid deployment, each service is deployed, sets policies, and adjusts IO resources in units of cgroups. The related technologies monitor IO resources based on statistical data at the single-process dimension or the entire-machine dimension, which is difficult to accurately analyze and thus cannot ensure the monitoring effect.

[0057] In view of this, an IO resource monitoring method is provided in an embodiment of the present application. For a target process group that implements a service function, it includes at least one service process related to the service. The target process group corresponds to a resource scheduling control structure, that is, a cgroup. At this time, taking the cgroup as a statistical unit, the IO resource usage data of each process managed by the cgroup is summarized. Thus, based on the summary result, the usage situation of the IO resources of the cgroup can be intuitively reflected, avoiding analysis errors caused by monitoring the IO resources of the entire machine or a single process, and improving the effect of IO resource monitoring.

[0058] The following takes Figure 1B as an example to illustrate the IO resource monitoring in an embodiment of the present application. As Figure 1B shown, it is a schematic diagram of IO resource monitoring provided in an embodiment of the present application. Figure 1B It schematically shows the processes of three different services and the situation of IO resource monitoring. For Service 1, the corresponding cgroup1 manages the IO resources of Service 1. Among them, the processes corresponding to Service 1 include Process 11, Process 12, and Process 13. For Service 2, the corresponding cgroup2 manages the IO resources of Service 2. Among them, the processes corresponding to Service 2 include Process 21, Process 22, and Process 23. For Service 3, the corresponding cgroup3 manages the IO resources of Service 3. Among them, the processes corresponding to Service 3 include Process 31, Process 32, and Process 33. Taking Service 1 as an example, when the kernel state collects the IO resource usage data of Process 11, Process 12, and Process 13 respectively, cgroup1 will not directly pass through the IO resource usage data of Process 11, Process 12, and Process 13 to the user state directly, but pre-statistics with cgroup1 as the unit, and summarize the IO resource usage data of Process 11, Process 12, and Process 13 managed by cgroup1. That is to say, for each service process in the target process group, when issuing the corresponding IO request to the hardware layer of the kernel for reading or writing, obtain the IO resource usage data of each preset path node in the issuing path to obtain the segmented data of the corresponding service process. Then, summarize the segmented data of each service process corresponding to Service 1, that is, taking cgourp1 as the statistical unit, summarize the segmented data of each service process, and upload the summarized segmented data to the user state. In this way, the user state can directly and accurately monitor and control the IO resources of cgroup1 based on the summarized segmented data.

[0059] The IO resource monitoring method provided in an embodiment of the present application can be applied to such as Figure 2In the application environment shown. Among them, the terminal 202 communicates with the server 204 through the network. The data storage system can store the data that the server 204 needs to process. The data storage system can be integrated on the server 204, or can be placed on the cloud or other servers.

[0060] In some embodiments, the terminal 202 is a terminal running a cloud disk client, and the server 204 provides a cloud download service for the cloud disk client. When the terminal 202 needs to read a file from the server 204, for example, picture or video data, the service process generates an IO request and sends the IO request to the server 204. The server 204 obtains the IO request. The service process belongs to the target process group, and the target process group corresponds to a resource scheduling control structure. The resource scheduling control structure is used to schedule and control the resources required by multiple service processes included in the target process group. The IO request is used to access the target disk. During the process of the IO request being sent to the target disk, the server 204 records the time-consuming information when the IO request reaches each preset path node on the sending path. Based on the time-consuming information recorded for multiple IO requests within a preset period, the IO resource usage data of the target process group on each disk is statistically analyzed to obtain a statistical result. According to the statistical result, in the resource scheduling control structure corresponding to the target process group, the server 204 adds an IO resource usage data structure, and monitors the IO resources used by multiple service processes included in the target process group according to the resource scheduling control structure.

[0061] In some other embodiments, taking the server 204 executing the IO resource monitoring method alone as an example for illustration. The server 204 deploys a corresponding operating system, including an application layer (deploying applications related to relevant services), a general block layer, and a hardware layer. For a target service, during the running of the corresponding service process, the service process (an instance of the application running) applies for an IO request and sends it to the general block layer. The general block layer sends the IO request to the disk in the hardware layer for corresponding reading or writing. During this process, the server 204 obtains the IO request from the service process. The service process belongs to the target process group, and the target process group corresponds to a resource scheduling control structure. The resource scheduling control structure is used to schedule and control the resources required by multiple service processes included in the target process group. The IO request is used to access the target disk. During the process of sending the IO request to the target disk, the server 204 records the time-consuming information when the IO request reaches each preset path node on the sending path. Based on the time-consuming information recorded for multiple IO requests within a preset period, the server 204 statistically calculates the IO resource usage data of the target process group on each disk to obtain a statistical result. According to the statistical result, in the resource scheduling control structure corresponding to the target process group, the server 204 adds an IO resource usage data structure, and monitors the IO resources used by multiple service processes included in the target process group according to the resource scheduling control structure.

[0062] Among them, the server 204 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 202 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto. Both the terminal 202 and the server 204 are provided with corresponding operating systems. The terminal 202 and the server 204 can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here. It can be understood that this method can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server.

[0063] In one embodiment, as Figure 3 shown, a method for monitoring IO resources is provided. In this embodiment, taking this method applied to the server 204 as an example for illustration, the method includes the following steps:

[0064] Step S302: Obtain an IO request from a service process. The service process belongs to a target process group, and the target process group corresponds to a resource scheduling control structure. The resource scheduling control structure is used to schedule and control the resources required by multiple service processes included in the target process group. The IO request is used to access the target disk.

[0065] Among them, the service process is a process regarding the target service. The target service can be an online service or an offline task, without specific limitation. Exemplarily, in the product interaction scenario, users and merchants achieve online product interaction through an e-commerce platform. There is a mixture of offline services and online services in this e-commerce platform. For example, the offline services that the e-commerce platform needs to process include order processing, inventory management, etc., and the online services that need to be processed include user payment, user product search, etc. Optionally, the execution of the target service involves at least one service process, and the set of at least one service process involved is determined as the target process group. That is to say, the target process group is a set of processes, including at least one service process. It can be understood that the service processes of different services are not mixed in one process group.

[0066] The resource scheduling control structure is a data structure. As described above, the resource scheduling control structure is regarded as a cgroup. The resource scheduling control structure is used to schedule and control the resources required by multiple service processes included in the target process group. In the embodiments of the present application, it means that the resource scheduling control structure is used to schedule and control the IO resources of the target process group. Optionally, one target service involves at least one resource scheduling control structure. Exemplarily, the target process group corresponding to the target service includes M processes, and the target service involves 2 resource scheduling control structures, for example, cgroup1 and cgroup2. Among them, cgroup1 manages m1 processes of the target service, and cgroup manages m2 processes of the target service, and the sum of m1 and m2 is M.

[0067] The IO request is a request initiated by the upper-layer application, used to request read or write operations to the corresponding disk. The upper-layer application includes various applications and services. The target disk is a kind of disk, which is a hardware device for storing data, and the disk is located at the hardware layer. During the running of each service process, at least one read or write operation is involved. For each read or write operation, the upper-layer application needs to initiate a corresponding IO request, and then enter the general block layer of the kernel to issue the IO request layer by layer until it is issued to the corresponding disk for corresponding read or write operations. It can be understood that each service process involves read or write operations on at least one disk.

[0068] Optionally, during the running of the service process on the server, when the service process needs to read or write data, an IO request is generated through the service process and the IO request is issued to the kernel of the server.

[0069] The above business process is also understood as an instance of a running application. Thus, in the case where the business process needs to read data from the target disk or write data to the target disk, an IO request for accessing the target disk is generated by the business process. Among them, the IO request includes a request type, and the request type includes a read type and a write type. Then, the business process issues the IO request to the general block layer of the kernel. The general block layer obtains the IO request from the business process.

[0070] Step S304, during the process of issuing the IO request to the target disk, record the time-consuming information when the IO request reaches each preset path node on the issuing path.

[0071] Among them, the process of issuing the IO request to the target disk refers to the process of the IO request being issued layer by layer from the general block layer of the kernel to the target disk in the hardware layer after entering the kernel. The issuing path refers to the process of the IO request being processed and issued layer by layer in the kernel.

[0072] The preset path node refers to a node on the issuing path. During the entire issuing path process, there are multiple preset path nodes involved. For example, the first path node where the issuing starts, the last path node where the issuing ends, and the intermediate nodes between the first path node and the last path node. The time-consuming information when reaching each preset path node refers to the duration consumed to complete this preset path node. This time-consuming information also reflects the latency situation at this preset path node. The longer the duration reflected by the time-consuming information, the more serious the latency. Further, the preset path node can be a node with time consumption or a node that is likely to cause serious latency. At this time, focus on recording the time-consuming information of multiple preset path nodes that may cause serious latency, and there is no need to record all the nodes in the issuing path, ensuring the efficiency and effectiveness of the recording.

[0073] Optionally, when the IO request starts to enter the kernel, the server determines the issuing path of the IO request and multiple preset path nodes of this issuing path. For each preset path node, after the server reaches this preset path node, it starts to count and record the time-consuming information of this preset path node.

[0074] Exemplarily, the issuing path of each IO request can be the same or different. The multiple preset path nodes involved in each issuing path can be the same or different. Exemplarily, determine the corresponding preset path nodes according to the disk type of the accessed disk.

[0075] Exemplarily, for each preset path node, when arriving at the preset path node, a time statistic function is called to statistically record the time-consuming information of the preset path node and synchronously record the basic information of the IO request. The basic information includes the size of the IO request, the request type, and the address of the target disk accessed. The time statistic function is used to statistically record the time at the preset path node.

[0076] In some embodiments, the time-consuming information when the IO request arrives at each preset path node on the distribution path is recorded, including: for each selected preset path node, recording the start time and the arrival time when the IO request arrives at the preset path node; according to the difference between the arrival time and the start time, determining the time-consuming information when the IO request arrives at the preset path node.

[0077] Among them, each path node corresponds to a distribution stage. Therefore, the distribution path of the IO request involves different distribution stages. For example, the distribution stage can be the stage of obtaining the IO request, and the distribution stage can be the stage of IO scheduling, etc. Therefore, the start time refers to the start time of performing this distribution stage, and the arrival time refers to the end time of completing this distribution stage. The time-consuming information includes the difference between the arrival time and the start time.

[0078] As described above, both the first path node and the last path node of the distribution path are preset path nodes. Therefore, after recording the start time and the arrival time when the IO request arrives at each preset path node, calculate the difference between the arrival time of the last path node and the start time of the first path node, and determine this difference as the total duration of the distribution path.

[0079] In this embodiment, for each selected preset path node, by recording the start time and the arrival time when the IO request arrives at the preset path node, the time-consuming information when the IO request arrives at the preset path node can be determined according to the difference between the arrival time and the start time, and thus the delay situation of each preset path node can be known. The larger the difference between the arrival time and the start time in the preset path node, the more serious the delay. Thus, when it is subsequently monitored that there is an abnormal time consumption in the target process group, the reason for the abnormality can be accurately queried based on the time-consuming information of each preset path node.

[0080] Before the above-mentioned IO request is distributed, if it is queried that the IO resource statistics function is enabled, then during the process of distributing the IO request to the target disk, the time-consuming information when arriving at each path node is recorded.

[0081] Step S306, based on the time-consuming information recorded for multiple IO requests within a preset time period, statistically calculate the IO resource usage data of the target process group on each disk to obtain a statistical result.

[0082] Among them, the preset time period is the time period for monitoring IO resources. Exemplarily, the preset time period can be the duration of the monitoring cycle of IO resource monitoring. For example, during the process from starting to count IO resources to ending the count of IO resources, multiple monitors are performed at a preset monitoring frequency, and the duration of each monitor is the preset time period.

[0083] The IO resource usage data of the disk refers to the usage data of IO resources when accessing the disk. Exemplarily, the IO resource usage data of the disk at least includes the elapsed time information of each IO request for accessing the disk. For example, the IO resource usage data of the disk includes the elapsed time information of each IO request for accessing the disk. Also for example, the IO resource usage data of the disk includes the elapsed time information of each IO request for accessing the disk and the basic information of each IO request. It should be noted that within this preset time period, at least one business process issues at least one IO request. The disks accessed by each IO request may be the same or different. Thus, determining the statistical result based on the IO resource usage data of the disk makes the statistical result clearer and more intuitive, that is, it can more clearly reflect which disks the target process group is involved in and the usage situation of the IO resources of each disk.

[0084] The statistical result is the result obtained by summarizing the IO resource usage data of each disk. Each disk refers to the disks involved in each business process in the target process group.

[0085] Optionally, the server determines at least one IO request that ends being issued within the preset time period, and the at least one IO request is issued by a business process belonging to the target process group. The server obtains the elapsed time information recorded by each of the at least one IO request, and determines the IO resource usage data of each disk based on the multiple elapsed time information, and summarizes the IO resource usage data of each disk to obtain the statistical result of the target process group within the preset time period.

[0086] Exemplarily, after the server detects that the IO resource statistics function is enabled at time t1, it starts timing, and also records the elapsed time information of each IO request issued within the preset time period. The server obtains multiple IO requests that end being issued within the preset time period. For the target process group, based on the mapping relationship between the process group and the IO request, it determines the IO requests corresponding to the target process group, and returns to the step of obtaining the elapsed time information recorded by each of the at least one IO request to continue execution.

[0087] Step S308, according to the statistical result, in the resource scheduling control structure corresponding to the target process group, add an IO resource usage data structure.

[0088] Among them, the IO resource usage data structure is a data structure. The IO resource usage data structure takes the resource scheduling control structure as a unit and summarizes the IO resource usage data of each process it manages. That is to say, the IO resource usage data structure reflects the overall IO resource usage situation of the target process group. Each target process group corresponds to at least one IO resource usage data structure.

[0089] Optionally, for the target process group, the server obtains the corresponding resource scheduling control structure and adds a new IO resource usage data structure corresponding to the target process group in the resource scheduling control structure. The newly added IO resource usage data structure contains this statistical result.

[0090] Optionally, for the IO requests issued by each business process in the target process group, determine the disks accessed by each IO request. For each disk, obtain the IO resource usage data corresponding to the disk from the statistical result, and in the resource scheduling control structure, generate an IO resource usage data structure corresponding to the disk according to the IO resource usage data corresponding to the disk. The number of IO resource usage data structures is the same as the number of disks.

[0091] Exemplarily, the target process group involves business process 1 and business process 2. During a preset period, business process 1 issues IO request 1, and business process 2 issues IO request 2 and IO request 3. At this time, IO request 1 accesses disk 1, IO request 2 accesses disk 2, and IO request 3 accesses disk 3. At this time, each disk corresponds to an IO resource usage data structure, that is, 3 new IO resource usage data structures are added.

[0092] Step S310, monitor the IO resources used by multiple business processes included in the target process group according to the resource scheduling control structure.

[0093] Among them, an IO resource usage data structure is newly added to the resource scheduling control structure mentioned here.

[0094] Optionally, after receiving a request to monitor the target process group, obtain the resource scheduling control structure containing the IO resource usage data structure, and parse the obtained resource scheduling control structure to obtain at least one IO resource usage data structure. Through the corresponding monitoring interface or monitoring tool, monitor the IO resources used by the business processes included in the target business process group based on at least one IO resource usage data structure.

[0095] Among them, the monitoring interface is an interface used to access and manipulate kernel information. Through this monitoring interface, the corresponding IO resource usage data structure can be directly obtained from the kernel to view the IO status of the target process group within a preset time period, for example, the number of reads and writes of each disk involved in the target process group.

[0096] The monitoring tool is a tool for monitoring IO resources based on the technology of running sandbox programs in a privileged context. The monitoring tool can perform secondary processing on at least one IO resource using a data structure and perform IO resource analysis. Once the target process group has an abnormal delay, the preset path node where the delay abnormality occurs in the delivery path can be checked to locate the cause of the delay abnormality.

[0097] Furthermore, when the server detects that the target process group has an abnormal delay within a preset period of time, the IO resource threshold managed by the resource scheduling control structure is increased. For example, while ensuring that the total amount of IO resources remains unchanged, the IO resource threshold corresponding to other process groups is reduced, that is, the IO resources managed by other resource scheduling control structures (that is, the resource scheduling control structures corresponding to other process groups) are reduced, or, according to the business types corresponding to other process groups, other process groups that are not currently important are deleted to increase the IO resources managed by the resource scheduling control structure corresponding to the target process group.

[0098] In the above IO resource monitoring method, by obtaining IO requests from business processes, the business processes belong to a target process group, and the target process group corresponds to a resource scheduling control structure. The resource scheduling control structure is used to schedule and control the resources required by multiple business processes included in the target process group. The IO requests are used to access a target disk. During the process of sending the IO requests to the target disk, by recording in real time the time-consuming information when the IO requests reach each preset path node on the sending path, the time-consuming situation of the IO requests at different IO preset path nodes in the sending path is further refined. In this way, once a high-latency problem occurs when the IO requests are sent, a quick response can be made based on the time-consuming information of each preset path node to accurately locate the IO preset path node that is most likely to have a problem. Based on the time-consuming information recorded for multiple IO requests within a preset period, the IO resource usage data of the target process group on each disk is statistically analyzed to obtain a statistical result. Then, according to the statistical result, an IO resource usage data structure is newly added to the resource scheduling control structure corresponding to the target process group. That is, before performing IO resource monitoring, the IO resource usage data belonging to the target process group is pre-statistically analyzed in the resource scheduling control structure, realizing the statistical analysis of IO resource usage data with the resource scheduling control structure corresponding to the target process group as the statistical unit. In this way, when performing IO resource scheduling, the IO resources used by multiple business processes included in the target process group can be directly monitored according to the resource scheduling control structure, avoiding analysis errors caused by monitoring IO resources for the entire machine or a single process and improving the effect of IO resource monitoring.

[0099] In some embodiments, the method further includes: during the process of sending the IO requests to the target disk, obtaining a set of path nodes of the sending path, where the set of path nodes includes multiple preset path nodes; and selecting a preset number of preset path nodes related to resource scheduling from the set of path nodes.

[0100] Among them, the set of path nodes is the set of path nodes involved in the sending path. The preset path nodes related to resource scheduling refer to the preset path nodes with a high degree of association with resource scheduling. Exemplarily, the preset path nodes related to resource scheduling can be path nodes with a possible delay or important path nodes that affect the sending of IO requests during the resource scheduling process.

[0101] Optionally, the server has pre-stored the set of path nodes of the sending path, and the set of path nodes involved in the sending of each IO request is the same. During the process of sending the IO requests to the target disk, the server obtains the set of path nodes and filters out multiple preset path nodes with a high degree of association between resource scheduling from the set of path nodes.

[0102] Exemplarily, the probability of each preset path node having a delay is statistically calculated, and the preset path nodes with a probability greater than the probability threshold are determined as the path nodes with a high degree of association with resource scheduling. For example, the server filters out preset path nodes such as obtaining an IO request, entering the IO scheduling algorithm queue, completing IO scheduling and entering the distribution queue to the hardware layer, and the hardware layer processing the IO request from the set of path nodes, and records the time-consuming information of the filtered preset path nodes.

[0103] In this embodiment, during the process of issuing an IO request to the target disk, a large number of preset path nodes are involved, and there are some preset path nodes that do not involve resource scheduling, that is, no delay will occur. Therefore, to ensure the monitoring efficiency, by selecting a preset number of preset path nodes related to resource scheduling from the set of path nodes, while ensuring the monitoring effect, the amount of data used for IO resource monitoring is reduced, and the rate of IO resource monitoring is improved.

[0104] In some embodiments, based on the time-consuming information recorded for multiple IO requests within a preset time period, the IO resource usage data of the target process group on each disk is statistically calculated to obtain a statistical result, including: filtering out the IO requests issued by the business processes in the target process group within the preset time period; determining the disks accessed by each of the filtered IO requests; classifying the recorded time-consuming information according to the accessed disks, and statistically calculating the IO resource usage data of the target process group on each disk according to the classification result to obtain a statistical result.

[0105] Optionally, the server obtains each IO request within a preset time period and obtains the mapping relationship between the business process group and the IO request. The server filters out the IO requests issued by the business processes in the target process group from each IO request within the preset time period according to this mapping relationship. The server determines the disk accessed by each of the filtered IO requests, and classifies the recorded time-consuming information of the IO requests accessing the same disk into one category to obtain a classification result. The classification result contains sets of time-consuming information of different categories, and each time-consuming information in each set of time-consuming information corresponds to the same disk, and each category corresponds to one disk. For each disk, the IO resource usage data of the disk is determined according to the set of time-consuming information of the category corresponding to the disk, and the IO resource usage data of each disk is summarized to obtain the statistical result of the target process group within the preset time period.

[0106] Exemplarily, after determining each set of time-consuming information, for each disk, the server counts the I / O statistical information of each I / O request accessing the disk and the merging quantity involved in the disk, so as to obtain the I / O resource usage data of the disk. The statistical information of each I / O request includes the corresponding basic information, the sector number to be accessed (i.e., the sector number within the disk to be accessed), and the time-consuming information of each preset path node. The merging quantity refers to how many I / O requests are merged within the disk. Exemplarily, when obtaining multiple I / O requests from a service process, if these multiple I / O requests all access the same disk, it is further determined whether the multiple I / O requests meet the merging conditions (for example, if the physical addresses accessed by each I / O request are consecutive and the data volumes read and written by each I / O request are small, then they are merged). If the multiple I / O requests meet the merging conditions, the I / O requests are merged, and the merged I / O requests are sent down. At this time, the quantity of the merged I / O requests is recorded.

[0107] Therefore, by summarizing the I / O resource usage data of each disk, the data of each disk is obtained. The data of each disk includes the quantity of read I / Os involved in the disk within a preset time period, the merging quantity of I / O requests, the sector numbers within the disk accessed by the I / O requests, and the quantity of write I / Os. Thus, by combining the data of each disk, the statistical result is obtained.

[0108] In this embodiment, first, by filtering out the I / O requests sent down by the service processes in the target process group within a preset time period, in this way, the disks accessed by the filtered I / O requests can be counted. Then, the recorded time-consuming information is classified according to the accessed disks, and based on the classification result, the I / O resource usage data of the target process group on each disk is counted. Thus, the statistical result is determined through the I / O resource usage data of the disks, making the statistical result clearer and more intuitive, that is, it can more clearly reflect which disks the target process group is involved in and the usage situation of the I / O resources of each disk, facilitating the subsequent effective monitoring of the I / O resources of the target process group.

[0109] In some embodiments, the statistical result includes the I / O resource usage data of each disk. According to the statistical result, in the resource scheduling control structure corresponding to the target process group, an I / O resource usage data structure is newly added, including: in the resource scheduling control structure corresponding to the target process group, determining the storage object of the I / O resource usage data structure corresponding to each disk; for each disk, storing the corresponding I / O resource usage data in the corresponding storage object to obtain the I / O resource usage data structure corresponding to the disk.

[0110] Among them, the storage object is a data structure to be filled, that is, a data structure without any data. Optionally, the server obtains the resource scheduling control structure corresponding to the target process group and the number of disks involved in the target process group. In the resource scheduling control structure, the server adds a corresponding storage object for each disk. For each disk, the server fills the corresponding IO resource usage data in the corresponding storage object to obtain the IO resource usage data structure corresponding to the disk.

[0111] Exemplarily, referring to the foregoing Figure 1B , taking Service 1 as an example, for the target process group of Service 1, it includes Process 11, Process 12, and Process 13. Cgroup1 manages the IO resources of the target process group. During a preset period, Process 11, Process 12, and Process 13 issued a total of n1 IO requests, which involved n2 disks. The storage objects of each disk are determined in cgroup1. For each disk, the IO resource usage data corresponding to the disk is stored in the corresponding storage object to obtain the corresponding dkstats structure (IO resource usage data structure). At this time, there are n2 mutually different dkstats structures in cgroup1. Subsequently, cgroup1 containing n2 dkstats structures is directly uploaded to the user space for monitoring of IO resources through corresponding monitoring tools or interfaces.

[0112] In this embodiment, in the resource scheduling control structure corresponding to the target process group, the storage objects of the IO resource usage data structures corresponding to each disk are determined in advance. For each disk, the corresponding IO resource usage data is stored in the corresponding storage object to obtain the IO resource usage data structure corresponding to the disk. In this way, the IO resource usage data belonging to the target process group is statistically analyzed in the resource scheduling control structure in advance. When monitoring in the user space later, effective monitoring of IO resources can be directly performed based on the resource scheduling control structure containing the IO resource usage data structure, avoiding analysis errors caused by monitoring IO resources for the entire machine or a single process, and improving the effect of IO resource monitoring.

[0113] In some embodiments, the method further includes: determining a process group with the target process group as the parent process group; summarizing the IO resource usage data structure corresponding to the target process group and the IO resource usage data structure corresponding to the process group with the target process group as the parent process group to obtain the recursive IO resource usage data structure of the target process group.

[0114] Among them, a process group with the target process group as the parent process group means that the target process group is the parent, and this process group is the child process group of the target process group. The resource scheduling control structure of the target process group can be understood as a tree-like hierarchical structure, that is, each level contains the resource scheduling control structure of the parent process group, and the next level of each level is the resource scheduling control structure of the corresponding child process group.

[0115] Optionally, when the server verifies that there are child process groups in the target process group, it obtains the recursive relationship of the target process group, which includes at least one corresponding child process group when the target process group is the parent process group. The server obtains at least one corresponding child process group of the target process group according to the recursive relationship of the target process group. The server obtains the resource scheduling control structure corresponding to the target process group and the resource scheduling control structures of each child process group. Each obtained resource scheduling control structure contains the corresponding IO resource usage data structure. The server aggregates the IO resource usage data structure of the target process group and the IO resource usage data structures corresponding to each child process group to obtain the recursive IO resource usage data structure of the target process group.

[0116] Exemplarily, before aggregation, in the resource scheduling control structure corresponding to the target process group, determine the storage object of the recursive IO resource usage data structure, and aggregate the IO resource usage data structures corresponding to the target process group and each child process group respectively to obtain the aggregation result, and store the aggregation result in the corresponding storage object to obtain the recursive IO resource usage data structure. Thus, it can be known that the resource scheduling control structure corresponding to the target process group contains the IO resource usage data structure and the recursive IO resource usage data structure.

[0117] Of course, if there is a process group with a child process group as the parent, that is, this process group is the grandchild process group of the target process group. At this time, after obtaining the target process group, the corresponding child process groups and grandchild process groups, aggregate the IO resource usage data structures corresponding to the target process group, each child process group and each grandchild process group respectively to obtain the recursive IO resource usage data structure of the target process group.

[0118] In this embodiment, after determining the process group with the target process group as the parent process group, aggregate the IO resource usage data structure corresponding to the target process group and the IO resource usage data structure corresponding to the process group with the target process group as the parent process group to obtain the recursive IO resource usage data structure of the target process group. In this way, it allows the user to monitor the IO situation of the resource scheduling control structure corresponding to the target process group in a recursive manner in a timely manner, improving the monitoring convenience.

[0119] In some embodiments, the method further includes: receiving an IO resource query request for a target process group; according to the IO resource query request, screening and parsing the IO resource usage data structures of each disk corresponding to the target process group, and checking whether there are IO requests with abnormal time consumption according to the time consumption records corresponding to each IO request in the parsing result; in the case where there are IO requests with abnormal time consumption, performing scheduling control on the IO resources of the target process group.

[0120] Optionally, the server receives an IO resource query request for a target process group, and the IO resource query request includes an identifier of the target process group. According to the identifier, the resource scheduling control structure of the target process group is screened out from the resource scheduling control structures of multiple process groups. The server parses the resource scheduling control structure to obtain a parsing result, and the parsing result includes the IO resource usage data structures of each disk, and according to the IO resource usage data structures of each disk, the time consumption information corresponding to each IO request is obtained.

[0121] The server determines the total duration of the distribution paths of each IO request according to the time consumption information of each IO request by calling a monitoring tool for abnormal latency monitoring, calculates the sum value of the total durations of each IO request, and in the case where the sum value is greater than the duration threshold, the server determines that there are IO requests with abnormal time consumption, and the server increases the IO resource threshold managed by the resource scheduling control structure.

[0122] Exemplarily, as Figure 4 shown, it is a schematic diagram showing the IO resource monitoring result in an embodiment. After calling the monitoring tool for abnormal latency monitoring, by obtaining the ID (Identity Document) of the cgroup corresponding to the target process group, the IO resource usage data structures of each disk in the cgroup are obtained, and thus, the basic information, the accessed sector number, the time consumption information, and the merge quantity corresponding to each IO request are obtained. As Figure 4 shown, based on the monitoring tool for abnormal latency monitoring, the IO statistical data of the croup can be intuitively seen, that is, the statistical data of the IO requests of the read type ( Figure 4 "read" in Figure 4 ): the number of requests of this request type in this cgroup ( Figure 4 "req_n" in Figure 4 is 18416, and the corresponding merge quantity is 5; the number of sectors used is ( Figure 4 "sectors_n" in Figure 4Statistics of IO requests for "write"): They are 18698, 11, 1197696, and 38425528 in sequence. As Figure 4 It shows the average data in different dimensions. For example, req_n / s, merge_n / s, and kbytes / s are the average values of the number of requests, the number of merges, and the number of sectors respectively (that is, the ratios of the number of requests, the number of merges, and the number of sectors to the monitoring interval). For example, under the read type, they are 2051.24, 0.56, and 65664.51 respectively, and under the write type, they are 2082.65, 1.23, and 66701.72 respectively; iss_time(us) / req and req_time(us) / req are the average values of the time consumed in the hardware and the sum value respectively (that is, the ratios of the time consumed in the hardware and the sum value to the total number of IO requests). For example, under the read type, they are 950.11 and 951.45 respectively, and under the write type, they are 2053.70 and 2055.06 respectively. Further, it can be determined that the IO time consumption of this cgroup is basically generated on the hardware side, and the software consumption is not much. It can be known that the first monitoring tool can filter the IO statistical data according to the required cgroup ID and perform secondary processing to intuitively display the usage of IO resources in the required cgroup.

[0123] In this embodiment, after receiving an IO resource query request for a target process group, according to the IO resource query request, filter and parse the IO resource usage data structure of each disk corresponding to the target process group. According to the time-consuming records corresponding to each IO request in the parsing result, it is possible to accurately and timely verify whether there are IO requests with abnormal time consumption. In the case of there being IO requests with abnormal time consumption, perform scheduling control on the IO resources of the target process group to ensure the effectiveness of the scheduling control.

[0124] In other embodiments, if there is no need to perform secondary processing on the IO resource usage data, the method further includes: receiving an IO resource query request for a target process group; according to the IO resource query request, call the monitoring interface, and filter and parse the IO resource usage data structure of each disk corresponding to the target process group.

[0125] Among them, the monitoring interface is a proc (process information virtual file system) interface, which is used to directly call the resource scheduling control structure of any process group.

[0126] Exemplarily, as Figure 5 shown, it is a schematic diagram of the display of IO resource monitoring results in another embodiment. After calling the monitoring interface, it can intuitively display the changes of each disk in the required cgroup within 3 seconds.Figure 5 In the first, second, third, and fourth rows, the data of Disk 1 (vdb1) and Disk 2 (vdb2) 3 seconds ago, and the data of Disk 1 (vdb1) and Disk 2 (vdb2) after 3 seconds are shown respectively. For each row, the first and second fields are the disk numbers, the third field is the disk name, the fourth to seventh fields are the IO quantities, sector numbers, and processing times corresponding to the read type respectively, and the eighth to eleventh fields are the IO quantities, sector numbers, and processing times corresponding to the write type respectively. The thirteenth field is the total processing time for reads and writes, and the remaining fields are reserved for temporary use. Therefore, it can be known that there is no access to Disk 2 before and after 3 seconds, and only access to Disk 1 is involved.

[0127] In this embodiment, by calling the monitoring interface, the IO resource usage data of different resource scheduling control structures is directly obtained from the kernel, so that the IO resources of the process group corresponding to the resource scheduling control structure can be quickly monitored, improving the efficiency of IO resource monitoring.

[0128] In some embodiments, the method further includes: determining the processing duration of each preset path node corresponding to the IO request and the total duration from the issuance of the IO request to the disk according to the duration information of each preset path node corresponding to the IO request with time-consuming anomalies; for each preset path node, calculating the ratio between the corresponding processing duration and the total duration, and determining the preset path node with abnormal duration on the issuance path based on the respective ratios of each preset path node.

[0129] Optionally, the server determines the IO requests with time-consuming anomalies, obtains the duration information of each preset path node in the IO requests with time-consuming anomalies, and determines the processing duration of each path node and the total duration from the issuance of the IO request to the disk.

[0130] For each preset path node, calculate the ratio between the corresponding processing duration and the total duration, filter out the ratios exceeding the ratio threshold, and determine the preset path node corresponding to the filtered ratio as the preset path node with abnormal duration.

[0131] Optionally, when there are multiple IO requests with time-consuming anomalies, after determining the processing duration of each path node and the total duration from the issuance of the IO request to the disk, for each preset path node, add up the processing durations of each IO request on this preset path node to obtain the first sum value of this preset path node. Add up the total durations of each IO request to obtain the second sum value. For each preset path node, calculate the ratio between the corresponding first sum value and the second sum value, filter out the ratios exceeding the ratio threshold, and determine the preset path node corresponding to the filtered ratio as the preset path node with abnormal duration.

[0132] In this embodiment, according to the time-consuming information of each preset path node corresponding to the IO request with time-consuming anomaly, the processing duration of each preset path node corresponding to the IO request and the total duration from the issuance of the IO request to the disk are determined. For each preset path node, the ratio between the corresponding processing duration and the total duration is calculated, and based on the respective ratios of the preset path nodes, the preset path nodes with abnormal time-consuming on the issuance path can be accurately queried. In this way, once the resource scheduling control structure corresponding to the target process group has a time-consuming anomaly, the preset path node causing the time-consuming anomaly can be timely and accurately located based on the time-consuming information of each preset path node, and thus the reason for the time-consuming anomaly can be accurately determined.

[0133] Regarding the aforementioned monitoring tool for abnormal delay monitoring, this type of monitoring tool essentially uses an IO resource usage data structure for IO resource monitoring. Further, the monitoring tool also includes a target monitoring tool that does not use the IO resource usage data structure for IO resource monitoring, that is, the data sources of the target monitoring tool and the monitoring tool for abnormal delay monitoring are different.

[0134] Among them, the target monitoring tool can be a first target monitoring tool with the ability to visually display the IO request delay distribution, or the target monitoring tool can also be a second target monitoring tool with the ability to monitor the issuance of a single IO request. For example, the first target monitoring tool includes a first delay distribution monitoring tool and a second delay distribution monitoring tool. The first delay distribution monitoring tool is used to visually display the overall delay histogram of the resource scheduling control structure. For example, based on the first delay distribution monitoring tool, the distribution of cgroup in different delay intervals can be visually displayed. The second delay distribution monitoring tool is used to visually display the delay histogram of different request types. The second target monitoring tool can perform a more microscopic analysis, that is, it can follow a single IO request and analyze the issuance situation of the single IO request and the time-consuming information at each preset path node.

[0135] Therefore, in some embodiments, before obtaining the IO request from the business process, it further includes: turning on the target monitoring tool and determining the temporary storage area corresponding to the target monitoring tool; the method further includes: recording the total duration of issuing the IO request to the target disk in the temporary storage area.

[0136] Among them, the temporary storage area is used to store the delay information of each IO request and the request type of the IO request, so that the target monitoring tool can obtain the corresponding statistical data from the temporary storage area. The delay information at least includes the total duration of the IO request issuance path. In other examples, if the second target monitoring tool is turned on, the delay information further includes the time-consuming information of each preset path node.

[0137] Exemplarily, when the server detects that the target monitoring tool is turned on, it determines the temporary storage area corresponding to the target monitoring tool. The server obtains the IO requests from the business process. After the IO requests are issued and end, it obtains the latency information of the IO requests, and the latency information includes the total duration of the IO request issuing path. Moreover, it stores the total duration and the request type of the IO request in the temporary storage area.

[0138] In some other examples, after the server obtains the IO requests from the business process, during the process of issuing the IO requests to the target disk, it stores the latency information of the IO requests in the temporary storage area. The latency information includes the elapsed time information of each preset path node and the total duration of the issuing path.

[0139] It should be noted that after using the monitoring interface for monitoring or turning on the monitoring tool for abnormal latency monitoring, the foregoing steps S302 to S310 are executed, that is, through the IO resource usage data structure in the resource scheduling control structure, the statistics of the IO resource usage data are realized with the resource scheduling control structure as the statistical unit. When using the target monitoring tool for monitoring, the data of the target monitoring tool comes from the real-time situation, that is, using the temporary storage area, the latency information of each IO request is stored in real time. Subsequently, when the target monitoring tool is monitoring, based on the obtained latency information of each IO request and the mapping relationship between the target process group and the IO request, it can also count the IO resource usage data of each disk corresponding to the target process group. The essence of the counting process here is also to count with the resource scheduling control structure as the statistical unit.

[0140] In another example, after determining to turn on the target monitoring tool and determining the temporary storage area, it can also be while executing steps S302 to S310, store the latency information of the IO requests in the temporary storage area. Subsequently, after turning off the target monitoring tool, when switching to the monitoring method of the monitoring interface or the monitoring tool for abnormal latency monitoring, it can directly perform IO resource scheduling based on the resource scheduling control structure containing the IO resource usage data structure.

[0141] In this embodiment, the target monitoring tool is turned on, and the temporary storage area corresponding to the target monitoring tool is determined. Thus, in the case where the IO requests are issued to the target disk, the total duration of issuing the IO requests to the target disk is recorded in the temporary storage area. In this way, subsequently, according to the identifier of the required resource scheduling control structure, it can also be monitored with the resource scheduling control structure as the statistical unit, thereby improving the effect of IO resource monitoring.

[0142] In some embodiments, the method further includes: receiving an IO resource query request for a target process group; according to the IO resource query request, obtaining the total duration of the IO requests of each disk corresponding to the target process group from the temporary storage area, and according to the total duration of each IO request, viewing the latency situation of the target process group through a target monitoring tool.

[0143] Optionally, after receiving an IO resource query request for a target process group, the server parses the IO resource query request to obtain the identifier corresponding to the target process group, determines the corresponding resource scheduling control structure according to the identifier, and according to the mapping relationship between the target process group and the IO requests, filters out the IO requests of each corresponding disk from the IO requests, and filters out the total duration of each IO request from the temporary storage area. The server views the latency situation of the target process group through the target monitoring tool according to the total duration of each IO request.

[0144] Exemplarily, when the target monitoring tool is the first latency distribution monitoring tool, as Figure 6 shown, it is a schematic diagram showing the IO resource monitoring results in another embodiment. The IO resource query request carries the identifiers of two target process groups, corresponding to cgroup1 and cgroup2 respectively. After the server obtains the total duration of each IO request corresponding to cgroup1 and the total duration of each IO request corresponding to cgroup2, through the first latency distribution monitoring tool, it statistically obtains the statistical results of each cgroup in different latency intervals. Figure 6 Both cgroup1 and croup2 involve one disk and the disks are the same. As Figure 6 shown, there are 14 latency intervals (unit: us), which are [0,1], [2,3], …, [1024,2047], [2048,4095], [4096,8191], [8192,16383] in sequence. According to the total duration of each IO request, the statistical results of each latency interval ( Figure 6 “count” in the figure), it can be known that there are no statistical values in the first 10 latency intervals, that is, the total duration of each IO request in cgroup1 is distributed in the last 4 latency intervals. Thus, visualizing the latency distribution in different latency intervals ( Figure 6 “distribution” in the figure), that is, the corresponding distribution histogram. Similarly, based on the total duration of each IO request corresponding to cgroup2, the corresponding distribution histogram is determined. According to Figure 6 shown, the latency situations of the two cgroups for the same disk are basically the same. Approximately 50% is distributed in the 2 - 4 ms (i.e., [2048,4095]) interval, 25% is in the 4 - 8 ms (i.e., [4096,8191]) interval, and the highest latency appears in the 8 - 16 ms (i.e., [8192,16383]) interval.

[0145] Exemplarily, in the case where the target monitoring tool is the second delay distribution monitoring tool, as Figure 7 shown, it is a schematic diagram showing the monitoring results of IO resources in another embodiment. The IO resource query request carries the identifier of a target process group, corresponding to cgroup3. After the server obtains the total duration of each IO request corresponding to cgroup3 and the type of each IO request from the temporary storage area, through the first delay distribution monitoring tool, it statistically analyzes the delay histograms of different types respectively. As Figure 7 shown, there are 14 delay intervals (unit: us), which are [0,1], [2,3], …, [1024,2047], [2048,4095], [4096,8191], [8192,16383] in sequence. Based on the total duration of each read-type IO request in cgroup3, the statistical results of each delay interval ( Figure 7 "count" in shown) can be known. It can be seen that there are no statistical values in the first 9 delay intervals, that is, the total duration of each read-type IO request in cgroup3 is distributed in the last 5 delay intervals. Thus, the delay distribution of the read type in cgroup3 ( Figure 7 "distribution" in shown), that is, the corresponding distribution histogram, is visualized. Similarly, based on the total duration of each write-type IO request in cgroup3, the statistical results of each delay interval ( Figure 7 "count" in shown) can be known. It can be seen that there are no statistical values in the first 11 delay intervals, that is, the total duration of the write-type IO requests in cgroup3 is distributed in the last 3 delay intervals. Thus, the delay distribution of the write type in cgroup3 is visualized. As Figure 7 can be known, in cgroup3, there are both read type and write type in total. Among them, the average delay of the read-type IO requests is slightly larger than that of the write-type IO requests. 80% of the write-type IO requests have a delay in the interval of 1~2ms (that is, [1024,2047]), while 30% of the read-type IO requests have a delay in the interval of 2~4ms (that is, [2048,4095]), and 60% are as high as 4~8ms (that is, [4096,8191]) interval. Therefore, if the service corresponding to this cgroup3 expects to optimize the overall service speed by reducing the IO delay situation. Then, for Figure 7 the reflected delay distribution histogram, more strategies prior to the read-type IO can be considered, and the scheduling effect will be relatively obvious.

[0146] Exemplarily, in the case where the target monitoring tool is the second target monitoring tool, as Figure 8As shown, it is a schematic diagram showing the monitoring results of IO resources in another embodiment. The IO resource query request carries the identifier of a target process group. For example, it is the ID of cgroup4. After the server obtains the time-consuming information of each IO request corresponding to cgroup4 at each preset path node and the types of each IO request from the temporary storage area, it uses the second target monitoring tool to count the statistical information of each IO request.

[0147] For each IO request, the statistical information sequentially includes Figure 8 the start time of the IO request at the start path node (corresponding to the column where "TIME (s)" is located in Figure 8 ), the cgroup ID to which it belongs (corresponding to the column where "CGID" is located in Figure 8 ), the IO request generation tool (corresponding to the column where "COMM" is located in Figure 8 ), the business process ID (corresponding to the column where "PID" is located in Figure 8 ), the hardware device (corresponding to the column where "DISK" is located in Figure 8 ), the request type (corresponding to the column where "T" is located in Figure 8 ), the IO request size (corresponding to the column where "SECTOR" is located in Figure 8 ), the sector number in the accessed disk (corresponding to the column where "SECS" is located in Figure 8 ), the bytes (corresponding to the column where "BYTES" is located in Figure 8 ), the time consumed by the IO in the queue (corresponding to "QUE (ms)" in Figure 8 ) and the time consumed in the hardware device (corresponding to the column where "LAT(ms)" is located in Figure 8 ). The data related to each IO request can be seen in the values of the corresponding columns in Figure 8 . As Figure 8 shown, the cgroup ID of each IO request is 97, the IO request generation tools are all generated by the fio (multi-threaded IO generation tool), and the types of hardware devices are all disk devices (i.e., vdb). The request types involved include read type ( Figure 8 "R" in Figure 8 ) and write type ( "W" in

[0148] ). The latency of read-type IO requests in the queue is relatively longer than that of write-type IO requests, and the processing time in the device is shorter than that of write IOs. Most read and write IOs can be completed within 1 - 2 ms, and the overall time consumption basically occurs in the hardware device. That is to say, through the second target monitoring tool, it is possible to further follow the issuance and completion status of each IO request. That is, when performing abnormal latency monitoring, it is also possible to use the second target monitoring tool to monitor whether there is a latency anomaly and locate the cause of the abnormal latency.In this embodiment, after receiving an IO resource query request for a target process group, according to the IO resource query request, the total duration of the IO requests of each disk corresponding to the target process group is obtained from the temporary storage area, and according to the total duration of each IO request, through the target monitoring tool, the latency situation of the target process group is checked. Thus, the IO resource usage of the corresponding resource scheduling control structure is monitored in real time. Thus, it is also possible to monitor with the resource scheduling control structure as the statistical unit, thereby improving the effect of IO resource monitoring.

[0149] The present application also provides an application scenario that applies the above-mentioned IO resource monitoring method. Specifically, the application of the IO resource monitoring method in this application scenario is as follows: In a product interaction scenario, users and merchants conduct product interactions on an interaction platform. Both online services and offline services are deployed on the interaction platform, and different services generate corresponding service processes. Since the issuance of IO requests is involved when executing the corresponding services, in order to ensure the normal operation of each service, it is necessary to monitor the IO resources of each service. At this time, the target process group with the highest sensitivity to latency is determined from different process groups. By using the IO resource monitoring method of the embodiment of the present application, it is possible to effectively monitor the IO resources of the target process group and adaptively adjust the thresholds of the IO resources of other process groups based on the monitoring results. Specifically, the server obtains an IO request from a service process, the service process belongs to the target process group, the target process group corresponds to a resource scheduling control structure, and the resource scheduling control structure is used to schedule and control the resources required by multiple service processes included in the target process group. The IO request is used to access the target disk; during the process of the IO request being sent to the target disk, the server records the time-consuming information when the IO request reaches each preset path node on the sending path; based on the time-consuming information recorded for multiple IO requests within a preset period, the IO resource usage data of the target process group on each disk is statistically analyzed to obtain a statistical result; the server adds an IO resource usage data structure to the resource scheduling control structure corresponding to the target process group according to the statistical result; and the IO resources used by multiple service processes included in the target process group are monitored according to the resource scheduling control structure.

[0150] Of course, it is not limited to this. The IO resource monitoring method provided by this application can also be applied to other application scenarios. For example, in a video processing scenario, involving a video platform, where online services (such as real-time video streams) and offline services (such as video transcoding and compression) are mixedly deployed on the video platform. At this time, different services involve IO request operations. Therefore, the IO resource monitoring method provided by this application can be used to monitor the IO resources of different services. Another example is in a medical image processing scenario, where online services (such as real-time medical data collection and processing) and offline services (such as medical image storage and management) are mixedly deployed on the same machine or the same server, and the IO resource monitoring method provided by this application can also be used to monitor the IO resources of different services.

[0151] The above application scenarios are only illustrative explanations. It can be understood that the application of the IO resource monitoring method provided by each embodiment of this application is not limited to the above scenarios.

[0152] In a specific embodiment, as Figure 9 shown, it is a schematic diagram of the IO resource monitoring steps in an embodiment. Specifically as follows:

[0153] Step 1: After the server detects that the IO resource statistics function is enabled, obtain the monitoring method of the IO resources. If the monitoring method is to monitor based on a monitoring tool for abnormal delay monitoring or to monitor using a monitoring interface, then execute Steps 2 to 6. If the monitoring method is to monitor based on a target monitoring tool, then execute Steps 7 to 8.

[0154] Step 2: When a monitoring tool for abnormal delay monitoring is enabled, or when the user selects to monitor using a monitoring interface, the server obtains the IO requests from the business processes. The business processes belong to the target process group, and the target process group corresponds to a resource scheduling control structure, which is used to schedule and control the resources required by multiple business processes included in the target process group. The IO requests are used to access the target disk.

[0155] Step 3: During the process of sending the IO requests to the target disk, obtain the set of path nodes of the sending path. The set of path nodes includes multiple preset path nodes; select a preset number of preset path nodes related to resource scheduling from the set of path nodes.

[0156] Exemplarily, as Figure 9 shown, the server filters out path nodes such as obtaining IO requests, IO scheduling algorithms, IO sending queues, and IO request completion from the set of path nodes. For each preset path node filtered out.

[0157] Step 4: For each selected preset path node, the server records the start time and arrival time when the IO request reaches the preset path node. According to the difference between the arrival time and the start time, the time-consuming information of the IO request reaching the preset path node is determined.

[0158] Step 5: The server filters out the IO requests issued by the business processes in the target process group within the preset time period. Determine the disks accessed by each of the filtered IO requests. Classify the recorded time-consuming information according to the accessed disks, and based on the classification results, count the IO resource usage data of the target process group on each disk to obtain the statistical results. The statistical results include the IO resource usage data of each disk. The server determines the storage objects of the IO resource usage data structures corresponding to each disk in the resource scheduling control structure corresponding to the target process group according to the statistical results. For each disk, store the corresponding IO resource usage data in the corresponding storage object to obtain the IO resource usage data structure corresponding to the disk. At this time, the resource scheduling control structure includes the IO resource usage data structure.

[0159] If there is a subclass process group in the target process group, then determine the process group with the target process group as the parent process group. And, summarize the IO resource usage data structure corresponding to the target process group and the IO resource usage data structure corresponding to the process group with the target process group as the parent process group to obtain the summary result, and store the summary result in the storage object of the recursive IO resource usage data structure of the target process group to obtain the recursive IO resource usage data structure. At this time, the resource scheduling control structure includes the IO resource usage data structure of each disk and the recursive IO resource usage data structure.

[0160] Step 6: After receiving the IO resource query request for the target process group, the server filters and parses the IO resource usage data structures corresponding to each disk of the target process group according to the IO resource query request through the monitoring tool for abnormal delay monitoring or the monitoring interface.

[0161] Furthermore, if it is necessary to further verify whether there is a time-consuming anomaly, further, obtain the time-consuming records corresponding to each IO request from the IO resource usage data structures of each disk, and through the monitoring tool for abnormal delay monitoring, determine the total duration of the issuing path of each IO request according to the time-consuming information of each IO request, calculate the sum value of the total durations of each IO request, and in the case where the sum value is greater than the duration threshold, the server determines that there is an IO request with a time-consuming anomaly, and the server increases the IO resource threshold managed by the resource scheduling control structure.

[0162] Step 7: After the server determines that the target monitoring tool is enabled, it determines the temporary storage area corresponding to the target monitoring tool. During the process of issuing the IO request to the target disk, the time-consuming information of the IO request at each preset path node and the total duration involved in the issuing path are both recorded in real time in the temporary storage area. For example, the time-consuming information for obtaining the IO request, the IO scheduling algorithm, the IO issuing queue, and the completion of the IO request are recorded in real time.

[0163] Step 8: The server receives an IO resource query request for the target process group; according to the IO resource query request, it obtains the total duration of the IO requests for each disk corresponding to the target process group from the temporary storage area, and based on the total duration of each IO request, it views the latency situation of the target process group through the target monitoring tool.

[0164] In this embodiment, by obtaining the IO requests from the service processes, where the service processes belong to the target process group, and the target process group corresponds to a resource scheduling control structure, which is used to schedule and control the resources required by multiple service processes included in the target process group, and the IO requests are used to access the target disk. During the process of sending the IO requests to the target disk, by recording in real time the time-consuming information when the IO requests reach each preset path node on the sending path, the time-consuming situation of the IO requests at different IO preset path nodes in the sending path is further refined. In this way, once a high-latency problem occurs when the IO requests are sent, it can quickly respond based on the time-consuming information of each preset path node to accurately locate the IO preset path node that is most likely to have problems. Based on the time-consuming information recorded for multiple IO requests within a preset period, the IO resource usage data of the target process group on each disk is statistically analyzed to obtain a statistical result. Then, according to the statistical result, an IO resource usage data structure is newly added to the resource scheduling control structure corresponding to the target process group. That is, before monitoring the IO resources, the IO resource usage data belonging to the target process group is pre-statistically analyzed in the resource scheduling control structure, realizing the statistical analysis of the IO resource usage data with the resource scheduling control structure corresponding to the target process group as the statistical unit. In this way, when scheduling the IO resources, it can directly monitor the IO resources used by multiple service processes included in the target process group according to the resource scheduling control structure, avoiding analysis errors caused by monitoring the IO resources of the entire machine or a single process, and improving the effect of IO resource monitoring. Therefore, it can improve the efficiency of IO resource monitoring during the business co-location process and is also conducive to analysis and decision-making. It can effectively reduce the business operation cost, save the power consumption of the IDC (Internet Data Center) data center, and reduce carbon emissions. In addition, based on the time-consuming information of each preset path node, it can accurately locate the preset path node that causes the abnormal delay in the case of abnormal delay. In addition, according to the user's needs, the monitoring method can be flexibly set to obtain the corresponding monitoring result, improving the user experience. And based on using different monitoring tools for monitoring, it can be realized to monitor from the macroscopic or microscopic observation dimension.

[0165] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0166] Based on the same inventive concept, an embodiment of the present application further provides an IO resource monitoring device for implementing the above-mentioned IO resource monitoring method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the IO resource monitoring device provided below can refer to the limitations on the IO resource monitoring method in the above text, and will not be repeated here.

[0167] In one embodiment, as Figure 10 shown, an IO resource monitoring device 1000 is provided, including: an acquisition module 1002, a recording module 1004, a statistics module 1006, a new addition module 1008, and a monitoring module 1010, where:

[0168] The acquisition module 1002 is configured to acquire IO requests from a service process. The service process belongs to a target process group, and the target process group corresponds to a resource scheduling control structure. The resource scheduling control structure is used to schedule and control the resources required by multiple service processes included in the target process group. The IO request is used to access a target disk;

[0169] The recording module 1004 is configured to record the time-consuming information when the IO request reaches each preset path node on the sending path during the process of sending the IO request to the target disk;

[0170] The statistics module 1006 is configured to statistically calculate the IO resource usage data of the target process group on each disk based on the time-consuming information recorded for multiple IO requests within a preset period, and obtain a statistical result;

[0171] The new addition module 1008 is configured to add an IO resource usage data structure to the resource scheduling control structure corresponding to the target process group according to the statistical result;

[0172] The monitoring module 1010 is configured to monitor the IO resources used by multiple service processes included in the target process group according to the resource scheduling control structure.

[0173] In some embodiments, the apparatus further includes a selection module configured to, during the process of sending an IO request to a target disk, obtain a set of path nodes of the sending path, where the set of path nodes includes a plurality of preset path nodes; and select a preset number of preset path nodes related to resource scheduling from the set of path nodes.

[0174] In some embodiments, the recording module 1004 is configured to, for each selected preset path node, record the start time and the arrival time when the IO request arrives at the preset path node; and determine the time-consuming information of the IO request arriving at the preset path node according to the difference between the arrival time and the start time.

[0175] In some embodiments, the statistics module 1006 is configured to filter out the IO requests sent by the service processes in the target process group within a preset time period; determine the disks accessed by the filtered IO requests respectively; classify the recorded time-consuming information according to the accessed disks, and statistically obtain the IO resource usage data of the target process group on each disk according to the classification result, so as to obtain a statistical result.

[0176] In some embodiments, the statistical result includes the IO resource usage data of each disk. The new addition module 1008 is configured to determine the storage object of the IO resource usage data structure corresponding to each disk in the resource scheduling control structure corresponding to the target process group; for each disk, store the corresponding IO resource usage data in the corresponding storage object to obtain the IO resource usage data structure corresponding to the disk.

[0177] In some embodiments, the apparatus further includes a summarization module configured to determine the process groups with the target process group as the parent process group; and summarize the IO resource usage data structure corresponding to the target process group and the IO resource usage data structures corresponding to the process groups with the target process group as the parent process group to obtain the recursive IO resource usage data structure of the target process group.

[0178] In some embodiments, the apparatus further includes a scheduling module configured to receive an IO resource query request regarding the target process group; filter and parse the IO resource usage data structures corresponding to each disk of the target process group according to the IO resource query request, and check whether there are any IO requests with abnormal time consumption according to the time-consuming records corresponding to the IO requests in the parsing result; and perform scheduling control on the IO resources of the target process group in the case of the existence of IO requests with abnormal time consumption.

[0179] In some embodiments, the apparatus further includes a determination module configured to determine, according to the duration information of each preset path node corresponding to the IO request with time-consuming anomaly, the processing duration of each preset path node corresponding to the IO request and the total duration from the issuance of the IO request to the disk; for each preset path node, calculate the ratio between the corresponding processing duration and the total duration, and determine the preset path nodes with abnormal duration on the issuance path based on the respective ratios of the preset path nodes.

[0180] In some embodiments, the determination module is further configured to activate a target monitoring tool and determine the temporary storage area corresponding to the target monitoring tool; the recording module is further configured to record the total duration from the issuance of the IO request to the target disk in the temporary storage area.

[0181] In some embodiments, the apparatus further includes a viewing module configured to receive an IO resource query request regarding a target process group; according to the IO resource query request, obtain the total duration of the IO requests of each disk corresponding to the target process group from the temporary storage area, and view the latency situation of the target process group through the target monitoring tool based on the total duration of each IO request.

[0182] Each module in the above IO resource monitoring apparatus can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.

[0183] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 11 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals through a network connection. When the computer program is executed by the processor, it implements an IO resource monitoring method.

[0184] Those skilled in the art can understand, Figure 11The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0185] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0186] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0187] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0188] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0189] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0190] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0191] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. An IO resource monitoring method, characterized in that, The method includes: Obtaining an IO request from a service process, where the service process belongs to a target process group, the target process group corresponds to a resource scheduling control structure, the resource scheduling control structure is used to schedule and control the resources required by multiple service processes included in the target process group, and the IO request is used to access a target disk; During the process of sending the IO request to the target disk, recording the time-consuming information when the IO request reaches each preset path node on the sending path; Based on the time-consuming information recorded for multiple IO requests within a preset period, statistically analyzing the IO resource usage data of the target process group on each disk to obtain a statistical result; According to the statistical result, adding an IO resource usage data structure in the resource scheduling control structure corresponding to the target process group; Monitoring the IO resources used by multiple service processes included in the target process group according to the resource scheduling control structure.

2. The method according to claim 1, characterized in that, The method further includes: During the process of sending the IO request to the target disk, obtaining a set of path nodes of the sending path, where the set of path nodes includes multiple preset path nodes; Selecting a preset number of preset path nodes related to resource scheduling from the set of path nodes.

3. The method according to claim 2, characterized in that, The recording the time-consuming information when the IO request reaches each preset path node on the sending path includes: For each selected preset path node, recording the start time and arrival time when the IO request reaches the preset path node; Determining the time-consuming information when the IO request reaches the preset path node according to the difference between the arrival time and the start time.

4. The method according to claim 1, characterized in that The statistically analyzing the IO resource usage data of the target process group on each disk based on the time-consuming information recorded for multiple IO requests within a preset period to obtain a statistical result includes: Filtering out the IO requests sent by the service processes in the target process group within the preset period; Determining the disks accessed by the filtered IO requests respectively; Classifying the recorded time-consuming information according to the accessed disks, and statistically analyzing the IO resource usage data of the target process group on each disk according to the classification result to obtain a statistical result.

5. The method according to claim 1, wherein The statistical result includes the IO resource usage data of each disk, and the adding an IO resource usage data structure in the resource scheduling control structure corresponding to the target process group according to the statistical result includes: In the resource scheduling control structure corresponding to the target process group, determining the storage object of the IO resource usage data structure corresponding to each disk; For each disk, storing the corresponding IO resource usage data in the corresponding storage object to obtain the IO resource usage data structure corresponding to the disk.

6. The method according to claim 1, wherein The method further includes: Determining the process groups with the target process group as the parent process group; Summarizing the IO resource usage data structure corresponding to the target process group and the IO resource usage data structures corresponding to the process groups with the target process group as the parent process group to obtain the recursive IO resource usage data structure of the target process group.

7. The method according to claim 1, wherein The method further includes: Receive an IO resource query request for a target process group; According to the IO resource query request, filter and parse the IO resource usage data structures of each disk corresponding to the target process group, and check whether there are IO requests with abnormal time consumption according to the time consumption records corresponding to each IO request in the parsing result; In the case where there are IO requests with abnormal time consumption, perform scheduling control on the IO resources of the target process group.

8. The method according to claim 7, characterized in that, The method further includes: According to the time consumption information of each preset path node corresponding to the IO request with abnormal time consumption, determine the processing duration of each preset path node corresponding to the IO request and the total duration of the IO request sent to the disk; For each preset path node, calculate the ratio between the corresponding processing duration and the total duration, and based on the respective ratios of each preset path node, determine the preset path nodes with abnormal time consumption on the sending path.

9. The method according to claim 1, wherein Before obtaining the IO request from the service process, it further includes: Start a target monitoring tool and determine the temporary storage area corresponding to the target monitoring tool; The method further includes: Record the total duration of the IO request sent to the target disk in the temporary storage area.

10. The method according to claim 9, characterized in that, The method further includes: Receive an IO resource query request for a target process group; According to the IO resource query request, obtain the total duration of the IO requests of each disk corresponding to the target process group from the temporary storage area, and view the latency situation of the target process group through the target monitoring tool according to the total duration of each IO request.

11. An IO resource monitoring device, characterized in that, The device includes: An acquisition module, configured to acquire an IO request from a service process, where the service process belongs to a target process group, the target process group corresponds to a resource scheduling control structure, the resource scheduling control structure is used to schedule and control the resources required by multiple service processes included in the target process group, and the IO request is used to access a target disk; A recording module, configured to record the time consumption information when the IO request reaches each preset path node on the sending path during the process of sending the IO request to the target disk; A statistics module, configured to perform statistics on the IO resource usage data of the target process group on each disk based on the time consumption information recorded for multiple IO requests within a preset period to obtain a statistical result; An adding module, configured to add an IO resource usage data structure to the resource scheduling control structure corresponding to the target process group according to the statistical result; A monitoring module, configured to monitor the IO resources used by multiple service processes included in the target process group according to the resource scheduling control structure.

12. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 10.

14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Storage IO multi-path shunting method and system based on Cgroup traceability

    CN122093310A

  • A Storage I / O Multipath Splitting Method and System Based on Cgroup Source Tracing

    CN122093310B