Performance Monitoring in Distributed Storage Systems

By generating probe requests and analyzing responses in a distributed storage system, the problem of difficulty in accurately monitoring system performance in the prior art is solved, and non-invasive performance evaluation and SLA compliance evaluation are realized, providing an accurate understanding of system performance.

CN114217948BActive Publication Date: 2025-07-22GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111304498.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2015-11-13
Filing Date
2016-09-27
Publication Date
2025-07-22
Estimated Expiration
2036-09-27

AI Technical Summary

Technical Problem

In distributed storage systems, prior art is difficult to accurately monitor system performance, especially end-to-end performance and SLA compliance without affecting client request processing.

Method used

By generating probing requests based on client requests, the distributed storage system prepares without accessing actual data, monitoring system performance, including generating probing requests and analyzing responses to calculate performance metrics.

Benefits of technology

Non-invasive performance monitoring of distributed storage systems is implemented, providing accurate evaluation of end-to-end performance and SLA compliance, reducing the impact on client requests, and providing statistical information on system hotspots and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114217948B_ABST
    Figure CN114217948B_ABST
Patent Text Reader

Abstract

The present disclosure relates to performance monitoring in a distributed storage system. Methods and systems for monitoring performance in a distributed storage system are described. An example method includes: identifying requests sent by a client to the distributed storage system, each request including a request parameter value of a request parameter; generating a probe request based on the identified requests, the probe request including a probe request parameter value of a probe request parameter representing a statistical sample of the request parameters included in the identified requests; sending the generated probe request to the distributed storage system via a network, wherein the distributed storage system is configured to perform preparatory work to serve each probe request in response to receiving the probe request; receiving a response to the probe request from the distributed storage system; and outputting at least one performance metric value for measuring the current performance state of the distributed storage system based on the received response.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Division Explanation

[0002] This application is a divisional application of Chinese Patent Application No. 201680047730.4 with a filing date of September 27, 2016. Technical Field

[0003] This specification generally relates to monitoring performance in a distributed storage system. Background Art

[0004] In a distributed system, various performance metrics can be tracked to determine the overall health of the system. For example, the amount of time (i.e., latency) it takes for the system to respond to a client request can be monitored to ensure that the system responds in a timely manner. Summary of the Invention

[0005] Generally, one aspect of the subject matter described in this specification can be embodied as a system and a method performed by a data processing device. The method includes the following actions: identifying requests sent by a client to a distributed storage system, each request including a request parameter value of a request parameter; based on the identified requests, generating probe requests, each probe request including a probe request parameter value of a probe request parameter representing a statistical sample of the request parameters included in the identified requests; sending the generated probe requests to the distributed storage system over a network, where the distributed storage system is configured to perform preparatory work in response to receiving the probe requests to serve each probe request; receiving responses to the probe requests from the distributed storage system; and outputting at least one performance metric value for measuring the current performance state of the distributed storage system based on the received responses.

[0006] One or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and potential advantages of the subject matter will be apparent from the description, the drawings, and the claims. Brief Description of the Drawings

[0007] Figure 1 is a schematic diagram of an exemplary environment for monitoring performance in a distributed storage system.

[0008] Figure 2 is a swimlane diagram of an exemplary process for processing client requests in a distributed storage system.

[0009] Figure 3 is a swimlane diagram of an exemplary process for monitoring performance in a distributed storage system.

[0010] Figure 4 is a flowchart of an exemplary process for monitoring performance in a distributed storage system.

[0011] Figure 5 is a schematic diagram of a computing device that can be used to implement the systems and methods described in this document.

[0012] In the various figures, the same reference numerals and markings denote the same units. Detailed implementation

[0013] In a distributed storage system, there are many factors that affect system performance. For example, a client can access the system through a public network such as the Internet. In this case, from the perspective of each client, performance issues anywhere in the public network may affect the performance of the distributed storage system. Issues with the distributed storage system itself, such as hardware failures, internal network failures, software errors, or other problems, may also affect the performance of the system as perceived by the client.

[0014] In some cases, the relationship between a distributed storage system and its clients may be governed by a service level agreement (SLA). An SLA typically includes performance goals that the provider of the distributed storage system agrees to meet in servicing client requests. For example, the SLA can indicate that the provider of the distributed storage system ensures that the latency does not exceed 10 ms in processing requests. In some cases, the SLA can include actions to be taken when the performance goals are not met, such as the provider refunding the client. Such an agreement can also include the stipulation that the distributed storage system provider is not responsible for performance issues caused by circumstances outside its control (e.g., public network outages, client network outages, client device problems).

[0015] Even in the absence of an SLA, the provider of a distributed storage system still desires to monitor the health of the system in a non-invasive manner that does not affect the performance of the system when servicing client requests.

[0016] Accordingly, the present disclosure describes techniques for monitoring performance in a distributed database system by profiling client requests. An exemplary method includes identifying requests sent by a client to a distributed storage system. The requests can include request parameters such as a request type, concurrency parameters, and a request target indicating the data involved in the request. In some cases, the requests can be identified based on information sent from the client to the distributed storage system. Based on the identified requests, probe requests are generated, the probe requests including probe request parameters representing statistical samples of the request parameters included in the identified requests. The generated probe requests are sent to the distributed storage system, which responds by performing preparatory work to service each probe request but does not actually access the data indicated by the request target. For example, in response to a probe request with a request type of "read", the distributed storage system can prepare to read the data indicated by the request target (e.g., locate the data, queue the request, etc.), but can read fields associated with data that is not accessible to the client. This allows the system to be profiled in a comprehensive manner without interfering with the processing of regular client requests. When a response to a probe request is received from the distributed storage system, performance metrics can be calculated.

[0017] In some cases, the process can be performed by a computing device co-located with the distributed storage system (e.g., on the same internal network) in order to measure the performance of the distributed storage system alone, such that the process does not indicate issues outside of the provider's control (e.g., public network outages).

[0018] The techniques described herein can provide the following advantages. By profiling actual client requests to generate probe requests, the techniques allow monitoring of a distributed storage system using an approximation of the client requests currently being served by the distributed storage system. Additionally, by monitoring the health of the distributed storage system without accessing data accessible to the client, the performance impact of such monitoring on actual client requests can be minimized. Further, by monitoring the distributed storage system in isolation from other factors outside the control of the system provider, a more accurate view of the true performance of the system can be obtained, where the true performance includes the end-to-end performance, such as latency, experienced by the client when requesting service from the distributed storage system. This information can be useful in determining compliance with an SLA. The techniques can also provide statistical information related to the distribution of the client requests themselves (e.g., whether a particular data item is more popular than other data items), which can enable the provider and the client to understand whether they should adjust their workloads (e.g., avoid hot spots). Additionally, the techniques allow profiling of different aspects of the distributed storage system such as queue times, actual processing times, server location times, and the current CPU and memory utilization, queue lengths, and other metrics of the servers. Further, the techniques may allow easier management of SLA / SLO compliance since the profiling is entirely under the control of the provider of the distributed storage system.

[0019] Figure 1 is a schematic diagram of an exemplary environment for monitoring performance in a distributed database system. As shown, environment 100 includes a distributed storage system 110 that includes a plurality of servers 112, each of which manages a plurality of data groups 114. In operation, a client 120 sends a request 122 to the distributed storage system 110. The distributed storage system 110 processes the request 122 and sends a response 124 to the client 120. The client 120 sends request information 126 to a probe 130. The request information 126 includes information related to the requests sent by a particular client 120 to the distributed storage system 110, such as the request type, request parameters, and request target for each request. The probe 130 receives the request information 126 and sends a probe request 132 to the distributed storage system 110 based on the request information 126. The distributed storage system 110 processes the probe request and returns a response 134 to the program 130. The probe 130 analyzes the response and outputs a performance metric 140 indicating the current performance of the distributed storage system 110, a particular server 112 within the distributed storage system 110, a particular data group 114 within the distributed storage system 110, or other components within the distributed storage system 110.

[0020] The distributed storage system 110 can be a distributed system that includes multiple servers 112 connected by a local or private network (not shown). In some cases, the local or private network can be entirely within a single facility, while in other cases, the local or private network can cover a large area and interconnect multiple facilities. The servers 112 can communicate with each other to service client requests 122 by storing, retrieving, and updating data requested by the clients 120. In some cases, the distributed storage system 110 can be a distributed database, a distributed file system, or other types of distributed storage. The distributed storage system 110 can also include components for managing and organizing the operation of the servers 112 within the system.

[0021] Within the distributed storage system 110, each server 112 can be a computing device that includes a processor and a storage device, such as a hard disk drive, for storing data managed by the distributed storage system 110. In some cases, data can be distributed to different servers 112 according to a distribution policy. For example, the distribution policy can specify that a particular table or file within the distributed storage system 110 must be stored on a particular number of servers 112 or in order to maintain redundancy. The distribution policy can also specify that data must be stored in multiple different locations to maintain geographical redundancy. In some cases, the servers 112 can use an external storage device or system, such as a distributed file system, instead of a directly connected permanent storage.

[0022] Each server 112 manages one or more data groups 114. The data groups 114 can include portions of the total data set managed by the distributed storage system 110. Each data group 114 can represent data that includes a portion of a table from a distributed database, one or more files from a distributed file system, or other partitions of data within the distributed storage system 110. In operation, the distributed storage system 110 can analyze each request 122 and probe request 132 to determine the specific data group 114 involved in the request or probe request based on the request target. Thereafter, the distributed storage system can route the request or probe request to the specific server 112 that manages the specific data group 114.

[0023] In some cases, the client 120 can be a user of the distributed storage system 110. The client 120 can also be an entity (such as a website or an application) that uses the distributed storage system 110 to store and retrieve data. Each client 120 can record information related to each request 122 it sends to the distributed storage system 110. In some cases, each client 120 can store a record of the entire request sent to the distributed storage system 110. Each client can also store a summary of the requests 122 sent to the distributed storage system, such as, for example, to store a count of requests with the same set of request parameters that have been sent. For example, the client 120 can record the fact that 5 requests of request type "read", concurrency parameter "stale", and request target a table named "customer" have been sent. Each client 120 can send the request information 126 to the probe 130, for example, at fixed intervals. The client 120 can send the request information 126 to the probe 130 via a public network such as the Internet. In some cases, the request information 126 can be collected by a software process or library running on the client that collects information related to the requests sent by the client. In some cases, the software library can be provided by the provider of the distributed storage system.

[0024] The probe 130 can analyze the request information received from the client 120 and generate a probe profile representing a statistical approximation of the requests described by the request information 126. For example, the probe 130 can analyze request information 126 that includes 10,000 requests of type "read" and 5000 requests of type "write", and generate a probe profile that indicates that 1000 probe requests of type "read" and 500 probe requests of type "write" should be sent in order to simulate and determine the performance of the distributed system in processing the original requests 122. In some cases, the probe 130 can select the number of probe requests 132 to generate such that the number is large enough to represent the requests 122 sent by the client 120, but small enough to minimize the impact on the performance of the distributed storage system 110.

[0025] Based on the probe profile, the probe 130 sends a probe request 132 to the distributed storage system 110. In some cases, the format of the probe request 132 can be the same as the format of the request 122 sent by the client 120, but can include an indication that they are probe requests rather than requests from the client 120. The distributed storage system 110 can receive the probe requests 132 and process them in the same way as it processes client requests 122, except that the distributed storage system cannot access the actual data indicated by the request target of each probe request. In some cases, the distributed storage system 110 can alternatively access data specifically allocated to the probe 130, such as special fields, columns, metadata values, or other values associated with the data indicated by the request target. In this way, the performance of several aspects of the distributed storage system 110 can be profiled without interfering with the processing of client requests 122. For example, a probe request including a concurrency parameter that would cause the requested data to be locked may instead cause the distributed storage system 110 to lock probe-specific data so as not to interfere with the processing of client requests. Such a feature allows the concurrency characteristics of the distributed storage system 110 to be profiled without affecting the processing of client requests.

[0026] As shown, the probe 130 generates one or more performance metrics 140 based on the probe requests 132 sent and the responses 134 received. In some cases, the performance metrics can include an overall system latency measured by the average amount of time taken by the distributed storage system 110 to respond to the probe requests. The performance metrics can also include availability as measured by the ratio of failed probe requests to successful probe requests. The performance metrics can also include local network latency, server queue latency (e.g., the average amount of time each probe request 132 waits on the server before being processed), disk or memory latency, or other performance metrics.

[0027] Figure 2It is a swimlane diagram of an exemplary process for processing client requests in a distributed storage system. At 205, client 120 sends a request to distributed storage system 110. At 210, distributed storage system 110 prepares to serve the request. For example, the preparations made by distributed storage system 110 may include: parsing the request; determining a specific server that stores the requested data; sending the request to the specific server; generating an execution plan for the specific request (e.g., determining which tables to access and how to manipulate the data to complete the request); or other operations. At 215, distributed storage system 110 accesses the data indicated by each request. This is in contrast to the processing of probe requests by distributed storage system 110, in which the data indicated by the request cannot be accessed (as described below). At 220, distributed storage system 110 sends a response to the request to the client. For example, if the client has requested to read specific data from distributed storage system 110, the response may include the requested data.

[0028] Figure 3 It is a swimlane diagram of an exemplary process for monitoring performance in a distributed database system. At 305, client 120 sends a request to distributed storage system 110, and distributed storage system 110 responds to the request at 310. At 315, client 120 provides information related to the sent request to probe 130. At 320, probe 130 generates a probe profile based on the information related to the request sent by the client. In some cases, the probe may receive information related to the requests sent by multiple clients and generate a probe profile based on this information. At 325, probe 130 sends a probe request to distributed storage system 110 based on the probe profile. At 330, distributed storage system 110 prepares to serve each probe request. At 335, distributed storage system 110 accesses probe-specific metadata associated with the data indicated by the request target of each probe request. At 340, distributed storage system 110 sends a response to the probe request to probe 130. At 345, the probe calculates a performance metric based on the response to the probe request.

[0029] Figure 4 It is a flowchart of an exemplary process for monitoring performance in a distributed database system. At 405, requests sent by clients to a distributed storage system are identified, and each request includes request parameters. In some cases, the request parameters include a request type, a concurrency parameter, and a request target indicating the data in the distributed storage system involved in the request.

[0030] At 410, a probe request is generated based on the recognized request. The probe request includes probe request parameters representing a statistical sample of the request parameters included in the recognized request. In some cases, generating the probe request includes: generating a number of probe requests that is less than the number of recognized requests. Generating the probe request may include generating a number of probe requests that is proportional to the number of recognized requests, the probe request including a specific request type, specific concurrency parameters, and a specific request target, and the recognized request including the specific request type, the specific concurrency parameters, and the specific request target.

[0031] At 415, the generated probe request is sent via a network to the distributed storage system. The distributed storage system is configured to perform preparatory work in response to receiving the probe request to serve each probe request. In some cases, the distributed storage system is configured not to read or write any client-accessible data when performing preparatory work in response to receiving the probe request to serve each probe request. In some implementations, the data in the distributed storage system includes probe fields that are not client-accessible, and wherein the distributed storage system is configured to access the probe fields associated with the request target in the probe request in response to receiving the probe request.

[0032] At 420, a response to the probe request is received from the distributed storage system. At 425, at least one performance metric of the distributed storage system is output based on the received response. In some cases, outputting the at least one performance metric includes: outputting a weighted average of at least one performance metric for a specific data set of the distributed storage system based on the response to the probe request, the probe request including a request target identifying data in the specific data set. The at least one performance metric may include at least one of the following: availability, disk latency, queuing latency, request preparation latency, or internal network latency. In some cases, the performance metric may include a weighted average, where the weights used in the calculation are derived from the request information 126. For example, if the number of requests with stale concurrency is 10 times the number of requests with strong concurrency, then the performance data of the stale concurrency probe requests may be 10 times the performance data of the strong concurrency probes in subsequent performance metrics.

[0033] In some cases, process 400 includes comparing the at least one performance metric with a service level objective (SLO) including target values for at least one performance metric of the distributed storage system. The SLO may be included in a service level agreement (SLA) for the distributed storage system. Process 400 may also include determining that the at least one performance metric does not meet the target value and outputting an indication that the at least one performance metric does not meet the target value.

[0034] Figure 5FIG. 0 is a block diagram of computing devices 500, 550 that can be used as a client or as a server or multiple servers to implement the systems and methods described in this document. Computing device 500 is intended to represent various forms of digital computers such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 550 is intended to represent various forms of mobile devices such as personal digital assistants, cellular phones, smart phones, and other similar computing devices. Additionally, computing device 500 or 550 can include a Universal Serial Bus (USB) flash drive. The USB flash drive can store an operating system and other applications. The USB flash drive can include input / output components such as a wireless transmitter or a USB connector that can be inserted into a USB port of another computing device. The components shown here, their connections and relationships, and their functions are only illustrative and are not intended to limit the implementation of the invention described and / or claimed in this document.

[0035] Computing device 500 includes a processor 502, a memory 504, a storage device 506, a high-speed interface 508 connected to the memory 504 and a high-speed expansion port 510, and a low-speed interface 512 connected to a low-speed bus 514 and the storage device 506. Each of the components 502, 504, 506, 508, 510, 512 is interconnected using various buses and can be mounted on a common motherboard or otherwise as appropriate. Processor 502 can process instructions for execution within computing device 500, the instructions including instructions for graphical information stored in memory 504 or on storage device 506 to be displayed on an external input / output device, the external input / output device such as a display 516 coupled to the high-speed interface 508. In other implementations, multiple processors and / or multiple buses can be used in conjunction with multiple memories and memory types as appropriate. Additionally, multiple computing devices 400 can be connected to each device for providing portions of the necessary operations (e.g., as a server cluster, a blade server group, or a multi-processor system).

[0036] Memory 504 stores information within computing device 500. In one implementation, memory 504 is a volatile memory unit. In another implementation, memory 504 is a non-volatile memory unit. Memory 504 can also be other forms of computer-readable media such as a magnetic disk or an optical disk.

[0037] The storage device 506 can provide mass storage for the computing device 500. In one implementation, the storage device 506 can be a computer-readable medium such as a floppy disk device, a hard disk device, an optical disk device, or a magnetic tape device, a flash memory or other similar solid-state memory device, or an array of devices including a storage area network or other configured devices. The computer program product can be tangibly embodied in an information carrier. The computer program product can also contain instructions that, when executed, perform one or more methods such as those described above. The information carrier is a computer-readable medium or a machine-readable medium such as the memory 504, the storage device 506, or the memory on the processor 502.

[0038] The high-speed controller 508 manages the bandwidth-intensive operations of the computing device 500, while the low-speed controller 512 manages the low-bandwidth-intensive operations. This assignment of tasks is merely exemplary. In one implementation, the high-speed controller 508 is coupled to the memory 504, the display 516 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 510 that can accept various expansion cards (not shown). In this implementation, the low-speed controller 512 is coupled to the storage device 506 and a low-speed expansion port 514. The low-speed expansion port, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices such as a keyboard, a pointing device, a scanner or a networking device such as a switch or a router, for example, via a network adapter.

[0039] As shown, the computing device 500 can be implemented in a variety of different forms. For example, it can be implemented as a standard server 520 or implemented multiple times in a group of such servers. It can also be implemented as part of a rack server system 524. Additionally, it can be implemented in a personal computer such as a laptop computer 522. Alternatively, components from the computing device 500 can be combined with other components (not shown) in a mobile device such as the device 550. Each of such devices can include one or more computing devices 500, 550, and the entire system can be composed of multiple computing devices 500, 550 that communicate with each other.

[0040] The computing device 550 includes a processor 552, a memory 564, an input / output device such as a display 554, a communication interface 566, a transceiver 568, and other components. The device 550 can also have a storage device such as a microdrive or other devices to provide auxiliary storage. Each of the components 550, 552, 564, 554, 566, 568 is interconnected using various buses, and several components can be mounted on a common motherboard or otherwise as appropriate.

[0041] The processor 552 may execute instructions within the computing device 550, which include instructions stored in the memory 564. The processor may be implemented as a chipset including separate multiple analog and digital processors on a chip. Additionally, the processor may be implemented using any of a number of architectures. For example, the processor 510 may be a CISC (Complex Instruction Set Computer) processor, a RISC (Reduced Instruction Set Computer) processor, or a MISC (Minimal Instruction Set Computer) processor. The processor may provide, for example, cooperation with other components of the device 550, such as control of a user interface, applications running on the device 550, and wireless communications performed via the device 550.

[0042] The processor 452 may communicate with a user via a control interface 558 and a display interface 556 coupled to a display 554. The display 554 may be, for example, a TFT display (Thin Film Transistor Liquid Crystal Display) or an OLED (Organic Light Emitting Diode) display, or other suitable display technology. The display interface 556 may include appropriate circuitry for driving the display 554 to present graphics and other information to the user. The control interface 558 may receive commands from the user and convert them for submission to the processor 552. Additionally, an external interface 562 may be provided for communicating with the processor 552 so that the device 550 can perform near field communication with other devices. In some implementations, the external interface 562 may provide, for example, wired communication, or in some implementations, may provide wireless communication, and may also use multiple interfaces.

[0043] The memory 564 stores information within the computing device 550. The memory 564 may be implemented as one or more of a computer-readable medium or media, volatile memory units, or non-volatile memory units. An extended memory 574 may also be provided and the extended memory 574 may be connected to the device 550 via an extended interface 572 including, for example, a SIMM (Single In-line Memory Module) card interface. Such extended memory 574 may provide additional storage space for the device 550, or may also store applications or other information of the device 550. Specifically, the extended memory 574 may include instructions for performing or supplementing the above processing, and may also include security information. Thus, for example, the extended memory 574 may be provided as a security module for the device 550 and may be programmed with instructions that allow the device 550 to be used securely. Additionally, security applications may be provided via the SIMM card together with additional information, such as placing identification information on the SIMM card in a non-crackable manner.

[0044] The memory may include, for example, a flash memory and / or an MRAM memory as described below. In one implementation, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods such as those described above. The information carrier is a computer or machine-readable medium that can be received, for example, by transceiver 568 or external interface 562, such as memory 464, extended memory 474, or the memory on processor 452.

[0045] Device 550 can communicate wirelessly through a communication interface 566, which may include digital signal processing circuitry where necessary. Communication interface 566 can provide communication under various modes or protocols such as GSM voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS. Such communication can occur, for example, through radio frequency transceiver 568. Additionally, short-range communication can occur, for example, using Bluetooth, WiFi, or such other transceivers (not shown). Additionally, a GPS (Global Positioning System) receiver module 570 can provide additional wireless data related to navigation and positioning to device 550, which can be used by applications running on device 550 as appropriate.

[0046] Device 550 can also communicate audibly using an audio codec 560, which can receive the spoken information from the user and convert it into usable digital information. Audio codec 560 can similarly generate audible sounds for the user, for example, through a speaker in the handset of device 550. Such sounds can include sounds from a voice call, can include recorded sounds (e.g., voice messages, music files, etc.), and can also include sounds generated by applications operating on device 550.

[0047] As shown, computing device 550 can be implemented in many different forms. For example, it can be implemented as a cellular phone 580. It can also be implemented as part of a smart phone 582, a personal digital assistant, or other similar mobile devices.

[0048] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0049] These computer programs (also referred to as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in high-level procedural and / or object-oriented programming languages, and / or in assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0050] For providing interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0051] The systems and descriptions herein can be implemented on the following computing systems, which include backend components (such as data servers), or middleware components (such as application servers), or frontend components (such as client computers having a graphical user interface or a web browser through which users can interact with the implementations of the systems and techniques herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (such as a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), peer-to-peer networks (with self-organizing or static members), grid computing infrastructures, and the Internet.

[0052] The computing system can include clients and servers. The clients and servers are typically located far from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs running on the respective computers and having a client-server relationship with each other.

[0053] Further implementations are summarized in the following examples:

[0054] Example 1: A computer-implemented method for measuring performance metrics in a distributed storage system, executed by one or more processors, the method including: identifying requests sent by a client to the distributed storage system, each request including a request parameter value of a request parameter; based on the identified requests, generating probe requests, the probe requests including probe request parameter values representing statistical samples of the request parameters included in the identified requests; sending the generated probe requests to the distributed storage system through a network, wherein the distributed storage system is configured to perform preparatory work in response to receiving the probe requests to serve each probe request; receiving responses to the probe requests from the distributed storage system; and outputting at least one performance metric value measuring the current performance state of the distributed storage system based on the received responses.

[0055] Example 2: The method according to Example 1, wherein generating the probe requests includes generating a number of probe requests that is less than the number of the identified requests.

[0056] Example 3: The method according to Example 1 or 2, wherein the request parameters include request type, concurrency parameter, and a request target indicating data within the distributed storage system involved in the request.

[0057] Example 4: The method according to Example 3, wherein generating the probe requests includes generating a number of probe requests proportional to the number of the identified requests, the probe requests including a specific request type, specific concurrency parameters, and a specific request target, and the identified requests including the specific request type, the specific concurrency parameters, and the specific request target.

[0058] Example 5: The method according to any one of Examples 1 to 4, wherein outputting the at least one performance metric includes outputting a weighted average of at least one performance metric for a specific data set of the distributed storage system based on responses to the probe requests, the probe requests including a request target for identifying data in the specific data set.

[0059] Example 6: The method according to any one of Examples 1 to 5, wherein the at least one performance metric includes at least one of the following: availability, disk latency, queuing latency, request preparation latency, or internal network latency.

[0060] Example 7: The method according to any one of Examples 1 to 6, wherein the distributed storage system is configured not to read or write any data accessible to the client when performing preparatory work in response to receiving the probe requests to service each probe request.

[0061] Example 8: The method according to any one of Examples 1 to 7, wherein the data in the distributed storage system includes probe fields inaccessible to the client, and wherein the distributed storage system is configured to access the probe fields associated with the request target in the probe requests in response to receiving the probe requests.

[0062] Example 9: The method according to any one of Examples 1 to 8, further comprising comparing the at least one performance metric with a service level objective (SLO) including target values for the at least one performance metric for the distributed storage system.

[0063] Example 10: The method according to Example 9, further comprising: determining that the at least one performance metric does not meet the target value; and outputting an indication that the at least one performance metric does not meet the target value.

[0064] Example 11: The method according to Example 9 or 10, wherein the service level objective is included in a service level agreement (SLA) of the distributed storage system.

[0065] Example 12: The method according to any one of Examples 1 to 11, wherein the probe requests are a statistical representation of the identified requests.

[0066] Example 13: A non-transitory computer-readable medium for storing instructions that, when executed, are operable to cause at least one processor to perform operations for measuring performance metrics in a distributed storage system, the operations including: identifying requests sent by a client to the distributed storage system, each request including request parameters; generating, based on the identified requests, probe requests, the probe requests including probe request parameters representing a statistical sample of the request parameters included in the identified requests; sending the generated probe requests to the distributed storage system via a network, wherein the distributed storage system is configured to perform preparatory work to serve each probe request in response to receiving the probe requests; receiving responses to the probe requests from the distributed storage system; and outputting at least one performance metric of the distributed storage system based on the received responses.

[0067] Example 14: The computer-readable medium according to Example 13, wherein generating the probe requests includes generating a number of probe requests that is less than the number of the identified requests.

[0068] Example 15: The computer-readable medium according to Example 13 or 14, wherein the request parameters include a request type, a concurrency parameter, and a request target indicating data within the distributed storage system involved in the request.

[0069] Example 16: The computer-readable medium according to Example 15, wherein generating the probe requests includes generating a number of probe requests that is proportional to the number of the identified requests, the probe requests including a specific request type, a specific concurrency parameter, and a specific request target, and the identified requests include the specific request type, the specific concurrency parameter, and the specific request target.

[0070] Example 17: The computer-readable medium according to one of Examples 13 to 16, wherein outputting the at least one performance metric includes outputting a weighted average of at least one performance metric of a specific data set of the distributed storage system based on responses to probe requests that include a request target identifying data in the specific data set.

[0071] Example 18: The computer-readable medium according to one of Examples 13 to 17, wherein the at least one performance metric includes at least one of the following: availability, disk latency, queuing latency, request preparation latency, or internal network latency.

[0072] Example 19: A computer-readable medium according to one of Examples 13 to 18, wherein the distributed storage system is configured to not read or write any data accessible by the client when performing preparation work to service each probe request in response to receiving the probe request.

[0073] Example 20: A system for measuring performance metrics in a distributed storage system, the system comprising: a memory for storing data; and one or more processors operable to access the memory and perform the following operations, the operations including: identifying requests sent by a client to the distributed storage system, each request including request parameters; generating a probe request based on the identified requests, the probe request including probe request parameters representing a statistical sample of the request parameters included in the identified requests; sending the generated probe request to the distributed storage system via a network, wherein the distributed storage system is configured to perform preparation work to service each probe request in response to receiving the probe request; receiving a response to the probe request from the distributed storage system; and outputting at least one performance metric of the distributed storage system based on the received response.

[0074] Although several implementations have been described above, other modifications are possible. Additionally, the logical flow depicted in the figures does not require the particular order or sequential order shown to achieve the desired result. Other steps may be provided or some steps may be excluded from the flow, and other components may be added to the system or other components may be removed from the system. Accordingly, other implementations are within the scope of the appended claims.

Claims

1. A method for monitoring performance in a distributed storage system, comprising: At data processing hardware, receiving a client request to access a target resource of the distributed storage system, the distributed storage system including a plurality of data groups, the target resource indicating a corresponding data group among the plurality of data groups; In response to receiving the client request, generating, by the data processing hardware, a health probe request including a probe profile, the health probe request being configured to identify the availability of the target resource of the distributed storage system based on the probe profile and probe-specific data groups associated with the indicated corresponding data group, the probe-specific data groups being independent of the indicated corresponding data group; Passing, by the data processing hardware, the health probe request to the distributed storage system; Receiving, at the data processing hardware, a response to the health probe request, the response being generated without the distributed storage system performing any read or write operations on the target resource; And Generating, by the data processing hardware, a health performance metric based on the response according to the health probe request, the health performance metric identifying the availability of the target resource.

2. The method according to claim 1, wherein, The passing of the health probe request to the distributed storage system occurs while another client request requests access to the target resource of the distributed storage system.

3. The method according to claim 1, wherein, Generating the health performance metric based on the response according to the health probe request includes determining whether the health probe request fails or succeeds in identifying the availability of the target resource based on the probe-specific data groups.

4. The method according to claim 1, wherein, Generating the health performance metric based on the response according to the health probe request includes: Determining whether the health probe request fails or succeeds in identifying the availability of the target resource based on the probe-specific data groups; and Expressing the health performance metric identifying the availability of the target resource by a ratio of health probe request failures to health probe request successes.

5. The method according to claim 1, further comprising comparing, by the data processing hardware, the health performance metric with a service level objective SLO, the service level objective SLO including a target value for the health performance metric for the availability of the target resource.

6. The method according to claim 5, further comprising: Determining, by the data processing hardware, that the health performance metric does not meet the target value; And Passing, by the data processing hardware, an indication that the health performance metric does not meet the target value.

7. The method according to claim 5, wherein The service level agreement SLA of the distributed storage system includes the service level objective.

8. The method according to claim 1, wherein The data processing hardware is co-located with the distributed storage system.

9. The method according to claim 1, wherein The passing of the health probe request to the distributed storage system occurs at fixed intervals.

10. A system for monitoring performance in a distributed storage system, comprising: Data processing hardware; And Memory hardware that communicates with the data processing hardware and stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations, the operations including: Receiving a client request to access a target resource of a distributed storage system, the distributed storage system including a plurality of data groups, the target resource indicating a corresponding data group among the plurality of data groups; In response to receiving the client request, generating a health probe request including a probe profile, the health probe request being configured to identify the availability of the target resource of the distributed storage system based on the probe profile and probe-specific data groups associated with the indicated corresponding data group, the probe-specific data groups being independent of the indicated corresponding data group; Passing the health probe request to the distributed storage system; Receiving a response to the health probe request, the response being generated without the distributed storage system performing any read or write operations on the target resource; and Generating a health performance metric based on the response to the health probe request, the health performance metric identifying the availability of the target resource.

11. The system according to claim 10, wherein, The passing of the health probe request to the distributed storage system occurs while another client request requests access to the target resource of the distributed storage system.

12. The system according to claim 10, wherein Generating the health performance metric based on the response to the health probe request includes determining whether the health probe request fails or succeeds in identifying the availability of the target resource based on the probe-specific data groups.

13. The system according to claim 10, wherein, Generating the health performance metric based on the response to the health probe request includes: Determining whether the health probe request fails or succeeds in identifying the availability of the target resource based on the probe-specific data groups; and Representing the health performance metric identifying the availability of the target resource by a ratio of health probe request failures to health probe request successes.

14. The system according to claim 10, wherein, The operations further include comparing the health performance metric with a service level objective SLO, the service level objective SLO including a target value for the health performance metric for the availability of the target resource.

15. The system according to claim 14, wherein, The operations further include: Determining that the health performance metric does not meet the target value; and Transmitting an indication that the health performance metric does not meet the target value.

16. The system according to claim 14, wherein A service level agreement SLA of the distributed storage system includes the service level objective.

17. The system according to claim 10, wherein, The data processing hardware is co-located with the distributed storage system.

18. The system according to claim 10, wherein, The passing of the health probe request to the distributed storage system occurs at fixed intervals.

Citation Information

Patent Citations

  • Load test simulator

    US20050216234A1