Hyper-fusion system IO time delay observation method and device based on IO dyeing technology

By introducing IO coloring technology into the hyper-converged system, IO latency observation and tuning are achieved throughout the entire life cycle, solving the problem of inaccurate tracking caused by IO splitting and merging in traditional methods, and improving the accuracy and convenience of IO stability and performance display.

CN120743375APending Publication Date: 2025-10-03JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510854181.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

While existing hyper-converged systems improve IO performance, they cannot effectively guarantee IO quality, especially IO stability and latency stability. They lack systematic observation methods and tools for the entire life cycle, and traditional observation methods are prone to inaccurate tracking due to splitting and merging in the IO process.

Method used

A hyper-converged system IO latency observation method based on IO coloring technology is adopted. By adding specific tag information to IO, the IO operation process is tracked and analyzed within the system, including UI configuration, IO coloring, latency capture and analysis, to achieve latency observation and tuning throughout the entire life cycle.

Benefits of technology

It provides a system-level visualization of IO runtime latency status, improves the performance and competitiveness of hyper-converged products, ensures business stability, and maintains tracking accuracy and integrity during IO splitting and merging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743375A_ABST
    Figure CN120743375A_ABST
Patent Text Reader

Abstract

The invention relates to a hyper-fusion system IO time delay observation method and device based on an IO dyeing technology, and aims to solve the problems of inaccurate IO time delay observation and low efficiency in the prior art. The method comprises the steps that a user starts IO time delay observation and configures parameters through a UI configuration module; the IO dyeing module marks a unique mark on IO sent by the selected VM; the IO time delay capturing module captures IO carrying dyeing information at a key point of an IO life cycle and records key information; and the IO time delay analysis module obtains the total time delay of the IO by analyzing the IO track information generated by the IO capture module and backtracking the full life cycle of the IO, and decides whether to further backtrack the interlayer time delay and the intra-layer time delay of the IO full path or not according to the comparison between the total time delay and a configured time delay threshold value, and outputs the interlayer time delay and the intra-layer time delay to the UI interface. According to the method, the IO operation of the specific VM is accurately tracked through the IO dyeing technology, the key information of the IO operation is comprehensively obtained, the performance bottleneck is quickly positioned, and the observation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of virtualization technology and hyper-convergence system performance monitoring, and in particular to a method and device for observing IO latency of a hyper-convergence system based on IO staining technology. Background Art

[0002] In the cloud computing era, new technologies are constantly emerging, but hyperconvergence is gaining a growing market share thanks to its advantages, including high resource utilization, strong scalability, significant cost-effectiveness, simple deployment, and simplified management. As new technologies develop and mature, the performance of hyperconverged systems is also improving. However, as hyperconverged systems are increasingly applied in key sectors such as finance and healthcare, I / O jitter can cause significant fluctuations or even interruptions in critical business operations such as trading systems and databases. Simply improving I / O performance without guaranteeing I / O quality fails to meet the requirements of these businesses. Consequently, these key sectors have placed new demands on hyperconverged system performance: how to improve I / O performance while ensuring I / O quality. I / O quality refers to I / O stability, including both overall I / O per second (IOPS) stability and individual I / O latency stability. IOPS and stability are key competitive areas among hyperconverged vendors. Through continuous optimization and improvement, leading hyperconverged vendors have been able to meet the requirements of various businesses. However, these vendors have paid less attention to I / O quality and have focused solely on the hyperconverged software itself. Faced with increasingly stringent market requirements for hyper-convergence, we must focus on and address IO stability. However, we lack effective and convenient means and tools to identify high-latency IO among tens of thousands of IOs per second and accurately assess the time it takes to transfer between the various software and hardware layers in the hyper-convergence and virtualization systems. This places new demands on our hyper-convergence systems: how to efficiently measure IO latency while ensuring system IOPS. Furthermore, as market competition becomes increasingly competitive and technological development enters a new era, we must not only focus on the latency of the hyper-convergence system itself, but also address the quality risks caused by IO latency throughout the entire IO lifecycle, from virtual machines to kernels to hyper-convergence systems to physical disks.

[0003] Currently, most hyper-converged systems focus their optimization and measurement of IO latency solely on the hyper-converged system itself. However, the hyper-converged system itself is only a small component of IO latency throughout its lifecycle. Optimizing IO stability, or latency, should be systematic, as problems at any stage in the IO lifecycle can lead to significant fluctuations in IO latency. With the intensifying competition in the hyper-converged world and increasingly stringent business requirements, how to comprehensively, accurately, and meticulously measure IO latency throughout its lifecycle—and provide powerful tools and methods for optimizing IO stability and performance—is a challenge we must consider and confront. At the same time, business continuity requires that IO latency measurement not impact business operations. Summary of the Invention

[0004] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.

[0005] The present invention proposes a method for observing IO latency of a hyper-converged system based on IO coloring technology, which is applied to system-level IO latency observation and tuning scenarios in a hyper-converged system. By dynamically observing IO latency in the hyper-converged system, a visual display of the system-level IO runtime latency status is provided for the hyper-converged cluster, thereby improving the performance and competitiveness of the hyper-converged product.

[0006] Another object of the present invention is to propose an IO latency observation device for a hyper-converged system based on IO staining technology.

[0007] To achieve the above objectives, the present invention proposes a method for observing IO latency in a hyper-converged system based on IO staining technology, comprising:

[0008] The user enables IO latency monitoring and configures parameters through the UI configuration module, including the VM to be monitored, the latency threshold, and the observation duration or scheduled observation time.

[0009] The IO coloring module uniquely tags the IOs issued by the selected VM to complete the coloring;

[0010] The IO latency capture module captures IOs carrying coloring information at key points in the IO lifecycle and records the IO timestamp, coloring information, and other key information.

[0011] The IO latency analysis module analyzes the IO trace information generated by the IO capture module and traces back the entire IO lifecycle based on the coloring information of the completed IO to obtain the total IO latency. Based on the comparison of the total latency with the configured latency threshold, the module determines whether to further trace back the inter-layer and intra-layer latency of the entire IO path and output the results to the UI.

[0012] The method for observing IO latency in a hyper-converged system based on IO staining technology according to an embodiment of the present invention may also have the following additional technical features:

[0013] In one embodiment of the present invention, the observation duration or scheduled observation time supports the use of a default value.

[0014] In one embodiment of the present invention, IOs that do not carry a staining mark are not observed, and the staining information carried by each IO is different.

[0015] In one embodiment of the present invention, the key points are distributed at the entrances and exits of IO processing in the virtualization layer, file system layer, block device layer, kernel layer, hyper-converged system layer, and other core processing points.

[0016] In one embodiment of the present invention, when the total IO latency is less than the configured latency threshold, the IO layer and inter-layer latency are not obtained. When the total IO latency is greater than or equal to the configured latency threshold, the IO is backtracked to obtain the inter-layer and intra-layer latency of the entire IO path, sort them by latency, and finally output them to the UI for display.

[0017] In one embodiment of the present invention, only IO records within 15 seconds before the end of the IO are traced back.

[0018] In one embodiment of the present invention, after the observation time ends, IO coloring stops, and IO tracking stops. The IO delay analysis module clears IO records that do not meet the delay threshold during this tracking process.

[0019] To achieve the above-mentioned purpose, the present invention further proposes an IO latency observation device for a hyper-converged system based on IO staining technology, comprising:

[0020] The UI configuration module is used by users to enable IO latency observation and configure parameters, including the VM to be observed, the latency threshold, and the observation duration or scheduled observation time;

[0021] The IO coloring module is used to uniquely mark the IOs issued by the selected VM to complete the coloring;

[0022] The IO latency capture module is used to capture IOs carrying coloring information at key points in the IO lifecycle and record the IO timestamp, coloring information, and other key information.

[0023] The IO latency analysis module parses the IO trace information generated by the IO capture module, traces back the entire IO lifecycle based on the coloring information of the completed IO, obtains the total IO latency, and compares the total latency with the configured latency threshold to determine whether to further trace back the inter-layer and intra-layer latency of the entire IO path and output them to the UI.

[0024] The present invention provides a method and device for observing IO latency in a hyperconverged system based on IO coloring technology. While ensuring the stability and availability of hyperconverged user services, IO coloring—that is, adding specific tag information to IOs—enables internal tracking and analysis of IO operation flows. This technology addresses the issue of traditional IO latency observation methods, which often fail to track or provide inaccurate tracking after splitting or merging IO flows.

[0025] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0027] Figure 1 This is a flowchart of a method for observing IO latency in a hyper-converged system based on IO staining technology according to an embodiment of the present invention;

[0028] Figure 2 is an IO path diagram according to an embodiment of the present invention;

[0029] Figure 3 This is a logic diagram of a method for observing IO latency in a hyper-converged system based on IO staining technology according to an embodiment of the present invention;

[0030] Figure 4 This is a structural diagram of an IO latency observation device for a hyper-converged system based on IO staining technology according to an embodiment of the present invention. DETAILED DESCRIPTION

[0031] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0032] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0033] The following describes a method and apparatus for observing IO latency in a hyper-converged system based on IO staining technology according to an embodiment of the present invention with reference to the accompanying drawings.

[0034] First, the technical terms mentioned in this invention are introduced:

[0035] IO: input and output;

[0036] IOPS: the number of reads and writes per second;

[0037] VM: virtual machine;

[0038] eBPF: Extended Berkeley Packet Filter;

[0039] OCFS: shared disk file system.

[0040] Figure 1 FIG. 1 is a flow chart of a method for observing IO latency of a hyper-converged system based on IO staining technology according to an embodiment of the present invention. Figure 1 Shown, including:

[0041] S1: The user enables IO latency monitoring through the UI configuration module and configures parameters, including the VM to be monitored, the latency threshold, and the observation duration or scheduled observation time.

[0042] S2, the IO coloring module uniquely tags the IOs issued by the selected VM to complete the coloring;

[0043] S3, the IO latency capture module captures IOs carrying coloring information at key points in the IO lifecycle and records the IO timestamp, coloring information, and other key information;

[0044] In step S4, the IO latency analysis module analyzes the IO trace information generated by the IO capture module and traces back the entire IO lifecycle based on the coloring information of the completed IO to obtain the total IO latency. Based on the comparison of the total latency with the configured latency threshold, the module decides whether to further trace back the inter-layer and intra-layer latency of the entire IO path and outputs the results to the UI.

[0045] Specifically, the present invention mainly includes a UI interface, an IO coloring module, an IO delay capture module, and an IO delay analysis module. The UI interface provides users with an IO delay observation function configuration entry and result display. Users can select the VM that needs to be observed on the UI as needed. The IO coloring module will color the IO of the selected VM according to the configuration information issued by the UI interface: that is, by adding specific tag information to the IO, it can track and analyze the process of IO operations within the system. In order to improve usability and highlight key issues, users can configure the delay threshold from the UI, that is, only display IO that exceeds the threshold. If the user does not configure it, the system default threshold is used, and the default threshold is set to 0.5ms. The colored IO still carries the coloring information when it is split and merged on the IO path. The IO path is as shown in the attached figure. Figure 1As shown, the figure also shows the points where IO splitting and merging occur. Splitting and merging IO at any point will introduce technical challenges to the dynamic tracking of IO latency, resulting in uncertainty in the tracking results. We have solved the problem that IO cannot be dynamically tracked due to merging and splitting through IO dyeing technology. The IO latency capture module is responsible for capturing the IO latency of a specified VM. When capturing IO through traditional ebpf means, it is impossible to distinguish between the IO to be detected and the IO of other VMs. There are problems such as capturing too much IO, significant impact on system IO performance, and discarded capture events, which ultimately lead to inaccurate IO latency, IO latency cannot display the latency between layers, etc., and the ease of use is greatly reduced. After the introduction of IO dyeing, the IO latency capture module filters the dyed IO and only focuses on the IO issued by the specified VM. When the IO passes through the attached Figure 2 The key points of the path shown record timestamps, IO coloring information and other key information. Since the latency of most IOs cannot be known until the IO is completed to see whether it meets the threshold set by the user, the IO latency capture module will capture all IOs of the VM to which it belongs, but will release the IOs that do not meet the threshold conditions after the tracking is completed, in order to accelerate performance analysis and reduce the performance impact of system resources on user services. After the IO is captured by the latency capture module, a track information will be generated in the system. At the same time, the track information carries the timestamp of the IO in the flow. The IO latency parsing module parses the IO track information generated by the IO capture module to generate the total latency, inter-layer latency, and intra-layer latency of the IO, and sorts them according to the size of the latency, and finally outputs them to the UI interface for display. Through the close cooperation and processing of various modules, an intuitive and convenient way is provided for latency observation throughout the entire life cycle of IO.

[0046] Furthermore, the schematic diagram of the implementation process of the hyper-converged system IO delay observation method based on IO staining technology of the present invention is shown in the attached figure. Figure 3 The specific implementation process is as follows:

[0047] In S11, users enable IO latency observation and configure parameters through the UI configuration module. They mainly configure the VM to be observed and the latency threshold. They can also choose to observe the IO of all VMs. Considering the impact on user services, users can also set the observation duration or schedule observation at a specific time. The latency and observation time support the use of default values.

[0048] Specifically, users can enable the IO duration observation function through the graphical interface (UI) configuration module and set relevant parameters according to actual needs. During the configuration process, users can specify the virtual machines (VMs) that need to be observed, or they can choose to monitor the IO of all virtual machines. In addition, users can also set IO latency thresholds to define what degree of delay is considered abnormal. In order to improve ease of use, the system also supports quick configuration using default values. Taking into account that the observation operation may have a certain impact on the user's business, the system provides flexible time control options. Users can set specific observation durations as needed, or make appointments for observations during low-peak business periods, so as to minimize interference with normal business operations. This flexibility not only improves the user experience, but also helps to conduct performance analysis and troubleshooting more efficiently.

[0049] S12: After enabling IO latency observation, the IO coloring module will tag the IOs issued by the selected VM, completing the coloring process. It should be noted that only IOs carrying the coloring mark are those we need to observe; IOs without the coloring mark are not of interest. Furthermore, the coloring information carried by each IO is unique. This ensures that the present invention can: 1. When capturing IOs, only those carrying the coloring mark are focused on key services; 2. After the IO is completed, IO latency analysis can be performed using the IO coloring mark to trace the IO latency across the entire link; 3. Processing of IOs during the process, such as splitting and merging, will not result in loss of IO tracking.

[0050] Specifically, after turning on the IO latency observation function, the IO dyeing module in the system will mark the IO requests issued by the selected virtual machine (VM), which is called "IO dyeing". This process will attach specific identification information to the IO request when it is generated, so that these dyed IOs are traceable throughout the system. It should be noted that only IOs carrying dyeing marks will be included in the observation scope, and IOs that are not dyed will not be collected or analyzed by the system, thereby avoiding the interference of irrelevant data on performance analysis. The dyeing information carried by each IO is unique or distinguishable, and can correspond to a specific observation task or business context, which provides an accurate basis for our subsequent data analysis. Based on the IO dyeing mechanism, the present invention can achieve the following three key capabilities:

[0051] 1. Focus on the observation target: When capturing and collecting IO data, the system only focuses on IOs that carry color tags, effectively filtering IO traffic on non-critical paths, ensuring that resources are concentrated on the IO paths of core businesses or problems to be analyzed; 2. Full-link latency backtracking: After the IO is completed, by identifying the color tags it carries, the system can completely restore the flow process of the IO in the entire storage path, including the time consumption of each stage, thereby accurately analyzing the end-to-end IO latency and facilitating the location of performance bottlenecks; 3. Ensure tracking continuity: Even if splitting, merging, and other operations occur during IO transmission, the system can accurately track the flow and changes of the original IO based on the color tags, ensuring that tracking information is not lost due to changes in the IO structure, thereby ensuring the integrity and accuracy of data analysis.

[0052] S13: After IO latency observation is enabled, the IO latency capture module captures IOs carrying coloring information at key points in the IO lifecycle. These key points are located at the entry and exit points of IO processing in the virtualization layer, file system layer, block device layer, kernel layer, hyper-converged system layer, and other core processing points. When the IO flow passes through these key points, the IO capture module records the IO timestamp, coloring information, and other key information such as offset and length, as well as the trajectory information.

[0053] Specifically, after enabling the IO latency observation feature, the system's IO latency capture module begins accurately tracking IOs carrying coloring information. This module captures data at multiple key points in the IO lifecycle, covering all entry and exit points and core processing steps involved in IO processing, from the virtualization layer, file system layer, block device layer, to the kernel layer, as well as the hyperconverged system layer. By deploying collection logic at these key locations, the system comprehensively records the flow of IOs throughout the entire path, providing complete data support for subsequent performance analysis. Whenever a colored IO passes through one of these key nodes, the capture module automatically records key attributes such as its timestamp, coloring flag, offset, and data length. Based on this information, it constructs a trajectory of the IO within the system. This refined data collection method not only reconstructs the IO's transmission process between various layers, but also accurately identifies the time spent at each stage, enabling users to deeply analyze the composition of IO latency and identify potential performance bottlenecks. Furthermore, this mechanism ensures the continuity and accuracy of IO tracking, even in complex storage architectures or multi-path environments, ensuring reliable observation results.

[0054] In step S14, the IO latency analysis module analyzes the IO trace information generated by the IO capture module, traces the entire IO life cycle based on the coloring information of the completed IO, and obtains the total IO latency:

[0055] If the total IO latency is less than the configured latency threshold, the IO latency at each layer and between layers is not obtained.

[0056] If the total IO latency is greater than or equal to the configured latency threshold, the system backtracks the IO according to the IO coloring information to obtain the inter-layer and intra-layer latency of the entire IO path. The latency is sorted and displayed on the UI. Considering that OCFS file system heartbeat IO timeouts can cause service interruptions and considering typical hyperconvergence latency statistics, only IO records within the 15 seconds before the IO end are backtracked to accelerate the analysis process.

[0057] S15: After the observation time ends, the IO stops staining and tracking. The IO delay analysis module clears the IO records that do not meet the delay threshold during this tracking process. Other records and results are still retained and can be viewed later.

[0058] Therefore, the present invention focuses on core business IO through IO coloring technology, solves the problem of inaccurate tracking caused by IO splitting and merging in the IO process, and provides a system-level latency observation technology for the entire life cycle and all paths of IO.

[0059] The present invention adopts a rationally designed IO coloring mechanism, which can accurately mark the IO of the VM to be observed, and provide an accurate decision-making basis for IO latency full-path tracing.

[0060] The present invention designs a reasonable IO coloring mechanism, which ensures that the coloring information carried by each IO is different during each IO delay observation process and is repeatable in different observation cycles.

[0061] The present invention designs a reasonable IO capture mechanism, which accurately captures the IO of the VM to be observed without affecting the business IO, and generates a track record with timestamp, IO coloring information and other key information for the flow of IO in the entire path.

[0062] The present invention designs a reasonable IO delay analysis mechanism, which uses the total IO delay to trace the IO trace record according to the coloring information and analyze the IO delay of each layer and between layers on the entire IO path.

[0063] During IO backtracking, the present invention only backtracks the trajectory records 15 seconds before the IO end time point, accelerating IO delay analysis.

[0064] After the IO delay observation is completed, the present invention deletes the trace records of the IO total delay that does not meet the threshold condition, thereby releasing the occupation of system resources.

[0065] In summary, the present invention addresses the problem that hyper-converged systems are unable to quickly, accurately, and systematically observe IO latency when facing increasingly stringent requirements for IO latency and stability from core businesses. A hyper-converged system IO latency observation method and device based on IO coloring technology is proposed. Users can enable or disable the IO latency observation function through the UI configuration module, and can also configure parameters such as latency thresholds and observation time. After enabling the IO latency observation function, by coloring and marking the IOs issued by the VM to be observed, the problem that traditional observation methods cannot accurately observe after IO splitting and merging, and cannot distinguish between the IOs of the VM to be observed and the IOs issued by other VMs, resulting in excessive tracking of IOs, loss of tracking events, and complex analysis is solved. By filtering the colored IOs and generating a track with timestamps, coloring, and other key information for the IOs, detailed information on the IO life cycle is recorded. After the IO ends, by analyzing the total latency and tracing back the IO track based on the coloring information, detailed information such as the inter-layer latency, intra-layer latency, and total latency of each layer on the IO path is generated. Finally, an intuitive result is generated to show the time consumed by the IO at each stage of the path, greatly enhancing the competitiveness of hyper-converged products.

[0066] According to the hyper-converged system IO latency observation method based on IO dyeing technology in an embodiment of the present invention, it is possible to add specific tag information to IO through IO dyeing technology under the premise of ensuring the stability and availability of hyper-converged user services, so that the process of IO operations can be tracked and analyzed within the system. The IO dyeing technology can solve the problem that the traditional IO latency observation method cannot track or tracks inaccurately after splitting and merging in the IO process. At the same time, in order to improve accuracy and ease of use, IO layering technology is introduced and the latency of the entire IO life cycle is displayed through the UI. Based on IO dyeing and layering technology, the IO latency is analyzed, which can not only display the traditional IO total latency, but also display the flow latency of the IO entire life cycle at each layer and between layers through the UI. Of course, as an observation tool for IO latency and business operation status, the user can only turn on IO latency observation when needed. After turning on the hyper-converged system IO latency observation, the system automatically runs and statistically displays the IO latency, providing an intuitive display of the system business operation status.

[0067] In order to implement the above embodiment, Figure 4 As shown, this embodiment also provides a hyper-converged system IO latency observation device 10 based on IO staining technology, including:

[0068] The UI configuration module 100 is used by the user to enable IO latency observation and configure parameters, including the virtual machine VM to be observed, the latency threshold, and the observation duration or scheduled observation time;

[0069] The IO coloring module 200 is used to uniquely mark the IOs issued by the selected VM to complete the coloring;

[0070] IO latency capture module 300, used to capture IOs carrying coloring information at key points in the IO life cycle and record the IO timestamp, coloring information and other key information;

[0071] The IO latency analysis module 400 is used to analyze the IO trace information generated by the IO capture module, trace back the entire IO life cycle based on the coloring information of the completed IO, obtain the total IO latency, and compare the total latency with the configured latency threshold to determine whether to further trace back the inter-layer latency and intra-layer latency of the entire IO path and output them to the UI interface.

[0072] Furthermore, the observation duration or scheduled observation time supports the use of default values.

[0073] Furthermore, IOs that do not carry staining labels are not observed, and the staining information carried by each IO is different.

[0074] Specifically, the present invention mainly includes a UI interface, an IO coloring module, an IO delay capture module, and an IO delay analysis module. The UI interface provides users with an IO delay observation function configuration entry and result display. Users can select the VM that needs to be observed on the UI as needed. The IO coloring module will color the IO of the selected VM according to the configuration information issued by the UI interface: that is, by adding specific tag information to the IO, it can track and analyze the process of IO operations within the system. In order to improve usability and highlight key issues, users can configure the delay threshold from the UI, that is, only display IO that exceeds the threshold. If the user does not configure it, the system default threshold is used, and the default threshold is set to 0.5ms. The colored IO still carries the coloring information when it is split and merged on the IO path. The IO path is as shown in the attached figure. Figure 1 As shown, the figure also shows the points where IO splitting and merging occur. Splitting and merging IO at any point will introduce technical challenges to the dynamic tracking of IO latency, resulting in uncertainty in the tracking results. We have solved the problem that IO cannot be dynamically tracked due to merging and splitting through IO dyeing technology. The IO latency capture module is responsible for capturing the IO latency of a specified VM. When capturing IO through traditional ebpf means, it is impossible to distinguish between the IO to be detected and the IO of other VMs. There are problems such as capturing too much IO, significant impact on system IO performance, and discarded capture events, which ultimately lead to inaccurate IO latency, IO latency cannot display the latency between layers, etc., and the ease of use is greatly reduced. After the introduction of IO dyeing, the IO latency capture module filters the dyed IO and only focuses on the IO issued by the specified VM. When the IO passes through the attached Figure 2The key points of the path shown record timestamps, IO coloring information and other key information. Since the latency of most IOs cannot be known until the IO is completed to see whether it meets the threshold set by the user, the IO latency capture module will capture all IOs of the VM to which it belongs, but will release the IOs that do not meet the threshold conditions after the tracking is completed, in order to accelerate performance analysis and reduce the performance impact of system resources on user services. After the IO is captured by the latency capture module, a track information will be generated in the system. At the same time, the track information carries the timestamp of the IO in the flow. The IO latency parsing module parses the IO track information generated by the IO capture module to generate the total latency, inter-layer latency, and intra-layer latency of the IO, and sorts them according to the size of the latency, and finally outputs them to the UI interface for display. Through the close cooperation and processing of various modules, an intuitive and convenient way is provided for latency observation throughout the entire life cycle of IO.

[0075] According to an embodiment of the present invention, a hyper-converged system IO latency observation device based on IO coloring technology introduces IO coloring technology, while ensuring the stability and availability of user services in a hyper-converged environment. By adding specific tag information to each IO operation that needs to be observed, the system can track and analyze the entire process of these IOs within the entire system. This coloring mechanism not only maintains the consistency of its identification when IO traverses different system layers (such as the virtualization layer, file system layer, block device layer, etc.), but also effectively handles situations such as IO splitting and merging that are difficult to handle with traditional observation methods, thereby significantly improving the accuracy and completeness of IO tracking. With the help of IO coloring technology, the system can accurately identify and restore the complete life cycle of each key IO even in complex storage paths or concurrent environments. To further improve observation accuracy and user experience, the system also introduces IO layering technology, which divides the IO flow process in the entire system into multiple logical layers and clearly displays the IO flow path and corresponding latency at each layer and between layers through a graphical interface (UI). Based on this mechanism, when analyzing IO latency, not only can the traditional overall IO response time be presented, but also the specific time consumption of each step can be drilled down to help users quickly locate performance bottlenecks. As a lightweight, on-demand observation tool, the IO latency observation feature only operates when enabled by the user, without causing a sustained impact on normal business operations. Once enabled, the system automatically collects and analyzes data and generates visualizations, providing users with intuitive insights into IO performance and business operation status, helping operations personnel efficiently complete performance tuning and troubleshooting.

[0076] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0077] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

Claims

1. A method for observing IO latency in a hyper-converged system based on IO staining technology, characterized in that: The following steps are involved: The user enables IO latency monitoring and configures parameters through the UI configuration module, including the VM to be monitored, the latency threshold, and the observation duration or scheduled observation time. The IO coloring module uniquely tags the IOs issued by the selected VM to complete the coloring; The IO latency capture module captures IOs carrying coloring information at key points in the IO lifecycle and records the IO timestamp, coloring information, and other key information. The IO latency analysis module analyzes the IO trace information generated by the IO capture module and traces back the entire IO lifecycle based on the coloring information of the completed IO to obtain the total IO latency. Based on the comparison of the total latency with the configured latency threshold, the module determines whether to further trace back the inter-layer and intra-layer latency of the entire IO path and output the results to the UI.

2. The method according to claim 1, characterized in that The observation duration or scheduled observation time supports the use of default values.

3. The method according to claim 1, characterized in that IOs that do not carry staining labels are not observed, and the staining information carried by each IO is different.

4. The method according to claim 1, wherein The key points are distributed at the entrances and exits of IO processing and other core processing points in the virtualization layer, file system layer, block device layer, kernel layer, and hyper-convergence system layer.

5. The method according to claim 1, wherein If the total IO latency is less than the configured latency threshold, the latency of each layer and between layers is not obtained. If the total IO latency is greater than or equal to the configured latency threshold, the system backtracks the IO to obtain the inter-layer and intra-layer latency of the entire IO path. The latency is then sorted and displayed on the UI.

6. The method according to claim 5, characterized in that Only the IO records within 15 seconds before the end of the IO are traced back.

7. The method according to claim 1, characterized in that After the observation time ends, IO coloring and tracking stops, and the IO latency analysis module clears IO records that do not meet the latency threshold during this tracking process.

8. A hyper-converged system IO latency observation device based on IO staining technology, characterized in that: include: The UI configuration module is used by users to enable IO latency observation and configure parameters, including the VM to be observed, the latency threshold, and the observation duration or scheduled observation time; The IO coloring module is used to uniquely mark the IOs issued by the selected VM to complete the coloring; The IO latency capture module is used to capture IOs carrying coloring information at key points in the IO lifecycle and record the IO timestamp, coloring information, and other key information. The IO latency analysis module parses the IO trace information generated by the IO capture module, traces back the entire IO lifecycle based on the coloring information of the completed IO, obtains the total IO latency, and compares the total latency with the configured latency threshold to determine whether to further trace back the inter-layer and intra-layer latency of the entire IO path and output them to the UI.

9. The device according to claim 8, characterized in that The observation duration or scheduled observation time supports the use of default values.

10. The device according to claim 8, characterized in that IOs that do not carry staining labels are not observed, and the staining information carried by each IO is different.