Business system configuration method, electronic device and storage medium
Automatically monitoring and optimizing storage system performance through input and output tracking tools and pre-trained configuration models solves the problem of time-consuming and ineffective manual analysis, achieves performance automation and precise performance optimization, and improves system stability and efficiency.
Patent Information
- Application Number
- CN202511093959.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-06
AI Technical Summary
The performance optimization of existing storage systems relies on manual analysis, which is time-consuming and ineffective. Performance bottlenecks frequently occur, affecting the stability and continuity of business systems.
Use input and output tracking tools to monitor the read and write performance data of business systems, use pre-trained configuration models to automatically analyze and generate target configuration files, and dynamically adjust system performance.
It achieves fast and accurate performance optimization, reduces dependence on professional technicians, reduces analysis time costs, avoids the recurrence of performance bottlenecks, and improves system stability and efficiency.
Smart Images

Figure CN120596036B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of storage systems, and in particular to a configuration method, electronic device, and storage medium for a business system. Background Art
[0002] In today's rapidly evolving digital transformation landscape, storage systems, as a core component of enterprise IT infrastructure, have a significant impact on the service quality and operational efficiency of business systems. With the continued development of enterprise IT systems, data volumes are exploding, and data access frequency is increasing. This is placing increasing pressure on storage systems to handle input and output (IO). Driven by this trend, storage performance has become a key indicator of the core capabilities of business systems. Performance bottlenecks not only hinder business expansion and innovation, but also significantly reduce user experience, negatively impacting operational efficiency and market competitiveness.
[0003] However, relevant optimization solutions for storage performance have obvious shortcomings. On the one hand, they rely heavily on manual analysis, requiring professional technicians to evaluate storage carrying capacity based on the IO model of the enterprise's business type and achieve optimization by adjusting the host usage strategy and the operating status of the internal storage modules. This not only requires extremely high professional capabilities from the technicians, but also consumes a lot of time and costs, making it difficult to quickly respond to the performance requirements of the business system. On the other hand, when a storage performance bottleneck actually occurs, it is often only temporarily resolved by restarting the system. However, the performance improvement after the restart is limited and cannot fundamentally solve the problem, resulting in repeated performance bottlenecks, seriously affecting the stability and continuity of the business system. Summary of the Invention
[0004] The present application provides a configuration method, electronic device and storage medium for a business system to at least solve the problems in related technologies such as reliance on manual analysis, high professional and technical requirements, long time consumption, and poor reliance on restart when performance bottlenecks occur.
[0005] The present application provides a configuration method for a business system, comprising: monitoring the input and output operations of the business system through an input and output tracking tool to obtain read and write performance data; wherein the business system includes a host layer, a data transmission network layer, and a storage layer; analyzing the read and write performance data to obtain analysis results; determining a target configuration file based on the analysis results and a pre-trained configuration model; wherein the pre-trained configuration model is trained based on different analysis result samples and corresponding configuration file samples, and the analysis result samples include analysis results corresponding to different input and output operations in different scenarios; the target configuration file is used to adjust the performance of the business system; and the business system is configured based on the target configuration file.
[0006] This application also provides a configuration device for a business system, including:
[0007] The monitoring module is used to monitor the input and output operations of the business system through the input and output tracking tool to obtain read and write performance data. The business system includes the host layer, data transmission network layer and storage layer;
[0008] An analysis module is used to analyze the read and write performance data and obtain analysis results;
[0009] The model inference module is used to determine the target configuration file based on the analysis results and the pre-trained configuration model. The pre-trained configuration model is trained based on different analysis result samples and corresponding configuration file samples. The analysis result samples include analysis results corresponding to different input and output operations in different scenarios. The target configuration file is used to adjust the performance of the business system.
[0010] The configuration module is used to configure the business system based on the target configuration file.
[0011] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of the above-mentioned business system configuration method when executing the computer program.
[0012] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the configuration method of the above-mentioned business system are implemented.
[0013] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned business system configuration method when executed by a processor.
[0014] This application introduces an input and output tracking tool to automatically monitor the read and write performance data of each level of the business system, replacing the traditional manual collection and analysis process, greatly reducing the dependence on professional and technical personnel and the time cost. At the same time, through the application of pre-trained configuration models, the performance optimization experience of business systems under different preset scenarios is solidified into a model, which can quickly match and generate target configuration files based on real-time analysis results, avoiding the lag and subjectivity of manual evaluation, and realizing the rapid output of optimization solutions. Through continuous monitoring and analysis of full-link performance data, bottleneck links can be accurately located, and targeted adjustments can be made to the business system based on the target configuration file, achieving fundamental relief of performance bottlenecks rather than temporary avoidance. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0016] Figure 1 A schematic diagram of a specific hardware architecture on which execution of a configuration method for a business system provided in an embodiment of the present application relies;
[0017] Figure 2 A flowchart of a method for configuring a business system provided in an embodiment of the present application;
[0018] Figure 3 A schematic diagram of the structure of a storage layer of a business system provided in an embodiment of the present application;
[0019] Figure 4 A schematic diagram of input and output operations of a business system provided in an embodiment of the present application;
[0020] Figure 5 A schematic diagram of the structure of a configuration device for a business system provided in an embodiment of the present application;
[0021] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0023] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0024] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0025] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the configuration method of the business system depends, the specific application environment architecture or specific hardware architecture is described here.
[0026] like Figure 1 FIG. 1 is a schematic diagram of a specific hardware architecture on which the execution of the business system configuration method depends, including a third-party host and a business system.
[0027] The third-party host is used to run input and output tracking tools, analyze read and write performance data, load pre-trained configuration models, and generate target configuration files. The third-party host must be equipped with sufficient CPU cores, memory, local storage, and a network interface to the business system host layer.
[0028] The business system host layer consists of at least one or more application servers (physical or virtual machines), which run specific business systems (such as databases and trading systems) and generate I / O operations. These servers must support the deployment of I / O tracking tools and communicate with the storage layer through the driver layer. The hosts must be configured with an operating system and a corresponding I / O scheduler to receive configuration files from third-party hosts.
[0029] The business system data transmission network layer connects the business system host layer and the storage layer, supporting the transmission of I / O data and configuration instructions. It includes network equipment and transmission protocols.
[0030] The storage layer of the business system includes multi-level storage components that support hierarchical processing of I / O operations. The cache layer, composed of the storage array's cache, is used to temporarily store I / O data. The volume layer and the Redundant Array of Independent Disks (RAID) layer implement storage space virtualization through logical volume management or RAID controllers, supporting volume mapping, striping, and redundant protection. Back-end storage devices such as hard disk drives (HDDs), solid-state drives (SSDs), or just a bunch of disks (JBODs) provide physical storage space and must support I / O command interaction with the back-end driver layer. The storage controller is responsible for coordinating I / O processing at each layer and supports dynamic adjustment of configuration parameters.
[0031] The embodiments of the present application provide a configuration method for a business system, and the method is described in detail in conjunction with the execution flow of the configuration method for the business system.
[0032] like Figure 2 As shown, an embodiment of the present application provides a flow chart of a method for configuring a business system, the method comprising the following steps S201 to S204:
[0033] S201. Monitor the input and output operations of the business system through an input and output tracking tool to obtain read and write performance data.
[0034] Among them, the input and output tracking tool (IO trace) uses Python scripts to implement the IO monitoring service program. The business system includes the host layer, data transmission network layer and storage layer. The storage layer is as follows: Figure 3 As shown in the figure, it includes the front-end host driver layer, host protocol layer, cache layer, volume layer, disk array (RAID) layer, back-end protocol layer, and back-end disk driver layer. The back-end disk driver layer is the back-end disk host bus adapter (HBA) driver layer.
[0035] The input and output operations of the business system are as follows: Figure 4 As shown in the figure, a third-party host uses the IO trace tool to monitor the input and output operations of the business system and collect read and write performance data. The IO flow is captured from the host layer to the data transmission network layer and then to the storage layer.
[0036] Read / write performance data is a detailed record of business system input / output operations, presented in the form of structured logs, binary files, or specially formatted text. This data can include at least one of the following: the host layer's first input / output latency, the data transmission network's hardware quality signal, and the storage layer's second input / output latency. This data also includes the initiation time of the input / output operation, the logical block address (LBA), the operation type (read / write), the data size, and the associated process.
[0037] In some embodiments, the business system of the present application sets a client at the storage layer, which is implemented in Python and establishes a communication connection with a third-party host. The third-party host obtains read and write performance data through the client at the storage layer.
[0038] In some embodiments, an input / output tracking tool is deployed at the host layer of the business system to monitor multiple storage volumes at the host layer. Specifically, the LBA intervals of multiple storage volumes are recorded; based on the LBA intervals of the storage volumes, the business phase corresponding to each input / output is identified. Business phases include, but are not limited to, initialization phase, operation phase, and backup phase. For example, a database index file may be concentrated in LBA interval A, data files in interval B, and log files in interval C; the backup phase may focus on accessing intervals B and C.
[0039] For example, LBA interval 0-100000 is frequently accessed by the database process, and is mainly random read, and is marked as the database index query phase. LBA interval 100001-200000 is frequently written sequentially by the log process, and is marked as the log write phase.
[0040] The above embodiment records the LBA interval of the storage volume and the corresponding IO behavior through the input and output tracking tool, without the need to manually interpret the IO log and storage volume mapping relationship one by one; through preset rules, such as writing the order of log processes into the corresponding log stage, or by associating with historical data, the business stage label is directly output, which reduces the requirements for the professional skills of the operator; the input and output tracking tool monitors IO behavior in real time at the host layer. Once characteristic access to a specific LBA interval is found, the corresponding business stage can be immediately marked, avoiding the time-consuming process of manually extracting logs and comparing data across systems afterwards; through the binding of LBA intervals and business stages, scattered IO data is converted into structured stage and behavior association information, reducing the time cost of manual sorting and analysis; through the mapping of LBA intervals and business stages, the root cause of the problem can be traced, which is conducive to accurately locating performance bottlenecks.
[0041] In some embodiments, an I / O tracking tool is used to monitor the I / O request queues of multiple storage volumes at the host layer and record the backlog of I / O operations in the I / O request queues. If the backlog exceeds a preset upper limit, multiple I / O operations for adjacent logical block addresses are merged; if the backlog is less than a preset lower limit, the I / O operation is split into multiple sub-operations.
[0042] Optionally, the lengths of the I / O request queues awaiting processing in multiple storage volumes are recorded. The I / O request queue length reflects the degree of I / O backlog. A high backlog increases latency, while a low backlog may waste bandwidth. Input and output operations are then merged or split based on the I / O request queue length. If the I / O request queue length is less than a first length threshold, the I / O operations are merged; if the I / O request queue length is greater than a second length threshold, the I / O operations are split. The first length threshold is less than the second length threshold.
[0043] For example, multiple small IO requests with adjacent logical block addresses are merged into a large IO, which reduces the number of IO operations and improves throughput, which is suitable for mechanical disks. For example, multiple consecutive 4KB write requests can be merged into a single 64KB write request.
[0044] As another example, splitting a large I / O request into multiple smaller ones prevents the large I / O from monopolizing bandwidth and causing excessive latency for other requests. This approach is suitable for SSD or concurrent scenarios. For example, a 1MB write request can be split into four 256KB requests. This precise optimization for different scenarios fundamentally alleviates performance bottlenecks.
[0045] The above embodiment uses automated tools to monitor the storage volume's IO request queue in real time, automatically recording the queue backlog level (such as queue length) and key attributes of IO operations (such as data volume and runtime). This eliminates the need for manual collection and analysis of the storage volume's IO status, reducing the complexity of human intervention. IO operations are dynamically adjusted based on the IO request queue backlog level. When the backlog is too high, IO operations from adjacent LBAs are merged to improve throughput. When the backlog is too low, IO operations are split to avoid bandwidth waste. This entire process requires no manual intervention, shortening the time it takes to respond to and resolve issues.
[0046] Optionally, record the data volume, run time, and end time of I / O operations. The data volume of I / O operations is directly related to the storage system's processing efficiency. Run time is the time it takes for an I / O operation to run on the central processing unit (CPU) core, representing the time it takes from the application initiating the I / O operation to the CPU processing. End time is the time it takes for the I / O operation to be written to the kernel driver layer and is used to calculate the total duration of the I / O operation. By recording the basic attributes of I / O operations and the duration of key steps, the spatial characteristics and resource consumption of I / O operations can be intuitively reflected.
[0047] By recording key metrics such as the run time and end time of I / O operations, we can pinpoint the root causes of performance issues, such as large I / O monopolizing bandwidth and causing latency, or excessive small I / O affecting throughput. This allows for targeted policy adjustments, such as splitting overly large I / O operations to resolve concurrency blockages or merging small I / O operations to improve mechanical disk efficiency. This precise root cause identification and optimization avoids business interruptions and recurring issues caused by restarts, fundamentally improving system stability and performance.
[0048] When using I / O tracking tools to monitor the input and output operations of a business system, you can optionally record at least one of the following: the start event (S), completion event (C), run event (K), queue event (Q), merge event (M), or split event (P). By marking key actions around the event nodes of the I / O operation, you can determine the sequence of events, clarify the duration and dependencies of each step, and build an event flow for the lifecycle of the I / O operation.
[0049] For example, for the start event, the IO operation ID, LBA, and timestamp are recorded; for the queuing event, the queue ID and current queue length are recorded; for the merge event or split event, the number of IO operations and the LBA range before and after the merge / split are recorded; for the running event, the CPU cores and time occupied by the IO operation during the processing are recorded; for the completion event, the IO operation completion timestamp is recorded.
[0050] The above embodiment clearly restores the processing logic and dependencies of IO operations by recording the event nodes and process paths of IO operations. Specifically, an automated tool records key events such as the start, completion, execution, queuing, merging, and splitting of IO operations, and marks each event with core information such as timestamp, queue length, and CPU usage. This replaces the manual process of sifting through massive log files for key information, directly constructing a clear event chain and significantly reducing reliance on specialized technical expertise.
[0051] In some embodiments, the storage layer of the business system includes a front-end host driver layer, a host protocol layer, a cache layer, a volume layer, a disk array layer, a back-end protocol layer, and a back-end disk driver layer. In the process of monitoring the input and output operations of the business system using an input and output tracking tool to obtain read and write performance data, when the input and output operation obtains a logical block address at the front-end host driver layer, it is marked as the start time; when the input and output operation enters the host protocol layer, the cache layer, the volume layer, the disk array layer, and the back-end protocol layer, the corresponding timestamp is recorded at the entrance to each layer and associated with the layer identifier; when the input and output operation is written to the disk at the back-end disk driver layer, the completion time is marked; the start time, the timestamps corresponding to each layer, and the completion time are integrated into a tracking record.
[0052] Specifically, the time taken for each link in the IO operation path, from initiation at the host layer to storage layer write-to-disk, is recorded. The IO path is divided into multiple layers: the front-end host driver layer (H), host protocol layer (O), cache layer (A), volume layer (V), RAID layer (R), and back-end driver layer (D). When the IO operation obtains the LBA address at the front-end host driver layer, it is marked as the start time S. Subsequently, as it flows through each layer, the time it enters that layer is recorded. Finally, after the IO operation is written to disk, the end time C is marked. This time information is integrated into a complete tracking record, including the LBA address of the IO operation, the time taken at each layer, the total time taken, and other information, and saved in the form of a file.
[0053] For example, when the front-end host driver layer (H) receives an IO request and parses the LBA address, it records the start time S. As the IO enters the host protocol layer O, cache layer A, volume layer V, RAID layer R, and back-end driver layer D, it records the corresponding timestamp at the entrance to each layer and associates it with the layer identifier. The next step is to capture the completion time of the IO operation. When the IO operation is finally completed and written to disk, the completion time C is marked at the back-end driver layer, and the time records of the IO at each layer are retroactively associated. During the lifecycle of each IO operation from S to C, the timestamps of each layer, such as the time difference from S to H, the time from H to O, and the time from O to A, as well as information such as the LBA address range, IO type (read / write), and data size, are aggregated into a structured trace record containing the following fields: IO unique ID, LBA start, LBA length, S time, H time, O time, A time, V time, R time, D time, C time, and IO type. These records are then written to a file in chronological order for storage.
[0054] The above embodiment uses automated tools to record the entry time, start time, and completion time of IO operations at various levels, such as the front-end host driver layer, host protocol layer, and cache layer, and integrates them into structured tracking records containing fields such as unique ID, LBA information, and time consumption at each layer. This replaces the process of manually piecing together data from scattered system logs and driver information, reduces dependence on professional technology, and eliminates the need for manual interpretation of complex underlying interactions. By clarifying the time consumption at each level, such as the time difference from the front-end host driver layer to the host protocol layer, and the time consumption from the cache layer to the volume layer, performance bottlenecks can be quickly located and analysis time is shortened. Refined analysis based on full-path tracing can trace the root cause of performance problems.
[0055] S202: Analyze the read and write performance data to obtain analysis results.
[0056] The analysis results can include performance statistics tables, presented in the form of Excel tables and charts.
[0057] In some embodiments, the read / write performance data includes event nodes for input / output operations, namely, at least one of a start event, a completion event, a run event, a queue event, a merge event, and a split event. When executing step S202, the event nodes are first sorted in a red-black tree according to the logical block address of the input / output operation. When a start event for any input / output operation is detected, a target completion event is determined from the completion events in the red-black tree. The target completion event overwrites the queue event for this input / output operation. An event flow is then established for this input / output operation, from the start event to the queue event to the target completion event. This event flow is further parsed to determine the performance bottleneck of this input / output operation.
[0058] Specifically, the third-party host reads IO operation events one by one and sorts each event node in a red-black tree according to the logical block address of the IO operation. Leveraging the efficient interval query capability of the red-black tree, this system enables rapid indexing of large numbers of IO operations. When an IO operation start event S is captured, the host scans the recorded completion events C in the red-black tree to find a completion event that covers the current IO queue interval (Q). This establishes a full link between the IO operation, from initiation to queuing to completion.
[0059] For example, assume that two IO operations [0, 10] and [10, 20] with adjacent logical block address intervals are merged into a large IO operation [0, 20]. When the merged IO is completed, the completion status of the original two IOs can be inferred through the interval coverage relationship, thereby calculating their respective delays.
[0060] By analyzing the associated IO event stream, bottlenecks can be identified based on the time taken at different stages. If anomalies occur during the CPU scheduling phase (e.g., from K events to other phases), feedback can be provided to the host client to adjust the scheduling algorithm. If the queuing phase (Q events) takes too long, this could indicate an issue with the scheduling mechanism or the underlying transmission link. If the CPU processing takes too long to complete (K to C), this likely points to an issue with the underlying driver or the link from the host to the storage.
[0061] The above-described embodiment uses a red-black tree to sort the logical block addresses of IO operations, leveraging its efficient interval query capabilities to quickly index and associate a large number of IO events, replacing the process of manually sorting and comparing scattered event data. Through automated red-black tree sorting and interval coverage queries, a full-link event flow from start events to queued events to completed events is directly constructed, reducing reliance on specialized technology and eliminating the need for manual intervention in complex association logic. The efficient red-black tree-based index can quickly associate event nodes and directly locate bottlenecks by analyzing the time consumed by each stage in the event flow, shortening analysis time.
[0062] Optionally, the read-write performance data includes event nodes of input and output operations, namely at least one of the start event, completion event, run event, queue event, merge event and split event. When executing step S202, firstly, the events of the same IO operation are sorted in chronological order to obtain an event chain, the interval time between each event node is calculated, and the interval time is added to the analysis results of the read-write performance data. For example, the interval from the queue event to the merge event represents the queue delay, and the interval from the merge event to the run event represents the merge time. By clarifying the sequence of events and the time consumption of each link such as queue delay and merge time, the performance problem can be quickly located and the analysis time is shortened. It can be made clear through the event flow that if the queue delay of a certain IO operation is too high, it may be a problem with the queue scheduling strategy; if the run event occupies the CPU core time for too long, it may be improper resource allocation. Based on these specific conclusions, it is helpful to find a targeted configuration file.
[0063] In some embodiments, by analyzing the time difference between the start time S and each layer in the tracking record, it is determined at which stage the IO operation is delayed, thereby locating the performance bottleneck. The time contribution of the IO operation at each layer can be calculated, such as the protocol layer processing time, the cache to disk synchronization time, etc. For example, a long A to V time may indicate a bottleneck in the data transmission from the cache layer to the volume layer, while an abnormal R to D time may point to a problem with RAID verification or the back-end driver.
[0064] S203: Determine the target configuration file based on the analysis results and the pre-trained configuration model.
[0065] Pre-trained configuration models are stored in the model library. These models are trained based on different analysis result samples and corresponding configuration file samples. The analysis result samples used for training are obtained by analyzing different input and output operations in different scenarios. The target configuration files are used to adjust the performance of the business system.
[0066] The model library includes analysis results for various input and output operations under different pre-defined scenarios, along with configuration models trained using corresponding configuration files. Optimal configurations for different I / O operations vary in host scheduling, storage caching strategies, and RAID verification efficiency.
[0067] We simulate various read and write workloads in a lab environment and test performance under different configurations, such as adjusting host-side I / O queue depth, storage cache refresh policy, and RAID stripe size. We then identify the optimal parameter combinations for each I / O characteristic and build a model library. These models encompass not only storage-layer configurations (such as volume-layer block size and RAID level selection) but also host-layer best practices (such as driver parameters and CPU affinity settings).
[0068] The configuration model training process involves first setting up the training environment and collecting data. A test platform simulating real-world scenarios is constructed, including front-end hosts (configured with different operating systems and drivers), data transmission networks, storage layers, and performance monitoring tools. For each I / O size, such as 4K and 8K, individual and combined scenarios are designed. Key parameters are adjusted within each pre-defined scenario: At the host level, different queue depths and I / O scheduling algorithms can be tested; at the storage level, cache sizes, RAID stripe sizes (e.g., 64K and 128K), and back-end disk types can be tested. For each parameter combination, I / O workloads are continuously run and performance metrics are recorded to identify the optimal parameter combination for that scenario. Each I / O operation characteristic is associated with the corresponding optimal profile to form a structured pre-trained model. For example, a profile for a 4K random write scenario might include a host-side queue depth of 8, a storage cache dirty data flush threshold of 50%, and a RAID5 stripe size of 64K. A profile for a 1M sequential read scenario might include host-side read-ahead enabled, a storage cache prefetch size of 2M, and a RAID0 stripe size of 1M.
[0069] In some embodiments, the reasoning process of the configuration model includes: based on the analysis results and the pre-trained configuration model, first obtaining the feature information of the input and output operations based on the analysis results, then matching the closest configuration model from the model library based on the feature information, and then inputting the analysis results into the configuration model to obtain the corresponding target configuration file.
[0070] Optionally, the analysis results include the performance bottleneck location of input and output operations, and / or the interval time between each event node. First, the bottleneck type such as excessive queuing delay, read-write conflict, etc., the address range of the involved logical block, the associated event node chain and other features are extracted from the performance bottleneck location; the time value, time fluctuation range and proportion of key nodes such as start events and queue events, running events and completion events are extracted from the event interval time to form a structured feature vector. Then, by calculating the similarity between the feature vector of the scenario to be analyzed and the feature samples corresponding to each pre-trained model in the model library, the Euclidean distance, cosine similarity and other algorithms can be used to quantify the degree of matching, and the model with the highest similarity is selected as the target model. The complete analysis results, including the specific bottleneck details, the original data of the time consumption of each event, etc., are input into the matched configuration model. The model is based on the built-in mapping relationship between different features and optimal configuration parameters. By performing pattern recognition and logical reasoning on the input analysis results, it generates a target configuration file containing specific adjustment parameters, such as adjusting the read and write queue length thresholds, optimizing the logical block address allocation strategy, adjusting the event scheduling priority, etc., which is directly used for the performance parameter configuration of the business system to achieve targeted optimization.
[0071] By combining specific input and output operation performance analysis results with pre-trained configuration models, we can achieve accurate recommendations for business system configurations. Because the pre-trained configuration model is trained based on a large number of analysis result samples and corresponding configuration file samples from different scenarios, it can match the most suitable configuration solution for different performance bottlenecks and event durations. This effectively guides the performance adjustment of business systems, improves system operational efficiency in different scenarios, and optimizes the overall performance of input and output operations.
[0072] The above-mentioned embodiment, by building a test platform that simulates real-world scenarios, automatically traverses different I / O sizes, single scenarios, and combined scenarios, and adjusts key parameters such as the host and storage layers. This replaces the manual trial-and-error process, automatically records performance metrics, and selects the optimal combination. By linking I / O operation characteristics with optimal configuration files through pre-trained models, performance optimization time is significantly shortened. Precise model-based configuration can fundamentally improve performance.
[0073] S204: Configure the business system based on the target configuration file.
[0074] Import the corresponding configuration file into the business system to automatically adjust parameters. Optionally, use the IO trace tool to verify whether the actual performance meets the expectations of the pre-trained configuration model. If there is a deviation, call alternative configurations for similar scenarios from the model library and fine-tune them to ensure the ultimate performance optimization.
[0075] In summary, the configuration method of a business system provided by the embodiment of the present application automatically collects key data such as LBA intervals, event nodes, and timestamps of each level through input and output tracking tools, and realizes event association and sorting with the help of red-black tree sorting, event stream construction and other technologies, replacing the traditional reliance on manual interpretation of logs and complex process of combing logic, so that non-professionals can also locate problems through structured data and tool output. From real-time monitoring of IO request queues, dynamic adjustment of merge / split strategies, to tracing the time consumption of each link through event streams and full-path timestamps, and then directly matching the optimal parameter configuration with the pre-trained model, the performance bottleneck position is quickly locked, the efficiency of performance analysis and problem location is improved, and the time consumption is shortened. Through means such as associating business stages with LBA intervals, parsing the dependencies of each link with event streams, and analyzing the time consumption of the full-path level, the root causes of the problem such as resource conflicts in the business stage, processing delays at a certain level, and unreasonable parameter configuration are clarified, and the target configuration file is determined through the pre-trained configuration model to achieve a fundamental solution to the performance bottleneck.
[0076] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0077] like Figure 5 As shown, Figure 5 A schematic diagram of a configuration device for a business system provided in an embodiment of the present application, the device comprising:
[0078] Monitoring module 501, used to monitor the input and output operations of the business system through the input and output tracking tool to obtain read and write performance data; wherein the business system includes a host layer, a data transmission network layer, and a storage layer;
[0079] An analysis module 502 is used to analyze the read and write performance data to obtain analysis results;
[0080] Model inference module 503 is used to determine a target configuration file based on the analysis results and a pre-trained configuration model. The pre-trained configuration model is trained based on different analysis result samples and corresponding configuration file samples. The analysis result samples include analysis results corresponding to different input and output operations in different scenarios. The target configuration file is used to adjust the performance of the business system.
[0081] The configuration module 504 is used to configure the business system based on the target configuration file.
[0082] As an optional implementation method provided in an embodiment of the present application, the read and write performance data includes event nodes of input and output operations; the monitoring module 501 is specifically used to monitor the input and output operations of the business system through an input and output tracking tool, and record the event nodes of the input and output operations; wherein the event nodes include at least one of a start event, a completion event, a running event, a queue event, a merge event, and a split event.
[0083] As an optional implementation provided in an embodiment of the present application, the analysis results include the performance bottleneck location of the input and output operations; the analysis module 502 is specifically used to sort the event nodes in the red-black tree according to the logical block address of the input and output operations; when the start event of any input and output operation is monitored, the target completion event is determined from the completion events in the red-black tree, and the target completion event covers the queued event of any input and output operation; an event flow of any input and output operation from the start event to the queued event and then to the target completion event is established; the event flow is parsed to determine the performance bottleneck location of any input and output operation.
[0084] As an optional implementation provided in an embodiment of the present application, the analysis results include the interval time between each event node; the analysis module 502 is specifically used to sort the event nodes of the input and output operations in chronological order to obtain an event chain; based on the event chain, the interval time between each event node is calculated.
[0085] As an optional implementation method provided in an embodiment of the present application, the read and write performance data includes tracking records; the storage layer includes a front-end host driver layer, a host protocol layer, a cache layer, a volume layer, a disk array layer, a back-end protocol layer and a back-end disk driver layer; the monitoring module 501 is specifically used to mark the start time when the input and output operation obtains the logical block address in the front-end host driver layer; when the input and output operation enters the host protocol layer, the cache layer, the volume layer, the disk array layer, and the back-end protocol layer, the corresponding timestamp is recorded at the entrance to each layer, and the layer identifier is associated; when the input and output operation is written to the disk in the back-end disk driver layer, the completion time is marked; the start time, the timestamp corresponding to each layer and the completion time are integrated into a tracking record.
[0086] As an optional implementation method provided in an embodiment of the present application, an input / output tracking tool is deployed on a client of the host layer; the read / write performance data includes the business stage corresponding to the input / output operation; the monitoring module 501 is specifically used to monitor the input / output operations of multiple storage volumes of the host layer through the input / output tracking tool; record the logical block address ranges of multiple storage volumes; and identify the business stage corresponding to the input / output operation based on the logical block address ranges.
[0087] As an optional implementation provided in an embodiment of the present application, the device also includes: a merging and splitting module, which is used to monitor the input and output request queues of multiple storage volumes of the host layer through an input and output tracking tool; record the backlog level of input and output operations in the input and output request queue; when the backlog level is greater than a preset backlog upper limit, merge multiple input and output operations of adjacent logical block addresses; when the backlog level is less than a preset backlog lower upper limit, split the input and output operation into multiple sub-operations.
[0088] As an optional implementation provided in the embodiment of the present application, the model reasoning module 503 is specifically used to determine the characteristic information of the input and output operations based on the analysis results; match the pre-trained configuration model corresponding to the input and output operations from the model library based on the characteristic information; input the analysis results into the pre-trained configuration model to obtain the target configuration file.
[0089] For the description of the features in the embodiment corresponding to the configuration device of the business system, please refer to the relevant description of the embodiment corresponding to the configuration method of the business system, and no further details will be given here.
[0090] like Figure 6 As shown, an embodiment of the present application further provides an electronic device, including a memory 601 and a processor 602, wherein the memory 601 stores a computer program, and the processor 602 is configured to run the computer program to execute the steps in any of the above-mentioned business system configuration method embodiments.
[0091] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned business system configuration method embodiments when running.
[0092] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0093] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned business system configuration method embodiments are implemented.
[0094] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned business system configuration method embodiments.
[0095] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0096] The above is a detailed introduction to the configuration method, electronic device and storage medium of a business system provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A method for configuring a business system, characterized in that: include: Monitor the input and output operations of the business system through an input and output tracking tool to obtain read and write performance data; wherein the business system includes a host layer, a data transmission network layer, and a storage layer; Analyzing the read and write performance data to obtain analysis results; Determining a target configuration file based on the analysis results and a pre-trained configuration model; wherein the pre-trained configuration model is trained based on different analysis result samples and corresponding configuration file samples, the analysis result samples including analysis results corresponding to different input and output operations in different scenarios; the target configuration file is used to adjust the performance of the business system; Configuring the business system based on the target configuration file; The read and write performance data includes event nodes of the input and output operations; The step of monitoring the input and output operations of the business system using an input and output tracking tool to obtain read and write performance data includes: monitoring the input and output operations of the business system using the input and output tracking tool and recording event nodes of the input and output operations; wherein the event nodes include at least one of a start event, a completion event, a running event, a queue event, a merge event, and a split event; The analysis results include the performance bottleneck location of the input and output operations; The analysis of the read and write performance data to obtain analysis results includes: sorting the event nodes in a red-black tree according to the logical block addresses of the input and output operations; when a start event of any input and output operation is monitored, determining a target completion event from the completion events in the red-black tree, wherein the target completion event covers the queued events of any input and output operation; establishing an event flow for any input and output operation from the start event to the queued event and then to the target completion event; and parsing the event flow to determine the performance bottleneck location of any input and output operation.
2. The method according to claim 1, characterized in that The analysis results include the time taken between each event node; The analyzing the read and write performance data to obtain analysis results includes: Arrange the event nodes of the input and output operations in chronological order to obtain an event chain; According to the event chain, the time taken between each event node is calculated.
3. The method according to claim 1, characterized in that The read and write performance data includes tracking records; The storage layer includes a front-end host driver layer, a host protocol layer, a cache layer, a volume layer, a disk array layer, a back-end protocol layer and a back-end disk driver layer; The input and output operations of the business system are monitored by the input and output tracking tool to obtain read and write performance data, including: When the input / output operation obtains the logical block address at the front-end host driver layer, it is marked as the start time; When the input / output operation enters the host protocol layer, cache layer, volume layer, disk array layer, and backend protocol layer, a corresponding timestamp is recorded at the entrance to each layer and associated with a layer identifier; When the input / output operation is written to the disk at the back-end disk drive layer, a completion time is marked; The start time, the timestamps corresponding to each layer, and the completion time are integrated into a tracking record.
4. The method according to claim 1, wherein The input and output tracking tool is deployed on the client of the host layer; the read and write performance data includes the business phase corresponding to the input and output operation; The input and output operations of the business system are monitored by the input and output tracking tool to obtain read and write performance data, including: Monitoring the input and output operations of the plurality of storage volumes of the host layer by using the input and output tracking tool; Recording logical block address intervals of the multiple storage volumes; A service phase corresponding to the input / output operation is identified according to the logical block address interval.
5. The method according to claim 4, characterized in that The method further comprises: Monitoring the input and output request queues of the plurality of storage volumes of the host layer by using the input and output tracking tool; Recording the backlog of input and output operations in the input and output request queue; When the backlog level is greater than a preset backlog upper limit, merging multiple input and output operations of adjacent logical block addresses; When the backlog level is less than a preset backlog lower limit, the input and output operation is split into multiple sub-operations.
6. The method according to claim 1, characterized in that Determining a target configuration file based on the analysis results and the pre-trained configuration model includes: Determining characteristic information of the input and output operations according to the analysis results; Matching the pre-trained configuration model corresponding to the input and output operations from a model library according to the feature information; The analysis result is input into the pre-trained configuration model to obtain the target configuration file.
7. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the business system configuration method according to any one of claims 1 to 6 when executing the computer program.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the business system configuration method according to any one of claims 1 to 6 are implemented.