Method, apparatus, electronic device, and storage medium for expanding the space of a thin volume

By obtaining historical access data of streamlined volumes and using pre-trained models to predict space growth needs, the problem of unreasonable resource allocation in traditional technologies is solved, precise expansion decisions are achieved, and the performance and adaptability of the storage system are improved.

CN120255822BActive Publication Date: 2025-08-05INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510712595.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-05
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

The traditional streamlined volume space automatic expansion technology relies on fixed strategies and empirical parameters, and lacks dynamic predictions of future data storage needs, resulting in unreasonable resource allocation, affecting storage performance and unable to meet massive data storage needs.

Method used

By obtaining the historical access data of the streamlined volume, extracting multi-dimensional feature information and inputting the pre-trained prediction model, dynamically predicting the required parameters of space growth, and adjusting the expansion granularity for expansion.

Benefits of technology

It realizes accurate allocation of resources, avoids excessive or insufficient, improves storage performance and system forward-lookingness, and meets massive data storage needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255822B_ABST
    Figure CN120255822B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, electronic device and storage medium for expanding the space of a thin volume, which relates to the technical field of storage systems. It includes obtaining historical access data and extracting multi-dimensional feature information, breaking through the traditional single decision dimension that only relies on the current remaining space, and realizing the dynamic characterization of data storage requirements. Using a pre-trained prediction model to analyze the multi-dimensional feature information, dynamically predicting the space growth demand parameters, and adjusting the expansion granularity according to the space growth demand parameters for expansion. Compared with the traditional expansion method that relies on a fixed strategy, the present application can achieve the precise allocation of resources, avoid over-allocation or under-allocation, and thus solve the problems of the traditional technology relying on a fixed strategy, lacking dynamic prediction and foresight.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of storage systems, and particularly to a method, device, electronic device, and storage medium for expanding the space of a thin volume. Background Art

[0002] In the current context of the booming development of big data, cloud computing, and cloud disk technologies, the demand for high-performance and high-availability in massive data storage is increasing day by day. To meet this demand, storage systems widely use the technology of automatically expanding the space of thin volumes. This technology realizes the dynamic management of storage resources by monitoring the usage status of the volume capacity in real time, calculating the expansion requirements according to established policies, and then applying for and allocating storage space from the storage pool.

[0003] However, traditional technologies for automatically expanding the space of thin volumes mostly rely on fixed policies and empirical parameters, and only make decisions based on the current remaining space. They cannot dynamically predict and adjust future data storage requirements, which easily leads to over-allocation or insufficiency of space resources, wasting resources and affecting the performance of business processing. Summary of the Invention

[0004] This application provides a method, device, electronic device, and storage medium for expanding the space of a thin volume, so as to at least solve the problems in related technologies that rely on fixed policies and empirical parameters, only make decisions based on the current remaining space, lack dynamic prediction and foresight of future data storage and access requirements, easily cause unreasonable resource allocation, degradation of storage performance, and inability to meet the requirements of massive data storage.

[0005] This application provides a method for expanding the space of a thin volume, including: obtaining the historical access data of the thin volume to be expanded; extracting feature information from the historical access data, where the feature information includes at least one of time series features, spatial distribution features, multi-volume relationship features, system load status, and scenario features; inputting the feature information into a pre-trained prediction model to obtain the space growth demand parameters output by the prediction model; adjusting the expansion granularity according to the space growth demand parameters, and expanding the thin volume to be expanded based on the adjusted expansion granularity.

[0006] This application also provides a device for expanding the space of a thin volume, including:

[0007] An obtaining module, configured to obtain the historical access data of the thin volume to be expanded;

[0008] A feature extraction module, configured to extract feature information from the historical access data, where the feature information includes at least one of time series features, spatial distribution features, multi-volume relationship features, system load status, and scenario features;

[0009] A model prediction module, configured to input the feature information into a pre-trained prediction model to obtain the space growth demand parameters output by the prediction model;

[0010] An expansion module, configured to adjust the expansion granularity according to the space growth demand parameter, and expand the thin volume to be expanded based on the adjusted expansion granularity.

[0011] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above thin volume space expansion methods when executing the computer program.

[0012] This application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above thin volume space expansion methods are implemented.

[0013] This application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any of the above thin volume space expansion methods are implemented.

[0014] By obtaining historical access data and extracting multi-dimensional feature information, this application breaks through the traditional single decision dimension that only relies on the current remaining space, and realizes the dynamic characterization of data storage requirements; uses a pre-trained prediction model to analyze the multi-dimensional feature information, dynamically predicts the space growth demand parameter, and adjusts the expansion granularity according to the space growth demand parameter for expansion. Compared with the traditional expansion method that relies on a fixed policy, this application can achieve precise resource allocation, avoid over-allocation or under-allocation, and thus solve the problems of the traditional technology relying on a fixed policy, lacking dynamic prediction and foresight. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a schematic diagram of the software and hardware architecture on which the execution of the thin volume space expansion method provided by the embodiment of this application depends;

[0017] Figure 2 It is a schematic flowchart of the thin volume space expansion method provided by the embodiment of this application;

[0018] Figure 3 [[ID= 31]]It is a schematic structural diagram of the thin volume space expansion device provided by the embodiment of this application;

[0019] Figure 4 It is a schematic structural diagram of an electronic device provided by the embodiment of this application. Detailed implementation manners

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0021] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0022] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the technical terms required in the embodiments:

[0023] A thin volume is a storage virtualization technology that allows users to create logical volumes larger than the actual physical storage capacity to meet future growth needs; when creating a thin volume, all storage space is not immediately allocated, but storage resources are dynamically allocated from the storage pool according to the actual data writing requirements. Its advantages include avoiding space waste caused by over-allocation of traditional fully allocated volumes, supporting dynamic expansion to meet business growth needs, reducing initial hardware investment, and lowering storage management costs, and it is widely used in enterprise-level storage systems.

[0024] Thin volumes usually use space auto-expansion technology to achieve dynamic space allocation. Space auto-expansion is an intelligent storage management technology that detects the usage of storage volumes in real time through a monitoring module and is automatically triggered when the available space in the storage system is lower than a preset threshold. The system allocates additional space from the storage pool according to predefined policies, and the expansion process does not require suspending the business, ensuring high service availability.

[0025] For a thin volume, the size of the volume available to the user and the size of the actual occupied storage space are different. The size of the volume available to the user refers to the volume size specified by the user when creating the volume, and the size of the actual occupied space is the size of the physical storage space occupied, which is the space actually allocated by the storage pool to the thin volume. The currently used space is the space where data has been actually written, and the currently available space is the space actually allocated by the storage pool to the thin volume minus the used space.

[0026] In the field of artificial intelligence, large language models specifically refer to deep learning models with extremely large parameter sizes. The number of parameters of such models can range from millions to billions or even more. With such a large number of parameters, they can capture and learn complex data distributions and patterns.

[0027] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0028] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the thin volume space expansion method depends, the specific application environment architecture or specific hardware architecture is described herein.

[0029] like Figure 1 As shown, Figure 1 A diagram of the hardware and software architecture used to implement the streamlined volume expansion method, separating model inference from storage input / output (IO) paths.

[0030] At the hardware level, hard drives are the physical carriers of data storage. Redundant Array of Independent Disks (RAID) technology combines and manages multiple hard drives, improving data storage reliability and read / write performance. Storage pools, built on RAID, consolidate physical storage resources and provide flexible resource allocation units, facilitating subsequent storage space management and allocation. Thin volumes are logical storage units carved out of storage pools. Instead of pre-allocating fixed-size space, they dynamically allocate on demand, improving storage resource utilization.

[0031] At the software level, the intelligent scheduling module uses large models for model inference. The large model can analyze and predict data access patterns and space growth requirements based on historical data and real-time monitoring information. For example, it can predict which data will be frequently accessed (hot data) within a certain time period and which storage areas may face insufficient space. The intelligent scheduling module rationally allocates storage resources based on the results of model inference. The storage IO path is a storage IO request issued by the user host, along the thin volume, storage pool, RAID, and finally to the hard disk for data read and write operations. During this process, the storage IO path focuses on the actual transmission and storage operations of the data, and is separated from the model inference function.

[0032] In the above software and hardware architecture, the intelligent scheduling module performs dynamic resource allocation on the storage I / O path according to the model inference result. The model inference is separated from the storage I / O path, enabling the two to evolve and optimize independently. The storage I / O path focuses on efficiently and stably transmitting and storing data, reducing latency and increasing throughput; the intelligent scheduling module and the large model focus on accurate resource prediction and scheduling strategy formulation, without being affected by the busyness of the storage I / O and thus not affecting the inference performance.

[0033] An embodiment of the present application provides a method for expanding the space of a thin volume. In combination with the execution process of the method for expanding the space of a thin volume, the method is described in detail.

[0034] As Figure 2 shown, the method for expanding the space of a thin volume includes the following steps S201 to S204:

[0035] S201. Obtain the historical access data of the thin volume to be expanded.

[0036] The access data is the data corresponding to the data access requests received by the thin volume to be expanded within a preset period before the current moment.

[0037] S202. Extract feature information from the historical access data. The feature information includes at least one of time series features, spatial distribution features, multi-volume relationship features, system load status, and scenario features.

[0038] The feature information includes at least one of time series features, spatial distribution features, multi-volume relationship features, system load status, and scenario features. Among them, the time series features reflect the periodic law of data access. The multi-volume relationship features reflect the association relationships between volumes, such as backup relationships, snapshot relationships, storage pool ownership relationships (whether belonging to the same storage pool), host mapping relationships (whether mapped to the same host), etc. The system load status includes the usage rate of the Central Processing Unit (CPU), CPU temperature, memory usage rate, wear degree of the Solid State Drive (SSD) storage medium, error rate of the Hard Disk Drive (HDD), etc. The scenario features such as regional holidays, business types, etc.

[0039] The multi-dimensional feature information can more comprehensively reflect the state of the storage system and the context information of data access compared with the traditional method that only depends on the current volume capacity usage, providing a richer basis for the model to accurately predict.

[0040] S203. Input the feature information into a pre-trained prediction model to obtain the spatial growth demand parameters output by the prediction model.

[0041] In some embodiments, the prediction model is deployed locally to enhance real-time performance, reduce the risk of data leakage, and strengthen privacy protection. The model can be run on a model inference acceleration chip or a general-purpose CPU chip of a neural network engine. On the one hand, dedicated hardware acceleration is used to improve the inference speed, and on the other hand, the negative impact of model inference on the input / output processing performance of the storage system is reduced.

[0042] In some embodiments, the training steps of the prediction model include S301~S305:

[0043] S301. Construct a basic model based on a time series model or a language model.

[0044] The time series model is a model based on the Transformer architecture to capture the periodic patterns in the time series.

[0045] The language model is a large language model (LLM). It enables the basic model to have a powerful context modeling ability and better capture the access patterns related to the time series.

[0046] S302. Obtain sample data.

[0047] The sample data includes historical access logs and simulation data.

[0048] The historical access log is a record of the relevant attributes of the data access requests received by the thin volume during the historical period.

[0049] In some embodiments, simulation data is obtained using simulation input / output test software. The simulation test data can be created by creating a preset number of thin volumes on the storage system, mapping these thin volumes to the host, and performing simulation tests using simulation IO test software such as the virtual disk benchmark tool (vdbench). Among them, vdbench is a test tool developed by Oracle for simulating disk I / O loads, mainly used to evaluate the performance and stability of storage systems (such as disks, file systems, etc.).

[0050] The simulation test can be carried out in multiple rounds by changing different IO models to generate diverse simulation data. For example, it is necessary to change the ratio of read / write IO, the size of the operation data of a single IO, the randomness of read / write data, the number of IOs issued per second, the throughput, etc. Among them, the IO model refers to the combination method and characteristics of data read / write operations. These diverse simulation data can simulate the diversity of real business scenarios.

[0051] S303. Preprocess the sample data, including data cleaning, outlier removal, field extraction, and format conversion.

[0052] Clean the sample data to handle missing values and duplicate values in the sample data. Exemplarily, handle missing values by forward filling or interpolation. Check if there are exactly duplicate records in the sample data, and if so, delete the duplicates.

[0053] Remove outliers from the sample data. Outliers can be detected by statistical methods, machine learning methods, etc. If it is confirmed that an outlier is incorrect data, delete it; if an outlier is obviously incorrect, correct it. Statistical methods such as the Z-score method, the Interquartile Range (IQR) method, etc. Machine learning methods such as tree-based algorithms, density-based clustering algorithms, etc. Exemplarily, use the IQR method to detect and delete records with extremely long or short access durations; check and correct obviously incorrect timestamps.

[0054] Field extraction is to extract valuable information from the sample data, such as time information, text information, etc. Time information such as the data access period, the time difference from the first access to the current access, etc. Key information can be extracted from the text fields of the sample data using regular expressions.

[0055] Format conversion is to convert the sample data into a format suitable for the basic model, including data type conversion, standardization, normalization, variable encoding, etc.

[0056] The above preprocessing steps convert the original sample data into a structured dataset suitable for model analysis.

[0057] Divide the preprocessed sample data into a training set, a validation set, and a test set. Time series cross-validation (such as dividing the training / validation set in chronological order) can be adopted to ensure the model's ability to capture temporal dependencies.

[0058] S304. Extract sample feature information from the preprocessed sample data.

[0059] The sample feature information includes at least one of data access request information, statistical information, volume relationship, system load, and scenario information.

[0060] The data access request information includes access time, read / write type, target volume information (including identifier, logical address, physical address), access data size, current used capacity of the target volume, current remaining available capacity of the target volume, and space to be increased for the target volume;

[0061] The statistical information includes the read / write request ratio, the number of read / write requests, the access data size, and the space growth size within a preset time period. The preset time period can be, for example, 5s, 30s, 10mins, 30mins, 1h, etc.

[0062] Volume relationships include association relationships such as backup relationships, snapshot relationships, storage pool ownership relationships, and host mapping relationships between volumes. Extracting volume relationships is to enable the model to learn the access similarity between volumes.

[0063] System load includes CPU usage rate, CPU temperature, memory usage rate, wear level of SSD storage media, HDD error rate, etc.

[0064] Scenario information includes holiday information and business types.

[0065] Exemplarily, sample feature information can be extracted from the preprocessed sample data and integrated into a sample matrix. Each sample corresponds to the state of a certain volume at a certain time point. The sample matrix includes: (1) Time features: hour, week, holiday flag, business type; (2) Volume attributes: volume ID, storage pool ID, associated volume ID, space usage rate; (3) Historical statistics: number of write requests, average write size, space growth amount within each time window; (4) System status: CPU usage rate, memory usage rate, media health; (5) Labels: whether the volume will generate write requests in the future time period. If so, further predict the write address range, write data size, and space growth demand parameters.

[0066] In the process of extracting sample feature information from the preprocessed sample data, it includes dividing time windows, processing associated volume features, embedding scenario features, etc. When dividing time windows, exemplarily, the data is sliced at a fixed time interval such as 1 minute. Each window generates a sample, including the features at the start time of the window and the label of the next window, such as predicting write requests in the next 1 minute. When processing associated volume features, if the access patterns of the target volume and the associated volume are strongly correlated, the historical features of the associated volume, such as the write size in the past 10 minutes, can be used as cross features and input into the model. When embedding scenario features, an Embedding Layer is used to map high-dimensional categorical features such as business types and regions into low-dimensional vectors and fuse them with numerical features.

[0067] S305. Adjust the parameters of the basic model according to the sample feature information to obtain a prediction model.

[0068] The sample feature information integrates time series features, spatial distribution features, multi-volume relationship features, system load status, and scenario features, comprehensively describing the influencing factors of lean volume expansion, thereby improving prediction accuracy. The multi-dimensional sample feature information restricts each other during the model training process. For example, if the remaining space of the volume is insufficient, the space growth demand parameter is increased to correct the prediction result and avoid underestimation. Another example is that during holidays, the write request probability is adjusted according to historical write requests.

[0069] In the process of adjusting the parameters of the basic model according to the sample feature information to obtain the prediction model, input the sample feature information, and the prediction model outputs the write request probability in the middle. If the write request probability is greater than the preset threshold, the prediction model is triggered to predict the write address, data size, and spatial growth demand parameters.

[0070] During the training process, the evaluation metrics of the prediction model include, but are not limited to: precision, recall, F1 value, root mean square error (RMSE), and mean absolute error (MAE).

[0071] When the model is iterated, it can be trained in groups according to the volume type to improve the pertinence. The attention mechanism can be introduced to make the prediction model focus on key features, such as the time feature sequence.

[0072] The above embodiments construct a diversified and standardized training set by generating simulation data, covering diversified application scenarios. It is also beneficial to accelerate the model iteration cycle, improve the model training efficiency and the generalization ability of the model. The simulation data can accurately depict the underlying behavior of the storage system, enabling the model to capture the detailed features that are easily overlooked in the real environment, and improving the real-time performance and accuracy of the model prediction.

[0073] The prediction model also outputs the data access pattern within the target time period, including the target volume, the number of write requests, the location of the target volume, and the size of the written data. It is used to optimize resource allocation. The traditional resource allocation method relies on the current usage situation and preset policies, lacking foresight for future data access requirements, having obvious lag, being difficult to achieve pre-emptive resource optimization configuration, ultimately resulting in a decline in storage performance and being unable to fully meet the growing demand for massive data storage. While this application uses the prediction model to infer the data access pattern within the future target time period, improving the storage performance, the foresight and intelligent level of the resource allocation of the storage system.

[0074] In some embodiments, the actual data access pattern is determined within the target time period, and the prediction accuracy is calculated according to the actual data access pattern and the data access pattern output by the prediction model. In the case where the prediction accuracy is less than the preset threshold, the prediction model is updated. Thus, based on the actual data access pattern, the prediction model is updated to adapt to the change of the access pattern.

[0075] It can be understood that the actual data access pattern of the target time period and the data access pattern of the target time period predicted by the previous prediction model are matched to form a control sample. The accuracy of the model prediction is evaluated. If the prediction accuracy is less than the preset threshold, indicating a low accuracy, the prediction model is updated, including optimizing the model hyperparameters through methods such as grid search and Bayesian optimization, and updating metadata such as the model version number, training time, and evaluation metrics.

[0076] Based on the closed-loop mechanism, the model can dynamically evaluate its own performance according to the real-time data access pattern, automatically trigger an update when the accuracy drops, and continuously adapt to the changes in the access pattern of the storage system and business requirements, ensuring the reliability of space expansion prediction.

[0077] In some embodiments, based on the data access pattern of the target time period output by the prediction model, the address blocks with high probability of access are screened out, and the corresponding address blocks are loaded into the cache, thereby reducing latency and improving read / write performance.

[0078] In some embodiments, based on the data access pattern of the target time period output by the prediction model, it is determined whether there are abnormalities, including high-frequency access abnormalities, abnormal concentration of destination addresses, read / write ratio abnormalities, etc. Thus, early warnings of storage system failures or security threats can be given.

[0079] In some embodiments, based on the data access pattern of the target time period output by the prediction model, a heat level is defined for the data according to the access probability, and the data is divided into hot data, warm data, and cold data. Exemplarily, the access probability > 50% is hot data, the access probability 10% - 50% is warm data, and the access probability < 10% is cold data; according to the heat level, the corresponding data is migrated to the corresponding medium, for example, hot data is stored in SSD, warm data is stored in HDD, and cold data is stored in a tape library or cloud storage.

[0080] In some embodiments, it is detected whether the prediction model outputs a space growth demand parameter within a preset duration; if not, the thin volume to be expanded is expanded according to a preset expansion strategy. The preset duration starts counting after determining the initial expansion granularity, and is used to balance the expected inference time of the model and the business's tolerance for the timeliness of expansion. The preset expansion strategy can determine the expansion granularity according to a fixed ratio of the written data volume, or determine the expansion granularity according to the current usage of the thin volume. For example, when the usage rate of the thin volume exceeds 80%, the expansion granularity is a fixed size of 1GB.

[0081] Specifically, after determining the initial expansion granularity according to the preset expansion strategy, a countdown is started, and the output of the prediction model is continuously monitored. Before the countdown ends, if the prediction model outputs a space growth demand parameter, the countdown is immediately terminated, and the expansion granularity is adjusted according to this space growth demand parameter. If the countdown ends and the space growth demand parameter output by the prediction model is still not obtained, it indicates that the model inference takes too long or fails, then give up waiting and expand the thin volume to be expanded according to the initial expansion granularity according to the preset expansion strategy.

[0082] In the above embodiments, to address the possible failures or excessive time consumption in model inference and ensure timely space expansion, the fault tolerance mechanism of the thin volume module will work in coordination with the static expansion strategy and the countdown mechanism to build a fault tolerance mechanism, avoiding a decline in storage performance or a backlog of write requests caused by waiting for too long.

[0083] S204. Expand the thin volume to be expanded by adjusting the expansion granularity according to the space growth demand parameter.

[0084] In some embodiments, divide the space growth demand parameter by the used capacity of the thin volume to be expanded, round up to obtain the target multiple; then, multiply the target multiple by the minimum unit of the storage pool space to determine the expansion granularity. Further, first detect whether the free space in the storage pool meets this expansion granularity. If so, call the expansion interface to allocate storage space to the thin volume to be expanded according to this expansion granularity for expansion.

[0085] Among them, the minimum unit of the storage pool space is the smallest unit for managing the storage pool space, and its size is determined by the underlying architecture of the storage pool and can be queried and modified through a specific API interface, such as 128MB, 1GB. Determine the target multiple according to the space growth demand parameter. If the target multiple is not an integer, round up to determine an integer multiple. It can be understood that the space growth demand parameter is rounded up to an integer multiple of the minimum unit of the storage pool space.

[0086] Optionally, a minimum target multiple threshold is set. If the calculated target multiple is lower than the minimum target multiple threshold, adjust the target multiple to the minimum target multiple threshold. During the process of calculating the expansion granularity, if the calculated value is not an integer multiple of the minimum unit of the storage pool space, round up to obtain the final expansion granularity. During expansion, detect whether the free space in the storage pool meets the expansion granularity. It can be understood to check whether there is enough free space in the storage pool to meet the expansion strength requirement. If so, call the expansion interface to allocate storage space to the thin volume to be expanded according to the determined expansion granularity for expansion. If not, it means that the free space in the storage pool is insufficient, then issue an alarm and suspend the expansion operation, waiting for the administrator to supplement storage resources.

[0087] The above embodiments allocate in integer multiples, which can ensure that the space is fully utilized, avoid the generation of fragmented space, and improve the utilization rate of storage resources. In data storage and management, following this expansion granularity for space allocation is beneficial to shorten the expansion time and can also ensure the consistency of space usage among different thin volumes.

[0088] In some embodiments, after multiplying the target multiple by the minimum unit of the storage pool space to determine the expansion granularity, and before detecting whether the free space of the storage pool meets the expansion granularity, the following steps are further included: monitoring the system load, and determining whether the system load is less than or equal to a preset load threshold; if so, reducing the expansion granularity; if not, increasing the expansion granularity. Adjustment factors corresponding to different predefined load levels can be used to appropriately increase or decrease the expansion granularity based on the adjustment factors.

[0089] The above embodiments enable the system to reduce the expansion granularity to avoid space waste under low load; when the system is under high load, increase the expansion granularity, expand the thin capacity in a large amount at one time, reduce the IO processing waiting caused by insufficient space, and reduce the overhead brought by the expansion frequency.

[0090] In some embodiments, if the space capacity of the thin volume after expansion is greater than the user-available capacity, the expansion granularity is reduced to the minimum processing space of a single input / output. The user-available capacity is the upper limit of the available capacity of the volume set by the user, such as 10TB. The minimum processing space of a single input / output refers to the minimum space required to process a single IO, such as 64KB. This is to prevent space waste caused by excessive expansion of unused space. For example, if the currently used capacity of the volume is 9.5TB, the user upper limit is 10TB, and the remaining available capacity is 0.5TB, if the expansion is carried out according to the original expansion granularity, it will exceed the limit, then recalculate according to the minimum space of a single IO to ensure that the capacity after expansion does not exceed 10TB.

[0091] Based on the space growth demand parameters obtained by model prediction, the above embodiments dynamically adjust the expansion granularity in combination with the space management rules of the storage system, which not only meets the requirements of standardized management but also avoids space waste caused by excessive one-time expansion.

[0092] In summary, the present application provides a method for expanding the space of a thin volume. The method extracts multi-dimensional features of historical access data, analyzes the time distribution of historical access data through time series features to identify the growth timing pattern; anticipates the space demand of hot spots through spatial distribution features; uses the relationship between volumes to plan the expansion of associated volumes through multi-volume relationship features; analyzes the impact of historical access data on the system business performance through the system load status; and identifies business scenarios through scenario features to match different spatial growth patterns in different scenarios, making the expansion decision closer to the real business needs. By analyzing multi-dimensional feature information through a prediction model, the space growth demand parameters for future time periods are output, so as to adjust the expansion granularity as needed and intelligently drive the automatic expansion of the thin volume, solving the problems of resource waste and performance fluctuation caused by the lack of foresight in traditional technologies.

[0093] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0094] An embodiment of the present application also provides a device for expanding the capacity of a thin volume, as Figure 3 shown. The device includes:

[0095] An acquisition module 311, configured to acquire historical access data of the thin volume to be expanded;

[0096] A feature extraction module 312, configured to extract feature information from the historical access data, where the feature information includes at least one of time series features, spatial distribution features, multi-volume relationship features, system load status, and scenario features;

[0097] A model prediction module 313, configured to input the feature information into a pre-trained prediction model to obtain a spatial growth demand parameter output by the prediction model;

[0098] An expansion module 314, configured to adjust the expansion granularity according to the spatial growth demand parameter, and expand the thin volume to be expanded based on the adjusted expansion granularity.

[0099] As an optional implementation manner provided by an embodiment of the present application, the device further includes a training module for training the prediction model, including: constructing a basic model based on a time series model or a language model; acquiring sample data, where the sample data includes historical access logs and simulated data, and the simulated data is access logs generated by a simulated input-output test software; extracting sample feature information from the preprocessed sample data; and adjusting the parameters of the basic model according to the sample feature information to obtain a prediction model.

[0100] As an optional implementation manner provided by an embodiment of the present application, the prediction model also outputs a data access pattern for a target time period;

[0101] The device further includes an update module for updating the prediction model, including: determining an actual data access pattern in the target time period; calculating a prediction accuracy according to the actual data access pattern and the data access pattern output by the prediction model; and updating the prediction model when the prediction accuracy is less than a preset threshold.

[0102] As an optional implementation provided by the embodiments of the present application, the expansion module 314 is specifically configured to: divide the space growth demand parameter by the used capacity of the thin volume to be expanded, round up to obtain the target multiple; multiply the target multiple by the minimum unit of the storage pool space to determine the expansion granularity; detect whether the free space of the storage pool meets the expansion granularity; if so, call the expansion interface to allocate storage space to the thin volume to be expanded according to the expansion granularity for expansion.

[0103] As an optional implementation provided by the embodiments of the present application, after multiplying the target multiple by the minimum unit of the storage pool space to determine the expansion granularity and before detecting whether the free space of the storage pool meets the expansion granularity, the expansion module 314 is further configured to: monitor the system load; reduce the expansion granularity when the system load is less than or equal to the preset load threshold; increase the expansion granularity when the system load is greater than the preset load threshold.

[0104] As an optional implementation provided by the embodiments of the present application, the expansion module 314 is further configured to:

[0105] When the space capacity of the expanded thin volume is greater than the user available capacity, reduce the expansion granularity to the minimum processing space of a single input / output.

[0106] As an optional implementation provided by the embodiments of the present application, the expansion module is further configured to: detect whether the prediction model outputs a space growth demand parameter within a preset duration; when the prediction model does not output a space growth demand parameter within the preset duration, expand the thin volume to be expanded according to the preset expansion strategy.

[0107] The present application provides an expansion device for the thin volume space. The device extracts multi-dimensional features of historical access data, analyzes the time distribution of historical access data through time series features to identify the growth timing pattern; anticipates the space demand of hot spots through spatial distribution features; utilizes the relationship between volumes through multi-volume relationship features to plan the expansion of associated volumes; analyzes the impact of historical access data on the system business performance through the system load status; identifies business scenarios through scenario features to match the space growth pattern under different scenarios, making the expansion decision closer to the real business needs. By analyzing multi-dimensional feature information through a prediction model, the space growth demand parameter for the future time period is output, so as to adjust the expansion granularity as needed and intelligently drive the automatic expansion of the thin volume, solving the problems of resource waste and performance fluctuation caused by the lack of foresight in traditional technologies.

[0108] For the description of the features in the embodiments corresponding to the expansion device for the thin volume space, reference can be made to the relevant descriptions in the embodiments corresponding to the expansion method for the thin volume space, which will not be elaborated here one by one.

[0109] Embodiments of the present application also provide an electronic device, such as Figure 4As shown in the figure, it includes a memory 401 and a processor 402. A computer program is stored in the memory 401, and the processor 402 is configured to run the computer program to execute the steps in any of the above embodiments of the method for expanding the space of a thin volume.

[0110] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the method for expanding the space of a thin volume when running.

[0111] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), external hard drives, magnetic disks, or optical discs, etc., various media that can store computer programs.

[0112] An embodiment of the present application also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the method for expanding the space of a thin volume.

[0113] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the method for expanding the space of a thin volume.

[0114] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0115] The above has introduced in detail a method, device, electronic device, and storage medium for expanding the space of a thin volume provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for expanding thin volume space, characterized in that: include: Obtain historical access data for the thin volume to be expanded; extracting feature information from the historical access data, the feature information including at least one of a time series feature, a spatial distribution feature, a multi-volume relationship feature, a system load state, and a scene feature; Inputting the feature information into a pre-trained prediction model to obtain a spatial growth demand parameter output by the prediction model; Adjusting the expansion granularity according to the space growth requirement parameter, and expanding the thin volume to be expanded based on the adjusted expansion granularity; Among them, adjusting the expansion granularity according to the space growth demand parameter includes: dividing the space growth demand parameter by the used capacity of the thin volume to be expanded, rounding up to obtain a target multiple; multiplying the target multiple by the minimum unit of storage pool space to determine the expansion granularity.

2. The method according to claim 1, characterized in that The training process of the prediction model includes: Build a basic model based on time series model or language model; Obtaining sample data, the sample data including historical access logs and simulated data, the simulated data being access logs generated by simulated input and output testing software; extracting sample feature information from the preprocessed sample data; The prediction model is obtained by adjusting the parameters of the basic model according to the sample feature information.

3. The method according to claim 1, characterized in that The prediction model also outputs data access patterns for a target time period; The updating process of the prediction model includes: determining actual data access patterns during the target time period; Calculating prediction accuracy based on the actual data access pattern and the data access pattern output by the prediction model; When the prediction accuracy is less than a preset threshold, the prediction model is updated.

4. The method according to claim 1, wherein Expanding the thin volume to be expanded based on the adjusted expansion granularity includes: Checking whether the free space in the storage pool meets the expansion granularity; If so, the expansion interface is called to allocate storage space to the thin volume to be expanded according to the expansion granularity for expansion.

5. The method according to claim 1, wherein After multiplying the target multiple by the minimum unit of storage pool space to determine the expansion granularity, the method further includes: Monitor system load; When the system load is less than or equal to a preset load threshold, reducing the expansion granularity; When the system load is greater than the preset load threshold, the expansion granularity is increased.

6. The method according to claim 1, characterized in that After adjusting the expansion granularity according to the space growth requirement parameter and expanding the thin volume to be expanded based on the adjusted expansion granularity, the method further includes: When the capacity of the thin volume space after expansion is greater than the user's available capacity, the expansion granularity is reduced to the minimum processing space of single input and output.

7. The method according to claim 1, characterized in that After extracting feature information from the historical access data, the method further includes: Detecting whether the prediction model outputs space growth demand parameters within a preset time period; When the prediction model does not output the space growth requirement parameter within the preset time period, the thin volume to be expanded is expanded according to a preset expansion strategy.

8. A device for expanding thin volume space, characterized in that: include: An acquisition module is used to obtain historical access data of the thin volume to be expanded; a feature extraction module, configured to extract feature information from the historical access data, wherein the feature information includes at least one of a time series feature, a spatial distribution feature, a multi-volume relationship feature, a system load state, and a scene feature; A model prediction module, configured to input the feature information into a pre-trained prediction model to obtain a spatial growth demand parameter output by the prediction model; An expansion module, configured to adjust the expansion granularity according to the space growth requirement parameter, and expand the thin volume to be expanded based on the adjusted expansion granularity; The expansion module is specifically used to divide the space growth demand parameter by the used capacity of the thin volume to be expanded when adjusting the expansion granularity according to the space growth demand parameter, round up to obtain a target multiple; multiply the target multiple by the minimum unit of the storage pool space to determine the expansion granularity.

9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor is configured to implement the steps of the method for expanding thin volume space as claimed in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the method for expanding thin volume space according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Capacity expansion method, prediction model creation method and device, equipment and medium

    CN109885469A