Simple volume space expansion method and device, electronic equipment and storage medium
By obtaining historical access data of thin volumes and using pre-trained models to dynamically predict the space growth needs, the problem of unreasonable resource allocation in traditional technologies is solved, and more efficient storage performance and resource utilization are achieved.
Patent Information
- Application Number
- CN202510712595.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The traditional streamlined volume space automatic expansion technology relies on fixed strategies and empirical parameters, and lacks dynamic predictions of future data storage needs, resulting in unreasonable resource allocation, affecting storage performance and unable to meet massive data storage needs.
By obtaining the historical access data of the volume to be expanded, multi-dimensional feature information is extracted and pre-trained prediction model is input, dynamically predicting the required parameters of space growth, and adjusting the expansion granularity for expansion.
It realizes precise allocation of resources, avoids excessive or insufficient, improves storage performance and resource utilization, and meets massive data storage needs.
Smart Images

Figure CN120255822A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of storage systems, and particularly to a method, device, electronic device, and storage medium for expanding the space of a thin volume. Background Art
[0002] At present, with the booming development of big data, cloud computing, and cloud disk technologies, the demand for high-performance and high-availability in massive data storage is increasing day by day. To meet this demand, storage systems widely use the technology of automatically expanding the space of thin volumes. This technology realizes the dynamic management of storage resources by monitoring the usage status of the volume capacity in real time, calculating the expansion requirements according to established policies, and then applying for and allocating storage space from the storage pool.
[0003] However, traditional technologies for automatically expanding the space of thin volumes mostly rely on fixed policies and empirical parameters, and only make decisions based on the current remaining space. They cannot dynamically predict and adjust future data storage requirements, which easily leads to over-allocation or shortage of space resources, wasting resources and affecting business processing performance. Summary of the Invention
[0004] This application provides a method, device, electronic device, and storage medium for expanding the space of a thin volume, so as to at least solve the problems in related technologies that rely on fixed policies and empirical parameters, only make decisions based on the current remaining space, lack dynamic prediction and forward-looking of future data storage and access requirements, easily cause unreasonable resource allocation, decline in storage performance, and inability to meet the requirements of massive data storage.
[0005] This application provides a method for expanding the space of a thin volume, including: obtaining historical access data of the thin volume to be expanded; extracting feature information from the historical access data, where the feature information includes at least one of time series features, spatial distribution features, multi-volume relationship features, system load status, and scenario features; inputting the feature information into a pre-trained prediction model to obtain the space growth demand parameters output by the prediction model; adjusting the expansion granularity according to the space growth demand parameters, and expanding the thin volume to be expanded based on the adjusted expansion granularity.
[0006] This application also provides a device for expanding the space of a thin volume, including: An obtaining module, configured to obtain historical access data of the thin volume to be expanded; A feature extraction module, configured to extract feature information from the historical access data, where the feature information includes at least one of time series features, spatial distribution features, multi-volume relationship features, system load status, and scenario features; A model prediction module, configured to input the feature information into a pre-trained prediction model to obtain the space growth demand parameters output by the prediction model; An expansion module, configured to adjust the expansion granularity according to the space growth demand parameter, and expand the thin volume to be expanded based on the adjusted expansion granularity.
[0007] The present application further provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any one of the above thin volume space expansion methods when executing the computer program.
[0008] The present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above thin volume space expansion methods are implemented.
[0009] The present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any one of the above thin volume space expansion methods are implemented.
[0010] By obtaining historical access data and extracting multi-dimensional feature information, the present application breaks through the traditional single decision dimension that only depends on the current remaining space, and realizes the dynamic characterization of data storage requirements; uses a pre-trained prediction model to analyze the multi-dimensional feature information, dynamically predicts the space growth demand parameter, and adjusts the expansion granularity according to the space growth demand parameter for expansion. Compared with the traditional expansion method that depends on a fixed strategy, the present application can achieve precise allocation of resources, avoid over-allocation or under-allocation, and thus solve the problems of the traditional technology depending on a fixed strategy, lacking dynamic prediction and foresight. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 It is a schematic diagram of the software and hardware architecture on which the execution of an expansion method for a thin volume space provided by an embodiment of the present application depends; Figure 2 It is a schematic flowchart of an expansion method for a thin volume space provided by an embodiment of the present application; Figure 3 It is a schematic structural diagram of an expansion device for a thin volume space provided by an embodiment of the present application; Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the protection scope of the present application.
[0014] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0015] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the technical terms required in the embodiments: A thin volume is a storage virtualization technology that allows users to create logical volumes larger than the actual physical storage capacity to meet future growth needs; when creating a thin volume, all storage space is not immediately allocated, but storage resources are dynamically allocated from the storage pool according to the actual data writing requirements. Its advantages lie in avoiding space waste caused by over-allocation of traditional fully allocated volumes, supporting dynamic expansion to meet business growth needs, reducing initial hardware investment, and lowering storage management costs, and it is widely used in enterprise-level storage systems.
[0016] Thin volumes usually use space auto-expansion technology to achieve dynamic space allocation. Space auto-expansion is an intelligent storage management technology that real-time detects the usage of storage volumes through a monitoring module and is automatically triggered when the available space in the storage system is lower than a preset threshold. The system allocates additional space from the storage pool according to predefined policies, and the expansion process does not require pausing the business, ensuring high service availability.
[0017] For thin volumes, the size of the volume available to the user and the actual occupied storage space size are different. The size of the volume available to the user refers to the volume size specified by the user when creating the volume, and the actual occupied space size is the physical storage space size occupied, which is the space actually allocated by the storage pool to the thin volume. The currently used space is the space where data has actually been written, and the currently available space is the space actually allocated by the storage pool to the thin volume minus the used space. In the field of artificial intelligence, large language models specifically refer to those deep learning models with extremely large numbers of parameters. The number of parameters of such models can range from millions to billions or more. With the huge number of parameters, they can capture and learn complex data distributions and patterns.
[0018] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.
[0019] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the expansion method of the thin-provisioned volume space depends, the specific application environment architecture or specific hardware architecture will be described herein.
[0020] As Figure 1 shown, Figure 1 It is a schematic diagram of the software and hardware architecture on which the execution of the expansion method of the thin-provisioned volume space depends. The model inference is separated from the storage input / output (IO) path.
[0021] At the hardware level, the hard disk is the physical carrier for data storage. The Redundant Array of Independent Disks (RAID) technology combines and manages multiple hard disks, which can improve the reliability and read / write performance of data storage. The storage pool is built based on RAID, integrating physical storage resources, providing flexible resource allocation units, and facilitating subsequent management and allocation of storage space. The thin-provisioned volume is a logical storage unit divided on the basis of the storage pool. It does not pre-allocate a fixed-size space, but dynamically allocates space on demand, improving the utilization rate of storage resources.
[0022] At the software level, the intelligent scheduling module performs model inference with the help of a large model. The large model can analyze and predict data access patterns, space growth requirements, etc. based on historical data, real-time monitoring information, etc. For example, it can predict which data will be frequently accessed (hot data) within a certain time period, and which storage areas may face insufficient space. The intelligent scheduling module reasonably allocates storage resources according to the model inference results. The storage IO path is the storage IO request sent by the user host, which travels along the thin-provisioned volume, the storage pool, and RAID, and finally reaches the hard disk for data read / write operations. In this process, the storage IO path focuses on the actual transmission and storage operations of data, and is separated from the model inference function.
[0023] In the above software and hardware architecture, the intelligent scheduling module dynamically allocates resources on the storage IO path according to the model inference results. The separation of model inference and the storage IO path enables the two to evolve and optimize independently. The storage IO path focuses on efficiently and stably transmitting and storing data, reducing latency and increasing throughput; the intelligent scheduling module and the large model focus on accurate resource prediction and scheduling strategy formulation, and will not affect the inference performance due to the busyness of storage IO.
[0024] The embodiments of this application provide a method for expanding the space of a thin-provisioned volume. The method will be described in detail in combination with the execution process of the method for expanding the space of the thin-provisioned volume.
[0025] As Figure 2 shown, the method for expanding the space of a thin volume includes the following steps S201 to S204: S201. Obtain the historical access data of the thin volume to be expanded.
[0026] The access data is the data corresponding to the data access requests received by the thin volume to be expanded within a preset time period before the current moment.
[0027] S202. Extract feature information from the historical access data, where the feature information includes at least one of time series features, spatial distribution features, multi-volume relationship features, system load status, and scenario features.
[0028] The feature information includes at least one of time series features, spatial distribution features, multi-volume relationship features, system load status, and scenario features. Among them, the time series features reflect the periodic law of data access. The multi-volume relationship features reflect the association relationships between volumes, such as backup relationships, snapshot relationships, storage pool ownership relationships (whether they belong to the same storage pool), host mapping relationships (whether they are mapped to the same host), etc. The system load status includes the usage rate of the Central Processing Unit (CPU), CPU temperature, memory usage rate, wear degree of the Solid State Drive (SSD) storage medium, error rate of the Hard Disk Drive (HDD), etc. The scenario features are such as regional holidays, business types, etc.
[0029] The multi-dimensional feature information can more comprehensively reflect the state of the storage system and the context information of data access compared with the traditional method that only relies on the current volume capacity usage, providing a richer basis for the model to accurately predict.
[0030] S203. Input the feature information into a pre-trained prediction model to obtain the space growth demand parameters output by the prediction model.
[0031] In some embodiments, the prediction model is deployed locally, which enhances real-time performance, reduces the risk of data leakage, and strengthens privacy protection. The model can be run on a model inference acceleration chip or a general CPU chip of a neural network engine. On the one hand, it utilizes dedicated hardware acceleration to improve the inference speed, and on the other hand, it reduces the negative impact of model inference on the input / output processing performance of the storage system.
[0032] In some embodiments, the training steps of the prediction model include S301 to S305: S301. Build a basic model based on a time series model or a language model.
[0033] Time series models, such as models based on the Transformer architecture, are used to capture periodic patterns in time series.
[0034] Language models, such as large language models (LLMs), enable the base model to have powerful context modeling capabilities and be more capable of capturing access patterns related to time series.
[0035] S302. Obtain sample data.
[0036] The sample data includes historical access logs and simulated data.
[0037] The historical access log is a record of the relevant attributes of the data access requests received by the thin volume during a historical period.
[0038] In some embodiments, the simulated data is obtained using simulation input-output test software. The simulated test data can be obtained by creating a preset number of thin volumes on the storage system, mapping these thin volumes to the host, and performing simulation tests using simulation IO test software such as the virtual disk benchmark tool (vdbench). Among them, vdbench is a test tool developed by Oracle for simulating disk I / O loads, mainly used to evaluate the performance and stability of storage systems (such as disks, file systems, etc.).
[0039] The simulation test can perform multiple rounds of tests by varying different IO models to generate diverse simulated data. For example, it is necessary to vary the ratio of read-write IOs, the size of the operation data for a single IO, the randomness of read-write data, the number of IOs issued per second, the throughput, etc. Among them, the IO model refers to the combination method and characteristics of data read-write operations. These diverse simulated data can simulate the diversity of real business scenarios.
[0040] S303. Preprocess the sample data, including data cleaning, outlier removal, field extraction, and format conversion.
[0041] Perform data cleaning on the sample data to handle missing values and duplicate values in the sample data. Exemplarily, the missing values are processed by forward filling or interpolation. Check whether there are exactly duplicate records in the sample data, and if so, delete the duplicates.
[0042] To remove outliers from the sample data, statistical methods, machine learning methods, etc. can be used to detect outliers. If an outlier is confirmed to be incorrect data, it is deleted; if the outlier is clearly incorrect, it is corrected. Statistical methods such as the Z-score method and the Interquartile Range (IQR) method. Machine learning methods such as tree-based algorithms and density-based clustering algorithms. Exemplarily, the IQR method is used to detect and delete records with abnormally long or short access durations; check and correct clearly incorrect timestamps.
[0043] Field extraction is to extract valuable information from the sample data, such as time information, text information, etc. Time information such as the data access period, the time difference from the first access to the current access, etc. Regular expressions can be used to extract key information from the text fields of the sample data.
[0044] Format conversion is to convert the sample data into a format suitable for the basic model, including data type conversion, standardization, normalization, variable encoding, etc.
[0045] The above preprocessing steps convert the original sample data into a structured dataset suitable for model analysis.
[0046] The preprocessed sample data is divided into a training set, a validation set, and a test set. Time series cross-validation (such as dividing the training / validation set in chronological order) can be adopted to ensure the model's ability to capture temporal dependencies.
[0047] S304. Extract sample feature information from the preprocessed sample data.
[0048] The sample feature information includes at least one of data access request information, statistical information, volume relationships, system load, and scenario information.
[0049] The data access request information includes access time, read / write type, target volume information (including identifier, logical address, physical address), access data size, current used capacity of the target volume, current remaining available capacity of the target volume, and space to be increased for the target volume; The statistical information includes the read / write request ratio, the number of read / write requests, the access data size, and the space growth size within a preset time period. The preset time period can be, for example, 5s, 30s, 10mins, 30mins, 1h, etc.
[0050] The volume relationships include association relationships such as backup relationships, snapshot relationships, storage pool ownership relationships, and host mapping relationships between volumes. Extracting volume relationships is to enable the model to learn the access similarities between volumes.
[0051] The system load includes CPU usage rate, CPU temperature, memory usage rate, wear level of SSD storage media, HDD error rate, etc.
[0052] The scenario information includes holiday information and business type.
[0053] Exemplarily, sample feature information is extracted from the preprocessed sample data and can be integrated into a sample matrix. Each sample corresponds to the state of a certain volume at a certain time point. The sample matrix includes: (1) Time features: hour, week, holiday flag, business type; (2) Volume attributes: volume ID, storage pool ID, associated volume ID, space usage rate; (3) Historical statistics: number of write requests, average write size, space growth amount within each time window; (4) System status: CPU usage rate, memory usage rate, media health; (5) Label: Whether the volume will generate write requests in the future time period. If so, further predict the write address range, write data size, and space growth demand parameters.
[0054] In the process of extracting sample feature information from the preprocessed sample data, it includes dividing time windows, processing associated volume features, embedding scenario features, etc. When dividing time windows, exemplarily, the data is sliced at a fixed time interval such as 1 minute. Each window generates a sample, including the features at the start time of the window and the label of the next window, such as predicting write requests in the next 1 minute. When processing associated volume features, if the access pattern of the target volume is strongly correlated with that of the associated volume, the historical features of the associated volume such as the write size in the past 10 minutes can be used as cross features and input into the model. When embedding scenario features, the Embedding Layer is used to map high-dimensional categorical features such as business type and region into low-dimensional vectors and fuse them with numerical features.
[0055] S305. Adjust the parameters of the basic model according to the sample feature information to obtain a prediction model.
[0056] The sample feature information fuses time series features, spatial distribution features, multi-volume relationship features, system load status, and scenario features, comprehensively describing the influencing factors of thin volume expansion, thereby improving the prediction accuracy. The multi-dimensional sample feature information restricts each other during the model training process. For example, if the remaining space of the volume is insufficient, the space growth demand parameter is increased to correct the prediction result and avoid underestimation. Another example is that during holidays, the write request probability is adjusted according to historical write requests.
[0057] In the process of adjusting the parameters of the basic model according to the sample feature information to obtain a prediction model, the sample feature information is input, and the prediction model outputs the write request probability in the middle. If the write request probability is greater than the preset threshold, the prediction model is triggered to predict the write address, data size, and space growth demand parameters.
[0058] During the training process, the evaluation metrics of the prediction model include, but are not limited to: precision, recall rate, F1 value, root mean square error (RMSE), and mean absolute error (MAE).
[0059] When the model is iterated, it can be trained in groups according to the volume type to improve the pertinence. The attention mechanism can be introduced to make the prediction model focus on key features, such as the time feature sequence.
[0060] The above embodiments construct a diversified and standardized training set by generating simulation data, covering diversified application scenarios. It is also beneficial to accelerate the model iteration cycle, improve the model training efficiency and the generalization ability of the model. The simulation data can accurately depict the underlying behavior of the storage system, enabling the model to capture the detailed features that are easily overlooked in the real environment, and improving the real-time performance and accuracy of the model prediction.
[0061] The prediction model also outputs the data access pattern within the target time period, including the target volume, the number of write requests, the location of the target volume, and the size of the written data. It is used to optimize resource allocation. The traditional resource allocation method relies on the current usage status and preset policies, lacking foresight for future data access requirements, having obvious lag, and being difficult to achieve pre-emptive resource optimization configuration, ultimately resulting in a decline in storage performance and being unable to fully meet the growing demand for massive data storage. However, this application uses the prediction model to infer the data access pattern within the future target time period, improving the storage performance and the foresight and intelligence level of the storage system resource allocation.
[0062] In some embodiments, the actual data access pattern is determined within the target time period, and the prediction accuracy is calculated based on the actual data access pattern and the data access pattern output by the prediction model. When the prediction accuracy is less than the preset threshold, the prediction model is updated. Thus, based on the actual data access pattern, the prediction model is updated to adapt to the change of the access pattern.
[0063] It can be understood that the actual data access pattern in the target time period is matched with the data access pattern of the target time period predicted by the previous prediction model to form a control sample. The accuracy of the model prediction is evaluated. If the prediction accuracy is less than the preset threshold, indicating a low accuracy, the prediction model is updated, including optimizing the model hyperparameters through methods such as grid search and Bayesian optimization, and updating the metadata such as the model version number, training time, and evaluation metrics.
[0064] Based on the closed-loop mechanism, the model in the above embodiments can dynamically evaluate its own performance according to the real-time data access pattern, automatically trigger an update when the accuracy drops, and continuously adapt to the change of the access pattern of the storage system and the business requirements, ensuring the reliability of the space expansion prediction.
[0065] In some embodiments, based on the data access patterns of the target time period output by the prediction model, address blocks with high access probabilities are screened out, and the corresponding address blocks are loaded into the cache, thereby reducing latency and improving read / write performance.
[0066] In some embodiments, based on the data access patterns of the target time period output by the prediction model, it is determined whether there are anomalies, including high-frequency access anomalies, abnormal concentration of destination addresses, abnormal read / write ratios, etc. Thus, early warnings of storage system failures or security threats can be given.
[0067] In some embodiments, based on the data access patterns of the target time period output by the prediction model, a heat level is defined for the data according to the access probability, and the data is divided into hot data, warm data, and cold data. Exemplarily, the access probability > 50% is hot data, the access probability 10% - 50% is warm data, and the access probability < 10% is cold data; according to the heat level, the corresponding data is migrated to the corresponding medium. For example, hot data is stored in SSDs, warm data is stored in HDDs, and cold data is stored in tape libraries or cloud storage.
[0068] In some embodiments, it is detected whether the prediction model outputs a space growth demand parameter within a preset duration; if not, the thin volume to be expanded is expanded according to a preset expansion strategy. The preset duration starts counting after determining the initial expansion granularity, and is used to balance the expected inference time of the model and the business's tolerance for the timeliness of expansion. The preset expansion strategy can determine the expansion granularity according to a fixed ratio of the written data volume, or determine the expansion granularity according to the current usage of the thin volume. For example, when the utilization rate of the thin volume exceeds 80%, the expansion granularity is a fixed size of 1GB.
[0069] Specifically, after determining the initial expansion granularity according to the preset expansion strategy, a countdown is started, and the output of the prediction model is continuously monitored. Before the countdown ends, if the prediction model outputs a space growth demand parameter, the countdown is immediately terminated, and the expansion granularity is adjusted according to the space growth demand parameter. If the countdown ends and the space growth demand parameter output by the prediction model is still not obtained, indicating that the model inference takes too long or fails, then the waiting is abandoned, and the thin volume to be expanded is expanded according to the initial expansion granularity according to the preset expansion strategy. The above embodiments are to address the possible failures or excessive time consumption problems of model inference, ensure timely space expansion, and the fault tolerance mechanism of the thin volume module will work in coordination with the static expansion strategy and the countdown mechanism to build a fault tolerance mechanism to avoid a decline in storage performance or a backlog of write requests caused by waiting too long.
[0070] S204. Adjust the expansion granularity according to the space growth demand parameter to expand the thin volume to be expanded.
[0071] In some embodiments, the space growth demand parameter is divided by the used capacity of the thin volume to be expanded, and the ceiling value is obtained as the target multiple. Then, the target multiple is multiplied by the minimum unit of the storage pool space to determine the expansion granularity. Further, first, it is detected whether the free space in the storage pool meets the expansion granularity. If so, the expansion interface is called to allocate storage space to the thin volume to be expanded according to the expansion granularity for expansion.
[0072] Among them, the minimum unit of the storage pool space is the smallest unit for managing the storage pool space, and its size is determined by the underlying architecture of the storage pool and can be queried and modified through a specific API interface, such as 128 MB or 1 GB. The target multiple is determined according to the space growth demand parameter. If the target multiple is not an integer, the ceiling value is taken to determine an integer multiple. It can be understood that the space growth demand parameter is rounded up to an integer multiple of the minimum unit of the storage pool space.
[0073] Optionally, a minimum target multiple threshold is set. If the calculated target multiple is lower than the minimum target multiple threshold, the target multiple is adjusted to the minimum target multiple threshold. During the process of calculating the expansion granularity, if the calculated value is not an integer multiple of the minimum unit of the storage pool space, the ceiling value is taken to obtain the final expansion granularity. During expansion, it is detected whether the free space in the storage pool meets the expansion granularity. It can be understood that it is checked whether there is enough free space in the storage pool to meet the expansion strength requirement. If so, the expansion interface is called to allocate storage space to the thin volume to be expanded according to the determined expansion granularity for expansion. If not, it means that the free space in the storage pool is insufficient, and an alarm is issued and the expansion operation is suspended, waiting for the administrator to supplement storage resources.
[0074] The above embodiments allocate in integer multiples, which can ensure that the space is fully utilized, avoid the generation of fragmented space, and improve the utilization rate of storage resources. In data storage and management, following this expansion granularity for space allocation is beneficial to shortening the expansion time and ensuring the consistency of space usage among different thin volumes.
[0075] In some embodiments, after multiplying the target multiple by the minimum unit of the storage pool space to determine the expansion granularity, and before detecting whether the free space in the storage pool meets the expansion granularity, it further includes: monitoring the system load, and determining whether the system load is less than or equal to a preset load threshold; if so, reducing the expansion granularity; if not, increasing the expansion granularity. Adjustment factors corresponding to different load levels can be predefined to appropriately increase or decrease the expansion granularity based on the adjustment factors.
[0076] The above embodiments enable the system to reduce the expansion granularity to avoid space waste when the system load is low; when the system load is high, increase the expansion granularity, greatly expanding the thin capacity in a single time, reducing the IO processing waiting caused by insufficient space, and reducing the overhead brought by the expansion frequency.
[0077] In some embodiments, if the capacity of the thin volume after expansion is greater than the user-available capacity, the expansion granularity is reduced to the minimum processing space for a single input / output. The user-available capacity is the upper limit of the volume available capacity set by the user, such as 10TB. The minimum processing space for a single input / output refers to the minimum space required to process a single IO, such as 64KB. This prevents waste of space caused by excessive expansion of unused space. For example, if the currently used capacity of the volume is 9.5TB, the user upper limit is 10TB, and the remaining available capacity is 0.5TB, and expansion according to the original expansion granularity will exceed the limit, then it is recalculated according to the minimum space for a single IO to ensure that the capacity after expansion does not exceed 10TB.
[0078] Based on the space growth demand parameters obtained by model prediction in the above embodiments, in combination with the space management rules of the storage system, the expansion granularity is dynamically adjusted, which not only meets the requirements of standardized management but also avoids waste of space caused by excessive expansion at one time.
[0079] In summary, the present application provides a method for expanding the thin volume space. The method extracts multi-dimensional features of historical access data, analyzes the time distribution of historical access data through time series features to identify the growth timing pattern; anticipates the space demand of hot regions through spatial distribution features; plans the expansion of associated volumes by using the multi-volume relationship features; analyzes the impact of historical access data on the system business performance through the system load status; and identifies the business scenario to match the space growth pattern under different scenarios through scenario features, so that the expansion decision is closer to the real business needs. By analyzing the multi-dimensional feature information through a prediction model, the space growth demand parameters for the future time period are output, and thus the expansion granularity is adjusted as needed to intelligently drive the automatic expansion of the thin volume, solving the problems of resource waste and performance fluctuation caused by the lack of foresight in traditional technologies.
[0080] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0081] The embodiments of the present application also provide an apparatus for expanding the thin volume space, as Figure 3 shown. The apparatus includes: An acquisition module 311, configured to acquire historical access data of the thin volume to be expanded; A feature extraction module 312, configured to extract feature information from the historical access data, where the feature information includes at least one of time series features, spatial distribution features, multi-volume relationship features, system load status, and scenario features; A model prediction module 313, configured to input the feature information into a pre-trained prediction model to obtain the space growth demand parameters output by the prediction model; The expansion module 314 is used to adjust the expansion granularity according to the space growth demand parameter, and expand the thin volume to be expanded based on the adjusted expansion granularity.
[0082] As an optional implementation manner provided by an embodiment of the present application, the device further includes a training module for training a prediction model, including: constructing a basic model based on a time series model or a language model; obtaining sample data, where the sample data includes historical access logs and simulation data, and the simulation data is access logs generated by a simulation input-output test software; extracting sample feature information from the preprocessed sample data; and adjusting the parameters of the basic model according to the sample feature information to obtain a prediction model.
[0083] As an optional implementation manner provided by an embodiment of the present application, the prediction model further outputs a data access pattern for a target time period; The device further includes an update module for updating the prediction model, including: determining an actual data access pattern in the target time period; calculating a prediction accuracy according to the actual data access pattern and the data access pattern output by the prediction model; and updating the prediction model when the prediction accuracy is less than a preset threshold.
[0084] As an optional implementation manner provided by an embodiment of the present application, the expansion module 314 is specifically used for: dividing the space growth demand parameter by the used capacity of the thin volume to be expanded, rounding up to obtain a target multiple; multiplying the target multiple by the minimum unit of the storage pool space to determine the expansion granularity; detecting whether the free space of the storage pool meets the expansion granularity; if so, calling an expansion interface to allocate storage space to the thin volume to be expanded according to the expansion granularity for expansion.
[0085] As an optional implementation manner provided by an embodiment of the present application, after multiplying the target multiple by the minimum unit of the storage pool space to determine the expansion granularity, and before detecting whether the free space of the storage pool meets the expansion granularity, the expansion module 314 is further used for: monitoring the system load; reducing the expansion granularity when the system load is less than or equal to a preset load threshold; and increasing the expansion granularity when the system load is greater than the preset load threshold.
[0086] As an optional implementation manner provided by an embodiment of the present application, the expansion module 314 is further used for: When the space capacity of the expanded thin volume is greater than the user available capacity, reducing the expansion granularity to the minimum processing space of a single input-output.
[0087] As an optional implementation manner provided by an embodiment of the present application, the expansion module is further used for: detecting whether the prediction model outputs a space growth demand parameter within a preset duration; and expanding the thin volume to be expanded according to a preset expansion strategy when the prediction model does not output a space growth demand parameter within the preset duration.
[0088] The present application provides an expansion device for a thin volume space. The device extracts multi-dimensional features of historical access data, analyzes the time distribution of the historical access data through time series features to identify the growth timing pattern; anticipates the space demand of hot regions through spatial distribution features; utilizes the relationship between volumes through multi-volume relationship features to plan the expansion of associated volumes; analyzes the impact of historical access data on the system business performance through the system load status; and identifies business scenarios through scenario features to match different spatial growth patterns under different scenarios, enabling the expansion decision to be closer to the real business needs. By analyzing the multi-dimensional feature information through a prediction model, the space growth demand parameters for a future time period are output, thereby adjusting the expansion granularity as needed and intelligently driving the automatic expansion of the thin volume, solving the problems of resource waste and performance fluctuations caused by the lack of foresight in traditional technologies.
[0089] For the description of the features in the corresponding embodiments of the expansion device for the thin volume space, reference can be made to the relevant descriptions in the corresponding embodiments of the expansion method for the thin volume space, which will not be elaborated here one by one.
[0090] An embodiment of the present application further provides an electronic device, as Figure 4 shown, including a memory 401 and a processor 402. A computer program is stored in the memory 401, and the processor 402 is configured to run the computer program to execute the steps in any of the above embodiments of the expansion method for the thin volume space.
[0091] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the expansion method for the thin volume space when running.
[0092] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.
[0093] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the expansion method for the thin volume space are implemented.
[0094] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the expansion method for the thin volume space are implemented.
[0095] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered as exceeding the scope of this application.
[0096] The above has introduced in detail a method, device, electronic device, and storage medium for expanding the space of a thin volume provided by this application. Specific examples have been used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for expanding the capacity of a thin volume, characterized in that including: Obtain historical access data of the thin volume to be expanded; Extract feature information from the historical access data, where the feature information includes at least one of time series features, spatial distribution features, multi-volume relationship features, system load status, and scenario features; Input the feature information into a pre-trained prediction model to obtain the spatial growth demand parameters output by the prediction model; Adjust the expansion granularity according to the spatial growth demand parameters, and expand the thin volume to be expanded based on the adjusted expansion granularity.
2. The method according to claim 1, characterized in that, The training process of the prediction model includes: Construct a basic model based on a time series model or a language model; Obtain sample data, where the sample data includes historical access logs and simulated data, and the simulated data is access logs generated by simulated input / output test software; Extract sample feature information from the preprocessed sample data; Adjust the parameters of the basic model according to the sample feature information to obtain the prediction model.
3. The method according to claim 1, wherein The prediction model also outputs the data access pattern for the target time period; The update process of the prediction model includes: Determine the actual data access pattern during the target time period; Calculate the prediction accuracy according to the actual data access pattern and the data access pattern output by the prediction model; Update the prediction model when the prediction accuracy is less than a preset threshold.
4. The method according to claim 1, wherein The adjusting the expansion granularity according to the spatial growth demand parameters and expanding the thin volume to be expanded based on the adjusted expansion granularity includes: Divide the spatial growth demand parameter by the used capacity of the thin volume to be expanded, and round up to obtain the target multiple; Multiply the target multiple by the minimum unit of the storage pool space to determine the expansion granularity; Detect whether the free space of the storage pool meets the expansion granularity; If so, call the expansion interface to allocate storage space to the thin volume to be expanded according to the expansion granularity for expansion.
5. The method according to claim 4, wherein After multiplying the target multiple by the minimum unit of the storage pool space to determine the expansion granularity and before detecting whether the free space of the storage pool meets the expansion granularity, the method further includes: Monitor the system load; Reduce the expansion granularity when the system load is less than or equal to a preset load threshold; Increase the expansion granularity when the system load is greater than the preset load threshold.
6. The method according to claim 1, characterized in that, After adjusting the expansion granularity according to the spatial growth demand parameters and expanding the thin volume to be expanded based on the adjusted expansion granularity, the method further includes: When the space capacity of the expanded thin volume is greater than the user available capacity, reduce the expansion granularity to the minimum processing space of a single input / output.
7. The method according to claim 1, characterized in that, After extracting the feature information from the historical access data, the method further includes: Detect whether the prediction model outputs spatial growth demand parameters within a preset duration; When the prediction model does not output spatial growth demand parameters within the preset duration, expand the thin volume to be expanded according to a preset expansion strategy.
8. An expansion device for a thin volume space, characterized in that, including: An acquisition module for acquiring historical access data of the thin volume to be expanded; A feature extraction module, configured to extract feature information from the historical access data, where the feature information includes at least one of time series features, spatial distribution features, multi-volume relationship features, system load status, and scenario features; A model prediction module, configured to input the feature information into a pre-trained prediction model to obtain spatial growth demand parameters output by the prediction model; An expansion module, configured to adjust the expansion granularity according to the spatial growth demand parameters, and expand the thin volume to be expanded based on the adjusted expansion granularity.
9. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the method for expanding the thin volume space according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program, when executed by a processor, implements the steps of the method for expanding the thin volume space according to any one of claims 1 to 7.
Citation Information
Patent Citations
Storage space distribution method and system for simplified volume and related components
CN109739778A
Capacity expansion method, prediction model creation method and device, equipment and medium
CN109885469A
Capacity expansion and contraction method and device, electronic equipment and readable storage medium
CN116244069A
Cloud hard disk data guarantee method and system based on adaptive prediction of usable time
CN117950583A
Image forming apparatus
JP2012011721A
Cited By
Simplified volume resource management method in storage system, electronic equipment, medium and product
CN120743804A
Thin volume resource management methods, electronic devices, media and products in storage systems
CN120743804B
Space management method of thin provisioning volume and electronic equipment
CN121635817A