Method and apparatus for determining the frequency of data access

By using machine learning models to predict the future performance of cloud platform block storage instances, the problem of inaccurate determination of data access frequency in existing technologies is solved, enabling more accurate resource allocation, reducing resource waste, and improving resource utilization.

CN114595063BActive Publication Date: 2025-10-28JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210231567.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-09
Publication Date
2025-10-28
Estimated Expiration
2042-03-09

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of determining the frequency of data access to block storage instances in cloud platforms through empirical or statistical methods is poor. This cannot accurately reflect the frequency of future accesses, leading to inaccurate allocation of physical resources and wasted resources.

Method used

By employing machine learning models and based on the attribute information of target block storage instances, future performance is predicted, and the frequency of data access in future periods is determined. The machine learning models are trained to train the characteristics of different images and tenants, thereby improving the accuracy of predictions.

Benefits of technology

It improved the accuracy of data access frequency, improved the accuracy of physical resource allocation, reduced resource waste, and increased resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114595063B_ABST
    Figure CN114595063B_ABST
Patent Text Reader

Abstract

This disclosure relates to a method and apparatus for determining the frequency of data access, and a computer-storable medium, relating to the field of computer technology. The method for determining the frequency of data access includes: acquiring attribute information related to a target block storage instance in a cloud platform; using a machine learning model based on the attribute information related to the target block storage instance to predict the future performance of the target block storage instance in a future period; and determining the frequency of data access in the target block storage instance in the future period based on the predicted future performance, wherein the determined frequency of access is used to allocate physical resources to the target block storage instance. According to this disclosure, the accuracy of determining the frequency of data access in a cloud platform can be improved, thereby improving the accuracy of physical resource allocation, increasing physical resource utilization, and reducing physical resource waste.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to methods and apparatus for determining the frequency of data access, and computer-storable media. Background Technology

[0002] In cloud computing scenarios, determining the frequency of data access in block storage instances within a cloud platform (also known as a cloud computing platform) is of great guiding significance for the rational allocation of physical resources by the cloud platform.

[0003] In related technologies, users determine the frequency of data access in block storage instances on cloud platforms based on experience, or determine the frequency of data access in block storage instances over a historical period based on the historical performance of block storage instances over a historical period using statistical knowledge. Summary of the Invention

[0004] In related technologies, determining the frequency of data access based on experience is too subjective and relies excessively on the user's accumulated experience, resulting in poor accuracy. Furthermore, determining historical access frequency based on statistical knowledge of historical performance data cannot accurately reflect the future access frequency of data within a block storage instance.

[0005] To address the aforementioned technical issues, this disclosure proposes a solution that can improve the accuracy of determining the frequency of data access in a cloud platform, thereby improving the accuracy of physical resource allocation, increasing physical resource utilization, and reducing physical resource waste.

[0006] According to a first aspect of this disclosure, a method for determining the access frequency of data is provided, comprising: acquiring attribute information related to a target block storage instance in a cloud platform; using a machine learning model based on the attribute information related to the target block storage instance to predict the future performance of the target block storage instance in a future period; and determining the access frequency of data in the target block storage instance in the future period based on the predicted future performance, wherein the determined access frequency is used to allocate physical resources to the target block storage instance.

[0007] In some embodiments, the future period includes multiple future moments, and the future performance includes future performance values ​​of at least one performance of the target block storage instance at each future moment. Determining the access frequency of data in the target block storage instance within the future period based on the predicted future performance includes: obtaining the maximum performance value of each performance of the target block storage instance; for each performance, filtering out the future performance peak corresponding to each performance from the multiple future performance values ​​corresponding to the multiple future moments, wherein the future performance peak reflects the general performance of each performance within the future period; and for the target block storage instance, determining the access frequency of data in the target block storage instance within the future period based on the maximum performance value of the at least one performance and the corresponding future performance peak.

[0008] In some embodiments, for the target block storage instance, determining the access frequency of data in the target block storage instance in the future period based on the maximum performance value of the at least one performance and the corresponding future performance peak includes: for each performance, determining the ratio of the corresponding future performance peak to the maximum performance value as a reference ratio; when the reference ratios corresponding to all performances are less than a first reference ratio threshold, determining the access frequency of data in the target block storage instance in the future period as a first level; when there is a reference ratio corresponding to at least one performance that is greater than or equal to the first reference ratio threshold, and the reference ratios corresponding to all performances are less than a second reference ratio threshold, determining the access frequency of data in the target block storage instance in the future period as a second level, where the second reference ratio threshold is greater than the first reference ratio threshold and the second level is higher than the first level; when there is a reference ratio corresponding to at least one performance that is greater than or equal to the second reference ratio threshold, determining the access frequency of data in the target block storage instance in the future period as a third level, where the third level is higher than the second level.

[0009] In some embodiments, the cloud platform includes at least one target image, and the machine learning model includes a first machine learning model corresponding to each target image. The first machine learning model corresponding to each target image is trained based on relevant data of a reference block storage instance mounted on a compute node created based on each target image. Based on attribute information related to the target block storage instance, the machine learning model is used to predict the future performance value that the target block storage instance will achieve in a future period. This includes: when the image used when the compute node mounted on the target block storage instance is created belongs to the at least one target image, based on attribute information related to the target block storage instance, using the first machine learning model corresponding to the image corresponding to the target block storage instance, predicting the future performance value that the target block storage instance will achieve in a future period.

[0010] In some embodiments, for each image in the cloud platform, if the ratio of the total capacity of all block storage instances mounted under the compute nodes created based on each image to the total capacity of all block storage instances in the cloud platform is greater than a first capacity ratio threshold, then each image is a target image.

[0011] In some embodiments, the cloud platform includes at least one target image. If the tenant of the block storage instance mounted on the compute node created based on each target image belongs to a preset tenant, the machine learning model includes a second machine learning model corresponding to each preset tenant. The second machine learning model corresponding to each preset tenant is trained based on relevant data from a reference block storage instance of each preset tenant. Based on attribute information related to the target block storage instance, the machine learning model predicts the future performance value of the data in the target block storage instance within a future period. This includes: if the image used when the compute node mounted on the target block storage instance was created belongs to the at least one target image, determining whether the tenant to which the target block storage instance belongs belongs to a preset tenant; and if the tenant to which the target block storage instance belongs belongs to a preset tenant, predicting the future performance value of the data in the target block storage instance within a future period based on attribute information related to the target block storage instance and using the second machine learning model corresponding to the tenant to which the target block storage instance belongs.

[0012] In some embodiments, for each tenant in the cloud platform, if the ratio of the total capacity of all block storage instances of each tenant to the total capacity of all block storage instances mounted under the compute node to which each tenant belongs is greater than a second capacity ratio threshold, then each tenant is a preset tenant.

[0013] In some embodiments, the machine learning model further includes a first machine learning model corresponding to each target image. The first machine learning model corresponding to each target image is trained based on the relevant data of the reference block storage instance mounted on the computing node corresponding to each target image. Predicting the future performance value of the data in the target block storage instance in the future period based on the attribute information related to the target block storage instance and using the machine learning model further includes: when the tenant to which the target block storage instance belongs does not belong to a preset tenant, predicting the future performance value of the data in the target block storage instance in the future period based on the attribute information related to the target block storage instance and using the first machine learning model corresponding to the image corresponding to the target block storage instance.

[0014] In some embodiments, the method for determining the access frequency of data further includes: if the image used when the compute node mounted on the target block storage instance is created does not belong to the at least one target image, determining the access frequency of the target block storage instance in the future period to a preset level.

[0015] In some embodiments, the duration of the future period is the time required for the target block storage instance to migrate a data volume equal to the capacity of the target block storage instance.

[0016] In some embodiments, the future performance value is a performance value related to read operations; and / or the performance value related to read operations is measured using at least one of the following performance metrics: throughput related to read operations and IOPS related to read operations.

[0017] In some embodiments, the relevant data of the reference block storage instance includes the attribute information of the reference block storage instance, the attribute information of the compute nodes attached to the reference block storage instance, and the historical performance values ​​of the data in the reference block storage instance over a historical period; and / or the attribute information related to the target block storage instance includes the attribute information of the target block storage instance and the attribute information of the compute nodes attached to the target block storage instance.

[0018] In some embodiments, the method for determining the frequency of data access further includes: for each target image, using relevant data of a reference block storage instance mounted under a compute node created based on the target image, training a basic machine learning model to obtain a first machine learning model; and / or for each preset tenant, using relevant data of the reference block storage instance of the preset tenant, training a first machine learning model corresponding to the target image corresponding to the preset tenant to obtain a second machine learning model.

[0019] According to a second aspect of this disclosure, an apparatus for determining the access frequency of data is provided, comprising: an acquisition module configured to acquire attribute information related to a target block storage instance in a cloud platform; a prediction module configured to predict the future performance of the target block storage instance in a future period using a machine learning model based on the attribute information related to the target block storage instance; and a determination module configured to determine the access frequency of data in the target block storage instance in the future period based on the predicted future performance, wherein the determined access frequency is used to allocate physical resources to the target block storage instance.

[0020] According to a third aspect of this disclosure, an apparatus for determining the frequency of data access is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute the method for determining the frequency of data access as described in any of the above embodiments based on instructions stored in the memory.

[0021] According to a fourth aspect of this disclosure, a computer-storeable medium is provided having computer program instructions stored thereon, which, when executed by a processor, implement the method for determining the frequency of data access as described in any of the above embodiments.

[0022] In the above embodiments, the accuracy of determining the frequency of data access in the cloud platform can be improved, thereby improving the accuracy of physical resource allocation, increasing physical resource utilization, and reducing physical resource waste. Attached Figure Description

[0023] The accompanying drawings, which form part of this specification, illustrate embodiments of this disclosure and, together with the specification, serve to explain the principles of this disclosure.

[0024] This disclosure will become clearer with reference to the accompanying drawings and the following detailed description, wherein:

[0025] Figure 1 This is a flowchart illustrating a method for determining the frequency of data access according to some embodiments of the present disclosure;

[0026] Figure 2 This is a block diagram illustrating an apparatus for determining the frequency of data access according to some embodiments of the present disclosure;

[0027] Figure 3 This is a block diagram illustrating an apparatus for determining the frequency of data access according to other embodiments of the present disclosure;

[0028] Figure 4 This is a block diagram illustrating a computer system for implementing some embodiments of the present disclosure. Detailed Implementation

[0029] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0030] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0031] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.

[0032] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0033] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.

[0034] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0035] Figure 1 This is a flowchart illustrating a method for determining the frequency of data access according to some embodiments of the present disclosure.

[0036] like Figure 1 As shown, the method for determining the access frequency of data includes: step S1, obtaining attribute information related to a target block storage instance in a cloud platform; step S2, using a machine learning model based on the attribute information related to the target block storage instance, predicting the future performance of the target block storage instance in a future period; and step S3, determining the access frequency of data in the target block storage instance in the future period based on the predicted future performance. For example, the method for determining the access frequency of data is performed by a device for determining the access frequency of data.

[0037] In the above embodiments, machine learning models are used to learn the features of attribute information related to the target block storage instance. Based on these learned features, the future performance of the target block storage instance in future periods is predicted. Then, based on the predicted future performance, the frequency of data access in future periods is determined. This not only enables advance prediction but also allows for a more objective determination of access frequency, improving the accuracy of data access frequency assessment. Accurately determining the frequency of data access is crucial for rationally allocating physical resources to block storage instances, thereby improving the accuracy of physical resource allocation, increasing physical resource utilization, and reducing physical resource waste.

[0038] In step S1, attribute information related to the target block storage instance in the cloud platform is obtained. For example, the block storage instance is a cloud disk instance. The data in the cloud disk instance is actually stored on the physical resources allocated to the cloud disk instance. A block storage instance is a storage instance that stores data in the form of data blocks.

[0039] In step S2, based on the attribute information related to the target block storage instance, a machine learning model is used to predict the future performance of the target block storage instance in future periods.

[0040] In some embodiments, the duration of the future period is the time required for the target block storage instance to migrate a data volume equal to the capacity of the target block storage instance. For example, if the target block storage instance has a capacity of 800GB, a single-disk data migration rate of 10MB / s, and a daily data migration volume per disk of 10MB / s × 3600s × 24h ≈ 800GB / day, then the time required for the target block storage instance to migrate a data volume equal to the capacity of the target block storage instance is 800GB ÷ 800GB / day = 1 day. Therefore, the future period is 1 day. By considering the daily migration volume per disk, it can be ensured that after determining the access frequency of the data within the target block storage instance, the data can be migrated to newly allocated physical resources within the future period, further improving the utilization rate of physical resources.

[0041] In some embodiments, the attribute information related to the target block storage instance includes the attribute information of the target block storage instance itself and the attribute information of the compute nodes mounted on the target block storage instance. For example, the attribute information of the target block storage instance itself includes at least one of the target block storage instance's specifications and capacity. The attribute information of the compute nodes mounted on the target block storage instance includes at least one of the compute node's specifications, number of CPUs, amount of memory, and network bandwidth configuration information. Generally, cloud providers offer different cloud disk service specifications to cloud disk users based on the storage medium and storage server type, such as performance-based and capacity-based. Generally, cloud providers offer different cloud computing node specifications to compute node users based on the compute server's CPU, memory ratio, CPU model, etc., such as general-purpose, memory-optimized, and compute-optimized.

[0042] In some embodiments, the cloud platform includes at least one target image, and the machine learning model includes a first machine learning model corresponding to each target image. The first machine learning model corresponding to each target image is trained based on data related to a reference block storage instance mounted on a compute node created based on each target image. When the image used to create the compute node mounted on the target block storage instance belongs to at least one target image, the future performance value of the target block storage instance in a future period is predicted using the first machine learning model corresponding to the image corresponding to the target block storage instance, based on attribute information related to the target block storage instance. The image used to create the compute node includes an operating system, application software, and related configurations within the operating system. Compute nodes that typically provide the same type of application service are created using the same image.

[0043] In some embodiments, the relevant data for the reference block storage instance includes the attribute information of the reference block storage instance, the attribute information of the compute nodes mounted on the reference block storage instance, and the historical performance values ​​of the data in the reference block storage instance over a historical period. For example, the attribute information of the target block storage instance itself includes at least one of the target block storage instance's specifications and capacity. The attribute information of the compute nodes mounted on the target block storage instance includes at least one of the compute node's specifications, number of CPUs, amount of memory, and network bandwidth configuration information.

[0044] In the above embodiments, the I / O operation behavior of block storage instances under different images in the cloud platform varies greatly. By using data related to each target image to train a first machine learning model that conforms to the corresponding characteristics of each target image, the accuracy of predicting future performance values ​​is improved, thereby further improving the accuracy of determining the frequency of data access, which in turn can further improve the accuracy of physical resource allocation, further improve physical resource utilization, and further reduce physical resource waste.

[0045] In addition, during the training process, not only the attribute information of the reference block storage instance and the historical performance values ​​of the data stored therein over a historical period are considered, but also the attribute information of the computing nodes it is mounted on. This can further improve the accuracy of the first machine learning model, thereby further improving the accuracy of determining the frequency of data access, which in turn can further improve the accuracy of physical resource allocation, further improve physical resource utilization, and further reduce physical resource waste.

[0046] In some embodiments, for each image in the cloud platform, if the ratio of the total capacity of all block storage instances mounted under the compute nodes created based on that image to the total capacity of all block storage instances in the cloud platform is greater than a first capacity ratio threshold, then that image is a target image. For example, the first capacity ratio threshold is 1%. Alternatively, different first capacity ratio thresholds can be set according to actual business needs. By training the first machine learning model on images with a relatively large total storage capacity, training costs can be reduced and training efficiency improved.

[0047] In some embodiments, the cloud platform includes at least one target image. If the tenants of the block storage instances mounted on the compute nodes created based on each target image belong to preset tenants, the machine learning model includes a second machine learning model corresponding to each preset tenant. The second machine learning model corresponding to each preset tenant is trained based on relevant data from the reference block storage instance of each preset tenant.

[0048] In some embodiments, the future performance values ​​that the data in the target block storage instance will reach in future periods can be predicted by using a machine learning model based on attribute information related to the target block storage instance.

[0049] First, if the image used when the compute node mounted on the target block storage instance was created belongs to at least one target image, determine whether the tenant to which the target block storage instance belongs belongs to the preset tenant.

[0050] Then, if the tenant to which the target block storage instance belongs is a preset tenant, based on the attribute information related to the target block storage instance, the second machine learning model corresponding to the tenant to which the target block storage instance belongs is used to predict the future performance value of the data in the target block storage instance in the future period.

[0051] In the above embodiments, the relevant characteristics (including I / O operation behavior characteristics) of block storage instances of different tenants under the same image are quite different. By training a second machine learning model that is more consistent with the characteristics of block storage instances at the tenant granularity for all or some tenants, the accuracy of predicting future performance values ​​can be further improved, thereby further improving the accuracy of determining the frequency of data access, which in turn can further improve the accuracy of physical resource allocation, further improve physical resource utilization, and further reduce physical resource waste.

[0052] In some embodiments, for each tenant in the cloud platform, if the ratio of the total capacity of all block storage instances in each tenant to the total capacity of all block storage instances mounted under the compute node to which each tenant belongs is greater than a second capacity ratio threshold, each tenant is a preset tenant. For example, the second capacity ratio threshold is 1%. Alternatively, different second capacity ratio thresholds can be set according to actual business needs. By training the first machine learning model on tenants with a relatively large total storage capacity under a certain image, training costs can be reduced and training efficiency improved.

[0053] In some embodiments, taking the example that the machine learning model also includes a first machine learning model corresponding to each target image, when the tenant to which the target block storage instance belongs does not belong to the preset tenant, the future performance value of the data in the target block storage instance in the future period is predicted based on the attribute information related to the target block storage instance and using the first machine learning model corresponding to the image corresponding to the target block storage instance.

[0054] In some embodiments, future performance values ​​are performance values ​​related to read operations. For example, read operation-related performance values ​​are measured using at least one of read operation-related throughput and read operation-related IOPS (Input / Output Per Second). Performance values ​​change over time, and the performance values ​​achieved by the block storage instance are measured or quantified using performance metrics. Read operation-related performance includes read IOPS performance and read throughput performance. Read IOPS performance is measured by the number of read requests that can be processed per second. Read throughput performance is measured by the read throughput bandwidth that can be processed per second.

[0055] In the above embodiments, read operations are closely related to the frequency of data access in the block storage instance. Using performance values ​​related to read operations to determine the frequency of data access can further improve the accuracy of determining the frequency of data access, thereby further improving the accuracy of physical resource allocation, further improving physical resource utilization, and further reducing physical resource waste.

[0056] In step S3, based on the predicted future performance, the access frequency of data in the target block storage instance over the future period is determined. The determined access frequency is used to allocate physical resources to the target block storage instance.

[0057] In some embodiments, the future period includes multiple future moments, and the future performance includes the future performance value of at least one performance characteristic of the target block storage instance at each future moment. Step S3 above can be implemented as follows.

[0058] First, obtain the maximum performance value for each performance metric for the target block storage instance.

[0059] Then, for each performance level, the peak performance value corresponding to each performance level is selected from multiple future performance values ​​at multiple future time points. The peak performance value reflects the general performance of each performance level over future periods. In some embodiments, the Nth largest future performance value among multiple future performance values ​​at future time points is selected as the peak performance value for that performance level. N is an integer greater than 0 and less than the total number of multiple future performance values. For example, the 95th percentile future performance value for each performance level in future periods is selected as the peak performance value, arranged in descending order.

[0060] Finally, for the target block storage instance, the frequency of data access in the target block storage instance in the future period is determined based on the maximum performance value of at least one performance parameter and the corresponding future performance peak.

[0061] In the above embodiments, the future performance peak can reflect the general situation of each performance in the future period. By determining the access frequency of data through the future performance peak and the maximum performance value, the generality and representativeness of the determined access frequency of data can be improved, and the impact of extreme data on the accuracy of determining the access frequency of data can be avoided. This further improves the accuracy of determining the access frequency of data, which in turn can further improve the accuracy of physical resource allocation, further improve the utilization rate of physical resources, and further reduce the waste of physical resources.

[0062] In some embodiments, the frequency of data access in a target block storage instance over a future period can be determined based on the maximum performance value of at least one performance parameter and the corresponding future performance peak.

[0063] First, for each performance level, determine the ratio of the corresponding future peak performance to the maximum performance value as a reference ratio.

[0064] Secondly, if the reference ratios for various performance metrics are all less than the first reference ratio threshold, the frequency of data access in the target block storage instance over future periods is determined as the first level. For example, the first level is low.

[0065] Then, if at least one reference ratio corresponding to a performance characteristic is greater than or equal to a first reference ratio threshold, and all reference ratios corresponding to various performance characteristics are less than a second reference ratio threshold, the access frequency of data in the target block storage instance in future periods is determined as a second level. The second reference ratio threshold is greater than the first reference ratio threshold. The second level is higher than the first level. In some embodiments, the second level is a moderate level. For example, the first reference ratio threshold is 20%. For example, the second reference ratio threshold is 60%. For example, other first and second reference ratio thresholds can also be set according to actual needs.

[0066] Finally, if at least one reference ratio corresponding to a performance level is greater than or equal to a second reference ratio threshold, the access frequency of data in the target block storage instance over future periods is determined to be at the third level. The third level is higher than the second level. For example, the third level is considered high.

[0067] In some embodiments, if the image used when the compute nodes mounted on the target block storage instance are created does not belong to at least one target image, the access frequency of the target block storage instance in the future period is determined to be a preset level. For example, for compute nodes whose total number of mounted block storage instances is less than a quantity threshold and whose ratio of the total capacity of the mounted block storage instances to the total capacity of all block storage instances on the cloud platform is less than a first capacity ratio threshold, the image used when the compute nodes are created does not belong to the target image, and the access frequency of the target block storage instances mounted under it in the future period can be preset to a first level.

[0068] In some embodiments, physical resources are allocated to the target block storage instance based on the frequency of access to data in the target block storage instance over a future period. The block storage instance in the cloud platform points to the allocated physical resources and stores data on those resources.

[0069] In some embodiments, if the data in the target block storage instance is accessed with a first degree of frequency in the future period, the data in the target block storage instance is considered cold data in the future period, and low-performance physical resources are allocated to the data in the target block storage instance.

[0070] In some embodiments, if the data in the target block storage instance is accessed at a second frequency in the future period, the data in the target block storage instance is considered warm data in the future period, and medium-performance physical resources are allocated to the data in the target block storage instance.

[0071] In some embodiments, if the access frequency of data in the target block storage instance is at level three in the future period, the data in the target block storage instance is considered hot data in the future period, and high-performance physical resources are allocated to the data in the target block storage instance.

[0072] In some embodiments, for each target image, the relevant data of the reference block storage instance mounted under the compute node created based on each target image can also be used to train the basic machine learning model to obtain the first machine learning model.

[0073] In some embodiments, for each preset tenant, relevant data from the reference block storage instance of each preset tenant can be used to train a first machine learning model corresponding to the target image of each preset tenant, thereby obtaining a second machine learning model. By retraining the first machine learning model, the accuracy of determining the frequency of data access in the block storage instance of the preset tenant can be further improved, which in turn can further improve the accuracy of physical resource allocation, further improve physical resource utilization, and further reduce physical resource waste.

[0074] Figure 2 This is a block diagram illustrating an apparatus for determining the frequency of data access according to some embodiments of the present disclosure.

[0075] like Figure 2 As shown, the device 2 for determining the frequency of data access includes an acquisition module 21, a prediction module 22, and a determination module 23.

[0076] Module 21 is configured to retrieve attribute information related to the target block storage instance in the cloud platform, such as performing actions like... Figure 1 Step S1 is shown.

[0077] Prediction module 22 is configured to predict the future performance of the target block storage instance over a future period based on attribute information associated with the target block storage instance and using a machine learning model, such as performing... Figure 1 Step S2 is shown.

[0078] Module 23 is configured to determine the access frequency of data in the target block storage instance over a future period based on predicted future performance. The determined access frequency is used to allocate physical resources to the target block storage instance, such as performing actions like... Figure 1 Step S3 is shown.

[0079] Figure 3 This is a block diagram illustrating an apparatus for determining the frequency of data access according to other embodiments of the present disclosure.

[0080] like Figure 3 As shown, the apparatus 3 for determining the frequency of data access includes a memory 31 and a processor 32 coupled to the memory 31. The memory 31 is used to store instructions for performing a method for determining the frequency of data access corresponding to an embodiment. The processor 32 is configured to perform a method for determining the frequency of data access in any of the embodiments of this disclosure based on the instructions stored in the memory 31.

[0081] Figure 4 This is a block diagram illustrating a computer system for implementing some embodiments of the present disclosure.

[0082] like Figure 4 As shown, the computer system 40 can be represented in the form of a general computing device. The computer system 40 includes a memory 410, a processor 420, and a bus 400 connecting different system components.

[0083] The memory 410 may include, for example, system memory, non-volatile storage media, etc. The system memory may store, for example, an operating system, application programs, a boot loader, and other programs. The system memory may include volatile storage media, such as random access memory (RAM) and / or cache memory. The non-volatile storage media may store, for example, instructions for performing at least one of the methods for determining the frequency of data access in a corresponding embodiment. Non-volatile storage media include, but are not limited to, disk storage, optical storage, flash memory, etc.

[0084] The processor 420 can be implemented using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete hardware components such as discrete gates or transistors. Accordingly, each module, such as the decision module and the determination module, can be implemented by executing instructions in the central processing unit (CPU) memory to perform the corresponding steps, or by implementing dedicated circuitry to perform the corresponding steps.

[0085] Bus 400 can use any of the various bus architectures. For example, bus architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, and Peripheral Component Interconnect (PCI) bus.

[0086] The computer system 40 may also include an input / output interface 430, a network interface 440, and a storage interface 450. These interfaces 430, 440, and 450, as well as the memory 410 and processor 420, can be connected via a bus 400. The input / output interface 430 provides a connection interface for input / output devices such as a monitor, mouse, and keyboard. The network interface 440 provides a connection interface for various networked devices. The storage interface 450 provides a connection interface for external storage devices such as floppy disks, USB flash drives, and SD cards.

[0087] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations thereof, can be implemented by computer-readable program instructions.

[0088] These computer-readable program instructions are provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable device to produce a machine, such that execution of the instructions by the processor produces means for implementing the functions specified in one or more boxes of the flowchart and / or block diagram.

[0089] These computer-readable program instructions may also be stored in a computer-readable storage medium. These instructions cause a computer to work in a particular manner to produce an article of manufacture, including instructions that implement the functions specified in one or more boxes in a flowchart and / or block diagram.

[0090] This disclosure may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.

[0091] The method and apparatus for determining the frequency of data access in the above embodiments, as well as the computer-storable medium, can improve the accuracy of determining the frequency of data access in a cloud platform, thereby improving the accuracy of physical resource allocation, increasing physical resource utilization, and reducing physical resource waste.

[0092] This concludes the detailed description of the methods, apparatus, and computer-storable media for determining the frequency of data access according to this disclosure. To avoid obscuring the concept of this disclosure, some details known in the art have not been described. Those skilled in the art will fully understand how to implement the technical solutions disclosed herein based on the above description.

Claims

1. A method for determining the frequency of data access, comprising: Obtain attribute information related to a target block storage instance in a cloud platform, wherein the cloud platform includes at least one target image; Based on the attribute information related to the target block storage instance, a machine learning model is used to predict the future performance of the target block storage instance in the future period. The machine learning model includes a first machine learning model corresponding to each target image. The first machine learning model corresponding to each target image is trained based on the relevant data of the reference block storage instance mounted on the computing node created based on each target image. Based on the predicted future performance, the access frequency of data in the target block storage instance within the future period is determined, and the determined access frequency is used to allocate physical resources to the target block storage instance. The step of predicting the future performance value of the target block storage instance in a future period using a machine learning model based on attribute information related to the target block storage instance includes: If the image used when the compute node mounted on the target block storage instance is created belongs to at least one of the target images, the future performance value that the target block storage instance will achieve in the future period is predicted based on the attribute information related to the target block storage instance and using the first machine learning model corresponding to the image corresponding to the target block storage instance.

2. The method for determining the frequency of data access according to claim 1, wherein, The future period includes multiple future moments, and the future performance includes future performance values ​​of at least one performance characteristic of the target block storage instance at each future moment. Determining the access frequency of data in the target block storage instance within the future period based on the predicted future performance includes: Obtain the maximum performance value for each performance metric of the target block storage instance; For each performance, a future performance peak corresponding to each performance is selected from multiple future performance values ​​corresponding to multiple future moments. The future performance peak reflects the general performance of each performance in the future period. For the target block storage instance, the access frequency of the data in the target block storage instance in the future period is determined based on the maximum performance value of the at least one performance and the corresponding future performance peak.

3. The method for determining the frequency of data access according to claim 2, wherein, For the target block storage instance, determining the access frequency of data in the target block storage instance within the future period based on the maximum performance value of the at least one performance characteristic and the corresponding future performance peak includes: For each performance level, determine the ratio of the corresponding future peak performance to the maximum performance value as a reference ratio; If the reference ratios corresponding to various performance parameters are all less than the first reference ratio threshold, the access frequency of the data in the target block storage instance within the future period is determined to be the first level. If at least one performance-related reference ratio is greater than or equal to the first reference ratio threshold, and all reference ratios for various performances are less than the second reference ratio threshold, the access frequency of the data in the target block storage instance in the future period is determined to be the second level, where the second reference ratio threshold is greater than the first reference ratio threshold, and the second level is higher than the first level. If at least one reference ratio corresponding to performance is greater than or equal to the second reference ratio threshold, the access frequency of data in the target block storage instance in the future period is determined to be a third level, which is higher than the second level.

4. The method for determining the frequency of data access according to claim 1, wherein, For each image in the cloud platform, if the ratio of the total capacity of all block storage instances mounted under the compute nodes created based on each image to the total capacity of all block storage instances in the cloud platform is greater than a first capacity ratio threshold, then each image is a target image.

5. The method for determining the frequency of data access according to claim 1, wherein, The cloud platform includes at least one target image. When the tenant of the block storage instance mounted on the compute node created based on each target image belongs to a preset tenant, the machine learning model includes a second machine learning model corresponding to each preset tenant. The second machine learning model corresponding to each preset tenant is trained based on relevant data from the reference block storage instance of each preset tenant. Based on attribute information related to the target block storage instance, the machine learning model predicts the future performance values ​​that the data in the target block storage instance will achieve in future periods, including: If the image used when the compute node mounted on the target block storage instance is created belongs to at least one target image, determine whether the tenant to which the target block storage instance belongs belongs to a preset tenant; If the tenant to which the target block storage instance belongs is a preset tenant, the future performance value of the data in the target block storage instance is predicted in the future period based on the attribute information related to the target block storage instance and using the second machine learning model corresponding to the tenant to which the target block storage instance belongs.

6. The method for determining the frequency of data access according to claim 5, wherein, For each tenant in the cloud platform, if the ratio of the total capacity of all block storage instances of each tenant to the total capacity of all block storage instances mounted under the compute node to which each tenant belongs is greater than a second capacity ratio threshold, then each tenant is a preset tenant.

7. The method for determining the frequency of data access according to claim 5, wherein, The machine learning model also includes a first machine learning model corresponding to each target image. The first machine learning model corresponding to each target image is trained based on the relevant data of the reference block storage instance mounted on the computing node corresponding to each target image. Based on the attribute information related to the target block storage instance, the machine learning model is used to predict the future performance value of the data in the target block storage instance in the future period, which also includes: If the tenant to which the target block storage instance belongs does not belong to the preset tenant, the future performance value of the data in the target block storage instance in the future period is predicted based on the attribute information related to the target block storage instance and using the first machine learning model corresponding to the image corresponding to the target block storage instance.

8. The method for determining the frequency of data access according to claim 1, further comprising: If the image used when the compute node mounted on the target block storage instance is created does not belong to the at least one target image, the access frequency of the target block storage instance in the future period is determined to be a preset level.

9. The method for determining the frequency of data access according to claim 1, wherein, The duration of the future cycle is the time required for the target block storage instance to migrate a data volume equal to the capacity of the target block storage instance.

10. The method for determining the frequency of data access according to claim 1, wherein, The future performance value is a performance value related to read operations; and / or The performance values ​​related to read operations are measured using at least one of the following performance metrics: read operation-related throughput and read operation-related IOPS (IOps per second).

11. The method for determining the frequency of data access according to any one of claims 1-8, wherein, The relevant data of the reference block storage instance includes the attribute information of the reference block storage instance, the attribute information of the computing nodes attached to the reference block storage instance, and the historical performance values ​​of the data in the reference block storage instance over a historical period. and / or The attribute information associated with the target block storage instance includes the attribute information of the target block storage instance and the attribute information of the compute nodes mounted on the target block storage instance.

12. The method for determining the frequency of data access according to claim 7, further comprising: For each target image, a basic machine learning model is trained using relevant data from the reference block storage instance mounted under the compute node created based on each target image, to obtain the first machine learning model. and / or For each preset tenant, using the relevant data of the reference block storage instance of each preset tenant, a first machine learning model corresponding to the target image corresponding to each preset tenant is trained to obtain a second machine learning model.

13. An apparatus for determining the frequency of data access, comprising: The acquisition module is configured to acquire attribute information related to a target block storage instance in a cloud platform, wherein the cloud platform includes at least one target image; The prediction module is configured to predict the future performance of the target block storage instance in a future period based on attribute information related to the target block storage instance and using a machine learning model. The machine learning model includes a first machine learning model corresponding to each target image, which is trained based on relevant data of a reference block storage instance mounted on a compute node created based on each target image. The prediction module is configured to predict the future performance value of the target block storage instance in a future period based on attribute information related to the target block storage instance and using the first machine learning model corresponding to the image corresponding to the target block storage instance, when the image used when the compute node mounted on the target block storage instance is created belongs to at least one target image. The determination module is configured to determine the access frequency of data in the target block storage instance within the future period based on the predicted future performance, and the determined access frequency is used to allocate physical resources to the target block storage instance.

14. An apparatus for determining the frequency of data access, comprising: Memory; as well as A processor coupled to the memory, the processor being configured to perform a method for determining the frequency of data access as described in any one of claims 1 to 12, based on instructions stored in the memory.

15. A computer-storeable medium having stored thereon computer program instructions that, when executed by a processor, implement the method for determining the frequency of access to data as described in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Cold and hot data identification method for data hierarchical mixed storage

    CN113792772A