Storage resource management method and system and storage medium

By calculating the popularity and popularity threshold of data access for layered management, a data distribution mapping table is generated, and storage resources are allocated based on this, the problem that traditional storage systems cannot flexibly cope with complex business environments, and efficient, flexible and secure storage resource management is achieved.

CN120353397APending Publication Date: 2025-07-22INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510487455.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Traditional storage solutions are difficult to flexibly cope with complex and changing business environments, cannot achieve smooth expansion and rapid iteration of storage systems, and cannot meet the needs of modern efficient and flexible storage solutions.

Method used

By obtaining the user's historical log files, calculating the access popularity and popularity threshold of data, performing popularity hierarchy, generating a data distribution mapping table, and configuring a variety of storage resources based on the layered data and mapping tables, including migration of high-speed and low-speed storage media, dynamic adjustment of bandwidth and cache, and differentiated encryption strategies.

Benefits of technology

It realizes dynamic allocation, efficient utilization, flexible configuration and security protection of storage resources, can calmly deal with various complex application scenarios, and improves the performance and resource utilization efficiency of the storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353397A_ABST
    Figure CN120353397A_ABST
Patent Text Reader

Abstract

The invention discloses a storage resource management method and system and a storage medium, and relates to the technical field of storage resource management, and the method comprises the steps: calculating the access popularity and popularity threshold of data according to a historical log file, and carrying out the popularity layering of the data in the historical log file according to the two; the data distribution mapping table is generated based on the historical log file and the access popularity, and finally multiple storage resources are allocated based on the layered data and the data distribution mapping table, so that the problems that application scenes of related technologies are limited, and the performance of a storage system is difficult to flexibly deal with complex and changeable business environments are solved; the technical problem that smooth expansion and rapid iteration of a storage system cannot be achieved is solved, and the technical effects of dynamic allocation, efficient utilization, flexible configuration and safety protection of storage resources and capability of leisurely coping with various complex application scenes are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of storage resource management, and particularly to a storage resource management method, system, and storage medium. Background Art

[0002] Traditional ODM (Original Design Manufacturer) mainly customizes storage solutions through a fixed process, including three stages: requirements analysis, template customization, and simulation optimization. In the requirements analysis stage, in-depth communication is carried out to understand the specific requirements of customers for storage performance, and a corresponding storage medium combination strategy is formulated; in the template customization stage, a series of highly customized parameter templates are pre-built according to the requirements of specific industries; finally, in the simulation optimization stage, simulation tools are used to evaluate the performance and cost-effectiveness of different configuration schemes.

[0003] However, as the business environment becomes complex and changeable and business scenarios iterate rapidly, the related technical methods are difficult to flexibly respond to the changing business environment, the performance of the storage system is difficult to keep up with the rapid development of the business, and it is impossible to achieve smooth expansion and flexible adjustment, and it is difficult to meet the growing needs of modern industries for efficient and flexible storage solutions. Summary of the Invention

[0004] This application provides a storage resource management method, system, and storage medium to at least solve the problems in the related art that the application scenarios of related technologies are limited, the performance of the storage system is difficult to flexibly respond to the complex and changeable business environment, and the smooth expansion and rapid iteration of the storage system cannot be achieved.

[0005] This application provides a storage resource management method, including: obtaining the historical log file of a user; calculating the access heat and heat threshold of data according to the historical log file; performing heat stratification on the data in the historical log file according to the access heat and heat threshold, and generating a data distribution mapping table based on the historical log file and the access heat; and allocating a variety of storage resources based on the stratified data and the data distribution mapping table.

[0006] This application also provides a storage resource management system, including: a hardware layer, on which a variety of pluggable storage resources are deployed; a scheduling layer, the scheduling layer includes a memory for storing computer programs; and a processor for implementing the steps of the above storage resource management method when executing the computer programs.

[0007] This application also provides a computer-readable storage medium, in which a computer program is stored, and wherein the computer program, when executed by a processor, implements the steps of any one of the above storage resource management methods.

[0008] With this application, since the access heat and heat threshold of data can be calculated based on historical log files, the data in the historical log files can be heat-layered by the two to generate a data distribution mapping table. Finally, multiple storage resources are allocated based on the layered data and the data distribution mapping table, achieving efficient, flexible, and secure configuration of storage resources in different scenarios. Therefore, it solves the technical problems in the related art that the application scenarios are limited, the performance of the storage system is difficult to flexibly cope with complex and changeable business environments, and the smooth expansion and rapid iteration of the storage system cannot be achieved, and achieves the technical effects of dynamic allocation, efficient utilization, flexible configuration, and security protection of storage resources, and can calmly handle various complex application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To more clearly illustrate the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0010] Figure 1 FIG. is a flowchart of a storage resource management method provided according to an embodiment of the present application;

[0011] Figure 2 FIG. is a flowchart of generating a data distribution mapping table provided according to an embodiment of the present application;

[0012] Figure 3 FIG. is a data access table provided according to an embodiment of the present application;

[0013] Figure 4 FIG. is a data distribution mapping table provided according to an embodiment of the present application;

[0014] Figure 5 FIG. is a structural diagram of a storage resource management system provided according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0016] It should be noted that in the description of this application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such a process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0017] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.

[0018] Specifically, Figure 1 is a schematic flowchart of a storage resource management method provided by an embodiment of this application. As Figure 1 shown, this method includes the following steps:

[0019] In step S101, obtain the user's historical log file.

[0020] Among them, the user's historical log file includes business type log information, which includes log data such as the access timestamp, access times, storage location, and data popularity of the data.

[0021] It can be understood that the user's historical log file obtained in the embodiment of this application contains business type log information, and the business type log information includes log data such as the access timestamp, access times, storage location, and data popularity of the data.

[0022] In step S102, calculate the access popularity and popularity threshold of the data according to the historical log file.

[0023] Among them, the access popularity of the data refers to a value calculated based on data such as the number of times a certain data object is accessed within a certain time period, which can intuitively indicate the access frequency of the data. The specific calculation method will be described in detail below and will not be elaborated here; the popularity threshold is a dynamic threshold calculated based on relevant data, which refers to the critical value for distinguishing hot data and cold data.

[0024] In the embodiment of this application, calculating the access popularity and popularity threshold of the data according to the historical log file includes: extracting the log data from the historical log file; generating a data access table according to the log data, setting the statistical time window of the data; extracting the data in the statistical time window from the data access table; using the data in the data access table within the target time window to calculate the access popularity and popularity threshold.

[0025] Among them, extracting log data from historical log files means screening out business - type log information from the user's historical log files, and extracting log data such as the access timestamp, access count, storage location, and data popularity of the data from the business - type log information; the statistical time window of the data is specifically set according to actual needs and is not specifically limited here.

[0026] It can be understood that in the embodiment of this application, the access popularity and popularity threshold of data are calculated based on the historical log file. First, log data such as the access timestamp, access count, storage location, and data popularity of the data in the historical log file are extracted, a data access table recording the access situation of each data is generated based on these data, and a suitable statistical time window is set according to actual needs. Then, the data within the set statistical time window is extracted from the data access table, and these data are used to calculate the access popularity and popularity threshold of each data item to distinguish between popular and non - popular data.

[0027] In the embodiment of this application, calculating the access popularity and popularity threshold using the data in the data access table within the target time window includes: identifying the access count, total number of data, and access time of the data within the target time window; calculating the popularity threshold of the data according to the access count and the total number of data; calculating the access popularity according to the access count, the total number of data, and the access time.

[0028] Among them, calculating the popularity threshold of the data according to the access count and the total number of data is calculated by the formula T = μ ± k×σ, where μ is the average value of the data access frequency, and the calculation method is to extract the access count n of each data within the target time window from the data access table i (i represents the i - th data), the total number of data types is m, and the total access count is calculated Thus, the average value of the data access frequency μ = N / m is obtained; σ is the access count of all data within the time window, which is the standard deviation of the same data set as the average value μ, and its calculation method is to obtain the access counts n1, n2,..., n of each data in the data access table m (a total of m data), k is a coefficient used to adjust the tightness of the threshold, with an initial value set to 1. Observe the hit rate of hot data and the system load. If the hit rate is insufficient (such as the proportion of hot data is too low), then lower the value of k; if the load is too high, then raise the value of k; calculating the access popularity according to the access count, the total number of data, and the access time, through the formula is calculated, where n is the data access count; λ is the decay rate parameter that controls the weight decay speed of the historical access record and can be dynamically adjusted. For example, if the access records in the recent 7 days are defined as hot data and it is required that the access weight 7 days ago decays to 20%, then λ = -ln(0.2) / 7≈0.23; t currentis the current time point and is used to calculate the interval between the data access time and the current time; t i is the time stamp of the i-th access to the data.

[0029] It can be understood that the embodiments of the present application can prepare a data access table by identifying the number of accesses, the total number of data, and the access time of the data within the target time window. According to the number of accesses and the total number of data, the heat threshold of the data can be calculated through the above heat threshold calculation formula. At the same time, according to the number of accesses, the total number of data, and the access time, the access heat can be calculated through the above access heat calculation formula. By comparing the access heat of the data with the heat threshold, the activity degree of the data within the target time window can be quantified, which helps to accurately manage the data and optimize the storage strategy.

[0030] In step S103, the data in the historical log file is heat-layered according to the access heat and the heat threshold, and a data distribution mapping table is generated based on the historical log file and the access heat.

[0031] Among them, heat layering refers to the hierarchical management of data according to the data access heat, which will be described in detail below and will not be elaborated here; the data distribution mapping table contains information such as data type (taking a hospital as an example, the data type can include customer case data and staff information data), the number of accesses, the storage location, and the access heat.

[0032] It can be understood that the embodiments of the present application can perform hierarchical management of the data in the historical log file according to the access heat and the heat threshold calculated above, and generate a data distribution mapping table based on this. The mapping table contains information such as data type, the number of accesses, the storage location, and the access heat to help better optimize resource allocation.

[0033] In the embodiments of the present application, heat-layering the data in the historical log file according to the access heat and the heat threshold includes: if the access heat is greater than or equal to the heat threshold, the data is marked as hot data; if the access heat is less than the heat threshold, the data is marked as cold data.

[0034] Among them, hot data refers to the data that is frequently accessed; cold data refers to the data that is rarely accessed. In the embodiments of the present application, the hot data and the cold data are distinguished by the heat threshold.

[0035] It can be understood that, in the embodiments of the present application, data in the historical log file is heat-layered according to the access heat and the heat threshold. First, the access heat of each piece of data is compared with the heat threshold. If the access heat of the data is greater than or equal to the heat threshold, it is marked as hot data, indicating that it has a high access frequency. On the contrary, if the access heat is less than the heat threshold, it is marked as cold data, indicating that its access frequency is low and it may be more suitable for storage in a medium with lower cost. Through this layering method, the distinction between hot and cold data can be effectively achieved, thereby optimizing the storage resource allocation and system performance.

[0036] In step S104, a variety of storage resources are allocated based on the layered data and the data distribution mapping table.

[0037] Among them, the variety of storage resources includes different storage media.

[0038] It can be understood that, after obtaining the data distribution mapping table in the embodiments of the present application, according to the layered data, the data distribution mapping table is used to identify information such as the data type, access times, storage location, access heat, etc. of each piece of data, and the corresponding storage resources are allocated according to the information of the data, ensuring efficient data access performance and optimizing the usage efficiency of the overall storage resources, and achieving reasonable allocation of a variety of storage resources.

[0039] In the embodiments of the present application, allocating a variety of storage resources based on the layered data and the data distribution mapping table includes: identifying the current storage media for storing hot data and cold data; if the current storage media is inconsistent with the target storage media for hot data and cold data, then migrating the hot data and cold data to their respective target storage media; updating the data distribution mapping table after the data migration is completed, and using the data distribution mapping table to allocate bandwidth and cache for different data.

[0040] Among them, the storage media includes high-speed storage media and low-speed storage media. High-speed storage media refer to those storage devices that can provide fast data reading and writing speeds, such as SSD (Solid State Drive, solid-state drive), etc. Low-speed storage media refer to storage devices with relatively slower reading and writing speeds, but usually provide a larger storage capacity and a lower unit storage cost, such as HDD (Hard Disk Drive, hard disk drive), etc.

[0041] It can be understood that after obtaining the data distribution mapping table in the embodiments of the present application, various storage resources are allocated based on the stratified data and the data distribution mapping table. First, the storage media where the hot data and cold data are currently located are identified and compared with their respective target storage media. If a mismatch is found, that is, the hot data is stored on a low-speed storage media or the cold data is stored on a high-speed storage media, a migration operation needs to be performed to move the data to a more suitable target storage media, that is, migrate the hot data to the high-speed storage media and the cold data to the low-speed storage media. After the migration is completed, the data distribution mapping table also needs to be updated according to the new data after migration to reflect the new storage location of each data, so as to optimize the system performance and resource utilization efficiency, help ensure that data with high access frequency can be accessed quickly, and at the same time reduce the storage cost.

[0042] In the embodiments of the present application, after migrating the hot data and cold data to their respective target storage media, it further includes: if the remaining space of the target storage media of the hot data is lower than the elimination threshold, the data in the target storage media of the hot data is eliminated according to the access heat of the hot data.

[0043] Among them, the elimination threshold is specifically set according to the actual situation and will not be specifically limited here.

[0044] It can be understood that in the embodiments of the present application, after completing the migration of the hot data and cold data to their respective target storage media, if it is detected that the remaining space of the high-speed storage media where the hot data is located is lower than the preset elimination threshold, it means that the storage space of the high-speed storage media is insufficient and the data elimination mechanism needs to be started. The data elimination mechanism will sort according to the access heat of the hot data and remove the data with lower access heat in sequence, so as to release the storage space of the high-speed storage media, ensure that the performance-critical hot data can continuously obtain the best storage support, help maintain the efficient operation of the storage system, and optimize resource utilization.

[0045] In the embodiments of the present application, allocating bandwidth and cache for different data by using the data distribution mapping table includes: allocating different bandwidths and caches for the hot data and cold data according to the data distribution mapping table, where the bandwidth and cache allocated for the hot data are both greater than those allocated for the cold data; monitoring the queue depth of the cache and the current proportion of the hot data in the data distribution mapping table, adjusting the bandwidth according to the queue depth, and adjusting the cache according to the current proportion.

[0046] Among them, the bandwidth refers to the data transmission rate of the network or storage interface, which determines how fast the data can be read and written; the cache is a high-speed storage mechanism used to temporarily store frequently accessed data for quickly responding to requests; the queue depth refers to the number of tasks waiting to be processed, such as I / O (input / output requests), which is an indicator to measure the system load. A higher queue depth may mean that the system is facing greater pressure.

[0047] It can be understood that, according to the information in the data distribution mapping table in the embodiments of the present application, different bandwidths and cache resources are allocated to hot data and cold data. Among them, due to its high access frequency, hot data needs to be allocated more bandwidth and a larger cache space to ensure fast response. At the same time, continuously monitor the queue depth of the cache and the proportion of hot data in the total data volume. When it is found that the cache queue depth increases, it indicates that the system may be overloaded. At this time, the bandwidth allocation should be adjusted to relieve the pressure. For example, when it is monitored that the cold data cache queue is saturated, temporarily increase its bandwidth quota to 80% to ensure the normal operation of the cache. In addition, dynamically adjust the cache size based on the current proportion of hot data to ensure that hot data can obtain sufficient cache support. For example, when the proportion of hot data increases, automatically recycle resources from the cold data cache pool, thereby optimizing the performance and efficiency of the overall system, maintaining the stability and high-efficiency operation of the system, and maximizing resource utilization.

[0048] In the embodiments of the present application, after deploying multiple storage resources based on the stratified data and the data distribution mapping table, it further includes: identifying the access heat of the data in the data distribution mapping table; adopting different encryption policies for different data according to the access heat.

[0049] It can be understood that, after deploying storage resources based on the heat level of the data in the embodiments of the present application, the access heat of each piece of data in the data distribution mapping table is identified. According to this access heat information, different encryption policies are applied to different types of data. The specific encryption policies are as follows:

[0050] In the embodiments of the present application, different encryption policies are adopted for different data according to the access heat, including: if the current data is hot data, the first encryption algorithm is used to encrypt the hot data; if the current data is cold data, the second encryption algorithm is used to encrypt the hot data, where the encryption intensity of the second encryption algorithm is higher than that of the first encryption algorithm, and the key rotation period of the second encryption algorithm is higher than that of the first encryption algorithm; if the current data has undergone data migration, the encryption algorithm of the current data is updated.

[0051] Among them, the first encryption algorithm is a lightweight encryption algorithm, such as AES-GCM (AES-Galois / Counter Mode, an authenticated encryption mode based on AES) or ChaCha20 (a stream encryption algorithm); the second encryption algorithm is a high-strength encryption algorithm such as AES-256 (Advanced Encryption Standard) or national cipher SM4 (a cipher algorithm used for wireless local area network products issued by the State Cryptography Administration of China), and the encryption strength is higher than that of the first encryption algorithm; the key rotation period of the first encryption algorithm is a short period, such as rotating every 12 hours, to reduce the risk of key leakage; the key rotation period of the second encryption algorithm is a long period, such as rotating every 30 days.

[0052] It can be understood that after the embodiments of the present application complete the allocation of storage resources based on the data heat level, different encryption strategies are adopted for different data according to the access heat. For cold data, a high-strength encryption algorithm such as AES-256 or national cipher SM4 is used for encryption, a longer key rotation period is set, for example, rotating every 30 days, and these data are stored in a low-cost encrypted storage pool, and the global unified key pool management is used to reduce performance loss; for hot data, a more efficient encryption algorithm such as AES-GCM or ChaCha20 is used for encryption, and a shorter key rotation period is set, for example, rotating every 12 hours, to reduce the risk of key leakage. In addition, independent encryption domain division is implemented for different types of hot data. Taking a hospital as an example, customer case data and staff information data can be divided into independent sub-domains respectively, and each sub-domain uses an independent key management channel, so as to achieve more detailed and flexible security management. When hot data turns into cold data, due to the long storage time of cold data, in order to avoid potential risks caused by algorithm obsolescence, the encryption algorithm upgrade will be triggered and the data will be migrated to the cold data encryption domain; if cold data turns into hot data, the lightweight encryption protocol will be temporarily enabled, and only the running temporary encryption policy needs to be adjusted without modifying the encrypted data, avoiding service interruption. For hot data with high access heat, a higher-level encryption standard and more stringent security measures are adopted to protect it from unauthorized access, while for cold data with lower access heat, a relatively simple encryption method is applied to balance security and cost efficiency, which can not only protect the security of sensitive data, but also optimize the configuration of encryption resources and improve the overall system efficiency and security.

[0053] Through the description of the above implementation manners, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation manner.

[0054] The storage resource management method proposed according to the embodiments of the present application can calculate the access heat and heat threshold of data based on historical log files, perform heat stratification on the data in the historical log files through the two, and generate a data distribution mapping table. Finally, various storage resources are allocated based on the stratified data and the data distribution mapping table, realizing the efficient, flexible and secure configuration of storage resources in different scenarios, achieving the dynamic allocation, efficient utilization, flexible configuration and security protection of storage resources, and being able to calmly handle various complex application scenarios.

[0055] The storage resource management method will be further described below through a specific embodiment.

[0056] This embodiment proposes a storage resource management and data security optimization system based on data stratification. By introducing a modular hardware layer and an intelligent scheduling layer, the efficient, flexible and secure configuration of storage resources is realized. The main functions of each module are as follows:

[0057] Modular hardware layer:

[0058] (1) Pluggable storage media: Support the hybrid deployment of HDD / SSD / NVMe (Non-Volatile Memory Express), and customers can select combinations according to cost / performance requirements.

[0059] (2) Multi-protocol interfaces: Integrate interface protocols such as SATA (Serial Advanced Technology Attachment), NVMe, and SCSI (Small Computer System Interface) to adapt to the communication standards of different terminal devices.

[0060] Intelligent scheduling layer:

[0061] (1) Hot and cold data migration algorithm: Based on the ODM customer historical log information, calculate hot and cold data, and automatically migrate high-priority data to high-speed media.

[0062] (2) Dynamically adjust the I / O bandwidth and cache resources: Dynamically adjust the I / O bandwidth and cache resources based on the data distribution mapping table.

[0063] (3) Dynamically adjust the encryption policy encryption: Divide independent encryption domains according to the data sensitivity level, and use different encryption algorithms for data with different sensitivity levels. High-frequency access data uses low-latency algorithms, and low-frequency data enables high-security algorithms.

[0064] The specific implementation manner of the technical implementation of this embodiment is as follows:

[0065] (1) The modular hardware layer supports users to select storage media according to cost / performance requirements, including HDD / SSD / NVMe, and select interface protocols according to actual usage scenarios. This step is used to provide a basis for the intelligent scheduling layer. If users select the same type of storage media, the cold and hot data migration algorithm of the intelligent scheduling layer is not enabled, and only the encryption policy of cold and hot data is dynamically adjusted, and different I / O bandwidths and cache resources are allocated according to different data types.

[0066] (2) The intelligent scheduling layer first adopts the cold and hot data migration algorithm, supports ODM customers to upload previous log files, parses and calculates cold and hot data based on the log files, sorts these data types, and preferentially migrates data with high access frequency to high-speed storage media, such as NVMe disks, as Figure 2 Generate a flowchart of the data distribution mapping table. Migrate data with low access frequency to low-speed storage media, such as HDD disks. The method for calculating cold and hot data is as follows:

[0067] Step 1: According to the log information, filter out business-class log information, collect log information such as the access timestamp, access times, storage location, and data heat of the data, and generate a data access table, such as Figure 3 Data access table.

[0068] Step 2: According to the log information filtered out in Step 1, calculate the threshold T and access heat H of different data. If the storage log period is very long, it can be divided by month / quarter, and later, as the customer logs increase, data outside the time window is eliminated.

[0069] Step 3: Based on the time window, statistically calculate the access frequency, set a dynamic threshold T to distinguish cold and hot data. For example, taking a hospital as an example, with the past 7 days as the time window, the data with the top 20% of access times is hot data. Its calculation method: T = μ ± k × σ.

[0070] Among them:

[0071] μ: The average value of data access frequency, and its calculation method is as follows:

[0072] (1) Extract the access times n of each data within the time window from the "data access table" generated in Step 1 i (i represents the i-th data).

[0073] (2) Calculate the total access times and the total number of data types. Total number of data types: m. Total access times:

[0074] (3) The average value of data access frequency μ = N / m.

[0075] σ: The number of accesses to all data within the time window, which is the standard deviation of the same data set as the average value μ, and its calculation method is as follows:

[0076] (1) Data set: The number of accesses n1, n2,..., n of each data in the "data access table" generated in step 1 m (There are m data in total).

[0077] (2)

[0078] (3) k: Coefficient (adjust the tightness of the threshold, the initial value is set to 1, observe the hit rate of hot data and the system load. If the hit rate is insufficient (such as the proportion of hot data is too low), then lower the value of k; if the load is too high, then increase the value of k).

[0079] Step 4: Calculate the access heat H of different types of data, and its calculation method: Among them:

[0080] t current : The current time point, used to calculate the interval between the data access time and the current time;

[0081] t i : The timestamp of the i-th access to the data;

[0082] λ: Decay rate parameter, which controls the weight decay speed of historical access records. Dynamically adjust the value of λ, and its calculation method is as follows: The business defines "hot data" as the access records in the recent 7 days. It is required that the access weight 7 days ago decays to 20%, then λ = -ln(0.2) / 7 ≈ 0.23.

[0083] n: The total number of accesses within the statistical window.

[0084] Each time a new access record is added, only the contribution value of the new record to H needs to be calculated, avoiding full recalculation.

[0085] Step 5: If the data heat H ≥ T, mark it as hot data and migrate it to high-speed media; if H < T and it is stored in high-speed media, then migrate it to low-speed media. When the remaining space of the high-speed media is lower than the threshold (such as 5%), eliminate the last data according to the heat value ranking, and generate a data distribution mapping table, such as Figure 4 Data distribution mapping table.

[0086] For example: Assume that a user log window contains 365 days of data, the current μ = 120 times / day, σ = 25, take k = 1.5, then T = 120 + 1.5 × 25 = 157.5. If a certain data block has 3 new accesses today, and its timestamps are t1, t2, t3, then incrementally update the H value: ΔH = e^(-0.01×0) + e^(-0.01×0) + e^(-0.01×0) = 3 (the decay on the same day can be ignored). When H ≥ T finally, it triggers the migration of hot data.

[0087] Step 6: After each data migration, update the data distribution mapping table.

[0088] Step 7: Recalculate μ, σ, and H in full volume monthly / quarterly to avoid the accumulation of incremental calculation errors.

[0089] Secondly, allocate different I / O bandwidths and cache resources according to different data types to improve the utilization rate of storage resources.

[0090] Step 1: For hot data, adopt a fixed high-priority I / O bandwidth (accounting for 60%-70% of the total bandwidth). Memory cache: Allocate 60%-80% of the capacity. SSD cache: Preload associated data. For cold data, adopt elastic allocation of the remaining bandwidth (30%-40%). HDD cache: Only cache metadata (such as file directory structure) and load data blocks on demand.

[0091] Step 2: Dynamically adjust I / O bandwidth and cache resources based on the data distribution mapping table. The dynamic adjustment of I / O bandwidth is achieved by monitoring the SSD / HDD queue depth. When the SSD queue is saturated, temporarily increase its bandwidth quota to 80%. And reserve an emergency bandwidth channel for hot data (10% of the total bandwidth) to ensure low latency for burst access. The dynamic adjustment of cache resources is achieved by monitoring the proportion of hot data. When the proportion of hot data increases, automatically recycle resources from the cold data cache pool (such as a 10% memory expansion step).

[0092] Then, according to the data heat of the data distribution mapping table, adopt different encryption strategies for different types of data, reducing the impact on performance while ensuring the security of customer data.

[0093] Step 1: Adopt lightweight encryption algorithms (such as AES-GCM or ChaCha20) for hot data to balance security and access efficiency. At the same time, set short-period key rotation (such as rotating every 12 hours) to reduce the risk of key leakage. And conduct independent encryption domain division, dividing subdomains according to data types (taking a hospital as an example: customer case domain, staff information domain), and each subdomain uses an independent key management channel.

[0094] Step 2: Use high-strength algorithms (such as AES-256 or national cryptographic SM4) for cold data, and extend the key rotation period to 30 days. Store it in a low-cost encrypted storage pool and adopt global unified key pool management. The above reduces performance loss.

[0095] Step 3: Monitor the data distribution mapping table and upgrade the encryption algorithm for the data with changed hot and cold data. If hot data turns into cold data, since cold data has a long storage time, to avoid potential risks caused by outdated algorithms. At this time, trigger the encryption algorithm upgrade (AES-GCM → AES-256) and migrate it to the cold data encryption domain. Since the cold data key rotation period is long, synchronously update the key management policy (replace the global key pool) at this time, and ensure the consistency of the old and new key systems by re-encrypting all data. If cold data turns into hot data, temporarily enable the lightweight encryption protocol, and only need to adjust the runtime encryption policy without modifying the encrypted data to avoid service interruption.

[0096] Through the above solution, not only can hot and cold data be identified in real time, but also data security can be guaranteed and the response speed can be improved with the least impact on performance.

[0097] An embodiment of the present application also provides a storage resource management system. Figure 5 For the structural schematic diagram of the storage resource management system provided by the embodiment of the present application, as Figure 5 shown, the system includes: a hardware layer 201, on which a variety of pluggable storage resources 202 are deployed; a scheduling layer 203, the scheduling layer 203 includes a memory 204 for storing computer programs; a processor 205 for implementing the steps of the above storage resource management method when executing the computer programs.

[0098] Among them, the variety of pluggable storage resources include pluggable storage media, which can support the hybrid deployment of HDD / SSD / NVMe, and any combination thereof can be selected according to cost / performance requirements; it includes multi-protocol interfaces, integrating interface protocols such as SATA, NVMe, and SCSI to adapt to the communication standards of different terminal devices.

[0099] It can be understood that on the hardware layer 201 of the embodiment of the present application, a variety of pluggable storage resources 202 are deployed, which can support the hybrid deployment of HDD / SSD / NVMe, and at the same time include multi-protocol interfaces, integrating interface protocols such as SATA, NVMe, and SCSI, and can adapt to the communication standards of different terminal devices. The embodiment of the present application also includes a scheduling layer 203 for storing computer programs and implementing the steps of the above storage resource management method when executing the computer programs.

[0100] For the description of the features in the embodiment corresponding to the storage resource management system, reference can be made to the relevant description of the embodiment corresponding to the storage resource management method, which will not be elaborated here one by one.

[0101] The storage resource management system proposed according to the embodiments of the present application can, through the hardware layer and the scheduling layer, calculate the access heat and heat threshold of data based on historical log files, stratify the data in the historical log files according to the two, generate a data distribution mapping table, and finally allocate various storage resources based on the stratified data and the data distribution mapping table, achieving efficient, flexible, and secure configuration of storage resources in different scenarios, achieving the technical effects of dynamic allocation, efficient utilization, flexible configuration, and security protection of storage resources, and being able to calmly handle various complex application scenarios.

[0102] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. Among them, the computer program is set to execute the steps in any one of the above embodiments of the storage resource management method when running.

[0103] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.

[0104] An embodiment of the present application also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above embodiments of the storage resource management method.

[0105] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above embodiments of the storage resource management method.

[0106] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0107] The above has introduced in detail a storage resource management method, system and storage medium provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A storage resource management method, characterized in that, The method includes: Obtaining the historical log file of the user; Calculating the access heat and heat threshold of the data according to the historical log file; Performing heat stratification on the data in the historical log file according to the access heat and the heat threshold, and generating a data distribution mapping table based on the historical log file and the access heat; Allocating multiple storage resources based on the stratified data and the data distribution mapping table.

2. The storage resource management method according to claim 1, wherein The calculating the access heat and heat threshold of the data according to the historical log file includes: Extracting the log data in the historical log file; Generating a data access table according to the log data and setting a statistical time window for the data; Extracting the data in the data access table within the statistical time window; Calculating the access heat and heat threshold by using the data in the data access table within the target time window.

3. The storage resource management method according to claim 2, wherein The calculating the access heat and heat threshold by using the data in the data access table within the target time window includes: Identifying the access times, total number of data, and access time of the data within the target time window; Calculating the heat threshold of the data according to the access times and the total number of data; Calculating the access heat according to the access times, the total number of data, and the access time.

4. The storage resource management method according to claim 1, characterized in that, The performing heat stratification on the data in the historical log file according to the access heat and the heat threshold includes: If the access heat is greater than or equal to the heat threshold, marking the data as hot data; If the access heat is less than the heat threshold, marking the data as cold data.

5. The storage resource management method according to claim 1, wherein The allocating multiple storage resources based on the stratified data and the data distribution mapping table includes: Identifying the current storage media for storing hot data and cold data; If the current storage media are inconsistent with the target storage media for the hot data and the cold data, migrating the hot data and the cold data to their respective target storage media; After completing the data migration, updating the data distribution mapping table, and using the data distribution mapping table to allocate bandwidth and cache for different data.

6. The storage resource management method according to claim 5, wherein, After migrating the hot data and the cold data to their respective target storage media, it further includes: If the remaining space of the target storage media for the hot data is lower than the elimination threshold, eliminating the data in the target storage media for the hot data according to the access heat of the hot data.

7. The storage resource management method according to claim 5, characterized in that, The using the data distribution mapping table to allocate bandwidth and cache for different data includes: Allocating different bandwidth and cache for the hot data and the cold data according to the data distribution mapping table, where the bandwidth and cache allocated for the hot data are both greater than those allocated for the cold data; Monitoring the queue depth of the cache and the current proportion of the hot data in the data distribution mapping table, adjusting the bandwidth according to the queue depth, and adjusting the cache according to the current proportion.

8. The storage resource management method according to claim 1, characterized in that After allocating multiple storage resources based on the stratified data and the data distribution mapping table, it further includes: Identifying the access heat of the data in the data distribution mapping table; Adopting different encryption strategies for different data according to the access heat.

9. A storage resource management system, characterized in that, It includes: A hardware layer, on which multiple pluggable storage resources are deployed; The scheduling layer, the scheduling layer includes a memory for storing a computer program; A processor for implementing the steps of the storage resource management method according to any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the storage resource management method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Cited By

  • Data caching method, system and device and storage medium

    CN120583063A

  • A data caching method, system, device and storage medium

    CN120583063B

  • Large model training resource optimization method and system based on Hadoop ecology

    CN121070632A

  • Big model training resource optimization method and system based on hadoop ecosystem

    CN121070632B

  • Distributed hierarchical data storage system and method

    CN121143721A