Data full-flash storage optimization method and system based on cloud computing
By building a dynamic access popularity prediction model and generating sharded storage topology maps and cache preloading strategy matrix, the data storage distribution of all-flash storage clusters is optimized, the problem of low data reading and writing efficiency in the existing technology is solved, and more efficient data storage and access is achieved.
Patent Information
- Application Number
- CN202510312385.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The prior art is difficult to effectively optimize the data storage distribution of all-flash storage clusters, resulting in low data reading and writing efficiency and lack of in-depth analysis of data block characteristics and historical access conditions.
By collecting multi-dimensional performance data of all-flash storage clusters in real time, analyzing the storage operation request flow of user terminals, traversing the historical access record library to extract the historical access timing characteristics and physical storage location proximity of data blocks, building a dynamic access popularity prediction model, and generating a sharded storage topology diagram across nodes and a cache preload strategy matrix.
It realizes scientific storage distribution of data blocks, improves the read and write performance of data, reduces the degree of fragmentation of storage space, and significantly improves the system's response speed and user experience.
Smart Images

Figure CN120179176A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud computing, and in particular, to a method and system for optimizing all-flash storage of data based on cloud computing. Background Art
[0002] With the rapid development of cloud computing technology, the amount of data has shown an explosive growth, posing higher requirements for the performance, reliability, and efficiency of data storage. All-flash storage has gradually become an important choice for data storage in cloud computing environments due to its high-speed data read and write performance. However, in practical applications, all-flash storage systems face many challenges, and existing technical solutions are difficult to meet the increasing complex requirements.
[0003] In the current field of cloud computing data storage, for the optimization of all-flash storage clusters, when processing storage operation requests from user terminals, traditional methods usually simply allocate storage locations according to fixed rules, lacking in-depth analysis of the characteristics of data blocks in the requests and historical access situations. For example, the historical access timing characteristics of data blocks to be stored are not considered, resulting in a lack of scientific distribution of data storage and an inability to reasonably arrange according to the actual access frequency and pattern of data, thereby affecting the read and write efficiency of data. At the same time, existing technologies also rarely pay attention to the physical storage location proximity of associated data blocks, leading to a relatively scattered distribution of data in the storage cluster, increasing the seek time and network transmission overhead during data access. Summary of the Invention
[0004] In view of the problems mentioned above, in combination with the first aspect of the present invention, embodiments of the present invention provide a method for optimizing all-flash storage of data based on cloud computing, the method comprising: Real-time collecting a multi-dimensional performance data set of each node in the all-flash storage cluster, the multi-dimensional performance data set including the fragmentation rate of the node storage space, the depth of the input / output request queue, and the network link bandwidth utilization rate; Receiving a storage operation request stream sent by a user terminal, and parsing the set of identifiers of data blocks to be stored and the corresponding operation mode tags in the storage operation request stream; Based on the set of identifiers of data blocks to be stored, traversing the historical access record library, and extracting the historical access timing characteristics of the data blocks to be stored and the physical storage location proximity of associated data blocks; According to the multi-dimensional performance data set and the historical access timing characteristics, constructing a dynamic access heat prediction model for the data blocks to be stored within a preset time period; Based on the dynamic access heat prediction model and the physical storage location proximity, generating a cross-node sharded storage topology map and a cache preloading strategy matrix.
[0005] For example, traversing the historical access record library based on the set of identifiers of data blocks to be stored, and extracting the historical access time sequence features of the data blocks to be stored and the physical storage location proximity of associated data blocks, includes: According to each data block identifier in the set of identifiers of data blocks to be stored, matching the corresponding historical access record in the historical access record library, where the historical access record includes a set of access timestamps and a subset of associated data block identifiers; Extracting the access timestamp sequence in the historical access record, and calculating the periodic access interval distribution parameter and the set of timestamps of burst access events based on the access timestamp sequence; Filtering out from the subset of associated data block identifiers a set of associated data block identifiers that have a co - access count exceeding a cooperation threshold with the data blocks to be stored within a preset time window; Obtaining the set of physical storage coordinates of the set of associated data block identifiers in the all - flash storage cluster, where the set of physical storage coordinates includes storage node identifiers and logical unit addresses; Generating the physical storage location proximity matrix of the associated data blocks based on the topological connection relationship between the storage node identifiers and the distance difference between the logical unit addresses in the set of physical storage coordinates.
[0006] On the other hand, an embodiment of the present invention further provides a data all - flash storage optimization system based on cloud computing, including a processor and a machine - readable storage medium. The machine - readable storage medium is connected to the processor. The machine - readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine - readable storage medium to implement the above - mentioned method.
[0007] Based on the above aspects, after the embodiments of the present application collect the multi-dimensional performance data sets of each node in the all-flash storage cluster in real time, they receive the storage operation request stream of the user terminal and parse the set of to-be-stored data block identifiers and operation mode tags therein. At the same time, based on the set of to-be-stored data block identifiers, they traverse the historical access record library, extract the historical access timing characteristics and the physical storage location proximity of associated data blocks, deeply mining the historical access characteristics of the data blocks themselves and their spatial relationships with surrounding data blocks. Furthermore, according to the multi-dimensional performance data sets and the historical access timing characteristics, a dynamic access heat prediction model for the to-be-stored data blocks within a preset time period is constructed. This dynamic access heat prediction model integrates real-time performance data and historical access rules, can dynamically and accurately predict the access heat of data blocks in the future period of time, can better adapt to the variability and complexity of data access in the cloud computing environment, and effectively improves the ability to predict data access trends. Finally, based on the dynamic access heat prediction model and the physical storage location proximity, a cross-node sharded storage topology map and a cache preloading strategy matrix are generated, realizing the optimized allocation and efficient utilization of storage resources. The cross-node sharded storage topology map takes into account the access heat and physical storage location relationships of data blocks, reasonably plans the storage distribution of data among different nodes, effectively balances the load pressure of each node, reduces the degree of storage space fragmentation, and improves the read and write performance of the entire storage cluster. At the same time, the cache preloading strategy matrix preloads the data blocks that may be frequently accessed into the cache in advance according to the predicted access heat, greatly reducing the waiting time for data access and significantly improving the system response speed and user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 is a schematic execution flow diagram of the all-flash storage optimization method based on cloud computing provided by an embodiment of the present invention.
[0009] Figure 2 is a schematic diagram of exemplary hardware and software components of the all-flash storage optimization system based on cloud computing provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0010] The present invention will be specifically described below in conjunction with the accompanying drawings of the specification. Figure 1 is a schematic flow diagram of the all-flash storage optimization method based on cloud computing provided by an embodiment of the present invention. The all-flash storage optimization method based on cloud computing will be introduced in detail below.
[0011] Step S110, collect the multi-dimensional performance data sets of each node in the all-flash storage cluster in real time, where the multi-dimensional performance data sets include the node storage space fragmentation rate, the input / output request queue depth, and the network link bandwidth utilization rate.
[0012] Specifically, for the data center of a certain enterprise, an all-flash storage cluster can be deployed to store a large amount of business data, such as the enterprise's financial data, customer information, sales records, etc. The all-flash storage cluster contains multiple nodes, and each node is responsible for storing and processing a part of the data.
[0013] In this scenario, the process of collecting the fragmentation rate of the node storage space is as follows: As the enterprise continuously performs data storage, deletion, and modification operations, the storage space in the node will gradually become fragmented. For example, when the finance department frequently updates the financial statement data, many small pieces of free space may be formed in the storage space instead of a continuous large piece of free space. By calculating the proportion of these fragmented spaces, the fragmentation rate of the node storage space can be obtained.
[0014] Regarding the depth of the input / output request queue, assume that the sales department of the enterprise conducts centralized statistics and reporting of sales data at the end of each month. During this period, a large number of read / write requests will be sent to the all-flash storage cluster. Each node will receive these requests and place the requests in the input / output request queue in the order of arrival. Thus, the number of requests in this queue is monitored in real time, and this number is the depth of the input / output request queue. For example, during the peak period of sales data processing, the depth of the input / output request queue of a node may reach several hundred or even thousands of requests.
[0015] The collection of the network link bandwidth utilization rate is related to the transmission of data between nodes. Specifically, different departments of the enterprise may share data. When the customer service department retrieves customer information from the storage cluster to provide services to customers, the data needs to be transmitted between nodes through the network link. During this process, the ratio of the actual data transmission volume per unit time on the network link to the total bandwidth of the link can be monitored, and this is the network link bandwidth utilization rate. For example, if the total bandwidth of the network link is 1000 Mbps and the actual data transmission volume at a certain moment is 500 Mbps, then the network link bandwidth utilization rate at this time is 50%.
[0016] Step S120: Receive the storage operation request stream sent by the user terminal, and parse the set of data block identifiers to be stored and the corresponding operation mode tags in the storage operation request stream.
[0017] Continuing with the above example of the enterprise data center, the user terminal can be devices such as computers and servers used by employees in various departments within the enterprise. Assume that the marketing department of the enterprise wants to store a new market research report in the all-flash storage cluster. The employees of the marketing department send a storage operation request stream to the storage cluster through specific software (user terminal).
[0018] The storage operation request stream contains a set of identifiers of data blocks to be stored. For example, a market research report may be divided into multiple data blocks for storage, and each data block has a unique identifier. At the same time, the request stream also carries the corresponding operation mode label, and here the operation mode label may be "write", indicating that this is a storage (write) operation. Thus, after receiving this storage operation request stream, it will be parsed to accurately extract the set of identifiers of data blocks to be stored and the operation mode label "write" for subsequent processing based on this information.
[0019] Step S130, traverse the historical access record library based on the set of identifiers of data blocks to be stored, and extract the historical access time sequence characteristics of the data blocks to be stored and the physical storage location proximity of associated data blocks.
[0020] Still based on the enterprise data center scenario, for the storage operation of a market research report. For example, each data block identifier in the set of identifiers of data blocks to be stored can be matched in the historical access record library. Assume that the historical access record library records the access situations of all past enterprise data.
[0021] For the extraction of historical access time sequence characteristics, taking one of the data blocks as an example, if its past access records show that it was frequently accessed from 9 am to 10 am on each working day, this constitutes an access timestamp sequence. Based on this access timestamp sequence, the periodic access interval distribution parameter can be calculated. For example, it is found that there is an access peak every 7 days on average, and this is the periodic access interval distribution parameter. At the same time, if there are a large number of sudden accesses during a special event (such as during the release of a company's new product), the timestamp set of this sudden access event can be recorded.
[0022] For the physical storage location proximity of associated data blocks, from the subset of associated data block identifiers, filter out the set of associated data block identifiers that have been co-accessed more than the collaboration threshold (assumed to be 10 times) with the data blocks to be stored within a preset time window (such as the past month). For example, some market analysis data blocks related to the market research report have been co-accessed 15 times in the past month. Then obtain the set of physical storage coordinates of these associated data block identifiers in the all-flash storage cluster. Assume that the data block of the market research report is stored at the logical unit address 100 of node A, while the associated market analysis data blocks are stored at the logical unit addresses 105 of node A and 200 of node B. Based on the topological connection relationship between these storage node identifiers (node A and node B are connected by a high-speed network) and the distance difference between the logical unit addresses (the distance between logical unit addresses 100 and 105 is relatively close, and the distance from 200 is relatively far), generate the physical storage location proximity matrix of the associated data blocks.
[0023] Step S140: Construct a dynamic access popularity prediction model for the data block to be stored within a preset time period based on the multi-dimensional performance data set and the historical access timing characteristics.
[0024] Still taking the above scenario as an example, extract the periodic access peak intervals (such as from 9 am to 10 am every weekday), the timestamps of sudden access events (such as the access peak during the new product release period), and the statistical values of access interval distributions (such as an access peak every 7 days on average) from the historical access timing characteristics.
[0025] Suppose the operation mode label corresponding to the market research report is "write", but some of the data it contains may be frequently read later for market strategy adjustment. Identify the read-write operation ratio corresponding to the operation mode label. When the read operation ratio exceeds a preset threshold (assumed to be 60%), activate the hot data prediction flag. For example, after analysis, it is found that there is an 80% probability that the market share analysis part in the market research report will be read later, exceeding the preset threshold of 60%, so the hot data prediction flag is activated.
[0026] Combine the periodic access peak intervals and the timestamps of sudden access events to generate a time-dimensional access probability density function. For example, if the access probability from 9 am to 10 am is 0.3 and the access probability corresponding to the sudden access event during the new product release period is 0.2, construct a time-dimensional access probability density function based on this data.
[0027] Based on the statistical values of access interval distributions and the hot data prediction flag, correct the weight parameters of the time-dimensional access probability density function. Suppose the statistical values of access interval distributions indicate that the access probability gradually decreases between two access peaks, and adjust the weight parameters of the time-dimensional access probability density function according to this situation.
[0028] Couple the corrected time - dimension access probability density function with the network link bandwidth utilization rate in the multi - dimensional performance data set. For example, extract the bandwidth fluctuation data set of the network link bandwidth utilization rate within multiple historical time windows from the multi - dimensional performance data set. Assume that the peak bandwidth utilization rate interval from 9:00 to 10:00 am on weekdays is 80%, and the average transmission rate is 500 Mbps. Align and map the peak bandwidth utilization rate interval with the time axis of the time - dimension access probability density function to generate a time - synchronized bandwidth utilization rate distribution sequence. Based on the average transmission rate, calculate the bandwidth weight factor of the time - dimension access probability density function for each time unit. Weight - correct the access probability values of the time - dimension access probability density function for the corresponding time units according to the bandwidth weight factor to generate a bandwidth - aware access probability density distribution function. Perform time - dimension normalization processing on the bandwidth - aware access probability density distribution function, and generate the heat - level parameter of the dynamic access heat prediction model based on the maximum probability value after normalization and the slope change of the probability distribution curve. This heat - level parameter can be used to represent the access heat condition of the market research report data block within a preset future time period.
[0029] Step S150: Generate a cross - node sharding storage topology graph and a cache pre - loading policy matrix based on the dynamic access heat prediction model and the physical storage location proximity.
[0030] Specifically, for the storage of the market research report data block, determine the sharding redundancy threshold and the minimum number of replicas of the data block to be stored according to the heat - level parameter output by the dynamic access heat prediction model. Assume that the heat level is relatively high, determine the sharding redundancy threshold to be 3 (indicating that the data block can be divided into at most 3 shards for storage), and the minimum number of replicas to be 2 (each shard has at least 2 replicas).
[0031] Traverse the storage space fragmentation rate and the input / output request queue depth of all nodes in the all - flash storage cluster, and filter out a subset of candidate nodes that meet the shard capacity constraints. For example, if the storage space fragmentation rate of node A is relatively low, the input / output request queue depth is within an acceptable range, and there is enough space to store the shards of the market research report data block, then node A may be selected into the subset of candidate nodes.
[0032] Based on the physical storage location proximity, calculate the storage location association score of each node in the subset of candidate nodes. Assume that node A is close to the storage location of the associated data block and has a relatively high storage location association score. Construct a storage path weight table between nodes according to the storage location association score and the network link bandwidth utilization rate in the multi - dimensional performance data set. For example, if the network link bandwidth utilization rate from node A to node B is high and the storage location association score is high, then the storage path weight value between them is relatively high.
[0033] Generate a sharded storage topology map with redundant path cross - connections based on the storage path weight table and the minimum number of replicas. For example, according to the storage path weight values between nodes in the storage path weight table, filter out a set of candidate storage paths whose weight values exceed a preset weight threshold (assumed to be 0.6). Determine that the number of sharded replicas of the market research report data block is 2 based on the minimum number of replicas, and allocate initial storage paths for each sharded replica. Suppose the storage path from node A to node B is first allocated for the first sharded replica. Traverse the storage path weight values in the set of candidate storage paths and find that the path weight value from node A to node B is the highest, so it is used as the default storage path for the sharded replica. Detect the connection status between nodes of this default storage path. If it is found that there is only one connection path from node A to node B, there is a single - point failure risk. Then select the backup storage path with the second - highest weight value (assumed that the path weight value from node A to node C is the second - highest) from the set of candidate storage paths as the cross - redundant path. Connect the cross - redundant path and the default storage path bidirectionally to generate a sharded storage topology map with redundant path cross - connections. Verify whether the path redundancy of each sharded replica in the sharded storage topology map meets the redundancy constraint conditions corresponding to the minimum number of replicas. If not, re - select the backup storage path and update the cross - connection relationship of the sharded storage topology map.
[0034] Generate a cache pre - loading policy matrix based on the dynamic access heat prediction model and the physical storage location proximity. Identify a sequence of target data blocks that have a spatial association with the market research report data block according to the physical storage location proximity, such as the market analysis data block mentioned before. Based on the dynamic access heat prediction model, predict the concurrent access probability distribution of the sequence of target data blocks within a preset time period. Suppose the market analysis data block has a 60% probability of being accessed simultaneously with the market research report data block in the next week. Calculate the pre - loading priority coefficient for each target data block according to the concurrent access probability distribution and the node storage space fragmentation rate in the multi - dimensional performance data set. If the node storage space fragmentation rate is low, the pre - loading priority coefficient may be high. Based on the pre - loading priority coefficient, allocate different cache retention periods and compression level parameters for the sequence of target data blocks. For example, for the market analysis data block with a high pre - loading priority coefficient, obtain the current cache space occupancy rate and the historical cache replacement frequency of the edge nodes in the all - flash storage cluster. When the concurrent access probability distribution is higher than the first preset threshold (assumed to be 50%), allocate a fixed retention period (such as 3 days) for the corresponding data block and lock the cache space. When the concurrent access probability distribution is lower than the second preset threshold (assumed to be 30%), dynamically adjust the decay rate of the retention period based on the historical cache replacement frequency. Perform variable compression rate processing on low - priority data blocks according to the compression level parameters, and record the metadata verification information of the compressed data blocks.
[0035] Based on the above steps, after the embodiments of the present application collect the multi-dimensional performance data sets of each node in the all-flash storage cluster in real time, they receive the storage operation request stream of the user terminal and parse the set of to-be-stored data block identifiers and operation mode tags therein. At the same time, based on the set of to-be-stored data block identifiers, they traverse the historical access record library, extract the historical access timing characteristics and the physical storage location proximity of associated data blocks, deeply mine the historical access characteristics of the data blocks themselves and their spatial relationships with surrounding data blocks. Furthermore, according to the multi-dimensional performance data sets and the historical access timing characteristics, a dynamic access heat prediction model for the to-be-stored data blocks within a preset time period is constructed. This dynamic access heat prediction model integrates real-time performance data and historical access rules, can dynamically and accurately predict the access heat of data blocks in the next period of time, can better adapt to the variability and complexity of data access in the cloud computing environment, and effectively improves the ability to predict data access trends. Finally, based on the dynamic access heat prediction model and the physical storage location proximity, a cross-node sharding storage topology map and a cache preloading strategy matrix are generated, realizing the optimized allocation and efficient utilization of storage resources. The cross-node sharding storage topology map takes into account the access heat and physical storage location relationship of data blocks, reasonably plans the storage distribution of data among different nodes, effectively balances the load pressure of each node, reduces the degree of storage space fragmentation, and improves the read and write performance of the entire storage cluster. At the same time, the cache preloading strategy matrix preloads the data blocks that may be frequently accessed into the cache in advance according to the predicted access heat, greatly reducing the waiting time for data access and significantly improving the system response speed and user experience.
[0036] For example, in a possible implementation manner, step S130 includes: Step S131, according to each data block identifier in the set of to-be-stored data block identifiers, match the corresponding historical access record in the historical access record library, where the historical access record includes a set of access timestamps and a subset of associated data block identifiers.
[0037] In the enterprise data center scenario, during the process related to the storage operation of data blocks in market research reports, first, according to each data block identifier of the market research report, the corresponding historical access records are matched in the historical access record library. Since the historical access record library details the past access situations of all enterprise data, the records corresponding to the data blocks in the market research report can be accurately found. These historical access records contain a set of access timestamps and a subset of associated data block identifiers. For example, for a specific data block in the market research report, the set of access timestamps in its historical access record shows the specific time of each past access, which may include a series of time points after a certain large-scale marketing activity carried out by the enterprise, such as 10:15 on March 1, 2023, 14:30 on March 5, 2023, etc. The subset of associated data block identifiers contains the identifiers of other data blocks that may be involved when accessing this data block.
[0038] Step S132: Extract the sequence of access timestamps in the historical access record, and calculate the periodic access interval distribution parameter and the set of timestamps of burst access events based on the sequence of access timestamps.
[0039] For example, the sequence of access timestamps of the specific data block in the above market research report is 10:15 on March 1, 2023, 14:30 on March 5, 2023, etc. in chronological order. Calculate the periodic access interval distribution parameter based on this sequence of access timestamps. The calculation process is as follows: First, count the time intervals between adjacent access timestamps. For example, the time interval from 10:15 on March 1, 2023 to 14:30 on March 5, 2023 is 4 days, 4 hours, and 15 minutes. Perform such calculations for all adjacent access timestamps. Then analyze these time intervals, count the frequency of each time interval. If the time interval of 4 days, 4 hours, and 15 minutes appears 3 times, which is the highest frequency, then this time interval is used as an important reference for the periodic access interval. Through the statistics and analysis of all time intervals in this way, the periodic access interval distribution parameter is obtained, which can reflect the periodic pattern of the data block being accessed in different time periods. At the same time, in this process, the set of timestamps of burst access events can also be determined. For example, when the enterprise conducts an annual financial audit, it needs to refer to a large amount of data in the market research report, resulting in a large number of concentrated accesses from April 1, 2023 to April 5, 2023. The timestamps during this period constitute the set of timestamps of burst access events.
[0040] Step S133: Screen out from the subset of associated data block identifiers the set of associated data block identifiers whose common access times with the data block to be stored exceed the collaboration threshold within a preset time window.
[0041] Assume that the preset time window is the past 6 months and the collaboration threshold is set to 15 times. Check the co - access times corresponding to each identifier within the subset of associated data block identifiers. For example, there is a market analysis data block that has been co - accessed with a market research report data block 20 times in the past 6 months, which exceeds the collaboration threshold of 15 times. Then it is screened out. These screened - out data block identifiers form the set of associated data block identifiers.
[0042] Step S134, obtain the set of physical storage coordinates of the set of associated data block identifiers in the all - flash storage cluster. The set of physical storage coordinates includes storage node identifiers and logical unit addresses.
[0043] For example, for the screened - out market analysis data block, its storage node identifier in the all - flash storage cluster is node A and the logical unit address is 105. There may be other associated data blocks stored in node B with a logical unit address of 200, etc. These storage node identifiers and logical unit addresses form the set of physical storage coordinates.
[0044] Step S135, generate a physical storage location proximity matrix for the associated data blocks based on the topological connection relationship between the storage node identifiers and the distance difference between the logical unit addresses in the set of physical storage coordinates.
[0045] For example, regarding the topological connection relationship between storage node identifiers, if there is a high - speed network connection with low network latency between node A and node B, then their topological connection relationship is relatively close. For the distance difference between logical unit addresses, such as the distance difference between the logical unit address 105 of node A and another logical unit address 110 of node A is relatively small, while the distance difference from the logical unit address 200 of node B is relatively large. By comprehensively considering these factors, the relationship between each pair of associated data blocks is quantitatively evaluated, thereby constructing a physical storage location proximity matrix for the associated data blocks. This physical storage location proximity matrix can accurately reflect the proximity relationship of each associated data block in terms of physical storage location, providing an important basis for subsequent data storage and management strategies.
[0046] In a possible implementation, step S140 includes: Step S141, extract the periodic access peak interval, burst access event timestamps, and access interval distribution statistical values from the historical access timing characteristics.
[0047] Specifically, regarding the periodic access peak interval, by reviewing the historical access records, it is found that the access volume of this data block is relatively large between 9:00 am and 10:00 am on each working day, and this time period is the periodic access peak interval. Regarding the time stamps of sudden access events, for example, during the period when the enterprise adjusts its quarterly sales strategy, there are a large number of sudden accesses to the market research report data block from June 15, 2023 to June 20, 2023, and these time points are the time stamps of sudden access events. For the statistical values of access interval distributions, the time intervals between each access are statistically analyzed. For example, the interval between the first access and the second access is 3 days, and the interval between the second and the third access is 5 days, etc. Through statistical analysis of a large number of such interval data, the statistical values of access interval distributions are obtained.
[0048] Step S142, identify the read-write operation ratio corresponding to the operation mode label. When the proportion of read operations in the read-write operation ratio exceeds a preset threshold, activate the hot data prediction flag.
[0049] Specifically, the market research report data block has an operation mode label when stored. Suppose it is found through analysis that 70% of the operations are read operations. Set the preset threshold to 60%. Since the proportion of read operations is 70% which exceeds the preset threshold of 60%, the hot data prediction flag is activated.
[0050] Step S143, combine the periodic access peak interval and the time stamps of sudden access events to generate a time-dimensional access probability density function.
[0051] Illustrated by a simple calculation process, assume that the initial access probability in the periodic access peak interval (9:00 am to 10:00 am on working days) is set to 0.3, and the access probability during the period from June 15, 2023 to June 20, 2023 corresponding to the time stamps of sudden access events is set to 0.2. For other time periods, corresponding probability values are set according to factors such as historical access frequencies, thus constructing a preliminary time-dimensional access probability density function.
[0052] Step S144, based on the statistical values of access interval distributions and the hot data prediction flag, correct the weight parameters of the time-dimensional access probability density function.
[0053] Suppose the statistical values of access interval distributions indicate that the access probability gradually decreases between two access peaks. For example, in a period of time after the periodic access peak interval, the access probability decreases by 0.05 every day. Since the hot data prediction flag is activated, indicating that this data block may be hot data, for this situation, the weight parameters need to be adjusted to slow down the decay rate of the access probability after the peak. For example, the original daily decay of 0.05 is adjusted to a daily decay of 0.03, thereby correcting the weight parameters of the time-dimensional access probability density function.
[0054] Step S145: Couplingly calculate the corrected time - dimension access probability density function with the network link bandwidth utilization rate in the multi - dimensional performance data set to generate the dynamic access heat prediction model, and the dynamic access heat prediction model outputs the heat level parameter of the data block to be stored.
[0055] For example, step S145 includes: Step S1451: Extract the bandwidth fluctuation data set of the network link bandwidth utilization rate within multiple historical time windows from the multi - dimensional performance data set. The bandwidth fluctuation data set includes the bandwidth utilization rate peak interval and the average transmission rate.
[0056] Step S1452: Align and map the bandwidth utilization rate peak interval in the bandwidth fluctuation data set to the time axis of the time - dimension access probability density function to generate a time - synchronized bandwidth utilization rate distribution sequence.
[0057] Step S1453: Calculate the bandwidth weight factor of the time - dimension access probability density function for each time unit based on the average transmission rate in the bandwidth utilization rate distribution sequence.
[0058] Step S1454: Weight - correct the access probability value of the time - dimension access probability density function for the corresponding time unit according to the bandwidth weight factor to generate a bandwidth - aware access probability density distribution function.
[0059] Step S1455: Perform time - dimension normalization processing on the bandwidth - aware access probability density distribution function, and generate the heat level parameter of the dynamic access heat prediction model based on the maximum probability value and the slope change of the probability distribution curve of the normalized bandwidth - aware access probability density distribution function.
[0060] For example, select the past 10 working days as the historical time window and count the network link bandwidth utilization rate for each working day. Among them, the peak interval of the bandwidth utilization rate may occur from 10 am to 11 am, with the peak reaching 80%, and the average transmission rate being 500 Mbps. Align and map the peak interval of the bandwidth utilization rate in the bandwidth fluctuation dataset to the time axis of the time dimension access probability density function to generate a time-synchronized bandwidth utilization rate distribution sequence. Assume that the time unit of the time dimension access probability density function is hours, the access probability from 9 am to 10 am is 0.3, and the peak bandwidth utilization rate from 10 am to 11 am is 80%. The process of calculating the bandwidth weight factor of the time dimension access probability density function for each time unit is as follows: First, determine the influence degree of the relationship between the average transmission rate and the peak bandwidth utilization rate on the access probability. If the average transmission rate is higher and the peak bandwidth utilization rate is higher, it indicates that the data transmission demand is large and the network resources are fully utilized during this time period, and the bandwidth weight factor of this time unit should be larger. For example, in a simple calculation method, if the average transmission rate is 500 Mbps and the peak bandwidth utilization rate is 80%, and a base weight is set to 1, then the bandwidth weight factor calculated based on these two values may be 1.2 (this is just an example calculation method, and the actual calculation will be based on more complex logic). According to this bandwidth weight factor, the access probability value of the time dimension access probability density function in the corresponding time unit is weighted and corrected. For example, the original access probability from 9 am to 10 am is 0.3, and after weighted correction, it becomes 0.3×1.2 = 0.36, thus generating a bandwidth-aware access probability density distribution function.
[0061] Furthermore, the normalization process is to adjust the access probability values of all time units to a specific range, such as between 0 and 1. The calculation process is as follows: First, find the maximum probability value in the bandwidth-aware access probability density distribution function, assume it is 0.4. Then divide the access probability value of each time unit by this maximum probability value. For example, if the original access probability of a certain time unit is 0.2, after normalization, it becomes 0.2÷0.4 = 0.5. Based on the maximum probability value of the normalized bandwidth-aware access probability density distribution function and the slope change of the probability distribution curve, generate the heat level parameter of the dynamic access heat prediction model. If the maximum probability value is close to 1 and the slope change of the probability distribution curve is large, it indicates that the data block has a high access heat and obvious heat changes in certain time periods, and the heat level parameter may be set to a higher level, such as level 3 (here it is assumed that the heat level is divided into levels 1 to 5); if the maximum probability value is small and the slope change is gentle, the heat level parameter may be set to a lower level, such as level 1. This heat level parameter can accurately reflect the access heat situation of the market research report data block within the preset time period, providing an important basis for subsequent data storage management strategies.
[0062] In a possible implementation, after step S143, the method further includes: Step S210: Divide a plurality of data access regions according to the node topology of the all-flash storage cluster, and configure an independent access frequency monitor for each region.
[0063] The node topology of the all-flash storage cluster may be a complex network structure composed of multiple storage nodes. For example, the nodes storing enterprise financial data are divided into one region, the nodes storing sales data are divided into another region, and the nodes related to the storage of market research report data blocks are divided into a specific region, etc. An independent access frequency monitor is configured for each such region, and these monitors can accurately count the access conditions of each region.
[0064] Step S220: Collect the actual access frequency distribution data of each region within a historical time window through the access frequency monitor.
[0065] Taking the region where the market research report data block is located as an example, set the historical time window as the past month. Within this month, the access frequency monitor details the actual access times of this region every day or every specific time period. For example, in the first few days of the month, since each department of the enterprise has just formulated a work plan, the access times to the market research report data block are relatively few, perhaps only 10 times a day; during the marketing plan adjustment period in the middle of the month, the access times increase to 30 times a day; when approaching the end of the month for monthly summary, the access times change again, perhaps reaching 20 times a day, thus obtaining the actual access frequency distribution data of this region within this historical time window.
[0066] Step S230: Fit and verify the actual access frequency distribution data with the time-dimensional access probability density function, and adjust the curve shape of the probability density function.
[0067] To explain with a simple calculation process, assume that the access probability density function in the time dimension has an access probability of 0.1 set at the beginning of the month, while the actual access count distribution data shows that the actual access count at the beginning of the month is less. Calculate the difference between the actual access count and the access count predicted according to the probability density function. For example, according to the probability density function, it is predicted that there should be 15 accesses per day at the beginning of the month, but there are actually only 10 accesses, and the difference is 5 times. Perform such difference calculations for each time period within the entire historical time window. If it is found that the difference in a certain time period is large, it indicates that the prediction of the probability density function in this time period is inaccurate. For example, during the marketing plan adjustment period in the middle of the month, the access count corresponding to the access probability predicted by the probability density function is 25 times, while there are actually 30 accesses, and the difference is 5 times. Adjust the curve shape of the probability density function according to these differences. If the actual access count in a certain time period is more than the predicted access count, appropriately increase the access probability in this time period to make the curve adjust upward in this time period; otherwise, adjust it downward.
[0068] Step S240, based on the adjusted probability density function, update the calculation rule of the heat level parameter of the dynamic access heat prediction model.
[0069] Before adjusting the probability density function, the calculation rule of the heat level parameter may be calculated based on certain eigenvalue of the original probability density function, such as the maximum probability value, the slope of the probability distribution curve, etc. For example, the original rule is that when the maximum probability value is greater than 0.3 and the slope is greater than a certain specific value, the heat level parameter is level 3. After adjusting the probability density function, these eigenvalue change. Determine the new eigenvalue according to the adjusted probability density function again, such as the maximum probability value becomes 0.35 and the slope also changes. Update the calculation rule of the heat level parameter according to these new eigenvalue, for example, the new rule may become that when the maximum probability value is greater than 0.35 and the slope meets the new condition, the heat level parameter is level 3. This can enable the dynamic access heat prediction model to more accurately reflect the access heat of the market research report data block in different situations, so as to provide a more accurate basis for operations such as data storage, management, and prefetching.
[0070] In a possible implementation manner, step S150 includes: Step S151, determine the sharding redundancy threshold and the minimum number of replicas of the data block to be stored according to the heat level parameter output by the dynamic access heat prediction model.
[0071] Suppose the heat level parameter of the market research report data block is relatively high, indicating that it may have a high access frequency in the future. Based on this, the sharding redundancy threshold is determined to be 3, which means that the data block can be divided into at most 3 shards for storage to improve data availability and reliability. At the same time, the minimum number of replicas is determined to be 2, that is, each shard should have at least 2 replicas. The purpose of doing this is to ensure data accessibility in the face of node failures, data corruption, and other situations.
[0072] Step S152: Traverse the storage space fragmentation rates and input / output request queue depths of all nodes in the all-flash storage cluster, and filter out a subset of candidate nodes that meet the shard capacity constraints.
[0073] For example, there are multiple nodes in the all-flash storage cluster, such as node A, node B, node C, etc. For node A, its storage space fragmentation rate is 20%, and the input / output request queue depth is 50 requests at the current moment. The storage space fragmentation rate of node B is 30%, and the input / output request queue depth is 80 requests. Suppose the shard capacity constraint is that the storage space fragmentation rate does not exceed 30% and the input / output request queue depth does not exceed 100 requests. Then both node A and node B meet this shard capacity constraint and are selected into the subset of candidate nodes. If the storage space fragmentation rate of node C is 40% or the input / output request queue depth is 150 requests, which does not meet the shard capacity constraint, it will not be selected into the subset of candidate nodes.
[0074] Step S153: Based on the physical storage location proximity, calculate the storage location association scores of each node in the subset of candidate nodes.
[0075] Taking the association between the previously mentioned market analysis data block and the market research report data block as an example, suppose the market analysis data block is stored on node A, and a related data block of the market research report data block is stored at a location adjacent to the logical unit address of node A. Then the storage location association score of node A with the market research report data block is relatively high. If another node B is far from the storage location of the related data block of the market research report data block, its storage location association score is relatively low. By comprehensively considering these factors, an accurate storage location association score is calculated for each node in the subset of candidate nodes.
[0076] Step S154: Based on the storage location association scores and the network link bandwidth utilization rates in the multi-dimensional performance data set, construct a storage path weight table between nodes.
[0077] For example, the storage location association score between node A and node B is relatively high, and the network link bandwidth utilization rate between them is 70%, indicating that the data transmission between these two nodes is highly efficient. Based on these factors, a relatively high weight value is assigned to the storage path between node A and node B. If the storage location association score between node A and node C is relatively low and the network link bandwidth utilization rate is 50%, then the weight value of the storage path between them is relatively low. By evaluating the relationships between all nodes within the candidate node subset in this way, a complete storage path weight table between nodes is constructed.
[0078] Step S155, generate a sharded storage topology map including redundant path cross-connections based on the storage path weight table and the minimum number of replicas.
[0079] For example, step S155 includes: Step S1551, according to the storage path weight values between nodes in the storage path weight table, filter the set of candidate storage paths whose storage path weight values exceed the preset weight threshold.
[0080] Step S1552, determine the number of sharded replicas of the data block to be stored based on the minimum number of replicas, and allocate an initial storage path for each sharded replica.
[0081] Step S1553, traverse the storage path weight values in the set of candidate storage paths, and select the main storage path with the highest weight value as the default storage path for the sharded replica.
[0082] Step S1554, detect the connection status between nodes of the default storage path. If there is a single-node connection path, select the backup storage path with the second-highest weight value from the set of candidate storage paths as the cross-redundant path.
[0083] Step S1555, perform a two-way connection between the cross-redundant path and the default storage path to generate a sharded storage topology map including redundant path cross-connections.
[0084] Assume that the preset weight threshold is 0.6. In the storage path weight table, the storage path weight value from node A to node B is 0.7, and the storage path weight value from node B to node C is 0.8. Then the storage paths from node A to node B and from node B to node C are selected into the candidate storage path set. Based on the minimum number of replicas, the number of shard replicas of the data block to be stored is determined to be 2, and an initial storage path is assigned to each shard replica. For example, the storage path from node A to node B is assigned to the first shard replica. Then, by traversing the storage path weights in the candidate storage path set, it is found that the path weight value from node A to node B is the highest, so it is used as the default storage path for the shard replica. Detect the connection status between nodes of this default storage path. If it is found that there is only one connection path from node A to node B, there is a risk of single point of failure. At this time, the backup storage path with the second highest weight value (assuming the path weight value from node A to node C is the second highest) is selected from the candidate storage path set as the cross-redundant path. The cross-redundant path and the default storage path are connected bidirectionally to generate a shard storage topology diagram containing cross-connections of redundant paths. For example, in this shard storage topology diagram, for the first shard replica, the path from node A to node B is the primary storage path, and the path from node A to node C is the cross-redundant path, and the two are connected bidirectionally to ensure that when node A or node B fails, the data can still be accessed through node C.
[0085] Step S1556: Verify whether the path redundancy of each shard replica in the shard storage topology diagram meets the redundancy constraint conditions corresponding to the minimum number of replicas. If not, reselect the backup storage path and update the cross-connection relationship of the shard storage topology diagram.
[0086] Specifically, since the minimum number of replicas is 2, for each shard replica, its path redundancy should be at least 1, that is, in addition to the primary storage path, there is at least one backup storage path. Check the path situation of each shard replica in the shard storage topology diagram. If it is found that the path redundancy of a certain shard replica does not meet the requirements, for example, there is only the primary storage path and no backup storage path, then reselect the backup storage path and update the cross-connection relationship of the shard storage topology diagram. Assume that a certain shard replica originally only has the primary storage path from node A to node B, which does not meet the redundancy constraint conditions. Reselect the path from node A to node C as the backup storage path from the candidate storage path set, and update the shard storage topology diagram to establish the correct cross-connection relationship between node A to node B and node A to node C, so as to ensure the reliability and availability of the entire shard storage topology diagram, meet the requirements of the enterprise data center for storing the data block of the market research report, and improve the storage security and access efficiency of the data.
[0087] In a possible implementation manner, the method further includes: Step S310: Detect whether there is a single point of failure risk path in the sharded storage topology graph.
[0088] For example, for a shard replica of market research report data block, its main storage path is from node A to node B. If this is the only storage path and there is no other backup path connected to it, then this is a single point of failure risk path. Because once node A or node B fails, such as the power supply of node A has a problem or the storage medium of node B is damaged, then this shard replica will not be accessible, thus affecting the availability of the entire data block.
[0089] Step S320: If there is a single point of failure risk path, then based on the real-time performance data of the candidate node subset, dynamically insert backup storage nodes to form a ring redundancy path.
[0090] Suppose the candidate node subset includes nodes A, B, C, etc. Check the real-time performance data of these nodes. For example, the real-time storage space fragmentation rate of node C is relatively low, the input / output request queue depth is also within a reasonable range, and the network link bandwidth utilization rate is relatively high, with a good performance state. At this time, dynamically insert node C into the path with a single point of failure risk. For the previous path from node A to node B, expand it to a ring redundancy path from node A to node C and then to node B. In this way, even if node A or node B fails, the data can still be transmitted through node C, improving the reliability of the data. In this process, it is necessary to carefully consider the connection relationship and data flow between nodes to ensure the rationality and effectiveness of the ring redundancy path.
[0091] Step S330: According to the inter-node delay data of the ring redundancy path, optimize the replica synchronization priority in the sharded storage topology graph.
[0092] For example, by measuring the delay data between each pair of nodes in the ring redundancy path. For example, the delay from node A to node C is 5 milliseconds, the delay from node C to node B is 3 milliseconds, and the delay from node B to node A is 4 milliseconds. Determine the replica synchronization priority based on these delay data. If there is an update to a shard replica on node A, since the delay from node C to node B is relatively small, then the update can be preferentially synchronized to node C, then from node C to node B, and finally from node B back to node A to ensure data consistency and timeliness. This process requires precise calculation and weighing of the impact of delays on different paths on data synchronization to ensure both the improvement of data reliability and the efficient synchronization of data.
[0093] Step S340: Associate and map the optimized sharded storage topology graph with the cache preloading policy matrix to generate a combined storage policy configuration file.
[0094] The cache preloading policy matrix contains cache policy information for data blocks related to the market research report data blocks, such as preloading priority coefficients, cache retention periods, and compression level parameters. The optimized sharded storage topology map is associated and mapped with this cache policy information. For example, for the sharded replica stored on node A in the sharded storage topology map, if the preloading priority coefficient of its corresponding associated data block in the cache preloading policy matrix is high, then in the joint storage policy configuration file, it will be marked that node A where the sharded replica is located needs to perform cache preloading operations with higher priority. Through this association mapping, a comprehensive joint storage policy configuration file is generated, and this joint storage policy configuration file will guide how the entire data storage system stores, caches, and synchronizes data for the market research report data blocks and their related data blocks.
[0095] In one possible implementation, the method further includes: Step S410, creating a storage lifecycle tracking log for the data block to be stored, so as to record the creation timestamp of the sharded replica, migration events, and cache status change history through the storage lifecycle tracking log.
[0096] Specifically, when a sharded replica of a market research report data block is created at a certain moment, the storage lifecycle tracking log will record this creation timestamp, such as 10:15 on July 1, 2023. During the storage process of the data block, if due to performance issues of a certain node or the need for load balancing, a migration event occurs for the sharded replica, such as migrating from node A to node C, then this migration event will be detailedly recorded in the log, including information such as the source node, target node, and migration time of the migration. At the same time, the change history of the cache status will also be recorded. If the cache of a certain sharded replica changes from the initial uncached state to the cached state, or the cache retention period changes, this information will be recorded.
[0097] Step S420, analyzing the abnormal event sequence in the storage lifecycle tracking log to identify potential causes of storage performance bottlenecks.
[0098] For example, it is found in the storage lifecycle tracking log that within a certain period, the sharded replicas of the market research report data blocks frequently have migration events, and at the same time, the cache status is also unstable. By analyzing this abnormal event sequence, it may be found that due to the high fragmentation rate of the storage space of a certain node (such as node A), the storage and access efficiency of the data blocks is reduced, resulting in frequent migration events and unstable cache status. Or it is found that because the network link bandwidth utilization rate suddenly decreases at certain moments, affecting data transmission and caching, these are potential causes of storage performance bottlenecks.
[0099] Step S430: Adjust the parameter update frequency of the dynamic access heat prediction model and the calculation logic of the shard redundancy threshold according to the potential cause.
[0100] If it is found that the problem is caused by too high fragmentation rate of the storage space of the node, then the parameter update frequency of the dynamic access heat prediction model can be adjusted. For example, the original parameter update frequency was once a day. Since the change in the storage space fragmentation rate may have a greater impact on the data access heat, the parameter update frequency is increased to once an hour to more timely reflect the actual access situation of the data. For the calculation logic of the shard redundancy threshold, if it is found that the instability of the network link bandwidth utilization has a greater impact on data availability, then when calculating the shard redundancy threshold, the consideration weight of the network link bandwidth utilization will be increased. For example, previously the shard redundancy threshold was determined to be 3 only based on the heat level parameter. Now due to the influence of network factors, the shard redundancy threshold may be adjusted to 4 according to the fluctuation of the network link bandwidth utilization to improve data reliability.
[0101] Step S440: Synchronize the adjusted parameter update frequency and calculation logic to all management nodes of the all-flash storage cluster in real time.
[0102] Each management node in the all-flash storage cluster needs to obtain these adjusted information to ensure that the entire storage system stores and manages data according to the new policy. For example, management node A, management node B, etc. all need to receive and update these parameters, so that when storing the market research report data block and other data blocks, decisions can be made according to the parameter update frequency of the new dynamic access heat prediction model and the calculation logic of the shard redundancy threshold, ensuring the efficiency, reliability and data availability of the entire storage system.
[0103] In a possible implementation manner, step S150 further includes: Step S156: Identify a target data block sequence having a spatial association with the data block to be stored according to the physical storage location proximity.
[0104] Specifically, in the all-flash storage cluster, the market research report data blocks are stored at specific nodes and logical unit addresses. By analyzing the proximity of physical storage locations, such as examining the relationship between the storage node identifiers and the logical unit addresses, other data blocks that are stored adjacent to the market research report data blocks are determined. Suppose the market research report data block is stored at the logical unit address 100 of node A. Some auxiliary data blocks related to the market research report, such as the original data collection records and preliminary analysis results during the market research process, may be stored at the same node and have similar logical unit addresses (such as 101 - 105). These data blocks constitute a sequence of target data blocks that are spatially associated with the market research report data block.
[0105] Step S157: Based on the dynamic access heat prediction model, predict the concurrent access probability distribution of the sequence of target data blocks within the preset time period.
[0106] The dynamic access heat prediction model outputs information such as heat level parameters for the market research report data block. Based on this, the concurrent access probability distribution of the sequence of target data blocks is predicted. For example, according to the high access heat of the market research report data block from 9 am to 10 am on weekdays, it is speculated that the original data collection record data block that is spatially associated with it also has a relatively high concurrent access probability during this time period, perhaps reaching 0.4. For the preliminary analysis result data block, its concurrent access probability may be slightly lower, at 0.3. For other time periods, similarly, based on the access heat trend of the market research report data block and the degree of association between each target data block and the market research report data block, the concurrent access probability for each time period is calculated, thereby obtaining the concurrent access probability distribution of the sequence of target data blocks within the preset time period.
[0107] Step S158: According to the concurrent access probability distribution and the node storage space fragmentation rate in the multi-dimensional performance data set, calculate the preloading priority coefficient for each target data block.
[0108] Taking the data block of the original data collection record as an example, its concurrent access probability is 0.4. Suppose the storage space fragmentation rate of node A is 20%. The calculation process is as follows: If the storage space fragmentation rate is low, it means that the node has more space for caching operations, which has a positive impact on the preloading priority. A basic score can be set. For example, when the concurrent access probability is 0.4, it corresponds to 40 points, and then it is adjusted according to the storage space fragmentation rate. Since the storage space fragmentation rate is 20%, which is at a relatively low level, a certain bonus is given. Suppose 10 points are added, then the preloading priority coefficient of the data block of the original data collection record is 50 points. For the data block of the preliminary analysis result, the concurrent access probability is 0.3. Suppose the storage space fragmentation rate of the node where it is located is 30%. The fragmentation rate of 30% is relatively high and has a certain negative impact on the preloading priority. According to the same calculation method, its basic score is 30 points, and 5 points may be subtracted due to the high fragmentation rate. The preloading priority coefficient is 25 points. Through such a calculation method, the preloading priority coefficient is calculated for each target data block.
[0109] Step S159, based on the preloading priority coefficient, allocate different cache retention periods and compression level parameters for the target data block sequence, and integrate the cache retention periods and compression level parameters according to the node dimension to generate a multi-layer structure of the cache preloading policy matrix.
[0110] There are multiple nodes in the all-flash storage cluster. For each node, integrate the cache retention periods and compression level parameters of the target data blocks stored on that node. For example, on node A, the cache retention period of the data block of the original data collection record is 3 hours, and the compression level is the low compression level; the cache retention period of the data block of the preliminary analysis result is adjusted dynamically (suppose it is 1 hour currently), and the compression level is the high compression level. Organize this information according to the node dimension to form a multi-layer structure of the cache preloading policy matrix. This cache preloading policy matrix can clearly guide the cache preloading operations of the all-flash storage cluster for the target data block sequence on different nodes, including when to cache, how long to cache, and what compression ratio to adopt, so as to improve the data access efficiency and the utilization rate of storage resources.
[0111] Among them, step S159 includes: Step S1591, obtain the current cache space occupancy rate and historical cache replacement frequency of the edge nodes in the all-flash storage cluster.
[0112] Step S1592, when the concurrent access probability distribution is higher than the first preset threshold, allocate a fixed retention period for the corresponding data block and lock the cache space.
[0113] Step S1593, when the concurrent access probability distribution is lower than a second preset threshold, dynamically adjust the retention period decay rate based on the historical cache replacement frequency.
[0114] Step S1594, perform variable compression rate processing on low-priority data blocks according to the compression level parameter, and record the metadata verification information of the compressed data blocks.
[0115] Suppose the current cache space occupancy rate of the edge node is 60%, and the historical cache replacement frequency is once every 2 hours. For the original data acquisition record data block with a relatively high preloading priority coefficient, when its concurrent access probability distribution is higher than a first preset threshold (assumed to be 0.35), allocate a fixed retention period for the corresponding data block and lock the cache space. For example, allocate a fixed cache retention period of 3 hours. During these 3 hours, even if the cache space is tight, this data block will not be replaced. For the preliminary analysis result data block, its concurrent access probability distribution is lower than a second preset threshold (assumed to be 0.25), and the retention period decay rate is dynamically adjusted based on the historical cache replacement frequency. Since the historical cache replacement frequency is once every 2 hours, when the cache space is tight, its cache retention period will decay at a relatively fast rate. For example, the retention period is reduced by 10 minutes every 30 minutes to release the cache space to the high-priority data blocks that need the cache more quickly. Perform variable compression rate processing on low-priority data blocks according to the compression level parameter. For the low-priority data block such as the preliminary analysis result, according to its compression level parameter (assumed to be a relatively high compression level), use a relatively high compression rate for compression. During the compression process, record the metadata verification information of the compressed data block, such as the checksum, etc., for verifying the data integrity when accessing the data block subsequently.
[0116] In a possible implementation manner, the method further includes: Step S510, deploy a dynamic load balancing controller in the all-flash storage cluster, and collect the performance data fluctuation trend of each node in real time.
[0117] For example, for each node storing data blocks of market research report data and their related data blocks, the dynamic load balancing controller continuously monitors their performance status. It records the changes in various performance data of nodes such as Node A and Node B over time. Taking Node A as an example, it monitors the change in the storage space fragmentation rate of Node A at different time periods, such as from 15% in the morning to 18% at noon and then to 20% in the afternoon; at the same time, it monitors the fluctuation of the input / output request queue depth, such as the queue depth being 30 requests at the beginning of the business in the morning, increasing to 80 requests at the peak of the business in the morning as the market department frequently accesses the data blocks of the market research report and their associated data blocks, and then gradually falling back to 50 requests in the afternoon, etc. By recording these data at different time points, the fluctuation trend of the performance data of each node is analyzed.
[0118] Step S520: Identify potential overloaded nodes according to the performance data fluctuation trend, and trigger the data block replica migration warning mechanism.
[0119] For example, it is observed that the growth rate of the input / output request queue depth of Node A is relatively fast, increasing from 30 requests to 80 requests in the past hour. Calculate its growth rate (the number of increased requests divided by the time interval). Assume that this growth rate exceeds the pre-set queue depth warning line (for example, increasing by 50 requests per hour). At the same time, the change gradient of the storage space fragmentation rate of Node A is also relatively large, rising from 15% to 20% reaching the fragmentation threshold (assumed to be the change range of 15% - 20%). At this time, Node A is identified as a potential overloaded node, thus triggering the data block replica migration warning mechanism.
[0120] Step S530: Based on the redundant path cross-connection relationship in the sharded storage topology graph, select the migration target node and calculate the optimal delay parameter of the migration path.
[0121] In the sharded storage topology graph, node A has redundant path cross - connection relationships with other nodes (such as node B, node C, etc.). View the real - time performance data of these nodes, including storage space fragmentation rate, input / output request queue depth, network link bandwidth utilization, etc. Assume that node B has a lower storage space fragmentation rate, a smaller input / output request queue depth, and a higher network link bandwidth utilization, then node B is selected as the migration target node. Calculate the optimal delay parameter of the migration path from node A to node B. Consider each link segment between node A and node B, such as the intermediate nodes passed through, network devices, etc. Measure the delay data of each segment. For example, the delay from node A to the intermediate node is 3 milliseconds, and the delay from the intermediate node to node B is 2 milliseconds, and the total delay is 5 milliseconds. Also consider the impact of factors such as network congestion and data transfer rate that may be encountered during the migration process on the delay. By comprehensively analyzing this data, determine the optimal delay parameter of the migration path from node A to node B.
[0122] Step S540, during the execution of the migration operation, maintain the access availability of the original data block and synchronously update the joint storage policy configuration file.
[0123] When starting to migrate the sharded replicas of the market research report data block on node A to node B, it is necessary to ensure that during the migration process, the market department or other departments that need to access this data block can still access the data normally. This may require some technical means, such as temporary copies of data, multi - path access to data, etc. At the same time, during the migration process, synchronously update the joint storage policy configuration file. The joint storage policy configuration file contains information such as the sharded storage topology graph of the market research report data block, the cache pre - loading policy matrix, etc. Since the storage location of the data block has changed, relevant information needs to be updated in the joint storage policy configuration file. For example, modify the original storage information about this data block on node A to the storage information on node B, including updating the storage path of the sharded replicas on node B, cache policies, and other relevant information, to ensure that the policies of the entire storage system match the actual storage situation of the data.
[0124] Among them, step S520 includes: Step S521, monitor the growth rate of the input / output request queue depth and the change gradient of the storage space fragmentation rate of the potential overload node.
[0125] Step S522, when the growth rate exceeds the queue depth warning line and the change gradient reaches the fragmentation threshold, generate a migration task queue.
[0126] Step S523, according to the data block priority tags in the migration task queue, sort the migration execution order and allocate migration bandwidth resources.
[0127] For example, for node A, as described above, the significant increase in the depth of its input / output request queue and the obvious rise in the storage space fragmentation rate within a short period are recorded in detail. When the growth rate exceeds the queue depth warning line and the change gradient reaches the fragmentation threshold, a migration task queue is generated. Assume that there are multiple shard replicas of the market research report data block on node A, and these shard replicas are marked with different data block priority tags according to factors such as their importance. When generating the migration task queue, they are sorted according to these priority tags. For example, the shard replicas related to the core data of the market research report are marked as high priority, those related to auxiliary data are marked as medium priority, and other data block replicas with lower relevance are marked as low priority. According to this priority order, the high-priority data block replicas are placed at the front of the migration task queue first, followed by the medium-priority ones, and the low-priority ones last. At the same time, according to the current network bandwidth resource situation of the all-flash storage cluster, migration bandwidth resources are allocated to the data block replicas in the migration task queue. If the total bandwidth is 1000 Mbps, depending on factors such as the priority and size of the data block, 500 Mbps of bandwidth may be allocated to the high-priority data block replicas, 300 Mbps to the medium-priority ones, and 200 Mbps to the low-priority ones.
[0128] Step S524, after the migration is completed, verify the data consistency of the target node and update the node status identifier of the shard storage topology map.
[0129] After the data block replica is completely migrated from node A to node B, the data on node B needs to be verified for consistency. This may include comparing the checksum of the data block to see if the size, content, etc. of the data block are the same as when it was on node A. If the data is consistent, it indicates that the migration was successful. Then update the node status identifier of the shard storage topology map, modify the status identifier of this data block replica on node A to migrated, and modify the status identifier on node B to received and stored this data block replica. Through such operations, the shard storage topology map can accurately reflect the storage location and status of the data block replicas, providing an accurate basis for subsequent storage management operations.
[0130] In a possible implementation manner, the method further includes: Step S610, implement a data integrity verification loop in the all-flash storage cluster, regularly scan the checksum information of the shard replicas, and when a checksum anomaly is detected, initiate a replica repair request based on the redundant path cross-connection relationship in the shard storage topology map.
[0131] In this embodiment, taking the market research report data block as an example, the all-flash storage cluster will scan the checksum information of the shard replicas of the market research report data blocks stored on each node at a predetermined time interval, such as every hour or every day. The checksum is a value representing the content characteristics of the data block calculated through a specific algorithm. For each shard replica of the market research report data block, its checksum can be recalculated and then compared with the previously stored checksum information.
[0132] Suppose that during a scan, it is found that the checksum of a shard replica of the market research report data block stored on node A does not match the previously stored value, which indicates that the data of this shard replica may be damaged or incorrect. At this time, the redundant path cross-connection relationship in the shard storage topology map can be referred to. For example, the shard storage topology map shows that this shard replica has a primary storage path on node A and is connected to node B through a redundant path, and a redundant replica of this shard replica is stored on node B. Based on this, a replica repair request can be sent to node B, requesting node B to provide the correct shard replica data to repair the damaged replica on node A.
[0133] Step S620, according to the compression level parameter in the cache preloading policy matrix, perform re-compression and cache status refresh operations on the repaired data block, update the metadata mapping relationship in the joint storage policy configuration file, and feedback the repair result to the user terminal.
[0134] In the cache preloading policy matrix, each data block has a corresponding compression level parameter. Suppose the compression level parameter of this shard replica of the market research report data block is the medium compression level. After the repair is completed, the repaired data block can be recompressed according to this medium compression level. For the compression process, according to the compression algorithm, the duplicate data in the data block can be processed to reduce the storage space occupied by the data block. At the same time, since the content of the data block has changed (after repair and recompression), the cache status needs to be refreshed. If the data block was previously in the cached state in the cache and had a certain cache retention period and cache policy, then these cache-related states need to be updated according to the new situation. For example, it may be necessary to recalculate the cache retention period or adjust the cache priority, etc.
[0135] The combined storage policy configuration file contains various information about the storage of data blocks in market research report. Among them, the metadata mapping relationship records the associations between data blocks and various information such as storage locations, caching policies, compression levels, etc. Since the data blocks have undergone repair, recompression, and cache state refresh, these related information have all changed. For example, the location of a data block on a storage node may change due to repair operations, or a change in the compression level may cause a change in its relationship in the cache preloading policy. Therefore, it is necessary to update the metadata mapping relationship in the combined storage policy configuration file to ensure that the information in the file matches the actual storage and management situation of the data blocks. Finally, the repair result can be fed back to the user terminal. For users of data blocks in market research reports, such as employees in the market department of an enterprise, they can receive notifications of the repair result through the user terminal. If the repair is successful, the notification may show "The shard copy of the data block in the market research report on node A has been successfully repaired, data integrity has been restored, and it can be used normally"; if the repair fails, the notification may show "The repair of the shard copy of the data block in the market research report on node A has failed. Please contact the administrator for further inspection." In this way, relevant personnel in the enterprise can timely understand the status of the data blocks so as to take further measures when needed.
[0136] Figure 2 FIG. shows a schematic diagram of exemplary hardware and software components of a cloud computing-based all-flash storage optimization system 100 that can implement the ideas of the present application provided by some embodiments of the present application. For example, the processor 120 can be used on the cloud computing-based all-flash storage optimization system 100 and is used to execute the functions in the present application.
[0137] The cloud computing-based all-flash storage optimization system 100 can be a general-purpose server or a special-purpose server, both of which can be used to implement the cloud computing-based all-flash storage optimization method of the present application. Although only one server is shown in the present application, for convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0138] For example, the cloud computing-based all-flash storage optimization system 100 may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as disks, ROM, or RAM, or any combination thereof. Exemplarily, the cloud computing-based all-flash storage optimization system 100 may further include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The methods of the present application can be implemented according to these program instructions. The cloud computing-based all-flash storage optimization system 100 further includes an input / output (I / O) interface 150 between the computer and other input / output devices.
[0139] For ease of explanation, only one processor is described in the cloud computing-based all-flash storage optimization system 100. However, it should be noted that the cloud computing-based all-flash storage optimization system 100 in the present application may further include multiple processors. Therefore, the steps executed by one processor described in the present application may also be jointly executed or separately executed by multiple processors. For example, if the processor of the cloud computing-based all-flash storage optimization system 100 executes steps A and B, it should be understood that steps A and B may also be jointly executed by two different processors or separately executed in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor jointly execute steps A and B.
[0140] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above-mentioned cloud computing-based all-flash storage optimization method is implemented.
[0141] It should be noted that, in order to simplify the presentation of the disclosure of the present invention and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes multiple features are incorporated into one embodiment, drawing, or description thereof.
Claims
1. A data all-flash storage optimization method based on cloud computing, characterized in that: The method comprises: Collecting multi-dimensional performance data sets of each node in the all-flash storage cluster in real time, wherein the multi-dimensional performance data sets include node storage space fragmentation rate, input and output request queue depth, and network link bandwidth utilization; Receiving a storage operation request stream sent by a user terminal, parsing a set of identifiers of data blocks to be stored and corresponding operation mode tags in the storage operation request stream; Traversing the historical access record library based on the set of identifiers of the data blocks to be stored, extracting the historical access time sequence characteristics of the data blocks to be stored and the physical storage location proximity of the associated data blocks; Constructing a dynamic access heat prediction model of the data block to be stored within a preset time period according to the multi-dimensional performance data set and the historical access time series characteristics; Based on the dynamic access heat prediction model and the physical storage location proximity, a cross-node shard storage topology map and a cache preloading strategy matrix are generated.
2. The method for optimizing data all-flash storage based on cloud computing according to claim 1, characterized in that: The step of constructing a dynamic access heat prediction model for the data block to be stored within a preset time period according to the multi-dimensional performance data set and the historical access time series characteristics includes: Extracting periodic access peak intervals, burst access event timestamps and access interval distribution statistics from the historical access timing characteristics; Identify the read-write operation ratio corresponding to the operation mode tag, and activate the hot data prediction mark when the proportion of read operations in the read-write operation ratio exceeds a preset threshold; Generate a time dimension access probability density function by combining the periodic access peak interval and the burst access event timestamp; Based on the access interval distribution statistics and the hotspot data prediction mark, modifying the weight parameter of the time dimension access probability density function; The modified time dimension access probability density function is coupled with the network link bandwidth utilization in the multi-dimensional performance data set to generate the dynamic access heat prediction model, and the dynamic access heat prediction model outputs the heat level parameter of the data block to be stored.
3. The method for optimizing data all-flash storage based on cloud computing according to claim 2, characterized in that: After generating the time dimension access probability density function by combining the periodic access peak interval and the burst access event timestamp, the method further includes: Divide a plurality of data access areas according to the node topology of the all-flash storage cluster, and configure an independent access frequency monitor for each area; The access frequency monitor is used to collect the actual access frequency distribution data of each area within the historical time window; Fitting and verifying the actual access frequency distribution data with the time dimension access probability density function, and adjusting the curve shape of the probability density function; Based on the adjusted probability density function, the heat level parameter calculation rules of the dynamic access heat prediction model are updated.
4. The method for optimizing data all-flash storage based on cloud computing according to claim 1, characterized in that: Based on the dynamic access heat prediction model and the physical storage location proximity, a cross-node shard storage topology diagram is generated, including: Determine the shard redundancy threshold and the minimum number of copies of the data block to be stored according to the heat level parameter output by the dynamic access heat prediction model; Traversing the storage space fragmentation rates and input / output request queue depths of all nodes in the all-flash storage cluster, and screening a subset of candidate nodes that meet the shard capacity constraints; Calculating a storage location association score for each node in the candidate node subset based on the physical storage location proximity; Constructing a storage path weight table between nodes according to the storage location association score and the network link bandwidth utilization in the multi-dimensional performance data set; Based on the storage path weight table and the minimum number of replicas, a shard storage topology diagram including redundant path cross connections is generated.
5. The method for optimizing data all-flash storage based on cloud computing according to claim 4, characterized in that: The method further comprises: Detect whether there is a single point failure risk path in the shard storage topology diagram; If there is a single point failure risk path, dynamically insert a spare storage node based on the real-time performance data of the candidate node subset to form a ring redundant path; Optimizing the replica synchronization priority in the shard storage topology diagram according to the inter-node delay data of the ring redundant path; The optimized shard storage topology map is associated and mapped with the cache preloading strategy matrix to generate a joint storage strategy configuration file.
6. The method for optimizing data all-flash storage based on cloud computing according to claim 4, characterized in that: The method further comprises: Creating a storage lifecycle tracking log for the data block to be stored, so as to record the creation timestamp, migration events and cache status change history of the shard copy through the storage lifecycle tracking log; Analyzing abnormal event sequences in the storage lifecycle tracking log to identify potential causes of storage performance bottlenecks; Adjusting the parameter update frequency of the dynamic access heat prediction model and the calculation logic of the shard redundancy threshold according to the potential cause; The adjusted parameter update frequency and calculation logic are synchronized in real time to all management nodes of the all-flash storage cluster.
7. The data all-flash storage optimization method based on cloud computing according to claim 1 is characterized in that: Based on the dynamic access heat prediction model and the physical storage location proximity, a cache preloading strategy matrix is generated, including: Identifying a target data block sequence spatially associated with the data block to be stored according to the physical storage location proximity; Based on the dynamic access heat prediction model, predict the concurrent access probability distribution of the target data block sequence within the preset time period; Calculating a preloading priority coefficient of each target data block according to the concurrent access probability distribution and the node storage space fragmentation rate in the multidimensional performance data set; Based on the preloading priority coefficient, assigning differentiated cache retention period and compression level parameters to the target data block sequence; Integrate the cache retention period and compression level parameters according to the node dimension to generate a multi-layer structure of the cache preloading strategy matrix; The step of allocating differentiated cache retention periods and compression level parameters to the target data block sequence based on the preloading priority coefficient includes: Obtaining the current cache space occupancy rate and historical cache replacement frequency of the edge nodes in the all-flash storage cluster; When the concurrent access probability distribution is higher than a first preset threshold, a fixed retention period is allocated to the corresponding data block and a cache space is locked; When the concurrent access probability distribution is lower than a second preset threshold, dynamically adjusting the retention period decay rate based on the historical cache replacement frequency; Variable compression rate processing is performed on the low priority data block according to the compression level parameter, and metadata verification information of the compressed data block is recorded.
8. The method for optimizing data all-flash storage based on cloud computing according to claim 5, characterized in that: The method further comprises: Deploy a dynamic load balancing controller in the all-flash storage cluster to collect performance data fluctuation trends of each node in real time; Identify potential overloaded nodes according to the performance data fluctuation trend, and trigger a data block replica migration early warning mechanism; Based on the redundant path cross-connection relationship in the shard storage topology diagram, selecting a migration target node and calculating an optimal delay parameter of the migration path; During the migration operation, the access availability of the original data block is maintained and the joint storage policy configuration file is updated synchronously; The triggering of the data block replica migration early warning mechanism includes: Monitor the growth rate of the input and output request queue depth of the potential overloaded node and the change gradient of the storage space fragmentation rate; When the growth rate exceeds the queue depth warning line and the change gradient reaches the fragmentation threshold, a migration task queue is generated; sorting the migration execution order and allocating migration bandwidth resources according to the data block priority tags in the migration task queue; After the migration is completed, the data consistency of the target node is verified and the node status identifier of the shard storage topology map is updated.
9. The method for optimizing data all-flash storage based on cloud computing according to claim 5, characterized in that: The method further comprises: Implementing a data integrity verification cycle in the all-flash storage cluster, regularly scanning the checksum information of the shard replicas, and when a checksum anomaly is detected, initiating a replica repair request based on the redundant path cross-connection relationship in the shard storage topology diagram; According to the compression level parameters in the cache preloading strategy matrix, recompression and cache status refresh operations are performed on the repaired data blocks, and the metadata mapping relationship in the joint storage strategy configuration file is updated and the repair result is fed back to the user terminal.
10. A data all-flash storage optimization system based on cloud computing, characterized in that: The cloud computing-based data all-flash storage optimization system includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the cloud computing-based data all-flash storage optimization method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Data processing method and apparatus
CN107908653A
Data storage method, device and equipment of full-flash storage system and storage medium
CN111124281A
Multi-tile memory management for detecting cross tile access, providing multi-tile inference scaling, and providing optimal page migration
CN113424148A
Predictive storage optimization method and system for distributed storage system
CN117762345A
Storage data management system and method based on flash memory
CN117806555A
Cited By
Redundancy strategy adjustment method and device, computer equipment, storage medium and product
CN120371226A
Storage position adjusting method and device, equipment, storage medium and program product
CN120631276A
Flash memory management method and device based on SD NAND and storage medium
CN120762601A
A flash memory management method and device based on SD NAND and a storage medium
CN120762601B
Efficient memory management method and system based on SOC chip
CN120832332A