Distributed block storage system performance optimization method and system

By dynamically partitioning and tiering storage, and combining node load status and data access frequency, the distributed block storage system is optimized, solving the problem of mixed storage of high-frequency and low-frequency data, improving system performance and stability, and reducing storage costs.

CN120803377AActive Publication Date: 2025-10-17NEWLIXON TECH CO LTD +1

Patent Information

Application Number
CN202511317978.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-10-17
Estimated Expiration
2045-09-16

AI Technical Summary

Technical Problem

In existing technologies, data storage with fixed block sizes results in the mixed storage of high-frequency and low-frequency access data, affecting read efficiency; the master-slave node partitioning method is static and cannot respond to changes in node load and failures in real time, leading to master node overload or waste of slave node resources; and the data replication process cannot adapt to the real-time status of nodes, increasing storage costs and affecting system performance.

Method used

By analyzing data access frequency, dynamic block partitioning is performed to generate multiple data blocks. Master and slave nodes are dynamically distinguished based on node load status. Data is stored by combining access frequency and data type, and the master-slave node relationship is updated in real time. Incremental replication and data association fusion reading are used to optimize storage system performance.

Benefits of technology

This system concentrates high-frequency data on high-performance nodes and distributes low-frequency data to low-load nodes, avoiding resource waste, improving overall storage efficiency, relieving pressure on overloaded master nodes, quickly switching slave nodes to master nodes, reducing storage costs and network overhead, and improving system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803377A_ABST
    Figure CN120803377A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses a distributed block storage system performance optimization method and system.The distributed block storage system performance optimization method comprises the steps that in response to the data storage requirement of a user, the access frequency of data to be stored is analyzed to conduct dynamic blocking, and multiple pieces of block data are generated; dynamically distinguishing data storage nodes into master nodes and slave nodes according to the load state of each node in the storage system; according to the access frequency, writing each piece of block data into a corresponding master node, and copying a data increment corresponding to each master node into a corresponding slave node to obtain an updated master node and an updated slave node; in response to a data reading demand of a user, analyzing access relevance between data, and fusing the updated slave nodes to read corresponding data; according to the method and the device, the data is dynamically partitioned and hierarchically stored, the pressure of the overload main node can be shared, the load condition between the data storage nodes is balanced, and the system performance is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, more particularly to a distributed block storage system performance optimization method and system. BACKGROUND

[0002] At present, the fixed block size block method is adopted in the data storage process, which leads to the mixed storage of high-frequency access data and low-frequency access data, affecting the reading efficiency; the static configuration is adopted for the master-slave node division method, which cannot respond to the node load change and failure in real time, which may cause the overload of the master node or the waste of the slave node resources; the full replication or fixed ratio replication is adopted in the data replication process, which cannot adapt to the real-time state of the node for adjustment, increasing the data storage cost and affecting the system performance.

[0003] To solve at least one of the above problems, the present application provides a distributed block storage system performance optimization method and system. SUMMARY

[0004] In view of the deficiencies in the prior art, the purpose of the present application is to provide a distributed block storage system performance optimization method and system, which can effectively solve the problems in the background art. The specific technical solutions of the present application are as follows:

[0005] The distributed block storage system performance optimization method comprises:

[0006] In response to the user's data storage requirement, the access frequency of the data to be stored is analyzed to perform dynamic block, and a plurality of block data is generated;

[0007] According to the load state of each node in the storage system, the data storage nodes are dynamically divided into master nodes and slave nodes;

[0008] According to the access frequency, each block data is written into the corresponding master node, and the data increment corresponding to each master node is replicated into the corresponding slave node to obtain updated master nodes and updated slave nodes;

[0009] In response to the user's data reading requirement, the access correlation between the data is analyzed, the corresponding data of the updated slave node is read by fusion to optimize the performance of the distributed block storage system.

[0010] Specifically, in response to the user's data storage requirement, the access frequency of the data to be stored is analyzed to perform dynamic block, and a plurality of block data is generated, comprising:

[0011] In response to the user's data storage requirement, the data type and data content of the data to be stored are analyzed;

[0012] According to the data type, the access frequency of the data is predicted by a preset access frequency prediction model to obtain a predicted access frequency;

[0013] According to the predicted access frequency and the preset chunking strategy, the data content is chunked to obtain a plurality of block data.

[0014] Specifically, the data storage nodes are dynamically divided into master nodes and slave nodes according to the load states of the nodes in the storage system, including:

[0015] According to the load states of the nodes in the storage system, the load scores of each node are calculated through a preset node load evaluation model;

[0016] The nodes with the load scores less than a preset first load threshold are taken as initial master nodes, and the nodes with the load scores greater than or equal to the preset first load threshold are taken as initial slave nodes;

[0017] At least one initial slave node is allocated to each initial master node by analyzing the positional relationship between the nodes and the consistency of the data storage types, and a master-slave node relationship graph is constructed;

[0018] According to the real-time states of the data storage nodes, the master-slave node relationship graph is updated through a preset node updating mechanism to obtain real-time master nodes and slave nodes.

[0019] Specifically, the master-slave node relationship graph is updated according to the real-time states of the data storage nodes through a preset node updating mechanism to obtain real-time master nodes and slave nodes, including:

[0020] According to the real-time states of the data storage nodes, the load scores of each node are updated in real time to obtain updated load scores;

[0021] When the updated load score of the initial master node is greater than a preset second load threshold, the node with the smallest updated load score is selected from the corresponding initial slave nodes as an auxiliary master node;

[0022] When the initial master node fails or malfunctions, the node with the smallest updated load score is selected from the corresponding initial slave nodes as a replacement master node;

[0023] According to the auxiliary master node and the replacement master node, the master-slave node relationship graph is updated to obtain real-time master nodes and slave nodes.

[0024] Specifically, according to the access frequency, each block data is written into the corresponding master node, and the data increment corresponding to each master node is copied into the corresponding slave node to obtain updated master nodes and updated slave nodes, including:

[0025] According to the access frequency, the block data is divided into hot data blocks, warm data blocks and cold data blocks;

[0026] Based on the real-time state of each master node, a preset node state evaluation model is used to calculate a comprehensive state score of each master node;

[0027] According to the comprehensive state score and the type of the block data, each block data is matched to a corresponding master node and written with corresponding data to obtain an updated master node;

[0028] According to the data writing condition of the updated master node, data increment is copied to a corresponding slave node to obtain an updated slave node.

[0029] Specifically, according to the comprehensive state score and the type of the block data, each block data is matched to a corresponding master node and written with corresponding data to obtain an updated master node, including:

[0030] According to the comprehensive state score, the master nodes are layered to obtain a plurality of layers of master nodes, wherein the plurality of layers of master nodes include high-layer master nodes, middle-layer master nodes and low-layer master nodes;

[0031] For hot data blocks, a corresponding first master node is matched in the high-layer master nodes according to the comprehensive state score from high to low;

[0032] For warm data blocks, a corresponding second master node is matched in the middle-layer master nodes according to the comprehensive state score from high to low;

[0033] For cold data blocks, a corresponding third master node is matched in the low-layer master nodes according to the comprehensive state score from high to low;

[0034] Until each block data is matched to a corresponding master node, the block data is written in the corresponding master node to obtain an updated master node.

[0035] Specifically, according to the data writing condition of the updated master node, data increment is copied to a corresponding slave node to obtain an updated slave node, including:

[0036] According to the data writing condition of the updated master node and the state of the slave node, a corresponding slave node is matched through a preset master-slave node mapping relationship;

[0037] Data increment in the updated master node is copied to the corresponding slave node to obtain an updated slave node.

[0038] Specifically, in response to the user's data reading demand, the access correlation between data is analyzed, and the corresponding data of the updated slave node is read through fusion, including:

[0039] In response to the user's data reading demand, target data blocks to be read are obtained;

[0040] By analyzing the access correlation between data, corresponding update slave nodes associated with the target data block are fused to read corresponding data.

[0041] Specifically, the corresponding update slave nodes associated with the target data block are fused to read corresponding data by analyzing the access correlation between data, including:

[0042] According to the access correlation between data, the data correlation access probability of each update slave node is calculated through a preset data block correlation analysis model;

[0043] The update slave nodes with a data correlation access probability greater than a preset correlation access probability threshold are fused to obtain a fused slave node block;

[0044] According to the update slave nodes corresponding to the target data block, a preset number of update slave nodes are selected from the fused slave node block as a set of candidate access nodes in order of the data correlation access probability from high to low;

[0045] The corresponding target data block is read from the set of candidate access nodes.

[0046] A distributed block storage system performance optimization system is used to implement the distributed block storage system performance optimization method, including:

[0047] A data division module analyzes the access frequency of to-be-stored data to perform dynamic blocking in response to a user's data storage requirement, and generates a plurality of block data;

[0048] A data storage node division module dynamically divides data storage nodes into master nodes and slave nodes according to the load state of each node in the storage system;

[0049] A data writing module writes each block data into a corresponding master node according to the access frequency, and copies the data increment corresponding to each master node to a corresponding slave node to obtain update master nodes and update slave nodes;

[0050] A data reading module analyzes the access correlation between data in response to a user's data reading requirement, fuses the update slave nodes to read corresponding data, and optimizes the performance of the distributed block storage system.

[0051] The beneficial effects of the present application: dynamically block and hierarchical storage of data, and update the master-slave node relationship according to the real-time state of the system storage node, so that high-frequency data is concentrated in high-performance nodes, and low-frequency data is distributed to low-load nodes, avoiding resource waste and improving overall storage efficiency; can share the pressure of the overloaded master node, and quickly switch the slave node to the master node in case of failure, improve system performance, and reduce storage cost and network overhead while improving system performance through incremental replication, hot and cold data separation, and dynamic resource allocation. BRIEF DESCRIPTION OF DRAWINGS

[0052] Figure 1 The workflow diagram of the distributed block storage system performance optimization method in the embodiment of the present application;

[0053] Figure 2 The schematic diagram of the auxiliary master node updating the master-slave node relationship in the embodiment of the present application;

[0054] Figure 3 The schematic diagram of the replacement master node updating the master-slave node relationship in the embodiment of the present application;

[0055] Figure 4 The schematic diagram of the updated slave node fusing read data in the embodiment of the present application. DETAILED DESCRIPTION

[0056] The present application will be further described in detail below in conjunction with the drawings and embodiments. In the embodiments of the present application, the words such as "exemplary" or "for example" are used to mean example, illustration, or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words such as "exemplary" or "for example" are intended to present the relevant concept in a specific manner.

[0057] Hereinafter, the terms "first", "second", etc. are generally referred to as words, only for description purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features limited by "first", "second" can explicitly or implicitly include one or more features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "multiple" is two or more.

[0058] Reference Figure 1 The specific implementation of the distributed block storage system performance optimization method of the present application is shown, including:

[0059] S101, in response to the user's data storage demand, analyze the access frequency of the data to be stored to perform dynamic blocking, and generate a plurality of block data;

[0060] S102, according to the load state of each node in the storage system, the data storage node is dynamically distinguished into master node and slave node;

[0061] S103, according to the access frequency, each block data is written into the corresponding master node, and the data increment corresponding to each master node is copied into the corresponding slave node, to obtain the updated master node and the updated slave node;

[0062] S104, in response to the user's data reading demand, the access correlation between data is analyzed, the corresponding data of the updated slave node is read by fusion, so as to optimize the performance of the distributed block storage system.

[0063] The embodiment dynamically completes data blocking and node division in the system by analyzing the data access frequency, and stores the blocked data to the corresponding node. Through data writing and fusion reading operation, the read-write performance of the distributed block storage system is effectively improved. The method can reduce the data access delay, store high-frequency and low-frequency data in high-performance and low-performance nodes respectively, fully utilize the performance difference of various nodes, improve the utilization rate of storage and computing resources, and optimize the overall performance of the distributed block storage system.

[0064] In the embodiment, after receiving the user's data storage demand, the data is dynamically blocked according to the data type and access frequency of the data to be stored. Data with similar access frequencies are grouped, and data with different access frequencies are allocated different sizes of data blocks. The dynamic blocking method fully considers the influence of the difference in data access frequency on the performance of the storage system, and avoids the problem of interference in high-frequency data access caused by the mixed storage of high-frequency and low-frequency data in the traditional method. By distinguishing between high-frequency and low-frequency data, the interference on high-frequency data access is reduced, thereby improving the speed of locating and reading high-frequency data and improving the data access efficiency.

[0065] Specifically, in the data read-write process, the state of the system data storage node will change accordingly. According to the real-time load change of the node, the master node and the slave node are dynamically divided. The dynamic node division mechanism can adapt to the node load change with time and data access mode, overcome the node load imbalance and resource waste problem caused by traditional static master-slave division, so that each node can more evenly bear the storage and access tasks, improve the system stability and reliability. In addition, when the master node fails, the system can quickly switch to the slave node, enhancing the system's fault response capability and overall reliability.

[0066] Further, the data is written to the matched master node according to the data access frequency, and the master node data increment is synchronized to the corresponding slave node. By matching the appropriate master node for data with different access frequencies, the problem of writing high-frequency data to nodes with insufficient performance is avoided, thereby improving data write efficiency. The incremental replication mechanism reduces the amount of data required for transmission, improving backup and node update speed.

[0067] Specifically, in the data reading phase, after receiving the user's data reading request, the updated slave nodes with strong association are fused according to the access association between the target data block and the data, forming a fused slave node block, and the data is obtained directly from the corresponding fused slave node block each time. Through association analysis, fusion reading is realized, effectively reducing the reading frequency and time, and improving the data retrieval efficiency and overall system performance.

[0068] The present application dynamically blocks and stores data, and updates the master-slave node relationship according to the real-time state of the system storage node, so that high-frequency data is concentrated in high-performance nodes, and low-frequency data is allocated to low-load nodes, avoiding resource waste and improving overall storage efficiency; it can share the pressure of overloaded master nodes, and quickly switch slave nodes to master nodes in case of failure, improving system performance, and through incremental replication, hot and cold data separation, and dynamic resource allocation, the system performance is improved while the storage cost and network overhead are reduced.

[0069] Further, in response to the user's data storage demand, the access frequency of the data to be stored is analyzed for dynamic blocking, generating a plurality of block data, including:

[0070] S201, in response to the user's data storage demand, the data type and data content of the data to be stored are analyzed;

[0071] S202, according to the data type, the access frequency of the data is predicted by a preset access frequency prediction model, to obtain a predicted access frequency;

[0072] S203, combining the predicted access frequency and the preset blocking strategy, the data content is blocked to obtain a plurality of block data.

[0073] In the embodiment, when receiving the data storage requirement of the user, the data to be stored is first acquired, and the data type and content thereof are analyzed. Since different types of data usually have different access modes and frequencies, for example, the access frequency of system configuration files is low, and the access frequency of real-time monitoring data is high, based on the data type, by means of a preset access frequency prediction model, the model includes but is not limited to a neural network trained by a large amount of historical data, the future access frequency of the data can be predicted according to the learned mode, the predicted access frequency is obtained, the system can identify the high-frequency access and low-frequency access data in advance, the high-frequency data is stored in a high-performance storage device, and the low-frequency data is stored in a low-cost device, so as to realize the optimization of storage resource allocation.

[0074] Specifically, in combination with the predicted access frequency and the preset block strategy, the block operation is performed on the data content, and the block strategy is made based on the access frequency threshold, data size, business requirement and various factors. In the embodiment, the data is divided into hot data, warm data and cold data according to the predicted access frequency, and different block sizes are used accordingly: a small block is set for the hot data, such as 16KB-64KB, to support fast response of frequent access; a medium block is used for the warm data, such as 128KB-512KB, to balance the access efficiency and storage cost; and a large block is used for the cold data, such as 1MB-4MB, to reduce the metadata management overhead. Through the analysis of the data type and content, in combination with the access frequency prediction and reasonable block, the data layout can be optimized according to the actual access characteristics, the read-write performance and storage efficiency are improved, and at the same time, different types and access modes of data can be adapted, and the flexibility and scalability of the system are enhanced.

[0075] Further, according to the load state of each node in the storage system, the data storage nodes are dynamically divided into master nodes and slave nodes, including:

[0076] S301, according to the load state of the nodes in the storage system, the load score of each node is calculated by a preset node load evaluation model;

[0077] S302, the node with a load score less than a preset first load threshold is taken as an initial master node, and the node with a load score greater than or equal to the preset first load threshold is taken as an initial slave node;

[0078] S303, at least one initial slave node is allocated to each initial master node by analyzing the positional relationship between the nodes and the consistency of the data storage types, and a master-slave node relationship graph is constructed;

[0079] S304, according to the real-time state of the data storage nodes, the master-slave node relationship graph is updated by a preset node updating mechanism, and the real-time master nodes and slave nodes are obtained.

[0080] In this embodiment, first, the load state of each data storage node in the system is analyzed and evaluated by a preset node load evaluation model, and the load score of each node is calculated. A plurality of node load indicators are collected by a system monitoring tool or a special software, including CPU usage, memory usage, disk I / O read / write rate, and network bandwidth occupancy rate, etc. The node load evaluation model adopts a weighted evaluation model, and the corresponding weight is allocated according to the influence degree of each indicator on the node load, and the comprehensive load score of the node is calculated by weighting, so that the node load state is directly reflected in a quantitative form, which is convenient for comparison and analysis between nodes.

[0081] For example, the weight of CPU usage is set to 0.3, the weight of memory usage is set to 0.2, the weight of disk I / O read / write rate is set to 0.4, and the weight of network bandwidth occupancy rate is set to 0.1. If the CPU usage of a node is 60%, the memory usage is 50%, the disk I / O read / write rate is 80 MB / s, and the network bandwidth occupancy rate is 30%, the node load score is: 0.3×60%+0.2×50%+0.4×80+0.1×30%, and the final score is a dimensionless value, which is convenient for unified comparison.

[0082] Further, based on the obtained node load score and the preset first load threshold, the nodes are divided into initial master nodes and slave nodes. The first load threshold can be set according to the system performance and computing demand. If the node load score is lower than the first load threshold, it is marked as an initial master node; if it is greater than or equal to the first load threshold, it is marked as an initial slave node. Considering the node resource availability, the node with lower load is used as the master node to undertake the main data storage responsibility, and the node with higher load is used as the slave node for data backup and auxiliary reading, which can avoid data concentration in a few nodes, realize reasonable allocation of storage tasks among heterogeneous load nodes, and improve the overall performance and stability of the system.

[0083] After completing the initial division of master and slave nodes, at least one slave node is allocated to each master node according to the positional relationship between the nodes and the consistency of the stored data types. When allocating, the slave node with short physical distance from the master node, matching storage type and low load is preferred to improve data transmission efficiency and ensure data consistency. Considering the network delay and data affinity, it helps to improve the system response speed and cooperation efficiency.

[0084] In addition, during the system operation, the master-slave node roles and corresponding relationships are dynamically adjusted according to the real-time state of the nodes and through a preset node update mechanism, and the master-slave node relationship diagram is updated in real time, which can ensure that it always reflects the actual state of the system. Through real-time role adjustment and relationship optimization, the system can adapt to node load fluctuations, maintain load balance, improve performance and stability, and quickly recover data storage and access capabilities in the event of node failure, further enhancing the reliability of the system.

[0085] Further, according to the real-time state of the data storage node, the master-slave node relationship graph is updated through a preset node updating mechanism to obtain real-time master nodes and slave nodes, including:

[0086] S401, according to the real-time state of the data storage node, the load score of each node is updated in real time to obtain an updated load score;

[0087] S402, when the updated load score of the initial master node is greater than a preset second load threshold, the node with the minimum updated load score is selected from the corresponding initial slave nodes as an auxiliary master node;

[0088] S403, when the initial master node fails or malfunctions, the node with the minimum updated load score is selected from the corresponding initial slave nodes as a replacement master node;

[0089] S404, according to the auxiliary master node and the replacement master node, the master-slave node relationship graph is updated to obtain real-time master nodes and slave nodes.

[0090] In this embodiment, the system monitoring tool is used to continuously collect the key performance indicators of each data storage node, and the load score is updated in real time according to the dynamic changes of the node state, so that the updated value reflecting the current load state of the node can be obtained. Based on the accurate load state, the node role and relationship graph are adjusted in time, so as to ensure the continuous and efficient and stable operation of the system, and the dynamic role allocation mechanism can quickly respond to the changes of the node load and protect the system performance.

[0091] As shown in Figure 2 , the updated load score of each initial master node is compared with a preset second load threshold, wherein the second load threshold can be set according to the system performance and actual demand. If the updated load score of a certain initial master node exceeds the second load threshold, it indicates that the load is too high and it is not suitable to continue to act as a master node. At this time, the initial slave node with the lowest updated load score is selected from the initial slave nodes belonging to the initial master node as an auxiliary master node, and the data writing task is shared by the original master node and the auxiliary master node. By introducing the auxiliary master node to disperse the load, the overall load distribution of the system can be optimized, and the performance and stability can be improved.

[0092] As shown in Figure 3 , in the distributed storage system, the initial master node may fail due to hardware or software failure. Once the failure is detected, the initial slave node with the lowest updated load score is selected from the initial slave nodes corresponding to the initial master node as a replacement master node to replace the original failed initial master node. By quickly replacing the master node, the data continuous availability can be maximized, the system interruption time is shortened, and the selection of low-load nodes in priority helps to improve the resource utilization efficiency.

[0093] Further, based on the determined auxiliary master node and replacement master node, the roles and connection relationships of the corresponding nodes in the relationship graph are updated, including establishing auxiliary master node related connections and adjusting the corresponding relationship between the replacement master node and the initial slave node; the updated master-slave node relationship graph is stored in the system storage module as the basis for subsequent data storage and access operations. By updating the relationship graph in real time to be consistent with the actual system state, data errors or operation abnormalities caused by inaccurate master-slave relationships can be avoided, ensuring the correct execution of data storage, backup and reading processes.

[0094] Further, according to the access frequency, each block data is written to the corresponding master node, and the data increment corresponding to each master node is copied to the corresponding slave node to obtain updated master nodes and updated slave nodes, including:

[0095] S501, according to the access frequency, the block data is divided into hot data block, warm data block and cold data block;

[0096] S502, based on the real-time state of each master node, the comprehensive state score of each master node is calculated through a preset node state evaluation model;

[0097] S503, combining the comprehensive state score and the type of block data, each block data is matched to the corresponding master node and written to the corresponding data to obtain updated master nodes;

[0098] S504, according to the data writing of the updated master node, the data increment is copied to the corresponding slave node to obtain updated slave nodes.

[0099] In this embodiment, the block data is divided into hot data block, warm data block and cold data block according to the access frequency, the hot data block is accessed frequently and needs to respond quickly, the cold data block is accessed less and has lower response speed requirements, and the warm data block is between the hot data block and the cold data block in terms of access frequency and performance requirements. According to the data type, it is stored in the corresponding data node, and each data block is allocated to the corresponding storage node.

[0100] At the same time, the storage performance and state of the master node are evaluated through a preset node state evaluation model, and the comprehensive state score of each master node is calculated. The node state evaluation model includes but is not limited to a weighted model, based on the collected master node state indicators, including CPU usage, memory usage, disk remaining space and network bandwidth utilization, etc. According to the influence of each indicator on the performance of the node, the corresponding weight is allocated, and the comprehensive state score is calculated by weighting. The comprehensive state score can comprehensively reflect the actual state of the master node, overcome the limitations of single indicator evaluation, and provide a basis for the matching of data blocks and nodes to ensure the efficiency and stability of data writing.

[0101] Specifically, the master nodes are layered according to the comprehensive state scores, different types of data blocks are matched with the nodes of the corresponding levels, and the master nodes with performance suitable for each data block are allocated and writing is performed, and the nodes become updated master nodes. Through the node matching mechanism, hot data blocks are written into high-performance master nodes to ensure access speed, and cold data blocks are stored in lower-performance nodes to save storage costs, while avoiding data centralized writing to a few nodes, achieving relatively balanced load and improving overall system performance.

[0102] Further, after the master nodes complete data writing, the data increments on the master nodes are replicated to the corresponding slave nodes. By monitoring the writing operation of the master nodes in real time, the data changes are recorded to obtain data increments, and then the incremental data is replicated to the corresponding slave nodes according to the preset master-slave node mapping relationship. After the replication is completed, the slave node becomes an updated slave node. Through the incremental replication mechanism, data backup can be realized to guarantee the recoverability and availability of data in the event of master node failure. In addition, the slave nodes process part of the read requests to reduce the pressure on the master nodes. Only the changed part of the data is replicated to reduce the transmission volume and storage overhead, thereby improving the system efficiency.

[0103] Further, in combination with the comprehensive state scores and the types of block data, each block data is matched to the corresponding master node and written to the corresponding data to obtain updated master nodes, including:

[0104] S601, layering the master nodes according to the comprehensive state scores to obtain a plurality of layers of master nodes, wherein the plurality of layers of master nodes include high-layer master nodes, middle-layer master nodes, and low-layer master nodes;

[0105] S602, for hot data blocks, matching to the corresponding first master nodes in the high-layer master nodes in descending order of the comprehensive state scores;

[0106] S603, for warm data blocks, matching to the corresponding second master nodes in the middle-layer master nodes in descending order of the comprehensive state scores;

[0107] S604, for cold data blocks, matching to the corresponding third master nodes in the low-layer master nodes in descending order of the comprehensive state scores;

[0108] S605, until each block data is matched to the corresponding master node, and the block data is written into the corresponding master node to obtain updated master nodes.

[0109] In the embodiment, the master nodes are stratified based on the calculated comprehensive state scores, and the stratification thresholds are preset according to actual performance requirements of the system. For example, the master nodes with a comprehensive state score greater than or equal to 80 points are classified as high-layer master nodes, the master nodes with a comprehensive state score greater than 40 points and less than 80 points are classified as middle-layer master nodes, and the master nodes with a comprehensive state score less than or equal to 40 points are classified as low-layer master nodes. Through the master node stratification mechanism, different performance nodes can be clearly distinguished, so that the data can be stored in the most suitable node according to the data characteristics, and the overall performance of the storage system is improved.

[0110] For the hot data blocks with high access frequency and strict performance requirements, the high-layer master nodes are arranged in descending order of comprehensive state scores, and the corresponding high-layer master nodes are selected from the head of the sorted list as the first master nodes for each hot data block in sequence. The high-layer master nodes have excellent performance and can meet the needs of the hot data blocks for high-speed reading and writing, effectively reducing the access delay.

[0111] For the warm data blocks with moderate access frequency and performance requirements, the middle-layer master nodes are sorted in descending order of comprehensive state scores, and the corresponding middle-layer master nodes are selected from the head of the sequence as the second master nodes for each warm data block in sequence. The performance of the middle-layer master nodes is moderate, which can meet the storage needs of the warm data blocks and avoid occupying high-layer node resources or affecting the access performance due to storage in low-layer nodes.

[0112] For the cold data blocks with low access frequency and relaxed performance requirements, the low-layer master nodes are arranged in descending order of comprehensive state scores, and the corresponding low-layer master nodes are selected from the front end of the sequence as the third master nodes for each cold data block in sequence. The low-layer master nodes have low comprehensive state scores and relatively economical costs, and are used for storing cold data to save storage costs while ensuring basic storage needs.

[0113] After matching the corresponding master nodes for each data block, the data is written into the corresponding nodes to complete the storage. During the writing process, measures such as verification are taken to ensure data integrity and consistency, and the state information of the master nodes is updated after writing, including disk usage and load state, so that the updated master nodes are obtained, data persistent storage is realized, and safety and availability are guaranteed.

[0114] Further, according to the data writing of the updated master nodes, the data increment is replicated to the corresponding slave nodes to obtain updated slave nodes, including:

[0115] S701, according to the data writing of the updated master nodes and the state of the slave nodes, the corresponding slave nodes are matched through a preset master-slave node mapping relationship;

[0116] S702, the data increment in the updated master node is replicated to the corresponding slave node to obtain the updated slave node.

[0117] In the embodiment, when the master node completes data writing, the corresponding data increment is synchronized to the corresponding slave node. The system continuously monitors the writing operation of the master node, records the data amount, time and data block position and the like, and obtains the state information of the slave node. According to the preset master-slave node mapping relationship, in combination with the data distribution and load state of the slave node, the slave node with sufficient resources and appropriate load is selected for data replication, so as to avoid replication to a node with insufficient resources or high load, thereby improving the resource utilization efficiency of the system.

[0118] Specifically, after determining the slave node corresponding to the master node, the data increment changed in the master node is replicated to the selected slave node. The data increment to be replicated is identified based on the writing information recorded by the master node. The high-efficiency data transmission protocol such as NFS, iSCSI or the like is used to transmit the incremental data to the corresponding slave node. The slave node writes the incremental data to the local storage device after receiving the incremental data, and updates the state information after completing the replication, thereby becoming an updated slave node. Through the incremental replication mechanism, the data transmission amount and replication time are effectively reduced, the data synchronization efficiency is improved, and the system load is reduced.

[0119] Further, in response to the data reading demand of the user, the access correlation between the data is analyzed, and the corresponding updated slave node associated with the target data block is fused to read the corresponding data.

[0120] S801, in response to the data reading demand of the user, the target data block to be read is obtained.

[0121] S802, by analyzing the access correlation between the data, the corresponding updated slave node associated with the target data block is fused to read the corresponding data.

[0122] In the embodiment, after the user initiates a data reading request to the distributed block storage system, the system responds and determines the target data block to be read, and locates the slave node storing the data block.

[0123] Specifically, based on the access correlation between the target data block and other data, the data block associated therewith is identified, and the updated slave node where the data block is located is determined for fusion. By pre-fetching the associated data that the user may access, the data reading speed can be improved. According to the correlation characteristics between the data accesses, the related data block information is obtained in advance, so as to effectively shorten the user waiting time, thereby improving the data reading efficiency and the overall response performance of the system.

[0124] Further, by analyzing the access correlation between the data, the corresponding updated slave node associated with the target data block is fused to read the corresponding data, including:

[0125] S901, according to the access correlation between the data, the data correlation access probability of each updated slave node is calculated through a preset data block correlation analysis model.

[0126] S902, merging the updated slave nodes whose data association access probability is greater than a preset association access probability threshold to obtain a fused slave node block;

[0127] S903. Select a preset number of update slave nodes from the fused slave node block as a set of candidate access nodes according to the update slave nodes corresponding to the target data block and in descending order of data association access probability;

[0128] S904: Read the corresponding target data block from the candidate access node set.

[0129] In this embodiment, based on the access correlation between data, a preset data block association analysis model is used, including but not limited to the Apriori algorithm, to analyze the data access correlation and calculate the associated access probability between the data blocks stored in each update slave node and the target data block. The model is trained based on a large amount of historical data and can evaluate the importance and access possibility of each update slave node data in the current reading operation.

[0130] like Figure 4 As shown in the figure, the associated access probability threshold is set according to the actual system conditions and experience, and the update slave nodes whose data associated access probability exceeds the associated access probability threshold are screened out for fusion. These update slave nodes that are closely associated with the target data block and have a higher access probability are combined into a fused slave node block. Through node fusion, we can focus on nodes that are important for the current read operation and improve the reading targeting.

[0131] Specifically, after forming a fused slave node block, the update slave node where the target data block is located is determined, and the update slave nodes in the fused slave node block are sorted from high to low based on the probability of data association access. A preset number of top nodes are selected from the update slave node sorting to form a set of candidate access nodes. This prioritizes the nodes with the highest association, increasing the probability of successfully acquiring the target data and reducing unnecessary node accesses.

[0132] Furthermore, the corresponding target data block is read from the set of alternative access nodes. When the user requests related data, the data is read directly from the set of alternative access nodes. This can avoid re-retrieval of the next read target from the node during the full update, thereby improving the data reading speed, reducing the reading time and resource consumption, and achieving performance optimization of the distributed block storage system by centralized reading of the nodes.

[0133] A distributed block storage system performance optimization system is used to implement a distributed block storage system performance optimization method, including:

[0134] The data division module analyzes the access frequency of the data to be stored to dynamically divide the data into a plurality of block data in response to the user's data storage requirement.

[0135] The data storage node division module dynamically divides the data storage nodes into master nodes and slave nodes according to the load state of each node in the storage system.

[0136] The data writing module writes each block data into the corresponding master node according to the access frequency, and replicates the data increment of each master node to the corresponding slave node to obtain updated master nodes and updated slave nodes.

[0137] The data reading module analyzes the access correlation between the data to perform a fusion reading of the corresponding data from the updated slave nodes in response to the user's data reading requirement, so as to optimize the performance of the distributed block storage system.

[0138] In this embodiment, the data division module dynamically divides the data to be stored into a plurality of data blocks in response to the user's data storage requirement. Based on the prediction of the data access frequency, the resource allocation in the subsequent storage and reading process is optimized, thereby improving the storage efficiency and data access performance. The data storage node division module dynamically divides the nodes into master nodes and slave nodes according to the load state of each node in the storage system, and constructs and updates the master-slave node relationship graph. By reasonably dividing the nodes, the storage system load balancing is achieved, and the system performance and stability are improved. At the same time, the slave nodes are used as data backups of the master nodes to enhance the data reliability and availability.

[0139] Specifically, the data writing module writes the block data into the matching master node according to the data access frequency and the real-time state of the master node, and replicates the data increment to the corresponding slave node, thereby optimizing the storage resource allocation, improving the data writing efficiency and reliability. The data reading module responds to the user's data reading request, performs a fusion reading operation on the updated slave nodes by analyzing the access correlation between the data, accurately locates the slave nodes associated with the target data block, preferentially reads the data from the nodes with close correlation, reduces the reading times and time, improves the data reading efficiency and accuracy, and further optimizes the performance of the distributed block storage system.

[0140] The above is only the preferred embodiment of the present application, and the protection scope of the present application is not limited to the above-mentioned embodiments. Any technical solutions falling within the scope of the present application should be considered as falling within the protection scope of the present application. It should be noted that some improvements and refinements made by those skilled in the art without departing from the principles of the present application should also be considered as falling within the protection scope of the present application.

Claims

1. A method for optimizing the performance of a distributed block storage system, characterized in that: include: In response to the user's data storage needs, the access frequency of the data to be stored is analyzed to perform dynamic block division and generate multiple block data; According to the load status of each node in the storage system, the data storage nodes are dynamically divided into master nodes and slave nodes; According to the access frequency, each block of data is written to the corresponding master node, and the data increment corresponding to each master node is copied to the corresponding slave node to obtain an updated master node and an updated slave node; In response to the user's data reading requirements, the access association between the data is analyzed, and the corresponding data is read from the update slave node in a fusion manner to optimize the performance of the distributed block storage system.

2. The distributed block storage system performance optimization method according to claim 1, characterized in that: The method of analyzing the access frequency of the data to be stored to dynamically divide the data into blocks in response to the user's data storage demand and generating a plurality of block data includes: In response to the user's data storage requirements, analyzing and obtaining the data type and data content of the data to be stored; According to the data type, the access frequency of the data is predicted using a preset access frequency prediction model to obtain a predicted access frequency; The data content is divided into blocks based on the predicted access frequency and the preset block strategy to obtain multiple block data.

3. The distributed block storage system performance optimization method according to claim 1, characterized in that: The method of dynamically distinguishing data storage nodes into master nodes and slave nodes according to the load status of each node in the storage system includes: Based on the load status of the nodes in the storage system, the load score of each node is calculated using the preset node load evaluation model; The node whose load score is less than the preset first load threshold is used as the initial master node, and the node whose load score is greater than or equal to the preset first load threshold is used as the initial slave node; By analyzing the position relationship between nodes and the consistency of data storage types, each initial master node is assigned at least one initial slave node to build a master-slave node relationship graph; According to the real-time status of the data storage node, the master-slave node relationship diagram is updated through a preset node update mechanism to obtain the real-time master node and slave node.

4. The distributed block storage system performance optimization method according to claim 3, characterized in that: The master-slave node relationship diagram is updated according to the real-time status of the data storage node through a preset node update mechanism to obtain the real-time master node and slave node, including: According to the real-time status of the data storage node, the load score of each node is updated in real time to obtain an updated load score; When the update load score of the initial master node is greater than the preset second load threshold, the node with the smallest update load score is selected from the corresponding initial slave nodes as the auxiliary master node; When the initial master node fails or fails, the node with the smallest update load score is selected from the corresponding initial slave nodes as the replacement master node; According to the auxiliary master node and the replacement master node, the master-slave node relationship diagram is updated to obtain the real-time master node and slave node.

5. The distributed block storage system performance optimization method according to claim 1, characterized in that: According to the access frequency, each block of data is written to the corresponding master node, and the data increment corresponding to each master node is copied to the corresponding slave node to obtain an updated master node and an updated slave node, including: According to the access frequency, the block data is divided into hot data blocks, warm data blocks and cold data blocks; Based on the real-time status of each master node, the comprehensive status score of each master node is calculated through the preset node status evaluation model; Combined with the comprehensive status score and the type of block data, each block data is matched to the corresponding master node and the corresponding data is written to obtain an updated master node; According to the data writing status of the updated master node, the data increment is copied to the corresponding slave node to obtain the updated slave node.

6. The distributed block storage system performance optimization method according to claim 5, characterized in that: Combined with the comprehensive status score and the type of block data, each block data is matched to the corresponding master node and the corresponding data is written to obtain an updated master node, including: According to the comprehensive status scores, the master nodes are layered to obtain multi-layer master nodes, wherein the multi-layer master nodes include high-layer master nodes, middle-layer master nodes and low-layer master nodes; For hot data blocks, they are matched to corresponding first master nodes in the high-level master nodes according to the comprehensive status scores from high to low; For warm data blocks, they are matched to corresponding second master nodes in the middle-level master nodes according to the comprehensive status scores from high to low; For cold data blocks, they are matched to corresponding third master nodes in the lower-level master nodes according to the comprehensive status scores from high to low; Until each block data is matched to the corresponding master node, the block data is written to the corresponding master node to obtain the updated master node.

7. The distributed block storage system performance optimization method according to claim 5, characterized in that: According to the data writing status of the updated master node, the data increment is copied to the corresponding slave node to obtain the updated slave node, including: According to the data written to the updated master node and the status of the slave node, the corresponding slave node is matched through the preset master-slave node mapping relationship; The data increment in the updated master node is copied to the corresponding slave node to obtain the updated slave node.

8. The distributed block storage system performance optimization method according to claim 1, characterized in that: The step of responding to a user's data reading demand, analyzing access relevance between data, and fusing the update slave nodes to read corresponding data includes: Responding to a user's data reading request, obtaining a target data block to be read; By analyzing the access association between data, the corresponding update slave nodes associated with the target data block are merged to read the corresponding data.

9. The distributed block storage system performance optimization method according to claim 8, characterized in that: The step of analyzing the access association between the data and fusing the corresponding updates associated with the target data block from the node to read the corresponding data includes: According to the access correlation between data, the data association access probability of each update slave node is calculated through the preset data block association analysis model; Merging the update slave nodes whose data association access probability is greater than a preset association access probability threshold to obtain a fused slave node block; According to the update slave nodes corresponding to the target data block, a preset number of update slave nodes are selected from the fused slave node block in descending order of the data association access probability as a set of candidate access nodes; Read the corresponding target data block from the candidate access node set.

10. A distributed block storage system performance optimization system, characterized in that: A method for optimizing the performance of a distributed block storage system according to any one of claims 1 to 9, comprising: A data partitioning module, in response to a user's data storage needs, analyzes the access frequency of the data to be stored to perform dynamic block division and generate multiple block data; The data storage node division module dynamically divides the data storage nodes into master nodes and slave nodes according to the load status of each node in the storage system; A data writing module writes each block of data to the corresponding master node according to the access frequency, and copies the data increment corresponding to each master node to the corresponding slave node to obtain an updated master node and an updated slave node; The data reading module responds to the user's data reading needs, analyzes the access correlation between the data, and integrates the updated slave nodes to read the corresponding data to optimize the performance of the distributed block storage system.

Citation Information

Patent Citations

  • File block storage method and device

    CN111291009A

  • Method for copying data and system for copying updated data from node

    CN118259824A

  • Financial data storage path optimization method based on big data

    CN118377432A

  • Data processing method and device, electronic equipment and storage medium

    CN118540334A

  • Distributed account book remote data synchronization method and system

    CN119449829A

Cited By

  • Processing system for improving data storage speed of server

    CN121334164A