Distributed system data management method and device, electronic equipment and storage medium

By introducing physical capacity awareness and system pre-check mechanisms, combined with metadata comparison and cleanup priority calculation, the problem of inconsistency between logical capacity and physical capacity in distributed storage systems during expansion is solved, achieving more efficient resource management and data consistency, and improving system stability and resource utilization.

CN120803710APending Publication Date: 2025-10-17JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510904658.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

During the expansion process, existing distributed storage systems suffer from inconsistencies between logical and physical capacity, leading to storage operation failures, system overloads, and data consistency risks. Furthermore, they lack effective node heterogeneity processing, impacting system stability and resource utilization.

Method used

By introducing a physical capacity awareness mechanism, scanning node metadata in real time, and dynamically calculating the actual physical available space, the wooden barrel effect model is used to calculate the system available capacity based on the lowest node capacity. Combined with the system pre-check mechanism, physical capacity verification is performed before capacity expansion. During the data migration process, data integrity is confirmed through metadata comparison, data blocks with low access frequency are cleared first, and audit logs of the clearing operations are recorded.

Benefits of technology

It improves the accuracy of capacity management, enhances the stability of system expansion, improves the efficiency of storage space recovery, ensures data consistency, enhances the system's maintainability and adaptability to node heterogeneity, and significantly improves the system's stability, reliability and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803710A_ABST
    Figure CN120803710A_ABST
Patent Text Reader

Abstract

The invention relates to a distributed system data management method and device, and the method comprises the steps: scanning node metadata in real time and dynamically calculating a physical available space, carrying out the physical capacity verification before the execution of storage operation, pausing the operation and triggering a space recovery task if the space is insufficient, responding to a data block reading event in a data migration process, and carrying out the storage operation. And marking the to-be-cleaned state, comparing and confirming the integrity through metadata after the storage of the target node is completed, and executing the atom deletion operation. Meanwhile, the cleaning priority is determined according to the access frequency of the data blocks, the space recovery efficiency is improved, and an audit log is recorded to ensure the operation traceability. The method can be applied to a server in a hyper-converged architecture, realizes data balance and storage space management during node capacity expansion, and is suitable for the fields of cloud computing, storage pooling, distributed storage systems and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computers, in particular to a distributed storage system, and more particularly to a distributed system data management method and device, an electronic device, and a storage medium. BACKGROUND

[0002] As an important infrastructure of modern data centers, distributed storage systems are widely used in cloud computing, virtualization, and large-scale data processing scenarios. In related technologies, through the collaborative work of computing, storage, and network resources, a distributed storage system under a hyper-converged architecture is constructed. Specifically, this system covers the entire process from data sharding, node balancing, to replica management, including logical capacity allocation, physical storage scheduling, data migration, and consistency guarantee. With the increase in the number of storage nodes and the dynamic changes in business loads, the system has higher requirements for the real-time and accuracy of capacity management, and traditional methods have been difficult to meet the needs of high availability and efficient resource utilization.

[0003] However, in the existing distributed storage capacity management method, logical capacity is directly used as the basis for resource allocation, and the occupation of migrated replicas in physical space is not fully considered, which may cause the system to still accept new storage requests in the case of insufficient physical space, leading to write failures or system overload. In addition, the traditional "clean up after migration" mechanism does not release the source node space before migration is completed, causing a "virtual high" phenomenon of storage capacity, affecting the scheduling efficiency and resource utilization of the system. At the same time, the existing technology lacks effective identification and processing of node heterogeneity, which may cause local node space to be exhausted during the expansion process, thereby affecting the stability and data consistency of the overall system. SUMMARY

[0004] The present application aims to at least partially solve one of the technical problems in the related art.

[0005] To this end, the first object of the present application is to propose a distributed system data management method, particularly suitable for storage resource management in hyper-converged infrastructure, aiming to solve the problems of storage operation failure, system overload, and data consistency risk caused by the inconsistency between logical capacity and physical capacity, the failure to release the space occupied by migrated replicas in time, and the low data cleaning efficiency during system expansion.

[0006] The second object of the present application is to propose a distributed system data management device.

[0007] The third object of the present application is to propose an electronic device.

[0008] The fourth object of the present application is to propose a computer-readable storage medium.

[0009] To achieve the above object, the first aspect of the present application provides a distributed system data management method, comprising: in response to an event that a new node joins a cluster, real-time scanning metadata of each node, dynamically calculating actual physical available space of each node, deducting space occupied by a to-be-deleted migration copy, and generating a node available capacity benchmark value; calculating a storage capacity that can be safely allocated by the system as a whole based on the node available capacity benchmark value, and before performing a storage operation, performing physical capacity verification according to a required storage capacity of the current operation, when detecting that the physical available space is lower than the required space, suspending the current operation and triggering a space recovery task, saving operation site information, continuously monitoring storage state changes, and after detecting that the space is released, automatically retrying; in a data migration process, in response to an event that a data block is read, marking the data block as a to-be-cleaned state in a metadata management module, when the data block is stored on a target node, confirming data integrity through metadata rapid comparison, and after confirming that the target node has completely saved the data, triggering an atomic deletion operation; calculating a cleaning priority according to an access frequency of the data block, preferentially cleaning the data block with low access frequency, so as to accelerate the release of storage space, and when performing the cleaning operation, recording an audit log of the cleaning operation.

[0010] In an embodiment of the present application, the capacity calculation model based on the wood barrel effect is used to take the lowest node capacity in the system as a benchmark to calculate the storage capacity that can be safely allocated by the system as a whole.

[0011] In an embodiment of the present application, before performing a LUN creation or expansion operation, the required storage capacity of the operation is calculated and compared with the storage capacity that can be safely allocated by the system as a whole, if the storage capacity that can be safely allocated by the system as a whole is less than the required storage capacity of the operation, the current operation is suspended and a space recovery task is triggered.

[0012] In an embodiment of the present application, in a data migration process, in response to an event that a data block is read by a target node, the data block is updated from a to-be-cleaned state to a read state in a metadata management module, so as to optimize the execution efficiency of a subsequent cleaning process.

[0013] In an embodiment of the present application, before triggering an atomic deletion operation, the storage state of a data block marked as a read state on a target node is preferentially verified, if the verification is passed, the cleaning operation is immediately performed, if the verification fails, a data retransmission process is automatically started, so as to ensure data consistency.

[0014] To achieve the above object, the second aspect of the present application provides a distributed system data management device, comprising: a data acquisition module, configured to scan metadata of each node in real time, dynamically calculate actual physical available space of each node, deduct space occupied by a to-be-deleted migration copy, and generate a node available capacity benchmark value in response to an event that a new node joins a cluster; a system pre-checking module, configured to calculate a storage capacity that can be safely allocated by the system as a whole based on the node available capacity benchmark value, and perform physical capacity checking based on a required storage capacity of a current operation before executing the storage operation, suspend the current operation and trigger a space recovery task when detecting that the physical available space is lower than the required space, save operation site information, continuously monitor storage state changes, and automatically retry after detecting that space is released; and a data cleaning module, configured to mark a data block as a to-be-cleaned state in a metadata management module in response to an event that the data block is read during data migration, confirm data integrity through metadata rapid comparison when the data block is stored on a target node, and trigger an atomic deletion operation after confirming that the target node has completely saved the data.

[0015] In an embodiment of the present application, the data acquisition module is further configured to calculate the storage capacity that can be safely allocated by the system as a whole based on a capacity calculation model based on the wood bucket effect, taking the lowest node capacity in the system as a benchmark.

[0016] In an embodiment of the present application, the system pre-checking module is further configured to calculate a required storage capacity of a current operation before executing a LUN creation or expansion operation, and compare the required storage capacity with the storage capacity that can be safely allocated by the system as a whole, and if the storage capacity that can be safely allocated by the system as a whole is less than the required storage capacity, suspend the current operation and trigger a space recovery task.

[0017] To achieve the above object, the third aspect of the present application provides an electronic device, comprising: a processor; a memory storing executable instructions; and the processor executes the instructions to implement the distributed system data management method of the first aspect of the present application.

[0018] To achieve the above object, the fourth aspect of the present application provides a computer readable storage medium storing a computer program, and the program is executed by a processor to implement the distributed system data management method of the first aspect of the present application.

[0019] The distributed system data management method, device, electronic equipment and computer readable storage medium of the embodiment of the application have the following technical effects and advantages by introducing a physical capacity sensing mechanism, a system pre-checking mechanism and a three-stage fast cleaning mechanism

[0020] Beneficial effects:

[0021] 1. Improve the accuracy of capacity management: by scanning metadata in real time and actively deducting the space occupied by migration copies, the system can accurately reflect the actual physical available capacity of each node, avoiding resource allocation errors caused by inconsistencies between logical capacity and physical capacity.

[0022] 2. Enhance the stability of system expansion: perform physical capacity verification before executing LUN creation or expansion operations to ensure that the system only performs operations when there is sufficient physical space, thereby avoiding operation failures or system overload due to insufficient space.

[0023] 3. Improve storage space recycling efficiency: start the cleaning process during data migration, and preferentially clean data blocks with low access frequency, significantly shortening the space recycling period and improving the utilization of storage resources.

[0024] 4. Ensure data consistency: after data migration is complete, confirm that the target node has saved the data completely through metadata comparison, and only perform atomic deletion operations after consistency verification is passed, ensuring the integrity and consistency of the data migration process.

[0025] 5. Enhance system maintainability and traceability: record detailed information of each cleaning operation through audit logs, facilitating quick problem location and taking appropriate measures to improve system operation efficiency and security.

[0026] 6. Adapt to node heterogeneity: use a capacity calculation model based on the bucket effect, taking the lowest node capacity in the system as the benchmark to avoid local storage space depletion problems caused by differences in node hardware configurations.

[0027] In summary, the application effectively solves the problems of capacity misjudgment, resource waste, and difficulty in ensuring data consistency in traditional methods during distributed system expansion, significantly improving the stability, reliability, resource utilization, and data security of the system, and has good technical advancement and practical application value.

[0028] Additional aspects and advantages of the application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS

[0029] The above and / or additional aspects and advantages of the application will become apparent and be readily appreciated from the following description, including the appended drawings, wherein:

[0030] Figure 1 is a flow chart of a distributed system data management method of an embodiment of the present application.

[0031] Figure 2 is a structural schematic diagram of a distributed system data management apparatus of an embodiment of the present application. DETAILED DESCRIPTION

[0032] Embodiments of the present application are described in detail below with reference to the attached drawings, of which examples of embodiments are shown, wherein the same or similar notations are used to denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by reference to the drawings are exemplary and are intended to explain the present application, and cannot be understood as limiting the present application.

[0033] In addition, the terms "first", "second", etc. are used only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined with "first", "second", etc. can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise explicitly and specifically limited.

[0034] Any process or method descriptions or any other descriptions herein can be understood as representing embodiments of the present application comprising an executable code of one or more modules, segments or portions for accomplishing a specific logical function or process, and the scope of the preferred embodiments of the present application includes additional implementation in which the functions are accomplished by additional means, in which the order of execution of the functions can be changed, including substantially simultaneously, and in which the functions can be performed in reverse order, as will be understood by those skilled in the art of the embodiments of the present application.

[0035] Figure 1 is a flow chart of a distributed system data management method of an embodiment of the present application.

[0036] As shown in Figure 1 The present application provides a distributed system data management method, which solves the problem of misjudgment of storage resources caused by inconsistency between logical capacity and physical capacity in the process of node expansion in traditional distributed storage systems through three core mechanisms of physical capacity awareness, system pre-inspection and rapid cleaning, and improves the efficiency of data migration and cleaning.

[0037] In step S101, in response to the event of adding a new node to the cluster, the metadata of each node is scanned in real time, the actual physical available space of each node is dynamically calculated, and the space occupied by the to-be-deleted migration copy is deducted to generate a node available capacity benchmark value.

[0038] Specifically, the prior art only relies on logical capacity when calculating available capacity, without considering the physical occupation of migrated copies, resulting in capacity misjudgment. For example, the traditional scheme does not deduct redundant copy space during migration, which may cause the system to misjudge that the capacity is sufficient, but the actual physical space is insufficient. The present application solves this problem by actively deducting the occupied space of migrated copies through real-time scanning of metadata.

[0039] In the embodiments of the present application, the physical capacity awareness module is responsible for dynamically calculating the actual physical available space of the newly added nodes in the cluster to ensure that the storage resources are reasonably and safely allocated during system expansion. The system scans the metadata of each node in real time through the data acquisition module, dynamically calculates the available capacity of each node, and adjusts the expansion strategy of the system based on this information.

[0040] The system first scans each node in the cluster in real time through the data acquisition module to obtain the metadata information of each node. By traversing the metadata of all nodes, the system can understand the effective available capacity Ci (unit: GB) of each node. These information provides a basis for the system to dynamically adjust the expansion scheme. The system uses storage pooling technology to manage nodes uniformly, and dynamically calculates the effective available capacity of each node through metadata scanning of the nodes. The effective available capacity refers to the actual available storage space of the node considering the redundant copies and space usage.

[0041] On this basis, the system uses the bucket effect model to determine the benchmark value SminS of the available capacity of the nodes during system expansion, taking the lowest node capacity as the benchmark to ensure that the physical space of all nodes in the cluster is evenly distributed, avoiding the risk of local storage depletion due to large differences in node capacity. The calculation formula is:

[0042] S min C1C2C3C i

[0043] Among them, C i represents the effective available capacity (unit: GB) of the i-th node, and S min is the benchmark value of the available capacity of the nodes for system expansion safety. This model takes the lowest node capacity as the benchmark to avoid the risk of local storage depletion due to node heterogeneity. By using the bucket effect model, the system avoids the problem of depletion of storage space of some nodes due to node heterogeneity (difference in capacity of different nodes).

[0044] Further, the system can dynamically calculate the overall available storage capacity Cavailable of the system in order to reasonably allocate storage resources during expansion. This capacity calculation considers the effective available space of all nodes in the system, including the total available capacity of the original nodes and the newly added nodes. The overall available storage capacity calculation formula is:

[0045] C available =S min ×N total

[0046] Where N total is the total number of existing storage nodes in the system, including the original nodes N original and the new nodes N new . This formula ensures that the system can dynamically adjust the resource allocation strategy according to the actual physical capacity when expanding, avoiding resource waste caused by false high logical capacity.

[0047] In the prior art, many systems only rely on logical capacity when calculating available capacity, without considering the physical occupation of storage replicas in nodes. The traditional solution may misjudge the capacity sufficiency of the system, resulting in not deducting the space occupied by redundant replicas during migration. For example, redundant replicas occupy a certain amount of physical storage space during migration, but these replica spaces are not deducted from the logical capacity, so the system may misjudge as sufficient capacity, but the actual physical space is insufficient.

[0048] Compared with the prior art, the present application solves this problem by actively deducting the space occupied by migration replicas through real-time scanning of metadata. The system can obtain the actual physical capacity of each node in real time and dynamically adjust the expansion and resource allocation strategy, thereby improving the accuracy and efficiency of system capacity utilization.

[0049] By introducing a physical capacity awareness module, the present application greatly improves the resource allocation efficiency and accuracy during system expansion. Compared with traditional technology, this solution not only improves the safety of expansion operation, but also avoids resource waste caused by false judgment of logical capacity.

[0050] Step S102, based on the node available capacity reference value, calculate the system overall safe allocation of storage capacity, and before executing the storage operation, according to the required storage capacity of the current operation, carry out physical capacity verification, when detecting that the physical available space is lower than the required space, suspend the current operation and trigger the space recycling task, save the operation site information, continuously monitor the storage state change, and after detecting the space release, automatically retry.

[0051] In one embodiment of the present application, based on the capacity calculation model of the wooden barrel effect, the lowest node capacity in the system is taken as the reference to calculate the system overall safe allocation of storage capacity.

[0052] Specifically, the prior art lacks a pre-check mechanism, resulting in still accepting storage requests when the physical space is insufficient, ultimately causing operation failure. The present application significantly reduces the operation failure rate through double verification (logical capacity and physical capacity) and automatic retry mechanism.

[0053] In the embodiments of the present application, before performing the LUN creation or expansion operation, the required storage capacity of the current operation is calculated and compared with the overall safe allocable storage capacity of the system. If the overall safe allocable storage capacity of the system is less than the required storage capacity of the current operation, the current operation is suspended and a space recycling task is triggered. Specifically, the system pre-checking module mainly functions to pre-check the physical capacity of the system before performing the LUN creation or expansion operation, to ensure smooth execution of the operation. In traditional storage systems, real-time checking of the physical capacity is often ignored, resulting in continued execution of the operation in the case of insufficient physical space, and thus causing write failure or system overload.

[0054] The system pre-checking module will first calculate the required storage capacity Crequire of the current operation. This calculation is based on the size of the LUN to be expanded, the capacity of the newly added storage node, the data migration amount, and other factors, to ensure that the storage requirements of the system are accurately evaluated.

[0055] Further, after calculating the required storage capacity, the system pre-checking module will obtain the current available physical space Cavailable of the system in real time, and compare it with the required storage capacity Crequire. If the current physical space Cavailable is greater than or equal to the required storage capacity Crequire, it indicates that the system has sufficient space to perform the operation, and the operation can continue. If the current physical space Cavailable is less than the required storage capacity Crequire, it indicates that the physical capacity is insufficient, and the system will suspend the current operation and start a space recycling task.

[0056] When the system detects insufficient physical space, it will automatically trigger a space recycling task. The core of this task is to recycle the storage space in the current system that is not fully utilized, such as cleaning up redundant copies or data blocks that are no longer needed. By recycling unnecessary storage resources, the system releases space for subsequent operations, ensuring that the needs of LUN expansion or creation can be met.

[0057] In addition, after starting the space recycling task, the system will continuously monitor the storage status. When it detects that the available space of the system has recovered to a sufficient level, the system will automatically retry the pre-checking. Before each retry, the system will re-verify the current physical capacity, to ensure that sufficient physical space is available before performing the storage operation again.

[0058] It should be noted that in the prior art, many distributed storage systems lack a pre-check mechanism for physical capacity, resulting in the system still accepting storage requests and attempting to perform expansion operations in the case of insufficient physical space. As a result, it may cause storage write failures, even system overload, affecting the stability and reliability of the entire storage environment. Compared with the prior art, the present application can effectively avoid the problem of insufficient physical capacity by introducing a double-check mechanism for logical capacity and physical capacity. By monitoring the physical space in real time, the system can automatically suspend operations when the space is insufficient and trigger a space recycling task to ensure that the space requirement is met.

[0059] In addition, the present application also designs an automatic retry mechanism. When the physical space of the system is restored, the retry mechanism can ensure that the storage operation continues without the need for manual intervention. This mechanism greatly improves the fault tolerance of the operation and the stability of the system.

[0060] Step S103, in the data migration process, in response to the event that the data block is read, the data block is marked as a to-be-cleaned state in the metadata management module, when the data block is stored on the target node, the data integrity is confirmed through the metadata quick comparison, and after confirming that the target node has completely saved the data, the atomic deletion operation is triggered.

[0061] In an embodiment of the present application, in the data migration process, in response to the event that the data block is read by the target node, the data block is updated from the to-be-cleaned state to the read state in the metadata management module, to optimize the execution efficiency of the subsequent cleaning process.

[0062] In an embodiment of the present application, before triggering the atomic deletion operation, the storage state of the data block marked as the read state on the target node is preferentially verified, if the verification is passed, the cleaning operation is immediately executed, if the verification fails, the data retransmission process is automatically started to ensure data consistency.

[0063] Specifically, in the possible implementation mode, the marking state can be realized by a distributed lock mechanism or version number management to further improve the concurrent processing capability.

[0064] In the embodiment of the present application, when the system triggers the data migration operation, the data cleaning module will mark the state of the data block to be migrated through the metadata management module. This operation is a key link to ensure that the data migration and cleaning process is efficient and orderly, especially in a large-scale distributed storage environment.

[0065] Specifically, during the data migration process, the system records the relevant information of each data block to be migrated in the metadata management module, such as data block ID, current migration status, access frequency, etc. In the initial state, all data blocks to be migrated are marked as "to be cleaned" status. This state indicates that the data block is in the state of migration, and the migration operation has not been completed.

[0066] Once the data block in migration is successfully read and saved by the new node, the system further marks the data block as "read" status. At this time, the system confirms that the target node has completely received and saved the data block, and is ready to enter the next cleaning operation.

[0067] Through this marking mechanism, the system can ensure that the data is deleted on the source node only after the migration is completed, avoiding the risk of data loss or inconsistency.

[0068] This marking mechanism provides a basis for subsequent cleaning processes, ensuring data consistency and integrity during data migration. Only after the target node reads and verifies the data integrity can the deletion operation be safely performed, avoiding data loss.

[0069] In addition to using the metadata management module for data block marking, the present application can also use other technical solutions to further optimize and improve the concurrent processing capability and data consistency:

[0070] 1) Distributed lock mechanism: To avoid multiple nodes accessing the same data block simultaneously, which may cause data consistency problems, the system can introduce a distributed lock mechanism. During the data migration process, the system can lock each data block to be migrated, ensuring that only one node can read or modify the data block at the same time. This can ensure that there is no concurrent conflict during data migration, improving the processing efficiency of the system.

[0071] 2) Version number management: Another alternative solution is to use version number management. Each data block can be attached with a version number, and the system updates the version number during migration to record the migration status of the data block. Through the version number, the system can clearly distinguish the migration status of each data block, avoiding inconsistency problems caused by concurrent access. At the same time, the version number can provide more accurate management and tracking functions.

[0072] By using the metadata management module to mark the status of the data block, the present application can ensure that the status of each data block during migration is timely and accurately tracked. This marking mechanism can effectively avoid data loss or inconsistency problems during migration, ensuring data integrity and consistency. In addition, by supporting distributed lock mechanism and version number management, etc. Technical alternatives can further improve the concurrent processing capability of the system, so as to more efficiently perform data migration tasks in large-scale distributed environments.

[0073] In addition, in the embodiments of the present application, the active verification module is responsible for verifying the data integrity in the data migration process, ensuring that the target node successfully receives and saves the migrated data blocks. The introduction of the verification process makes the data migration not only rely on the migration operation of the data blocks itself, but also ensures the reliability and consistency of the data through the verification mechanism.

[0074] After the data migration is completed, the system will quickly compare the migration data block states of the source node and the target node through the metadata management module. The verification process first checks whether the target node has completely received all data blocks, and compares them with the data blocks of the source node to verify their consistency.

[0075] If the verification is passed, that is, the data block integrity verification of the target node is successful, the atomic deletion operation is triggered. At this time, the system can safely delete the data copy on the source node and release the storage space. If the verification fails, that is, the target node fails to correctly receive some data blocks, the system will start the data retransmission process. This process ensures that any failed part of the migration can be retransmitted to the target node in time, avoiding data loss or inconsistency.

[0076] The verification process uses a hash verification algorithm (such as SHA-256) to ensure data consistency. By calculating the hash values of the data blocks of the source node and the target node, the system can effectively confirm whether the data has changed during the migration process. Hash value matching means data consistency, otherwise the data needs to be retransmitted.

[0077] In the prior art, the data migration operation often lacks a verification mechanism after migration. In the traditional scheme, the cleaning operation after data migration is performed immediately after the migration of the data blocks on the target node is completed, but no effective verification is performed during this process. This approach may cause data consistency problems, especially when data loss or errors occur during the migration process, the system cannot discover and correct them in time.

[0078] The present application verifies the data integrity of the target node immediately after the data migration is completed by introducing an active verification mechanism, ensuring data consistency. If the verification fails, the system ensures complete data migration through the automatic retransmission process, thereby avoiding system failures caused by data loss or inconsistency. Compared with the prior art, the present application greatly improves the reliability and accuracy of the data migration process.

[0079] By adopting the hash check algorithm, the application provides an efficient and reliable check method, ensuring that the data of the target node is completely consistent after data migration. At the same time, the data retransmission mechanism further ensures that the data loss that may occur during migration can be repaired in time. This design significantly improves the integrity and security of data migration, avoids potential problems caused by inconsistent data, and ensures data consistency in distributed systems.

[0080] In step S104, the cleaning priority is calculated according to the access frequency of the data block, and the data block with low access frequency is preferentially cleaned to speed up the release of storage space, and an audit log of the cleaning operation is recorded when the cleaning operation is performed.

[0081] Specifically, in the embodiments of the present application, the cleaning priority calculation module is responsible for calculating the cleaning priority according to the access frequency of the data block. This step ensures that the system can effectively and orderly clean up the storage space, and preferentially release the data block that is not frequently accessed, thereby optimizing the utilization efficiency of storage resources.

[0082] The system records the access times of each data block through the access frequency statistical module, and calculates its access frequency according to the time window. The access frequency reflects the usage of the data block in a certain period of time, and the frequently accessed data block is preferentially retained, and the infrequently accessed data block can be preferentially cleaned.

[0083] In order to dynamically adjust the cleaning priority of the data block, the system adopts a sliding window algorithm. The sliding window algorithm can perform weighted average on the access frequency of the data block in a time period, and dynamically update the priority with time. This enables the system to respond to changes in data access patterns in time and optimize the cleaning strategy.

[0084] The cleaning priority calculation formula is:

[0085]

[0086] Where P i is the cleaning priority of data block i, f i is the access frequency of data block i (unit: times / unit time), and α is the priority adjustment coefficient (default value is 1.0). The formula ensures that the data block with low access frequency has a higher cleaning priority through an inverse proportional relationship, so the system can preferentially clean the data block with low access frequency and long time of non-use, thereby quickly releasing the storage space.

[0087] According to the priority calculated by the formula, the system sorts the data blocks to ensure that the data block with low access frequency is preferentially cleaned. Preferentially cleaning these data blocks not only improves the utilization rate of storage space, but also reduces the impact of cleaning operations on high-frequency access data, avoiding unnecessary impact on frequently used data.

[0088] To further optimize the cleaning process, the system can introduce a machine learning model to predict the future access patterns of data blocks. By analyzing historical access data, the machine learning model (such as random forest) can predict the future access trends of data blocks, thus more accurately adjusting the cleaning priority. This can help the system dynamically adapt to changes in data access behavior, avoid misjudgment and excessive cleaning of high-frequency access data blocks, and improve overall performance and efficiency. After introducing the machine learning model, the system can predict the future usage of certain data blocks by continuously learning and training, and further dynamically adjust their cleaning priorities. For example, if a data block has a high probability of future access, even if the current access frequency is low, the system will lower its cleaning priority based on the prediction results to avoid mis-cleaning.

[0089] In the prior art, data cleaning is usually based on fixed rules without a mechanism for dynamically adjusting cleaning priorities. Most systems may clean data at fixed time intervals or other static rules, but this cannot efficiently cope with changes in access frequency. Especially when facing large-scale data storage, fixed cleaning rules may cause frequently used data to be mis-cleaning, affecting system performance.

[0090] The present application can dynamically adjust the cleaning priority according to the actual access frequency of the data block and the future usage prediction by combining the sliding window algorithm and the machine learning model. This method not only ensures better protection of high-frequency data blocks, but also efficiently releases storage space, improving the response capability of the system and the utilization rate of storage resources.

[0091] By introducing dynamic cleaning priority calculation, the present application realizes efficient management of storage resources. The sliding window algorithm and machine learning optimization scheme ensure that the system can adapt to changes in access patterns in real time and release storage space without affecting frequently used data. Compared with traditional methods, the present application can significantly improve the recycling efficiency of storage space while ensuring data security, reducing the storage cost of redundant data.

[0092] In the embodiments of the present application, the audit log module is responsible for recording detailed operation logs when the system performs cleaning operations. Log recording ensures the transparency and traceability of the data cleaning process, facilitating subsequent operation review, problem positioning and troubleshooting.

[0093] At each execution of a cleaning operation, the system records the following key information through the audit log module:

[0094] Operation time: records the specific time when each operation occurs, used to track the time nodes of cleaning activities.

[0095] Data block ID: identifies the data block being cleaned or processed, ensuring that each data block operation has a clear identification.

[0096] Cleanup Status: Indicates the current status of the data block, such as "Cleaned" indicating that the data block has been deleted, and "Pending Cleanup" indicating that the data block is still in the pending cleanup list.

[0097] Verification Result: Records the verification result of the cleaned data block, usually "Pass" indicating that the target node data block has been correctly migrated and cleaned, or "Fail" indicating that the verification has not passed and the data migration has not been successful. The log record follows a standardized format to ensure the integrity and consistency of the data cleaning process. All log information is formatted according to the ISO 27001 standard, which not only ensures the accuracy of the information, but also meets international standards for data security and information management.

[0098] The audit log format is: Laudit = (Top, Dblock, Sstatus, Rresult)

[0099] Where, T op is the operation time, D block is the data block ID, S status is the cleanup status (such as "Cleaned", "Pending Cleanup"), and R result is the verification result (such as "Pass", "Fail").

[0100] All audit logs will be saved and can be queried and audited at any time. The storage of log files complies with the ISO 27001 standard, ensuring their integrity and security. Through this audit log record, administrators can review historical cleaning operations at any time, and in the event of any problems, can quickly trace back to specific operation details. This is crucial for system maintenance and problem troubleshooting.

[0101] In the prior art, many systems lack audit log recording functions when performing data cleaning operations. Lack of effective log recording will result in insufficient operation transparency, and once system failure or data problems occur, it may be very difficult to locate and trace the problem. At this time, system administrators cannot quickly obtain key details in the cleaning process, increasing the difficulty of troubleshooting, and cannot effectively evaluate the correctness and efficiency of the cleaning operation.

[0102] Compared with the prior art, the present invention significantly improves the maintainability and problem tracking ability of the system by introducing a structured audit log recording mechanism. By recording detailed operation logs, the system not only improves the transparency of the data cleaning process, but also ensures the traceability and verifiability of the entire cleaning process. This is particularly important for subsequent operation review and data security monitoring.

[0103] By introducing the audit log recording mechanism, the present application significantly improves the transparency, traceability and maintainability of the data cleaning operation. Compared with the traditional technology, this function not only improves the fault diagnosis efficiency of the system, but also enhances the control ability of the system to the data cleaning process.

[0104] The distributed system data management method of the embodiment of the present application has the following technical effects and beneficial effects by introducing the physical capacity awareness mechanism, the system pre-check mechanism and the three-stage quick cleaning mechanism:

[0105] 1. Improve the accuracy of capacity management: By scanning the metadata in real time and actively deducting the space occupied by the migration copy, the system can accurately reflect the actual physical available capacity of each node, avoiding resource allocation errors caused by the inconsistency between logical capacity and physical capacity.

[0106] 2. Enhance the stability of system expansion: Perform physical capacity verification before executing LUN creation or expansion operation, ensure that the system performs operation only when the physical space is sufficient, thereby avoiding operation failure or system overload caused by insufficient space.

[0107] 3. Improve the efficiency of storage space recycling: Start the cleaning process during the data migration process, preferentially clean the data blocks with low access frequency, significantly shorten the space recycling period, and improve the utilization rate of storage resources.

[0108] 4. Ensure data consistency: After the data migration is completed, confirm that the target node has saved the data completely through metadata comparison, and only perform atomic deletion operation after the consistency verification is passed, to ensure the integrity and consistency during the data migration process.

[0109] 5. Enhance system maintainability and traceability: Record the detailed information of each cleaning operation through audit log, which facilitates quick positioning of the cause and taking corresponding measures when problems occur, improving the operation and maintenance efficiency and security of the system.

[0110] 6. Adapt to node heterogeneity: Use the capacity calculation model based on the bucket effect, taking the lowest node capacity in the system as the benchmark, to avoid the problem of local storage space depletion caused by differences in node hardware configuration.

[0111] In summary, the present application effectively solves the problems of capacity misjudgment, resource waste, and difficulty in ensuring data consistency in the traditional method during the expansion process of the distributed system, significantly improves the stability, reliability, resource utilization rate and data security of the system, and has good technical advancement and practical application value.

[0112] Figure 2 is a structural schematic diagram of a distributed system data management device according to an embodiment of the present application.

[0113] AsFigure 2 As shown, the application provides a distributed system data management device, comprising:

[0114] The data collection module 11 is configured to, in response to an event of a new node joining the cluster, scan the metadata of each node in real time, dynamically calculate the actual physical available space of each node, deduct the space occupied by the to-be-deleted migration copy, and generate a node available capacity benchmark value.

[0115] The system pre-checking module 12 is configured to calculate the storage capacity that can be safely allocated by the system as a whole based on the node available capacity benchmark value, and perform physical capacity checking according to the required storage capacity of a current operation before executing the operation, suspend the current operation and trigger a space recycling task when detecting that the physical available space is lower than the required space, save operation site information, continuously monitor the storage state change, and automatically retry after detecting that the space is released.

[0116] The data cleaning module 13 is configured to, in response to an event of a data block being read during data migration, mark the data block as a to-be-cleaned state in the metadata management module, confirm the data integrity through metadata quick comparison when the data block is stored on a target node, and trigger an atomic deletion operation after confirming that the target node has completely saved the data.

[0117] The data cleaning module 13 is further configured to calculate a cleaning priority according to the access frequency of the data block, and preferentially clean the data block with low access frequency, so as to accelerate the release of the storage space.

[0118] The data cleaning module is further configured to record an audit log of the cleaning operation when the cleaning operation is performed.

[0119] In an embodiment of the application, the data collection module 11 is further configured to calculate the storage capacity that can be safely allocated by the system as a whole based on a capacity calculation model based on the wood bucket effect, taking the lowest node capacity in the system as a benchmark.

[0120] In an embodiment of the application, the system pre-checking module 12 is further configured to, before performing a LUN creation or expansion operation, calculate the required storage capacity of the operation, and compare the required storage capacity of the operation with the storage capacity that can be safely allocated by the system as a whole, and if the storage capacity that can be safely allocated by the system as a whole is less than the required storage capacity of the operation, suspend the current operation and trigger a space recycling task.

[0121] As to the device in the above embodiment, the specific manner in which each module performs the operation has been described in detail in the embodiment of the method, and will not be described in detail here.

[0122] It should be understood that various parts of the present application can be implemented in hardware, software, firmware or a combination thereof. In the above embodiments, various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any of the following technologies, known in the art, or their combination, can be used: discrete logic circuitry having logic gates for implementing logic functions upon data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.

[0123] In the present application, unless otherwise clearly specified and limited, the terms "mounting", "connecting", "connecting", and the like should be understood in a broad sense, for example, it can be fixedly connected, or it can be detachably connected, or it can be integrated; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium; it can be the internal communication of two elements or the interaction relationship between two elements, unless otherwise clearly limited. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0124] In the description of the present application, the description of the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In the present application, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in the present application and the features of the different embodiments or examples without contradiction.

[0125] Although the embodiments of the present application have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.

Claims

1. A distributed system data management method, characterized in that: The following steps are involved: In response to the event of a new node joining the cluster, the system scans the metadata of each node in real time, dynamically calculates the actual physical available space of each node, and deducts the space occupied by the replicas to be deleted to generate the node available capacity baseline value; The system calculates the safe allocation capacity of the entire system based on the node available capacity benchmark value. Before executing a storage operation, the system performs a physical capacity check based on the storage capacity required for the current operation. If the physical available space is detected to be lower than the required space, the system pauses the current operation and triggers a space reclamation task. The system also saves the operation site information, continuously monitors storage status changes, and automatically retries when space is released. During the data migration process, in response to the event that a data block is read, the data block is marked as pending cleanup in the metadata management module. When the data block is stored on the target node, the data integrity is confirmed through a quick metadata comparison. After confirming that the target node has completely saved the data, the atomic delete operation is triggered. The cleanup priority is calculated based on the access frequency of the data block, and data blocks with low access frequency are cleaned first to speed up the release of storage space. When the cleanup operation is performed, an audit log of the cleanup operation is recorded.

2. The distributed system data management method according to claim 1, characterized in that: Through a capacity calculation model based on the bucket effect, the system's overall securely allocated storage capacity is calculated with the lowest node capacity in the system as the benchmark.

3. The distributed system data management method according to claim 1, characterized in that: Before performing a LUN creation or expansion operation, the storage capacity required for this operation is calculated and compared with the storage capacity that can be safely allocated by the entire system. If the storage capacity that can be safely allocated by the entire system is less than the storage capacity required for this operation, the current operation is suspended and a space reclamation task is triggered.

4. The distributed system data management method according to claim 1, characterized in that: During the data migration process, in response to an event that a data block is read by a target node, the data block is updated from a pending cleanup state to a read state in the metadata management module to optimize the execution efficiency of subsequent cleanup processes.

5. The distributed system data management method according to claim 1, characterized in that: Before triggering the atomic delete operation, the storage status of the data block marked as read on the target node is checked first. If the check passes, the cleanup operation is performed immediately. If the verification fails, the data retransmission process will be automatically started to ensure data consistency.

6. A distributed system data management device, characterized in that: include: The data collection module is used to respond to the event of a new node joining the cluster, scan the metadata of each node in real time, dynamically calculate the actual physical available space of each node, and deduct the space occupied by the migration replicas to be deleted to generate the node available capacity baseline value; The system pre-check module is used to calculate the storage capacity that can be safely allocated by the entire system based on the node available capacity baseline value. Before executing a storage operation, it performs a physical capacity check based on the storage capacity required for the current operation. If it detects that the physical available space is lower than the required space, it pauses the current operation and triggers a space reclamation task. It also saves the operation site information, continuously monitors storage status changes, and automatically retries when it detects that space has been released. The data cleaning module is used to respond to the event of a data block being read during the data migration process and mark the data block as a pending cleanup state in the metadata management module. When the data block is stored on the target node, the data integrity is confirmed through a quick metadata comparison. After confirming that the target node has completely saved the data, an atomic delete operation is triggered; The data cleaning module is also used to calculate the cleaning priority according to the access frequency of the data block, giving priority to cleaning the data blocks with low access frequency to speed up the release of storage space; The data cleaning module is also used to record the audit log of the cleaning operation when performing the cleaning operation.

7. The distributed system data management device according to claim 6, characterized in that: The data acquisition module is also used to: Through a capacity calculation model based on the bucket effect, the system's overall securely allocated storage capacity is calculated with the lowest node capacity in the system as the benchmark.

8. The distributed system data management device according to claim 6, characterized in that: The system pre-check module is also used to: before performing a LUN creation or expansion operation, calculate the storage capacity required for this operation and compare it with the storage capacity that can be safely allocated by the system as a whole. If the storage capacity that can be safely allocated by the system as a whole is less than the storage capacity required for this operation, the current operation is suspended and the space recovery task is triggered.

9. An electronic device, characterized in that: include: processor; a memory storing executable instructions; When the processor executes the instruction, the distributed system data management method according to any one of claims 1 to 5 is implemented.

10. A computer-readable storage medium, characterized in that A computer program is stored, and when the program is executed by a processor, a distributed system data management method as claimed in any one of claims 1 to 5 is implemented.