Metadata synchronization method, storage cluster system, electronic equipment and storage medium

CN120804044APending Publication Date: 2025-10-17JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510911980.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-17

Smart Images

  • Figure CN120804044A_ABST
    Figure CN120804044A_ABST
Patent Text Reader

Abstract

The invention discloses a metadata synchronization method, a storage cluster system, electronic equipment and a storage medium, and relates to the field of computer storage systems.The method comprises the steps that a metadata synchronization request is sent to each storage cluster in a first storage cluster set through a main storage cluster in the storage cluster system; a first data snapshot of a target storage cluster is generated at the moment when a metadata synchronization request is obtained through the target storage cluster, metadata change content of the target storage cluster is determined according to the first data snapshot and a second data snapshot, and the metadata change content is sent to a main storage cluster; and updating the metadata in the main storage cluster according to the obtained metadata change content through the main storage cluster, and synchronizing the updated metadata to each storage cluster in the first storage cluster set. By the adoption of the scheme, the problem that metadata synchronization cannot be well carried out between storage clusters is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computer storage systems, in particular to a metadata synchronization method, a storage cluster system, an electronic device and a storage medium. BACKGROUND

[0002] With the rapid development of digital economy in recent years, and the digital transformation of industries such as finance, scientific research and medical treatment, user business production systems have generated a large amount of high-value data. The data will change from hot to cold in access demand and frequency over time. From the perspective of reducing operation and maintenance costs, these high-value cold data need to be migrated to low-cost and more durable cold storage media, and data centers in regions with lower operation and maintenance costs. Another application scenario is that large enterprises deploy storage clusters in multiple regional parks or data centers, and need to flow and share data distributed in different regions according to business needs. In summary, the demand for data distributed in multiple regions and efficient management is increasingly urgent.

[0003] Currently, data migration between two sets of storage clusters is usually performed by separately deploying an archiving software system. After the data is migrated from the storage cluster deployed in region A to the storage cluster in region B, if the business system initiates an access to the data in the storage cluster in region A, a message that the data has been migrated will be returned, and the user needs to manually migrate the required data from the storage cluster in region B back to the storage cluster in region A. The problem with this solution is that the storage system where the data is migrated is isolated from the original storage system, the migration and back migration of data require manual operation, and there is a lack of global data view. The data migrated to region B cannot be accessed on the management platform of the storage cluster in region A, and the management and access efficiency of the data is low.

[0004] To solve this problem, the metadata of the storage clusters in multiple regions needs to be managed in a unified file system. The global file system (GFS) is used to manage and maintain the full metadata information of the multiple storage clusters deployed in different regions. This involves the need for metadata synchronization between storage clusters in different sites. Since the storage clusters in each site can support read and write requests of the business system at the same time, the metadata information in each site also changes in real time as the data is written. Therefore, to realize the global file system function, metadata synchronization between storage clusters in each site is a key requirement.

[0005] In view of the problem in the related art that the metadata between storage clusters cannot be well synchronized, no effective solution has been proposed so far. SUMMARY

[0006] The application provides a metadata synchronization method, a storage cluster system, an electronic device and a storage medium to at least solve the problem that metadata cannot be well synchronized between storage clusters.

[0007] The application provides a metadata synchronization method applied to a storage cluster system, wherein the storage cluster system comprises a plurality of storage clusters, and the method comprises the following steps: sending, by a master storage cluster in the storage cluster system, a metadata synchronization request to each storage cluster in a first storage cluster set, wherein the first storage cluster set comprises storage clusters other than the master storage cluster in the plurality of storage clusters; generating, by a target storage cluster, a first data snapshot of the target storage cluster at a time point when the metadata synchronization request is acquired, determining metadata change content of the target storage cluster according to the first data snapshot and a second data snapshot, and sending the metadata change content to the master storage cluster, wherein the target storage cluster is a storage cluster in the first storage cluster set that acquires the metadata synchronization request, the first data snapshot comprises metadata information of the target storage cluster at the time point, and the second data snapshot comprises metadata information of the target storage cluster at a time point when the target storage cluster acquires the metadata synchronization request last time; and updating, by the master storage cluster, metadata in the master storage cluster according to the acquired metadata change content, and synchronously updating the updated metadata to each storage cluster in the first storage cluster set.

[0008] The application further provides a storage cluster system, wherein the storage cluster system comprises a plurality of storage clusters, and the storage cluster system comprises: a master storage cluster configured to send a metadata synchronization request to each storage cluster in a first storage cluster set, wherein the first storage cluster set comprises storage clusters other than the master storage cluster in the plurality of storage clusters; and a target storage cluster configured to generate a first data snapshot of the target storage cluster at a time point when the metadata synchronization request is acquired, determine metadata change content of the target storage cluster according to the first data snapshot and a second data snapshot, and send the metadata change content to the master storage cluster, wherein the target storage cluster is a storage cluster in the first storage cluster set that acquires the metadata synchronization request, the first data snapshot comprises metadata information of the target storage cluster at the time point, and the second data snapshot comprises metadata information of the target storage cluster at a time point when the target storage cluster acquires the metadata synchronization request last time; and the master storage cluster is further configured to update metadata in the master storage cluster according to the acquired metadata change content, and synchronously update the updated metadata to each storage cluster in the first storage cluster set.

[0009] The application further provides an electronic device comprising a memory configured to store a computer program and a processor configured to execute the computer program to implement the steps of any of the above metadata synchronization methods.

[0010] The application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program.

[0011] The application further provides a computer program product, which comprises a computer program.

[0012] In the application, the master storage cluster acts as a coordinator of metadata synchronization, periodically sends a metadata synchronization request to the first storage cluster set, triggers the clusters to generate a snapshot of the current time, and accurately identifies the metadata change content by comparing with the snapshot of the last time. This metadata incremental synchronization method based on the snapshot only transmits the changed metadata, avoids unnecessary data transmission, greatly improves the synchronization efficiency and does not affect the normal business of the cluster. After obtaining the metadata change content, the master storage cluster updates the local metadata, and then synchronizes the updated metadata to each storage cluster in the first storage cluster set in an incremental manner. By using the above scheme, the problem that the storage clusters cannot well synchronize the metadata is solved, the consistency of data is ensured, the influence of the synchronization operation on the performance of the cluster is reduced, and the continuity and efficiency of the business system are ensured. BRIEF DESCRIPTION OF DRAWINGS

[0013] In order to more clearly illustrate the embodiments of the application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0014] Figure 1 It is a hardware structure block diagram of a metadata synchronization method according to an embodiment of the application;

[0015] Figure 2 It is a flowchart of a metadata synchronization method according to an embodiment of the application;

[0016] Figure 3 It is a process diagram of metadata synchronization between multiple clusters according to an embodiment of the application;

[0017] Figure 4 It is a data read-write access flowchart according to an embodiment of the application;

[0018] Figure 5 It is a structure block diagram of a storage cluster system according to an embodiment of the application. DETAILED DESCRIPTION

[0019] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative work are within the scope of protection of this application.

[0020] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0021] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0022] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the metadata synchronization method depends, the specific application environment architecture or specific hardware architecture is described here.

[0023] The method embodiments provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 1 : is a hardware structure diagram of a metadata synchronization method according to an embodiment of the present application. Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data, wherein the above-mentioned server device may also include a transmission device 106 for communication functions and an input and output device 108. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above server device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0024] The memory 104 can be used to store computer programs, such as software programs of application software and modules, such as a computer program corresponding to the starting method of the operating system in the embodiments of the present application. The processor 102 performs various functional applications and data processing, i.e., implements the above method, by running the computer programs stored in the memory 104. The memory 104 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include a memory remotely arranged with respect to the processor 102, which can be connected to a server device through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0025] The transmission device 106 is used to receive or send data via a network. The specific examples of the above network can include a wireless network provided by a communication provider of a server device. In one example, the transmission device 106 includes a network adapter (NIC), which can be connected to other network devices through a base station so as to be able to communicate with the Internet.

[0026] For better understanding, in the related art, the global file system is used to realize unified management of full metadata of multiple storage clusters, realize a global unified view of data, and data of all other clusters can be accessed on any storage cluster. The metadata information includes file size, timestamp, storage path, owner, permission, and other information, and the storage path information therein can be used to determine the cluster where the file is located. When it is necessary to access the file across clusters, the information of the required file can be obtained through the global metadata information, and the file of the remote cluster is pulled to the local in the background, and the business layer is not aware of the data flow between clusters, i.e., a global view and unified management of data across regions and across clusters are realized.

[0027] According to the number of deployed storage clusters, there are some differences in the mechanism of metadata synchronization between multiple storage clusters. If two sets of clusters are deployed in A and B respectively, and are managed uniformly through the global file system, when the storage cluster has data writing or modification, the changed part is sent to the other site in real time through incremental synchronization, i.e., the metadata of the two sets of clusters in A and B are synchronized to each other in real time, so as to ensure that the clusters in A and B have consistent full metadata information.

[0028] If more than three clusters are deployed in multiple regions, one cluster needs to be selected as the master cluster and the other clusters as slave clusters, and then the metadata synchronization between the clusters is periodically performed. The specific steps are as follows: first, after entering the metadata synchronization period, the master cluster sends a request for updating metadata to each slave cluster. After receiving the request, each slave cluster suspends data landing, and writes data into the cache first, and suspends the operation request of modifying / deleting / writing of the file system in the cluster; second, each slave cluster sends the metadata to the master cluster through incremental synchronization. After receiving the metadata, the master cluster updates the metadata of the cluster. After the step is executed, the master cluster contains the full metadata information of the slave clusters in the current metadata synchronization period; third, the master cluster pushes the metadata information changed compared with the last synchronization period to each slave cluster through incremental synchronization. After receiving the metadata information, each slave cluster updates the metadata information of the cluster. At this time, each cluster has been updated to the consistent full metadata information in the current metadata synchronization period. Fourth, each slave cluster sends a response information of completing the metadata update to the master cluster, writes the data in the cache into the hard disk, and executes the suspended operation request. The current metadata synchronization period ends. Each cluster resumes the normal read / write request, and the local metadata information is updated.

[0029] However, the following problems exist in the prior art: (1) The master cluster and the slave cluster are synchronized according to the pre-set metadata synchronization period. The synchronization period can be set according to the business needs. However, in order to ensure the consistency of the metadata, the business modification / deletion / writing request is suspended, and the data landing is suspended, which affects the timeliness of the business request response. (2) The metadata of the master cluster and the slave cluster can be consistent only after one metadata synchronization period. However, as the clusters continuously process the business access request, the corresponding metadata is continuously updated. If the interval between two metadata synchronization periods is long, the metadata of the clusters will be greatly different, and the newly written data cannot be queried in the global view and cannot be accessed across the clusters. If the interval between the metadata periods of the clusters is shortened, the performance of the clusters and the timeliness of the request response will be reduced. (3) Only one master cluster is configured. When an abnormal situation such as a crash occurs, the metadata synchronization of the whole global file system will be interrupted.

[0030] To solve the above problems, a metadata synchronization method is provided in the embodiment. The method is applied to a storage cluster system, and the storage cluster system includes a plurality of storage clusters, Figure 2 The flowchart of the metadata synchronization method according to the embodiment of the application is shown in Figure 2 The method includes the following steps S202-S206:

[0031] Step S202: sending, by the master storage cluster in the storage cluster system, a metadata synchronization request to each storage cluster in a first storage cluster set, wherein the first storage cluster set comprises storage clusters other than the master storage cluster in the plurality of storage clusters;

[0032] Step S204: generating, by a target storage cluster, a first snapshot of the target storage cluster at a time when the metadata synchronization request is obtained, determining metadata change content of the target storage cluster according to the first snapshot and a second snapshot, and sending the metadata change content to the master storage cluster, wherein the target storage cluster is a storage cluster in the first storage cluster set that obtains the metadata synchronization request, the first snapshot has metadata information of the target storage cluster at the time, and the second snapshot has metadata information of the target storage cluster at a time when the target storage cluster obtains the metadata synchronization request last time;

[0033] Step S206: updating, by the master storage cluster, metadata in the master storage cluster according to the obtained metadata change content, and synchronously updating the updated metadata to each storage cluster in the first storage cluster set.

[0034] In the above steps, the master storage cluster acts as a coordinator of metadata synchronization, periodically sends a metadata synchronization request to the first storage cluster set, triggers these clusters to generate a snapshot at the current time, and accurately identifies the metadata change content by comparing with the snapshot at the last time. This snapshot-based metadata incremental synchronization method only transmits the changed metadata, avoids unnecessary data transmission, greatly improves the synchronization efficiency and does not affect the normal business of the cluster. After obtaining the metadata change content, the master storage cluster updates the local metadata, and then synchronizes the updated metadata to each storage cluster in the first storage cluster set in an incremental manner. By using the above scheme, the problem that the storage clusters cannot well synchronize the metadata is solved, the consistency of data is ensured, the influence of the synchronization operation on the performance of the cluster is reduced, and the continuity and efficiency of the business system are ensured.

[0035] In an exemplary embodiment, the above-mentioned synchronization of the updated metadata to each storage cluster in the first storage cluster set can be implemented by the following steps S11-S13:

[0036] Step S11: determining third snapshot data, wherein the third snapshot data is snapshot data generated at a completion time, the third snapshot data has metadata information of the master storage cluster at the completion time, and the completion time is a time when the master storage cluster lastly updates the metadata in the master storage cluster according to the obtained metadata change content;

[0037] It should be noted that after the master storage cluster completes the update of the metadata, the metadata information of the master storage cluster after the last completion of the metadata update is obtained, and the purpose is to provide a reference point for subsequent determination of the full metadata change content, that is, the difference between the current state of the master storage cluster and the third snapshot data can be compared to determine the metadata changes of all storage clusters since the last synchronization completion.

[0038] Step S12: determining the full metadata change content between the updated metadata and the third snapshot data;

[0039] It should be noted that after the third snapshot data is determined, the master storage cluster calculates the full metadata change content between the updated metadata and the third snapshot data. This part of the content reflects the net change of the metadata of all storage clusters since the last metadata synchronization completion, including newly added, modified and deleted metadata records. By calculating the full metadata change content, the master storage cluster can accurately identify the metadata part that needs to be synchronized to other clusters, avoiding unnecessary data transmission and improving the efficiency and accuracy of the synchronization process.

[0040] Step S13: sending the full metadata change content to each storage cluster in the first storage cluster set, wherein the multiple storage clusters have performed metadata synchronization after the completion time and before sending the full metadata change content.

[0041] It should be noted that the full metadata change content obtained by analysis is synchronized to each storage cluster in the first storage cluster set. In order to ensure the consistency of the metadata, this synchronization operation is based on all changes after the third snapshot data. By sending these changes to all slave clusters and backup clusters, each cluster can update its local metadata information, thereby maintaining the synchronization and consistency of the global metadata. In addition, since the multiple storage clusters have performed metadata synchronization after the completion time and before sending the full metadata change content, it means that the system is in a relatively stable state to synchronize the data change content, further ensuring the consistency of the data and the stability of the system.

[0042] It should be noted that through the above steps, the master storage cluster can efficiently synchronize the latest metadata information to all clusters, support real-time update of the global file system and cross-cluster data access, while reducing the occupation of network resources in the synchronization process and improving the overall performance of the system.

[0043] In one exemplary embodiment, after the primary storage cluster's metadata is updated, it prioritizes metadata synchronization with the backup storage cluster. The primary storage cluster's full metadata changes are sent to the backup storage cluster, which then completes the update. After this step is complete, both the primary and backup storage clusters maintain the latest full metadata at the time of the synchronization request.

[0044] After metadata synchronization between the primary and backup storage clusters is complete, the primary storage cluster immediately pushes all metadata changes to each secondary storage cluster. Upon receiving the changes, the secondary storage cluster completes the update. At this point, the primary, backup, and secondary storage clusters have all completed the update. Each storage cluster's metadata information includes the metadata of the other storage clusters at the time of synchronization, as well as the latest information for its own storage cluster.

[0045] In an exemplary embodiment, the method further includes: in the event of a failure of the primary storage cluster, a backup storage cluster in the storage cluster system takes over the work of the primary storage cluster and becomes the new primary storage cluster in the storage cluster system; and determining a new backup storage cluster from a second storage cluster set through the new primary storage cluster, wherein the second storage cluster set includes storage clusters in the first storage cluster set except the new primary storage cluster.

[0046] It should be noted that in the present application, a backup storage cluster is also set up in the storage cluster system to cope with the failure of the main storage cluster and ensure the high availability and data consistency of the global file system. When a failure of the main storage cluster is detected, the backup storage cluster in the system will automatically take over the role and task of the main storage cluster and quickly switch to become the new main storage cluster. The system needs to re-establish a redundant protection mechanism. At this time, the new main storage cluster will elect a new backup storage cluster from the second storage cluster set. The second storage cluster set is all the clusters remaining in the first storage cluster set after excluding the cluster that has been promoted to the main storage cluster. The new backup storage cluster can be elected based on certain strategies, such as the health status, load conditions, geographical distribution, etc. of the cluster, to ensure that the newly selected backup storage cluster can take over the task of taking over when the main storage cluster fails again, providing system-level redundant protection.

[0047] It should be noted that through the automatic switching of the backup storage cluster and the election mechanism of the new backup storage cluster, the disaster recovery capability of the global file system and the continuity of data access are significantly improved. When the primary storage cluster fails, the backup storage cluster can quickly take over the responsibilities of the primary storage cluster, avoiding the interruption of metadata synchronization, thereby ensuring the consistency of data and the uninterrupted service of the global file system. At the same time, the election of the new backup storage cluster increases the redundancy of the system, further enhancing the stability and reliability of the system. Even in the case of multiple failures, through the role conversion between clusters, the continuous metadata synchronization and service access can be ensured.

[0048] In an exemplary embodiment, the method further comprises steps S21-S23:

[0049] Step S21: sending, by each storage cluster in the plurality of storage clusters, cluster information of the storage cluster to the remaining storage clusters in the plurality of storage clusters, wherein the cluster information comprises performance indicators, health status, load condition, redundancy, geographic information, and resource usage of the storage cluster; the performance indicators comprise processor utilization, memory utilization, and network bandwidth utilization;

[0050] It should be noted that in a multi-storage cluster environment, each storage cluster actively broadcasts its cluster information to other clusters in the system, which covers the performance indicators (such as processor utilization, memory utilization, and network bandwidth utilization), health status, load condition, redundancy, geographic information, and resource usage of the cluster. Through this information sharing mechanism, all clusters can obtain the latest status of each member in the system, providing basic data for subsequent election decisions.

[0051] Step S22: selecting, by each storage cluster in the plurality of storage clusters, a primary storage cluster and a backup storage cluster from the plurality of storage clusters according to a preset election rule and the obtained cluster information of the storage cluster, and sending the identity of the selected primary storage cluster and the identity of the backup storage cluster to the remaining storage clusters in the plurality of storage clusters;

[0052] It should be noted that on the basis of cluster information sharing, each storage cluster participates in the process of selecting a primary storage cluster and a backup storage cluster according to a set of preset election rules, taking into account factors such as performance, health status, and geographic location. The design of the election rule aims to balance the load of the cluster, maximize the system performance, and enhance the redundancy protection.

[0053] Step S23: In a case where the reference storage cluster is determined to have the most votes for being elected as the primary storage cluster, the reference storage cluster determines itself as the primary storage cluster in the storage cluster system and broadcasts first broadcast information to the rest of the storage clusters in the plurality of storage clusters, where the first broadcast information is used to announce that the reference storage cluster is the primary storage cluster in the storage cluster system; and in a case where the reference storage cluster is determined to have the most votes for being elected as the standby storage cluster, the reference storage cluster determines itself as the standby storage cluster in the storage cluster system and broadcasts second broadcast information to the rest of the storage clusters in the plurality of storage clusters, where the second broadcast information is used to announce that the reference storage cluster is the standby storage cluster in the storage cluster system.

[0054] It should be noted that when a certain storage cluster (reference storage cluster) obtains the most votes for being elected as the primary storage cluster through the election, it will confirm itself as the new primary storage cluster and broadcast first broadcast information to other clusters to officially announce its primary storage cluster identity. Conversely, if the reference storage cluster wins in the election of standby storage clusters, it will confirm itself as a standby storage cluster and broadcast second broadcast information to announce the identity of the standby storage cluster. This mechanism ensures that each cluster can update its internal role mapping in a timely manner, whether in initial deployment or in dynamic role conversion, to quickly adapt to the new cluster architecture and maintain the high availability and consistency of the system.

[0055] It should be noted that the above steps are based on real-time cluster information to dynamically select the most suitable clusters to serve as primary and standby storage clusters, optimizing the performance and resource utilization of the system.

[0056] In an exemplary embodiment, the method further comprises: in a case where the first storage cluster in the storage cluster system obtains a read data request, determining, by the metadata of the first storage cluster, a storage location of target data corresponding to the read data request; in a case where the storage location of the target data is determined to be a second storage cluster in the storage cluster system, obtaining the target data from the second storage cluster through the global file system, and responding to the read data request according to the obtained target data.

[0057] It should be noted that in the present application, the storage cluster system provides a transparent data access layer for users through the global file system, so that even if the data is actually stored in different physical clusters, the user can read the operation as accessing local data. When the first storage cluster receives a read data request, it first determines the actual storage location of the corresponding data (target data) of the request by querying the local metadata information. If the target data is not stored in the first storage cluster receiving the request, but is located in the second storage cluster, the first storage cluster will not directly tell the user this information, but automatically initiate a data acquisition request to the second storage cluster through the global file system, seamlessly pulling the target data from the second storage cluster. This process is completely transparent to the user, and the user does not need to care about the actual location of the data, only needs to initiate a read request, and the system will automatically identify and schedule the data, and finally return the target data to the user, completing the response to the read data request.

[0058] It should be noted that in the present embodiment, through the intelligent scheduling of the global file system, the user can access the cross-cluster data without awareness, greatly improving the convenience of data access and user experience. In addition, it also avoids the manual data migration operation that the user may perform due to unknown data storage location, reduces the time delay of data access, and improves the efficiency of data reading.

[0059] In an exemplary embodiment, the above step S102 comprises: sending, by the master storage cluster, a metadata synchronization request to each storage cluster in the first storage cluster set at every preset time interval; and / or sending a metadata synchronization request to each storage cluster in the first storage cluster set in the case of obtaining prompt information of any one of the first storage cluster set, wherein the prompt information is used to prompt that the metadata update frequency of the storage cluster is greater than the preset frequency.

[0060] It should be noted that the master storage cluster will actively send a metadata synchronization request to each storage cluster in the first storage cluster set according to the preset time interval. This periodic synchronization mechanism can ensure that even in the absence of sudden data changes, the metadata information between the storage clusters can be periodically updated, thereby maintaining a consistent view of the global file system and supporting cross-cluster data access and management.

[0061] In addition, in addition to periodic synchronization, the master storage cluster can also respond to synchronization requests triggered by specific events. When the system detects that the metadata update frequency of any storage cluster in the first storage cluster set exceeds the preset frequency, that is, the cluster frequently performs data write or modification operations in a short period of time, the master storage cluster immediately sends a metadata synchronization request to all clusters. This mechanism is particularly suitable for data hotspot scenarios. When the metadata of a cluster changes frequently, the latest state of the cluster can be synchronized in time to avoid data access delay or error caused by asynchronous metadata, thereby enhancing the real-time response capability and data consistency of the global file system.

[0062] It should be noted that in the present embodiment, in combination with periodic synchronization and event-driven synchronization, the present application can dynamically adjust the synchronization strategy according to the system state and data changes, ensuring periodic updating of metadata and responding to high-frequency changes in time, thereby improving the intelligence and adaptability of the synchronization mechanism.

[0063] In an exemplary embodiment, the above step S102 can also be implemented in the following manner: predicting, by the target large model, future metadata change of the storage cluster in the storage cluster system according to historical metadata change of the storage cluster in the storage cluster system; determining the sending frequency of the metadata synchronization request according to the future metadata change of the storage cluster in the storage cluster system; and sending, by the master storage cluster, the metadata synchronization request to each storage cluster in the first storage cluster set according to the sending frequency.

[0064] It should be noted that the present application introduces a target large model to predict the future metadata change of the storage cluster in the storage cluster system, and dynamically adjusts the sending frequency of the metadata synchronization request based thereon. By analyzing the historical metadata change data of the storage cluster, the target large model can learn and predict the future data change trend, including the update frequency, change type and scale of the metadata. Based on these predictions, the system can intelligently determine the ideal synchronization frequency of each storage cluster, neither too frequent to affect performance, nor too long to cause data inconsistency.

[0065] For example, if the model predicts that the metadata of a cluster will be updated frequently in the future, the system will automatically increase the synchronization frequency of the cluster to ensure the real-time performance of the metadata; on the contrary, if the update activity is less, the synchronization frequency will be reduced accordingly to avoid unnecessary resource consumption.

[0066] It should be noted that in the present embodiment, by using the target large model to predict the future metadata change of the storage cluster, and intelligently adjusting the sending frequency of the metadata synchronization request, the resource utilization efficiency and data consistency level of the global file system in the multi-cluster environment are significantly improved. It can accurately match the data change dynamics, avoid resource waste caused by excessive synchronization, and at the same time ensure timely updating of metadata during data active period, thereby optimizing the real-time performance and user experience of cross-cluster data access.

[0067] In an exemplary embodiment, the method further comprises metadata synchronization rate adjustment based on network conditions, specifically comprising: monitoring the network conditions between the main storage cluster and each storage cluster in real time, including but not limited to network delay, bandwidth occupancy and data transmission rate; according to the network conditions, dynamically adjusting the data transmission rate and synchronization strategy when sending metadata synchronization requests to each storage cluster, the synchronization strategy including: using batch small data packet synchronization in high delay network environment; in the case of low network bandwidth, preferentially synchronizing key metadata changes to ensure the consistency of core data while reducing the occupation of network resources during synchronization.

[0068] In the present embodiment, through real-time monitoring of network conditions and strategy adjustment, the scheme effectively optimizes the network resource utilization in the metadata synchronization process, reduces the network burden, and at the same time guarantees the consistency of metadata, improves the overall performance and stability of the system.

[0069] In an exemplary embodiment, the method further comprises: the global file system monitors the access mode of the application layer to data, predicts possible high access areas and times; before the predicted high access time, the main storage cluster actively sends a request to the related storage cluster to pre-load metadata, and synchronizes the metadata to be accessed in advance to the local to reduce the delay of real-time data request; the main storage cluster dynamically adjusts the pre-loading strategy according to the real-time feedback of the application layer, to ensure that the pre-loaded metadata is consistent with the application demand, and to optimize the data access efficiency.

[0070] In the present embodiment, through application layer prediction and metadata pre-loading, the scheme effectively shortens the data access time, especially during peak period, significantly improves the response speed of the global file system, enhances the user experience, and at the same time reduces the frequency of real-time synchronization, reduces the overall resource consumption and network pressure of the system.

[0071] Obviously, the above-described embodiments are only a part of the embodiments of the present application, not all the embodiments. In order to better understand the above method, the above process is described in combination with the embodiments, but not used to limit the technical solutions of the embodiments of the present application, specifically:

[0072] The multiple cluster roles are divided into three types, master cluster, backup cluster and slave cluster. When the system is deployed, one cluster is selected as the master cluster, another cluster is selected as the backup cluster, and the remaining clusters are selected as slave clusters. The roles of the master and backup are set from the perspective of the overall system redundancy and safety of the global file system. The main task of the master cluster is to receive the metadata updates submitted by the backup cluster and each slave cluster, synthesize the full metadata information, and push the updated data to other backup / slave clusters. When the master cluster fails to access, the backup cluster will switch to the master cluster and take over the master cluster business, and then elect one from the slave clusters to switch to the backup cluster. The backup cluster periodically receives the metadata updates pushed by the master cluster and updates them. The main task of the slave cluster is to respond to the metadata synchronization request issued by the master cluster in a timely manner, submit its metadata update information to the master cluster, and receive the metadata pushed by the master cluster to complete local update.

[0073] After the roles of each cluster are set, the global file system is deployed, and a management relationship is established between the clusters. The master cluster acts as the master management site, which can manage and access all data in the managed clusters. The master cluster actively triggers metadata update requests and sends them to each cluster periodically.

[0074] The remaining clusters submit metadata updates to the master cluster: in the normal working state, each cluster can handle read and write requests of its own business system, so its metadata information will change in real time with data writing, and each cluster will have differences. After receiving the metadata synchronization request sent by the master cluster, the backup cluster and each slave cluster immediately generate a snapshot at this moment, compare it with the snapshot at the last moment, and submit the changed metadata information to the master cluster. After receiving the metadata updates submitted by the remaining clusters, the master cluster synthesizes them into a latest full metadata information. At this time, the metadata of the master cluster includes the full metadata of all clusters managed under the global file system.

[0075] Master-backup synchronization: after the metadata update of the master cluster is completed, the metadata synchronization with the backup cluster is performed first, and the update content of the full metadata of the master cluster is sent to the backup cluster. After receiving, the backup cluster completes the update of the cluster. After this step, the metadata of the master cluster and the backup cluster are both the latest full metadata at the time of the synchronization request execution.

[0076] Master-to-slave synchronization: after the metadata synchronization of the master and backup clusters is completed, the master cluster immediately pushes the update content of the full metadata to each slave cluster, and the slave cluster completes the update of the cluster after receiving. At this time, the master cluster, the backup cluster and the slave cluster have all completed the update, and the metadata information of each cluster includes the metadata information of the remaining clusters at the synchronization time and the latest information of the cluster.

[0077] The processing mechanism in abnormal situation: when the main cluster fails to access, the backup cluster will switch to the main cluster role and take over the task of the main cluster, start to be responsible for the scheduling of metadata update request, and the synthesis and update pushing of full metadata, maintain the interaction with other clusters, and support the normal operation of global file system. At the same time, one of the clusters is elected as a backup cluster.

[0078] It should be noted that, Figure 3 The metadata synchronization process between multiple clusters is shown, wherein ① is that the main cluster periodically sends metadata synchronization request to each cluster; ② is that the backup cluster and the slave cluster respond to the synchronization request, generate snapshot at this moment, judge the changed content compared with the last moment through the snapshot, and submit the metadata update content to the main cluster. The main cluster synthesizes the received metadata update content into full metadata and writes it into the metadata database; ③ is that the main cluster synchronizes the metadata update to the backup cluster; ④ is that the main cluster pushes the update content of the full metadata to each slave cluster, and the slave cluster receives and updates the local information.

[0079] That is, the metadata between each cluster managed by the global file system is periodically synchronized, and the synchronization request is initiated by the main cluster and sent to the backup cluster and the slave cluster. The backup cluster and the slave cluster respond immediately after receiving the synchronization request, generate the snapshot of the file system at this moment, and compare and analyze the changed metadata information through the snapshot, and submit the changed metadata content to the main cluster. After receiving the metadata update submitted by the rest of the clusters, the main cluster synthesizes a full metadata and updates the local metadata database. After the main cluster completes the update, it synchronizes to the backup cluster preferentially, synchronizes the update content of the full metadata to the backup cluster immediately, and the backup cluster receives the update data and completes the local metadata update. After the metadata synchronization of the main and backup clusters is completed, the main cluster pushes the update content of the metadata to each slave cluster, and the slave cluster receives the update data and completes the local metadata update. At this point, each cluster completes this metadata synchronization, and the metadata of each cluster includes the full metadata of the information of other clusters, which can support the global unified view of data and cross-cluster access. It should be noted that in the metadata synchronization process, since the snapshot method is used to obtain the metadata change information, it does not affect the access request of the business, and the data read and write of the cluster can be normally performed.

[0080] It should be noted that, Figure 4A data read-write access flow diagram is shown, wherein ① is that a business system initiates a write request to write data into a local cluster file system; ② is that after the writing is completed, the metadata information of the local cluster is synchronously updated; ③ is that the business system initiates a read request to a backup cluster; ④ is that the backup cluster receives the read request, queries the local metadata to find that the required file is in a remote cluster; ⑤ is that the backup cluster initiates a cross-cluster access request to migrate the required file back to the local backup cluster; and ⑥ is that the read file answers the read request, the read operation is completed, and the business system is not aware of the data flow between clusters.

[0081] Specifically, the business system initiates a write request, and after the data is written into the local cluster, the metadata information is synchronously updated; the business system initiates a read request to the backup cluster, and after the backup cluster receives the read request, it is queried in the local cluster whether the required file exists; since the metadata of the local cluster has been synchronized, it is full metadata containing information of other remote clusters, and the required file in the remote slave cluster can be queried; a file migration request is initiated to the slave cluster to migrate the required file back to the local; the file is read and the business read request is answered, and the read request is executed. The business system is not aware of the data migration between clusters. That is, the global file system supports global file view and cross-cluster data flow.

[0082] It should be noted that, by means of the master-slave role management and master-backup redundancy protection mechanism between multiple clusters, and the metadata incremental synchronization method based on snapshots, the periodic and timely synchronization of metadata between multiple clusters is realized, the global metadata management and unified view of the global file system are realized, and the automatic scheduling of cross-cluster data is realized. And in the metadata synchronization process, the normal operation of the cluster business is not affected, realizing an efficient cross-region and cross-cluster data management mechanism.

[0083] Through the description of the above implementation manner, those skilled in the art can clearly understand that the method according to the above embodiment can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better implementation manner.

[0084] The embodiment of the application also provides a storage cluster system, Figure 5 is a structural block diagram of a storage cluster system according to an embodiment of the application, as Figure 5 shown, the system comprises:

[0085] The main storage cluster 502 is configured to send a metadata synchronization request to each storage cluster in the first storage cluster set, wherein the first storage cluster set comprises storage clusters other than the main storage cluster in the plurality of storage clusters;

[0086] The target storage cluster 504 is configured to generate a first snapshot of the target storage cluster at the time when the metadata synchronization request is obtained, to determine metadata change content of the target storage cluster according to the first snapshot and a second snapshot, and to send the metadata change content to the master storage cluster, where the target storage cluster is a storage cluster in the first storage cluster set that obtains the metadata synchronization request, the first snapshot has metadata information of the target storage cluster at the time, and the second snapshot has metadata information of the target storage cluster at the time when the last metadata synchronization request is obtained.

[0087] The master storage cluster 502 is further configured to update the metadata in the master storage cluster according to the obtained metadata change content, and to synchronize the updated metadata to each storage cluster in the first storage cluster set.

[0088] In the above system, the master storage cluster acts as a coordinator of metadata synchronization, periodically sends a metadata synchronization request to the first storage cluster set, triggers the clusters to generate a snapshot at the current time, and accurately identifies the metadata change content by comparing the snapshot at the current time with a snapshot at the last time. This metadata incremental synchronization method based on snapshots only transmits the changed metadata, avoids unnecessary data transmission, greatly improves the synchronization efficiency, and does not affect the normal business of the cluster. After obtaining the metadata change content, the master storage cluster updates the local metadata, and then synchronizes the updated metadata to each storage cluster in the first storage cluster set in an incremental manner. The above scheme solves the problem that the storage clusters cannot well synchronize the metadata, guarantees the data consistency, reduces the impact of the synchronization operation on the performance of the cluster, and ensures the continuity and efficiency of the business system.

[0089] In an exemplary embodiment, the master storage cluster 502 is further configured to determine third snapshot data, where the third snapshot data is snapshot data generated at a completion time, the third snapshot data has metadata information of the master storage cluster at the completion time, and the completion time is a time when the master storage cluster last completes the update of the metadata in the master storage cluster according to the obtained metadata change content; to determine full metadata change content between the updated metadata and the third snapshot data; and to send the full metadata change content to each storage cluster in the first storage cluster set, where the multiple storage clusters have performed metadata synchronization after the completion time and before the full metadata change content is sent.

[0090] In an example embodiment, the target storage cluster includes a backup storage cluster in the storage cluster system, for taking over the work of the primary storage cluster in case of failure of the primary storage cluster, and becoming a new primary storage cluster in the storage cluster system, and determining a new backup storage cluster from a second storage cluster set, wherein the second storage cluster set includes the storage clusters in the first storage cluster set except the new primary storage cluster.

[0091] In an example embodiment, the storage cluster system includes: any storage cluster, the storage cluster being any one of the storage clusters in the storage cluster system, for sending cluster information of the storage cluster to the rest of the storage clusters in the plurality of storage clusters, wherein the cluster information includes: performance indicators, health status, load status, redundancy, geographic information, resource usage of the storage cluster; the performance indicators include: processor utilization, memory utilization, network bandwidth utilization; electing a primary storage cluster and a backup storage cluster from the plurality of storage clusters according to a preset election rule and the obtained cluster information of the storage cluster, and sending an identification of the elected primary storage cluster and an identification of the backup storage cluster to the rest of the storage clusters in the plurality of storage clusters; a reference storage cluster, for determining that the reference storage cluster itself is the primary storage cluster in the storage cluster system in case of determining that the reference storage cluster is elected as the primary storage cluster by the most votes, and broadcasting first broadcast information to the rest of the storage clusters in the plurality of storage clusters, wherein the first broadcast information is used to announce that the reference storage cluster is the primary storage cluster in the storage cluster system; determining that the reference storage cluster itself is the backup storage cluster in the storage cluster system in case of determining that the reference storage cluster is elected as the backup storage cluster by the most votes, and broadcasting second broadcast information to the rest of the storage clusters in the plurality of storage clusters, wherein the second broadcast information is used to announce that the reference storage cluster is the backup storage cluster in the storage cluster system.

[0092] In an example embodiment, the storage cluster system includes a first storage cluster, for determining a storage location of target data corresponding to a read data request through metadata of the first storage cluster in case of obtaining the read data request; obtaining the target data from a second storage cluster in the storage cluster system through a global file system in case of determining that the storage location of the target data is the second storage cluster, and responding to the read data request according to the obtained target data.

[0093] In an example embodiment, the primary storage cluster 502 is configured to send a metadata synchronization request to each storage cluster in the first storage cluster set every preset time interval, and / or send a metadata synchronization request to each storage cluster in the first storage cluster set in case of obtaining prompt information of any one of the storage clusters in the first storage cluster set, wherein the prompt information is used to prompt that the metadata update frequency of the storage cluster is greater than a preset frequency.

[0094] In an example embodiment, the master storage cluster 502 is configured to predict, by the target large model, future metadata change of the storage cluster in the storage cluster system according to historical metadata change of the storage cluster in the storage cluster system; determine the sending frequency of the metadata synchronization request according to the future metadata change of the storage cluster in the storage cluster system; and send, by the master storage cluster, the metadata synchronization request to each storage cluster in the first storage cluster set according to the sending frequency.

[0095] The features of the embodiments corresponding to the storage cluster system can be referred to the related descriptions of the embodiments of the metadata synchronization method, which will not be repeated here.

[0096] The embodiments of the present application also provide an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above metadata synchronization method embodiments.

[0097] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to perform the steps in any of the above metadata synchronization method embodiments when running.

[0098] In an example embodiment, the above computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0099] The embodiments of the present application also provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to perform the steps in any of the above metadata synchronization method embodiments.

[0100] The embodiments of the present application also provide another computer program product, which comprises a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to perform the steps in any of the above metadata synchronization method embodiments.

[0101] Those skilled in the art will further realize that the mere concepts, teachings, and embodiments described herein are merely meant to provide an enabling description of embodiments of the present application and that various modifications can be made thereto without departing from the scope of the present application. Accordingly, the appended claims are intended to embrace all such alterations, modifications, and variations of the embodiments described herein that are within the scope of this application, including all the preferred embodiments.

[0102] The above provides a kind of metadata synchronization method and device, electronic equipment, storage medium, computer program product provided in the present application in detail.The principle and implementation of the present application are described in this paper by applying specific examples, the above example is only used to help understand the method and its core idea of the present application.It should be pointed out that, for the ordinary skilled in the art, without departing from the principle of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of the claims of the present application.

Claims

1. A metadata synchronization method, characterized in that: Applied to a storage cluster system, the storage cluster system includes multiple storage clusters, including: Sending a metadata synchronization request to each storage cluster in a first storage cluster set through a primary storage cluster in the storage cluster system, wherein the first storage cluster set includes storage clusters in the multiple storage clusters except the primary storage cluster; Generate a first data snapshot of the target storage cluster at the moment the metadata synchronization request is obtained by the target storage cluster, determine metadata changes of the target storage cluster based on the first data snapshot and the second data snapshot, and send the metadata changes to the primary storage cluster, wherein the target storage cluster is the storage cluster in the first storage cluster set that obtains the metadata synchronization request, the first data snapshot has metadata information of the target storage cluster at the moment, and the second data snapshot has metadata information of the target storage cluster at the moment when the metadata synchronization request was last obtained; The metadata in the primary storage cluster is updated according to the acquired metadata change content through the primary storage cluster, and the updated metadata is synchronized to each storage cluster in the first storage cluster set.

2. The metadata synchronization method according to claim 1, characterized in that: Synchronizing the updated metadata to each storage cluster in the first storage cluster set includes: Determining third snapshot data, wherein the third snapshot data is snapshot data generated at a completion time, the third snapshot data includes metadata information of the primary storage cluster at the completion time, and the completion time is the time when the primary storage cluster last completed updating the metadata in the primary storage cluster based on the acquired metadata change content; Determining full metadata changes between the updated metadata and the third snapshot data; The full metadata changes are sent to each storage cluster in the first set of storage clusters, wherein metadata synchronization has been performed between the plurality of storage clusters after the completion time and before the full metadata changes are sent.

3. The metadata synchronization method according to claim 1, characterized in that: The method further comprises: In the event of a failure of the primary storage cluster, the backup storage cluster in the storage cluster system takes over the work of the primary storage cluster and becomes the new primary storage cluster in the storage cluster system; A new backup storage cluster is determined from a second storage cluster set using the new primary storage cluster, wherein the second storage cluster set includes storage clusters in the first storage cluster set except the new primary storage cluster.

4. The metadata synchronization method according to claim 1, wherein: The method further comprises: Sending cluster information of each storage cluster to the remaining storage clusters in the multiple storage clusters through each storage cluster in the multiple storage clusters, wherein the cluster information includes: performance indicators, health status, load status, redundancy, geographical information, and resource usage of the storage cluster; the performance indicators include: processor utilization, memory utilization, and network bandwidth utilization; Electing a primary storage cluster and a backup storage cluster from each of the multiple storage clusters according to a preset election rule and the acquired cluster information of the storage clusters, and sending the identifiers of the elected primary storage cluster and the backup storage cluster to the remaining storage clusters in the multiple storage clusters; When it is determined that the reference storage cluster has the most votes to be elected as the primary storage cluster, the reference storage cluster is determined to be the primary storage cluster in the storage cluster system, and first broadcast information is broadcast to the remaining storage clusters in the multiple storage clusters, wherein the first broadcast information is used to notify that the reference storage cluster is the primary storage cluster in the storage cluster system; and Through the reference storage cluster among the multiple storage clusters, when it is determined that the reference storage cluster has the largest number of votes to be elected as the backup storage cluster, the reference storage cluster itself is determined to be the backup storage cluster in the storage cluster system, and a second broadcast information is broadcast to the remaining storage clusters among the multiple storage clusters, wherein the second broadcast information is used to notify that the reference storage cluster is the backup storage cluster in the storage cluster system.

5. The metadata synchronization method according to claim 1, characterized in that: The method further comprises: When a first storage cluster in the storage cluster system obtains a read data request, determining a storage location of target data corresponding to the read data request through metadata of the first storage cluster; When it is determined that the storage location of the target data is the second storage cluster in the storage cluster system, the target data is obtained from the second storage cluster through the global file system, and the read data request is responded to according to the obtained target data.

6. The metadata synchronization method according to claim 1, characterized in that: Sending a metadata synchronization request to each storage cluster in the first storage cluster set through the primary storage cluster in the storage cluster system includes: Sending, via the primary storage cluster, a metadata synchronization request to each storage cluster in the first storage cluster set at a preset time interval; and / or When prompt information of any storage cluster in the first storage cluster set is obtained, a metadata synchronization request is sent to each storage cluster in the first storage cluster set, wherein the prompt information is used to prompt that the metadata update frequency of the storage cluster is greater than a preset frequency.

7. The metadata synchronization method according to claim 1, characterized in that: Sending a metadata synchronization request to each storage cluster in the first storage cluster set through the primary storage cluster in the storage cluster system includes: Predicting future metadata changes of the storage clusters in the storage cluster system based on historical metadata changes of the storage clusters in the storage cluster system using the target large model; Determining a frequency of sending metadata synchronization requests according to future metadata changes of the storage cluster in the storage cluster system; A metadata synchronization request is sent to each storage cluster in the first storage cluster set through the primary storage cluster according to the sending frequency.

8. A storage cluster system, characterized in that: The storage cluster system includes multiple storage clusters, including: A primary storage cluster is configured to send a metadata synchronization request to each storage cluster in a first set of storage clusters, wherein the first set of storage clusters includes storage clusters in the plurality of storage clusters except the primary storage cluster; a target storage cluster, configured to generate a first data snapshot of the target storage cluster at the moment of obtaining the metadata synchronization request, determine metadata changes of the target storage cluster based on the first data snapshot and the second data snapshot, and send the metadata changes to the primary storage cluster, wherein the target storage cluster is the storage cluster in the first storage cluster set that obtains the metadata synchronization request, the first data snapshot contains metadata information of the target storage cluster at the moment, and the second data snapshot contains metadata information of the target storage cluster at the moment of the last metadata synchronization request; The primary storage cluster is further configured to update the metadata in the primary storage cluster according to the acquired metadata change content, and synchronize the updated metadata to each storage cluster in the first storage cluster set.

9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the metadata synchronization method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the metadata synchronization method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Online file system error detection method and device

    CN121743090A

  • An online file system error detection method and device

    CN121743090B