Data synchronization method, system, device and equipment and computer medium

Through the cooperation of alternate metadata nodes and message middleware, the data cached by Alluxio is synchronized in real time, solving the problem of data synchronization efficiency in Alluxio cached in HDFS, and improving data processing efficiency and the stability of HDFS cluster.

CN120295984APending Publication Date: 2025-07-11TENCENT DIGITAL TIANJIN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410040544.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The data synchronization method cached in HDFS in Alluxio has problems such as time lag and increased client response time, resulting in low data processing efficiency.

Method used

Log entries are obtained through the alternate metadata node, and the target path is pushed to the cache device using message middleware. The cache device synchronizes data from the distributed file system, eliminates data nodes that take time to process read requests that meet preset conditions, and selects the appropriate target data node for synchronization.

Benefits of technology

Real-time data synchronization of cache devices is realized, data processing efficiency and the stability of HDFS clusters are improved, and user experience is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295984A_ABST
    Figure CN120295984A_ABST
Patent Text Reader

Abstract

The invention discloses a data synchronization method, system, device and equipment and a computer medium, which can be applied to various scenes such as a cloud technology, a distributed system, a distributed storage technology and data synchronization. The method comprises; obtaining a plurality of log entries; sending a target path of the corresponding operated file to a message middleware based on the plurality of log entries, so that the message middleware is pushed to a corresponding cache device, and the cache device performs data synchronization from the distributed file system; the method for performing data synchronization from the distributed file system by the cache device comprises the following steps: acquiring a first identifier list comprising an identifier of at least one first data node from the distributed file system based on a target path; removing the second identifier contained in the blacklist from the first identifier list to obtain a residual identifier, and determining a target data node based on the residual identifier; and obtaining the synchronous data from the target data node. The data processing efficiency based on the cache device can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of distributed systems, and in particular, to a data synchronization method, system, device, equipment, and computer medium. Background Art

[0002] Hadoop Distributed File System, abbreviated as HDFS, is a distributed file system used to store and manage large-scale data sets. NameNode, the metadata node of the HDFS file system, is responsible for managing the namespace of the file system and the data block mapping information, maintaining the file system metadata and the data block replication status, etc. It records the data node information where each block in each file is located, but it does not permanently store the block location information because this information is reconstructed by the data nodes when the system starts. DataNode: The data node of the HDFS file system, responsible for storing and managing data blocks and performing read and write operations on data blocks.

[0003] The data in HDFS is finally stored on the disks of DataNode nodes. Therefore, sometimes to accelerate client access to HDFS or reduce the pressure of HDFS access, an Alluxio cache cluster is added between the client and HDFS to cache HDFS. The introduction of the Alluxio cache requires ensuring the consistency between the data in the Alluxio cache and the HDFS data. That is to say, when the data in HDFS is modified, Alluxio needs to discard the cached data and re-obtain it from HDFS. Currently, one of the cache data synchronization methods adopted by Alluxio is periodic timed synchronization, or when an access request from the client is obtained, Alluxio performs cache data synchronization. Among them, periodic timed synchronization has a time lag, and synchronizing when the client accesses will increase the response time of the client. These solutions will all lead to poor cache effects of Alluxio, and further lead to the technical problem of low efficiency in processing data based on Alluxio. Summary of the Invention

[0004] The embodiments of this application provide a data synchronization method, system, device, equipment, and computer medium, which can improve the efficiency of processing data based on Alluxio.

[0005] On the one hand, a data synchronization method is provided, which is applicable to a standby metadata node in a distributed file system, and includes:

[0006] Obtain a plurality of log entries, each log entry including: the file name of the file being operated, the operation type of the file being operated, and the path where the file being operated is located;

[0007] Send the target path where the corresponding operated file is located to the message middleware based on the multiple log entries, so that the message middleware pushes the target path to the corresponding cache device, and the cache device synchronizes data from the distributed file system based on the target path. The cache device is used to provide the data in the distributed file system to the client;

[0008] Wherein, the cache device synchronizes data from the distributed file system based on the target path, including:

[0009] The cache device obtains a first identifier list including identifiers of at least one first data node from the main metadata node in the distributed file system based on the target path;

[0010] Eliminate the second identifiers included in the blacklist from the first identifier list to obtain the remaining identifiers, wherein the second identifiers stored in the blacklist are identifiers of data nodes whose time-consuming for processing read requests meets a preset condition;

[0011] Determine the target data node based on the remaining identifiers;

[0012] Obtain the corresponding data from the target data node for synchronization.

[0013] On the other hand, a data synchronization system is provided, including a standby metadata node and a cache device in the distributed file system; the standby metadata node is used to: obtain multiple log entries, each log entry including: the file name of the operated file, the operation type of the operated file, and the path where the operated file is located; send the target path where the corresponding operated file is located to the message middleware based on the multiple log entries, so that the message middleware pushes the target path to the corresponding cache device, and the cache device synchronizes data from the distributed file system based on the target path. The cache device is used to provide the data in the distributed file system to the client;

[0014] The cache device synchronizes data from the distributed file system based on the target path, including: the cache device obtains a first identifier list including identifiers of at least one first data node from the main metadata node in the distributed file system based on the target path; eliminate the second identifiers included in the blacklist from the first identifier list to obtain the remaining identifiers, wherein the second identifiers stored in the blacklist are identifiers of data nodes whose time-consuming for processing read requests meets a preset condition; determine the target data node based on the remaining identifiers; obtain the corresponding data from the target data node for synchronization.

[0015] On the other hand, a data synchronization device is provided, which is applicable to a standby metadata node in a distributed file system and includes:

[0016] An acquisition unit, configured to acquire a plurality of log entries, each log entry including: the file name of the file to be operated, the operation type of the file to be operated, and the path where the file to be operated is located;

[0017] A sending unit, configured to send the target path where the corresponding file to be operated is located to a message middleware based on the plurality of log entries, so that the message middleware pushes the target path to the corresponding cache device, and the cache device synchronizes data from the distributed file system based on the target path, and the cache device is used to provide data in the distributed file system to a client;

[0018] Wherein, the cache device synchronizes data from the distributed file system based on the target path, including:

[0019] The cache device acquires a first identifier list including identifiers of at least one first data node from a primary metadata node in the distributed file system based on the target path;

[0020] Eliminating a second identifier included in a blacklist from the first identifier list to obtain remaining identifiers, wherein the second identifier stored in the blacklist is an identifier of a data node whose time-consuming for processing a read request meets a preset condition;

[0021] Determining target data nodes based on the remaining identifiers;

[0022] Acquiring corresponding data from the target data nodes for synchronization.

[0023] On the other hand, a computer-readable storage medium is provided, where the computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the steps in the method described in any one of the above embodiments.

[0024] On the other hand, a computer device is provided, the computer device includes a processor and a memory, the memory stores a computer program, and the processor is configured to execute the steps in the method described in any one of the above embodiments by calling the computer program stored in the memory.

[0025] The solution applicable to the standby metadata node in the distributed file system provided by the embodiments of the present application includes: obtaining a plurality of log entries, each log entry including: the file name of the file being operated, the operation type of the file being operated, and the path where the file being operated is located; based on the plurality of log entries, sending the target path where the corresponding file being operated is located to the message middleware, so that the message middleware pushes the target path to the corresponding cache device, and the cache device synchronizes data from the distributed file system based on the target path, and the cache device is used to provide data in the distributed file system to the client; wherein, the cache device synchronizes data from the distributed file system based on the target path, including: the cache device obtains a first identifier list including the identifiers of at least one first data node from the primary metadata node in the distributed file system based on the target path; removing the second identifiers included in the blacklist from the first identifier list to obtain the remaining identifiers, wherein the second identifiers stored in the blacklist are the identifiers of the data nodes whose time-consuming for processing read requests meets a preset condition; determining target data nodes based on the remaining identifiers; and obtaining corresponding data from the target data nodes for synchronization. Among them, the messages in the message middleware are subscribed by the cache device. Through the solution of the present application, the introduction of the message middleware can push the target path in real time, so that the cache device can synchronize data from the distributed file system in real time, improving the efficiency of processing data based on the cache device. Description of the Drawings

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0027] Figure 1a It is a schematic diagram of the scenario where Alluxio provided by the embodiments of the present application performs data synchronization periodically and regularly;

[0028] Figure 1b It is a schematic diagram of the scenario where Alluxio provided by the embodiments of the present application performs data synchronization when the client accesses;

[0029] Figure 2a It is a schematic flowchart of the data synchronization method provided by the embodiments of the present application;

[0030] Figure 2b It is a schematic diagram of the scenario of the data synchronization method provided by the embodiments of the present application;

[0031] Figure 2cSchematic structural diagram of the data synchronization system provided by the embodiments of the present application;

[0032] Figure 2d Scenario schematic diagram of the data synchronization method provided by the embodiments of the present application;

[0033] Figure 2e Scenario schematic diagram of a data synchronization method provided by the present application;

[0034] Figure 2f Scenario schematic diagram for determining the time consumption of a read request provided by the embodiments of the present application;

[0035] Figure 2g Scenario schematic diagram of the data synchronization process provided by the embodiments of the present application;

[0036] Figure 2h Scenario schematic diagram for determining the time consumption of a read request provided by the embodiments of the present application;

[0037] Figure 3 Schematic structural diagram of the data synchronization device provided by the embodiments of the present application;

[0038] Figure 4 Schematic structural diagram of the computer device provided by the embodiments of the present application. Detailed implementation manners

[0039] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0040] The embodiments of the present application can be applied to various scenarios such as cloud technology, distributed systems, distributed storage technology, and data synchronization.

[0041] The embodiments of the present application provide a data synchronization method, system, device, equipment and computer medium. Specifically, the data synchronization method of the embodiments of the present application is applicable to a standby metadata node in a distributed file system. Specifically, it can be executed by a computer device including the standby metadata node. Among them, the computer device can be a terminal or a server, etc. The terminal can be a smart phone, a tablet computer, a laptop computer, a smart voice interaction device, a smart home appliance, a wearable smart device, an aircraft, a smart vehicle terminal, etc. The terminal can also include a client, and the client can be a video client, a browser client or an instant messaging client, etc. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0042] First of all, some terms in the embodiments of the present application are explained below to facilitate the understanding of those skilled in the art.

[0043] HDFS, short for Hadoop Distributed File System, is a distributed file system. HDFS adopts a master / slave architecture. An HDFS cluster usually consists of an Active NameNode and several DataNodes. To avoid the single point problem of the NameNode, a standby NameNode is usually set as a backup.

[0044] NameNode: The NameNode is a central server responsible for managing the namespace of the file system and client access, such as creating, closing, renaming files or directories. It is responsible for determining the mapping of data blocks to specific storage nodes. Under its consent and scheduling, data blocks are created, deleted, and replicated.

[0045] DataNode: The DataNode is the actual storage node of HDFS, responsible for managing the storage on the node where it is located; handling read and write requests from clients. And it regularly reports heartbeats and the storage locations of blocks. From the perspective of the HDFS system, the NameNode is mainly responsible for the storage and operation of metadata, and the DataNode is responsible for actual storage. Usually, a process is deployed on one machine for each DataNode, and these machines are distributed on multiple racks.

[0046] Alluxio is a memory-centric virtual distributed storage system. It unifies the way of data access and builds a bridge for upper-layer computing frameworks and underlying storage systems.

[0047] Metadata, also known as mediation data or relay data, is data about data. It mainly describes information about data attributes and is used to support functions such as indicating storage locations, historical data, resource searches, and file records. Metadata is a kind of electronic catalog. To achieve the purpose of cataloging, it is necessary to describe and collect the content or characteristics of the data, and then assist in data retrieval.

[0048] Block: For files on HDFS, internally, a file is actually divided into one or more data blocks for storage, and these data blocks are stored on a group of DataNodes.

[0049] The data in HDFS is finally stored on the disks of DataNode nodes. Therefore, sometimes to accelerate client access to HDFS or reduce the pressure of HDFS access, an Alluxio cache cluster is added between the client and HDFS for HDFS caching. The introduction of Alluxio caching requires ensuring the consistency between the data in Alluxio cache and the data in HDFS. That is to say, when the data in HDFS is modified, Alluxio needs to discard the cached data and retrieve it again from HDFS. Currently, one of the ways Alluxio adopts for cache data synchronization is periodic timed synchronization (as Figure 1a shown), or when an access request from the client is received, Alluxio performs cache data synchronization (as Figure 1b shown). Among them, periodic timed synchronization has a time lag, and synchronizing when the client accesses will increase the response time of the client and affect the user experience. These solutions will all result in poor caching effects of Alluxio, and further lead to the technical problem of low efficiency in processing data based on Alluxio.

[0050] HDFS is a distributed file system that provides file storage capabilities for big data services. The NameNode service in HDFS is responsible for allocating and querying the DataNode nodes corresponding to data blocks. When a client writes a data block, N (the number of replicas) DataNode nodes are allocated according to a policy for writing data; when a client reads a data block, data is read from a certain DataNode in the replicas according to a certain policy. In the operations of a distributed file system, read requests are often more than write requests, so the response of read requests is more important to users. When there are many requests and high pressure in a cluster, a DataNode may respond slowly to read requests for block data. For example, the client takes more than 20 ms to receive a response. In addition, due to the aging of hardware, the processing performance of some DataNodes will also decline. Currently, when reading a multi-replica data block, a read request is always sent to the first DataNode node in the replica list. If the first DataNode node takes more than 20 ms to process the read request for some reason at this time, but this DataNode node does not belong to an abnormal failure node (abnormal failure nodes will be kicked out of the read list), this DataNode node continues to receive requests, while other replica DataNode nodes of this block data can process read requests faster, providing a better read response experience for users. To solve this problem, this method can identify slow requests for read requests of replica DataNodes, ensure that the HDFS client can select appropriate DataNode nodes for data synchronization, and improve the overall read experience of the HDFS cluster. Currently, when the caching device where the HDFS client is located sends a read request to HDFS, the caching device obtains the DataNode replica list where the block1 data block is located from the NameNode. The read request is sent to the first DataNode node in the obtained replica list. And when the caching device synchronizes data from the distributed file system, it can specifically synchronize data based on the master metadata node in the distributed file system.

[0051] If the list of block replicas returned by the master metadata node is sorted according to the network distance between the HDFS client and the DataNode, the unavailable DataNode replica nodes are ranked at the end of the list. After obtaining the list, the caching device always selects the first DataNode replica to read the data. However, the sorting of the replica list does not consider the replica response latency factor. If the request response of the caching device exceeds 20 ms, data synchronization may fail. Moreover, the read request load during data synchronization is not reasonably balanced. The replica list of a block usually has multiple DataNode nodes (the default is three replicas), and the current rule always sends read requests to the first DataNode node in the replica list, while other replica nodes rarely have the opportunity to bear the read requests for this block. In some high-pressure situations, it will cause high pressure on some DataNode nodes, and further make the HDFS cluster unstable when the caching device synchronizes data from the distributed file system.

[0052] The solution proposed in this paper enables the caching device to actively sense the response latency of each DataNode node when sending read requests and receiving read request responses, and adaptively select the appropriate DataNode node in the replica list to send the next read request through rules, thereby improving the stability of HDFS, reducing the read response latency of the HDFS cluster, and enhancing the user experience.

[0053] First, this application proposes a data synchronization method, which is implemented based on the Alluxio subscription message mode. The Standby Namnode is used to parse the Editlog to accurately find the modified files, and then the message middleware is used to notify Alluxio to update the modified file data.

[0054] Figure 2a FIG. is a schematic flow chart of a data synchronization method provided by this application. This method is applicable to the standby metadata node in the distributed file system and at least includes the following steps S201-S202:

[0055] S201. Obtain multiple log entries, each log entry including: the file name of the file being operated on, the operation type of the file being operated on, and the path where the file being operated on is located;

[0056] Optionally, the aforementioned distributed file system is the HFDS file system, which may include a master metadata node, a standby metadata node, and multiple data nodes. Among them, the master metadata node refers to the Active NameNode, the standby metadata node refers to the Standby NameNode, and the data node refers to the DataNode.

[0057] Optionally, operations on files in the HDFS file system are first sent by the HDFS client to the Active NameNode. The Active NameNode generates an Editlog and synchronizes the Editlog to the standby namenode. The Standby NameNode will synchronize the updates of the Active NameNode through the Editlog. Among them, the Editlog refers to the file operation log, which can be a write-ahead log.

[0058] Optionally, when the HDFS client modifies file metadata and the Active NameNode fails, when the Active NameNode is restarted later, it can be restored through the Editlog.

[0059] Optionally, the foregoing multiple log entries can be stored in the file operation log.

[0060] S202. Send the target path where the corresponding operated file is located to the message middleware based on the multiple log entries, so that the message middleware pushes the target path to the corresponding cache device, and the cache device synchronizes data from the distributed file system based on the target path. The cache device is used to provide data in the distributed file system to the client;

[0061] Optionally, sending the target path where the corresponding operated file is located to the message middleware based on the multiple log entries includes: determining the target path where the target operated file with the operation type being the preset operation type is located based on the multiple log entries; sending the target path to the message middleware.

[0062] In some alternative embodiments of the present application, the foregoing preset operation type includes any one or more of the following: Delete (delete file), close (close file), rename (rename file), truncate (truncate file), append (append write to file).

[0063] Among them, the foregoing multiple log entries are log entries corresponding to multiple operated files, and the target operated file is a file among the foregoing multiple operated files with the operation type being the preset operation type.

[0064] Optionally, the foregoing target path is a path in the HDFS file system.

[0065] In some alternative embodiments of the present application, reference can be made to Figure 2bAs shown, the foregoing standby metadata node may include a file operation log, and a log parsing module in the standby metadata node may obtain and parse the foregoing multiple log entries. An intelligent module in the standby metadata node may execute to determine the target path where the target file to be operated with the operation type being a preset operation type based on the foregoing multiple log entries.

[0066] In some alternative embodiments of the present application, the foregoing sending the target path to the message middleware means generating a publish message based on the target path and then sending the publish message to the message middleware, wherein the content of the publish message includes the path name of the target path.

[0067] Optionally, the cache device in the present application refers to Alluxio.

[0068] Alluxio needs to obtain the foregoing publish message through subscription. Optionally, Alluxio in the present application may refer to an Alluxio cache cluster. In some application scenarios, there may be multiple Alluxio cache clusters as caches for the HDFS file system, and each Alluxio cache cluster corresponds to an application. The Alluxio cache cluster may cache the relevant data of its corresponding application.

[0069] In some alternative embodiments of the present application, the application corresponding to the Alluxio cache cluster may access the Alluxio cache cluster through its corresponding client.

[0070] Specifically, refer to Figure 2c as shown, Figure 2c Client 1, Client 2, and Client 3 in refer to the clients that can access the Alluxio cache cluster corresponding to it. Figure 2c Alluxio1, Alluxio2, and Alluxio3 in are 3 Alluxio cache clusters. The application corresponding to Alluxio1 accesses Alluxio1 through Client 1, the application corresponding to Alluxio2 accesses Alluxio2 through Client 2, and the application corresponding to Alluxio3 accesses Alluxio3 through Client 3.

[0071] Figure 2c The 3 clients in access different directories for use by the application program. The directories corresponding to Alluxio1, Alluxio2, and Alluxio3 are / path1, / path2, and / path3 respectively, and each Alluxio cache cluster only caches the file data of the corresponding directory.

[0072] Furthermore, refer to Figure 2dAs shown, the message middleware includes multiple subscription queues. The foregoing step of sending the target path to the message middleware includes the following S2031 - S2032:

[0073] S2031. Determine the corresponding target subscription queue based on the target path;

[0074] S2032. Send the target path to the message middleware, so that the message middleware stores the target path in the target subscription queue corresponding to the target path.

[0075] Optionally, the name of the target subscription queue corresponding to the target path may be the same as the path name of the target path.

[0076] Optionally, different target paths may correspond to different target subscription queues.

[0077] Optionally, different target paths may correspond to the same target subscription queue.

[0078] Optionally, different Alluxio cache clusters may subscribe to messages in different target subscription queues.

[0079] Figure 2d In, the messages in subscription queue 1 subscribed by Alluxio1, the messages in subscription queue 2 subscribed by Alluxio2, and the messages in subscription queue 3 subscribed by Alluxio1.

[0080] Through the multi - queue subscription notification mode in this application, each Alluxio cache cluster subscribes to the directory changes of a sub - relationship. By receiving the messages with the operation file path names published by the standby Namenode, it obtains the files that have been modified. Then, the Alluxio cache cluster deletes the file data and re - fetches it from the HDFS file system. The HDFS file system in this application may refer to the HDFS cluster.

[0081] Optionally, the method further includes the following S01 - S02:

[0082] S01. Obtain the setting information of the preset operation type set by relevant personnel;

[0083] Specifically, relevant personnel can set the foregoing setting information through the device for setting the preset operation type, and the setting information may specifically include the preset operation type.

[0084] S02. Determine the preset operation type based on the setting information.

[0085] Furthermore, Figure 2eA scenario schematic diagram of a data synchronization method provided by this application. Alluxio can specifically subscribe to messages in the message middleware through a subscription module and perform data synchronization through an HDFS client.

[0086] The solution applicable to the standby metadata node in the distributed file system provided by the embodiments of this application includes: obtaining a plurality of log entries, each log entry including: the file name of the file being operated, the operation type of the file being operated, and the path where the file being operated is located; based on the plurality of log entries, sending the target path where the corresponding file being operated is located to the message middleware, so that the message middleware pushes the target path to the corresponding cache device, and the cache device performs data synchronization from the distributed file system based on the target path, and the cache device is used to provide data in the distributed file system to the client. Among them, the messages in the message middleware are subscribed by the cache device. Through the solution of this application, introducing the message middleware can push the target path in real time, so that the cache device can synchronize data from the distributed file system in real time, improving the efficiency of processing data based on the cache device. This solution uses the Standby Namnode to parse the Editlog log to accurately locate the modified file, and then notifies the Alluxio system to update the modified file data through the message middleware, so as to complete the cache synchronization of Alluxio, thereby improving the cache hit of Alluxio and accelerating the access efficiency of user service data.

[0087] Optionally, the cache device performing data synchronization from the distributed file system based on the target path includes the following S41 - S44:

[0088] S41. The cache device obtains a first identifier list including the identifiers of at least one first data node from the primary metadata node in the distributed file system based on the target path;

[0089] Optionally, the cache device obtaining a first identifier list including the identifiers of at least one first data node from the primary metadata node in the distributed file system based on the target path includes:

[0090] The cache device determines the storage location of the data in the target path based on the target path and the preset correspondence;

[0091] Sending the storage location to the primary metadata node in the distributed file system;

[0092] Receiving a first identifier list composed of the identifiers of at least one first data node corresponding to the storage location feedback by the primary metadata node to the cache device based on the storage location.

[0093] Wherein, the aforementioned storage location refers to the identifier of a block, and the number of at least one first data node is the number of replicas in the HDFS cluster.

[0094] Optionally, the at least one first data node is a data node associated with the HDFS client in the aforementioned cache device.

[0095] Optionally, any one of the at least one first data node is the data node where the data in the target path is stored.

[0096] S42. Exclude the second identifiers included in the blacklist from the first identifier list to obtain the remaining identifiers, where the second identifiers stored in the blacklist are the identifiers of the data nodes whose time-consuming for processing read requests meets a preset condition;

[0097] Optionally, as can be seen Figure 2f As shown, when the obtained first identifier list includes the identifier dn1 of DataNode1, the identifier dn2 of DataNode2, and the identifier dn3 of DataNode3, if it is determined that dn3 is included in the blacklist, then dn3 is excluded to obtain the remaining identifiers dn1 and dn2.

[0098] S43. Determine the target data node based on the remaining identifiers;

[0099] Optionally, determining the target data node based on the remaining identifiers includes: taking the data node corresponding to the identifier with the earliest order in the remaining identifiers as the target data node. For example, Figure 2f in this case, dn1 is taken as the target data node.

[0100] S44. Obtain the corresponding data from the target data node for synchronization.

[0101] Optionally, the corresponding data obtained from the target data node is synchronization data.

[0102] In some alternative embodiments of the present application, the cache device is further configured to execute the following S21-S22:

[0103] S21. Obtain a second identifier list including the identifiers of at least one second data node from the master metadata node in the distributed file system;

[0104] Optionally, the second data node is a data node associated with the HDFS client.

[0105] S22. Determine the pending data node based on the second identifier list;

[0106] Optionally, the node pointed to by the identifier with the highest ranking in the second identifier list may be used as the pending data node.

[0107] S23. Obtain the time taken to process each read request when the pending data node processes a preset number of read requests;

[0108] S24. Determine whether the ratio of the number of read requests that exceed the first preset duration to the preset number of read requests in the preset number of read requests is greater than a preset threshold. If so, add the identifier of the pending data node to the blacklist.

[0109] Optionally, the aforementioned preset number may be 10.

[0110] Optionally, every 10 read requests can be regarded as an identification window. For the read requests within each identification window, calculate the response time taken by the pending data node to process the read requests.

[0111] Optionally, reference may be made to Figure 2g as shown in Figure 2g DataNode1, DataNode2, and DataNode3 in represent 3 data nodes. The aforementioned cache device can determine the time taken for read requests from the HDFS client and each DataNode node.

[0112] Optionally, for an identification window, every time a read request whose time taken meets the preset condition is found, control the parameter representing that the time taken meets the preset condition to be incremented by 1. When all the read requests within an identification window have been traversed, then count the number of read requests whose time taken meets the preset condition among the preset number.

[0113] In some alternative embodiments of the present application, when the time taken exceeds the first preset duration, it is regarded as the time taken meeting the preset condition.

[0114] Optionally, the aforementioned preset threshold may be 0.8.

[0115] Optionally, if it is determined that the ratio of the number of read requests whose time taken meets the preset condition to the preset number of read requests in the preset number of read requests is not greater than the preset threshold, then determine not to add the pending data node to the blacklist.

[0116] Optionally, after determining not to add the pending data node to the blacklist, the aforementioned S23 - S24 may be continued to be executed. Specifically, reference may be made to Figure 2g as shown in, and the present application does not make any limitations in this regard.

[0117] Optionally, each read request may be a preset read request, or a historical read request for data synchronization, or other read requests. The present application does not make any limitations in this regard.

[0118] In some alternative embodiments of the present application, the caching device is further configured to execute the following S041-S042:

[0119] S041. When the ratio of the number of read requests that exceed the first preset duration to the preset number is greater than a preset threshold, request a third identifier list including the identifiers of at least one third data node from the primary metadata node in the distributed file system based on the last read request that exceeds the first preset duration among the preset number of read requests;

[0120] Optionally, the last read request that exceeds the first preset duration may be the current request, and the third data node is the data node storing the data to be read corresponding to the read request.

[0121] S042. Determine whether the number of identifiers in the third identifier list that are not included in the blacklist is not less than 2. If so, determine to add the identifier of the to-be-determined data node to the blacklist.

[0122] In some alternative embodiments of the present application, the caching device is further configured to:

[0123] For each second identifier in the blacklist, determine whether the duration for which the second identifier has been added to the blacklist exceeds a second preset duration. If so, remove the second identifier from the blacklist.

[0124] In some other alternative embodiments of the present application, the caching device is further configured to: obtain an instruction from a relevant person for adjusting the blacklist, and adjust the blacklist. Specifically, the adjustment of the blacklist may include adding, modifying, deleting, etc. the second identifiers in the blacklist.

[0125] Optionally, the aforementioned second preset duration may be 300s or 30s, and the second preset duration may be stored in the caching device.

[0126] The solution of the present application proposes to perform fine-grained window recognition on the requests sent by the HDFS client, construct a temporary blacklist for the DataNode nodes with response latency exceeding the first preset duration, and use this to enable the caching device to adaptively and actively adjust to select a suitable replica DataNode node to process when reading a block next time, thereby improving the read response latency of the HDFS cluster and enhancing the user experience.

[0127] It should be noted that the process of synchronizing cached data in the present application specifically refers to the process in which the caching device synchronizes data through the HDFS client.

[0128] Each embodiment of the present application further provides a data synchronization system, including a standby metadata node and a cache device in a distributed file system; the standby metadata node is used to: obtain multiple log entries, each log entry including: the file name of the operated file, the type of operation on the operated file, and the path where the operated file is located; based on the multiple log entries, the target path where the corresponding operated file is located is sent to the message middleware, so that the message middleware pushes the target path to the corresponding cache device, so that the cache device synchronizes data from the distributed file system based on the target path, and the cache device is used to provide the data in the distributed file system to the client;

[0129] The cache device synchronizes data from the distributed file system based on the target path, including: the cache device obtains a first identifier list including an identifier of at least one first data node from a primary metadata node in the distributed file system based on the target path; removes second identifiers included in a blacklist from the first identifier list to obtain remaining identifiers, wherein the second identifiers stored in the blacklist are identifiers of data nodes whose time consumption for processing read requests meets a preset condition; determines a target data node based on the remaining identifiers; and obtains corresponding data from the target data node for synchronization.

[0130] Specifically, the aforementioned distributed file system is a HFDS file system, which may include a primary metadata node, a standby metadata node, and a plurality of data nodes, wherein the primary metadata node refers to an Active NameNode, the standby metadata node refers to a Standby NameNode, and the data node refers to a DataNode.

[0131] Optionally, the HDFS client sends the file operation in the HDFS file system to the Active NameNode first. The Active NameNode generates an Editlog and synchronizes the Editlog to the Standby NameNode. The Standby NameNode synchronizes the updates of the Active NameNode through the Editlog. The Editlog refers to the file operation log, which can be a write-ahead log.

[0132] Optionally, the aforementioned multiple log entries may be stored in a file operation log.

[0133] Optionally, sending the target path of the corresponding operated file to the message middleware based on the multiple log entries includes: determining the target path of the target operated file whose operation type is a preset operation type based on the multiple log entries; and sending the target path to the message middleware.

[0134] In some alternative embodiments of the present application, the foregoing preset operation types include any one or more of the following: Delete (delete file), close (close file), rename (rename file), truncate (truncate file), append (append write file).

[0135] Wherein, the foregoing multiple log entries are log entries corresponding to multiple files to be operated, and the foregoing target file to be operated is a file among the foregoing multiple files to be operated whose operation type is a preset operation type.

[0136] Optionally, the foregoing target path is a path in the HDFS file system.

[0137] In some alternative embodiments of the present application, the foregoing standby metadata node may include file operation logs, and a log parsing module in the standby metadata node may obtain and parse the foregoing multiple log entries. The intelligent module in the standby metadata node may execute to determine the target path where the target file to be operated with the operation type of the preset operation type is located based on the foregoing multiple log entries.

[0138] In some alternative embodiments of the present application, the foregoing sending the target path to the message middleware means generating a publish message based on the target path and then sending the publish message to the message middleware, wherein the content of the publish message includes the path name of the target path.

[0139] Optionally, the cache device obtains a first identifier list including the identifiers of at least one first data node from the primary metadata node in the distributed file system based on the target path, including: the cache device determines the storage location of the data in the target path based on the target path and a preset correspondence relationship; sends the storage location to the primary metadata node in the distributed file system; and receives a first identifier list composed of the identifiers of at least one first data node corresponding to the storage location fed back by the primary metadata node based on the storage location.

[0140] Wherein, the foregoing storage location refers to the identifier of a block, and the number of at least one first data node is the number of replicas in the HDFS cluster.

[0141] Optionally, the foregoing at least one first data node is a data node associated with the HDFS client in the foregoing cache device.

[0142] Optionally, any one of the foregoing at least one first data nodes is the data node where the data in the target path is stored.

[0143] The solution for the standby metadata node applicable to the distributed file system provided by the embodiments of the present application includes: obtaining a plurality of log entries, each log entry including: the file name of the file being operated, the operation type of the file being operated, and the path where the file being operated is located; based on the plurality of log entries, sending the target path where the corresponding file being operated is located to a message middleware, so that the message middleware pushes the target path to the corresponding cache device, and the cache device synchronizes data from the distributed file system based on the target path, and the cache device is used to provide the data in the distributed file system to the client; wherein, the cache device synchronizing data from the distributed file system based on the target path includes: the cache device obtaining a first identifier list including the identifiers of at least one first data node from the primary metadata node in the distributed file system based on the target path; removing the second identifiers included in the blacklist from the first identifier list to obtain the remaining identifiers, wherein the second identifiers stored in the blacklist are the identifiers of the data nodes whose time-consuming for processing read requests meets a preset condition; determining target data nodes based on the remaining identifiers; and obtaining corresponding data from the target data nodes for synchronization. Among them, the messages in the message middleware are subscribed by the cache device. Through the solution of the present application, the introduction of the message middleware can push the target path in real time, so that the cache device can synchronize data from the distributed file system in real time, improving the efficiency of processing data based on the cache device.

[0144] For the specific implementation manners corresponding to this embodiment, reference may be made to the foregoing content, which will not be elaborated herein.

[0145] Each embodiment of the present application provides a data synchronization device. Figure 3 As a schematic structural diagram of the data synchronization device, the device is applicable to the standby metadata node in the distributed file system and includes:

[0146] An obtaining unit 31, configured to obtain a plurality of log entries, each log entry including: the file name of the file being operated, the operation type of the file being operated, and the path where the file being operated is located;

[0147] A sending unit 32, configured to send the target path where the corresponding file being operated is located to a message middleware based on the plurality of log entries, so that the message middleware pushes the target path to the corresponding cache device, and the cache device synchronizes data from the distributed file system based on the target path, and the cache device is used to provide the data in the distributed file system to the client;

[0148] Wherein, the cache device synchronizing data from the distributed file system based on the target path includes:

[0149] The cache device obtains a first identifier list including identifiers of at least one first data node from the master metadata node in the distributed file system based on the target path;

[0150] Remove the second identifiers included in the blacklist from the first identifier list to obtain remaining identifiers, where the second identifiers stored in the blacklist are identifiers of data nodes whose time consumption for processing read requests meets a preset condition;

[0151] Determine target data nodes based on the remaining identifiers;

[0152] Obtain corresponding data from the target data nodes for synchronization.

[0153] In some alternative embodiments of the present application, when the foregoing device is used to send the target path where the corresponding operated file is located to the message middleware based on the multiple log entries, it is specifically used for:

[0154] Determine the target path where the target operated file with an operation type of a preset operation type is located based on the multiple log entries;

[0155] Send the target path to the message middleware.

[0156] In some alternative embodiments of the present application, the message middleware includes multiple subscription queues. When the foregoing device is used to send the target path to the message middleware, it is specifically used for:

[0157] Determine a corresponding target subscription queue based on the target path;

[0158] Send the target path to the message middleware, so that the message middleware stores the target path in the target subscription queue corresponding to the target path.

[0159] In some alternative embodiments of the present application, the device is further used for:

[0160] Obtain the setting information for setting the preset operation type by relevant personnel;

[0161] Determine the preset operation type based on the setting information.

[0162] In some alternative embodiments of the present application, the cache device is further used for:

[0163] Obtain a second identifier list including identifiers of at least one second data node from the master metadata node in the distributed file system;

[0164] Determine pending data nodes based on the second identifier list;

[0165] Obtain the time consumed for processing each read request when the to-be-determined data node processes a preset number of read requests;

[0166] Determine whether the ratio of the number of read requests that exceed a first preset duration among the preset number of read requests to the preset number is greater than a preset threshold. If so, add the identifier of the to-be-determined data node to the blacklist.

[0167] In some alternative embodiments of the present application, the caching device is further configured to:

[0168] When the ratio of the number of read requests that exceed the first preset duration to the preset number is greater than the preset threshold, request a third identifier list including the identifiers of at least one third data node from the primary metadata node in the distributed file system based on the last read request that exceeds the first preset duration among the preset number of read requests;

[0169] Determine whether the number of identifiers in the third identifier list that are not included in the blacklist is not less than 2. If so, determine to add the identifier of the to-be-determined data node to the blacklist.

[0170] In some alternative embodiments of the present application, the caching device is further configured to: for each second identifier in the blacklist, determine whether the duration for which the second identifier has been added to the blacklist exceeds a second preset duration. If so, remove the second identifier from the blacklist.

[0171] Optionally, the caching device is further configured to: for each second identifier in the blacklist, determine whether the duration for which the second identifier has been added to the blacklist exceeds a second preset duration. If so, remove the second identifier from the blacklist. All of the above technical solutions can be combined arbitrarily to form alternative embodiments of the present application, which will not be elaborated herein one by one.

[0172] Each unit in the above-mentioned apparatuses can be implemented in whole or in part by software, hardware, and their combination. Each of the above-mentioned units can be embedded in the processor in the computer device in hardware form or be independent of it, or be stored in the memory in the computer device in software form so that the processor can call and execute the operations corresponding to each of the above-mentioned units.

[0173] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the functions of the module or unit.

[0174] Optionally, the foregoing storage node determination device can be integrated into the device where the NameNode is located. The foregoing storage device can be integrated into the device where the DataNode is located.

[0175] Optionally, the present application also provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the foregoing method embodiments are implemented. Figure 4 It is a schematic structural diagram of the computer device provided by the embodiments of the present application. The computer device can be a terminal or a server. As Figure 4 shown, the computer device 600 can include: a communication interface 601, a memory 602, a processor 603, and a communication bus 604. The communication interface 601, the memory 602, and the processor 603 communicate with each other through the communication bus 604. The communication interface 601 is used for the computer device 600 to perform data communication with external devices. The memory 602 can be used to store software programs and modules. The processor 603 runs the software programs and modules stored in the memory 602, such as the software programs corresponding to the foregoing method embodiments.

[0176] Optionally, the processor 603 can call the software programs and modules stored in the memory 602 to perform the following operations:

[0177] Obtain a plurality of log entries, each log entry including: the file name of the file being operated, the operation type of the file being operated, and the path where the file being operated is located;

[0178] Based on the plurality of log entries, send the target path where the corresponding file being operated is located to the message middleware, so that the message middleware pushes the target path to the corresponding cache device, and the cache device synchronizes data from the distributed file system based on the target path. The cache device is used to provide the data in the distributed file system to the client.

[0179] Optionally, the processor 603 can also call the software programs and modules stored in the memory 602 to perform other functions.

[0180] The present application also provides a computer-readable storage medium for storing a computer program. The computer-readable storage medium can be applied to a computer device, and the computer program causes the computer device to execute the corresponding processes in the various methods in the embodiments of the present application. For the sake of brevity, details are not repeated here.

[0181] The present application also provides a computer program product, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the corresponding processes in the various methods in the embodiments of the present application. For the sake of brevity, details are not repeated here.

[0182] The present application also provides a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the corresponding processes in the various methods in the embodiments of the present application. For the sake of brevity, details are not repeated here.

[0183] It should be understood that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware decoding processor, or executed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0184] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.

[0185] It should be understood that the above memory is by way of example but not limitation. For example, the memory in the embodiments of the present application can also be a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate SDRAM (DDR SDRAM), an enhanced SDRAM (ESDRAM), a synch link DRAM (SLDRAM), and a direct rambus random access memory (DR RAM), etc. That is to say, the memory in the embodiments of the present application is intended to include but not be limited to these and any other suitable types of memory.

[0186] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0187] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0188] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0189] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0190] In addition, each functional unit in the embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0191] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server) to execute all or part of the steps of the methods described in each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0192] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A data synchronization method, characterized in that, Applicable to a standby metadata node in a distributed file system, including: Obtain multiple log entries, each log entry including: the file name of the file being operated, the operation type of the file being operated, and the path where the file being operated is located; Based on the multiple log entries, send the target path where the corresponding file being operated is located to a message middleware, so that the message middleware pushes the target path to the corresponding cache device, and the cache device synchronizes data from the distributed file system based on the target path, and the cache device is used to provide data in the distributed file system to a client; Wherein, the cache device synchronizes data from the distributed file system based on the target path, including: The cache device obtains a first identifier list including identifiers of at least one first data node from a primary metadata node in the distributed file system based on the target path; Exclude a second identifier included in a blacklist from the first identifier list to obtain remaining identifiers, wherein the second identifier stored in the blacklist is the identifier of a data node whose time-consuming for processing a read request meets a preset condition; Determine target data nodes based on the remaining identifiers; Obtain corresponding data from the target data nodes for synchronization.

2. The method according to claim 1, wherein Sending the target path where the corresponding file being operated is located to a message middleware based on the multiple log entries, including: Determine the target path where the target file being operated with an operation type being a preset operation type is located based on the multiple log entries; Send the target path to the message middleware.

3. The method according to claim 2, wherein The message middleware includes multiple subscription queues, and sending the target path to the message middleware includes: Determine a corresponding target subscription queue based on the target path; Send the target path to the message middleware, so that the message middleware stores the target path in the target subscription queue corresponding to the target path.

4. The method according to claim 2, characterized in that, The method further includes: Obtain setting information for setting the preset operation type by relevant personnel; Determine the preset operation type based on the setting information.

5. The method according to claim 1, characterized in that, The cache device is further used for: Obtain a second identifier list including identifiers of at least one second data node from a primary metadata node in the distributed file system; Determine pending data nodes based on the second identifier list; Obtain the time-consuming for each read request when the pending data nodes process a preset number of read requests; Determine whether the ratio of the number of read requests whose time-consuming exceeds a first preset duration to the preset number among the preset number of read requests is greater than a preset threshold. If so, add the identifier of the pending data node to the blacklist.

6. The method according to claim 5, wherein The cache device is further used for: When the ratio of the number of read requests whose time-consuming exceeds a first preset duration to the preset number is greater than a preset threshold, request a third identifier list including identifiers of at least one third data node from a primary metadata node in the distributed file system based on the last read request that exceeds the first preset duration among the preset number of read requests; Determine whether the number of identifiers in the third identifier list that are not included in the blacklist is not less than 2. If so, determine to add the identifier of the to-be-determined data node to the blacklist.

7. The method according to claim 1, wherein The cache device is further configured to: for each second identifier in the blacklist, determine whether the duration for which the second identifier has been added to the blacklist exceeds a second preset duration. If so, remove the second identifier from the blacklist.

8. A data synchronization system, characterized in that, It includes a standby metadata node and a cache device in a distributed file system; the standby metadata node is configured to: obtain a plurality of log entries, each log entry including: the file name of the file being operated, the operation type of the file being operated, and the path where the file being operated is located; based on the plurality of log entries, send the target path where the corresponding file being operated is located to a message middleware, so that the message middleware pushes the target path to the corresponding cache device, and the cache device synchronizes data from the distributed file system based on the target path, and the cache device is configured to provide the data in the distributed file system to a client; The cache device synchronizes data from the distributed file system based on the target path, including: the cache device obtains a first identifier list including the identifiers of at least one first data node from the primary metadata node in the distributed file system based on the target path; removes the second identifiers included in the blacklist from the first identifier list to obtain remaining identifiers, where the second identifiers stored in the blacklist are the identifiers of data nodes whose time-consuming for processing read requests meets a preset condition; determines target data nodes based on the remaining identifiers; and obtains corresponding data from the target data nodes for synchronization.

9. A data synchronization device, characterized in that, Applicable to a standby metadata node in a distributed file system, including: An obtaining unit, configured to obtain a plurality of log entries, each log entry including: the file name of the file being operated, the operation type of the file being operated, and the path where the file being operated is located; A sending unit, configured to send the target path where the corresponding file being operated is located to a message middleware based on the plurality of log entries, so that the message middleware pushes the target path to the corresponding cache device, and the cache device synchronizes data from the distributed file system based on the target path, and the cache device is configured to provide the data in the distributed file system to a client; Wherein, the cache device synchronizes data from the distributed file system based on the target path, including: The cache device obtains a first identifier list including the identifiers of at least one first data node from the primary metadata node in the distributed file system based on the target path; Removes the second identifiers included in the blacklist from the first identifier list to obtain remaining identifiers, where the second identifiers stored in the blacklist are the identifiers of data nodes whose time-consuming for processing read requests meets a preset condition; Determines target data nodes based on the remaining identifiers; Obtains corresponding data from the target data nodes for synchronization.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded by a processor to execute the steps in the method according to any one of claims 1-7.

11. A computer device, characterized in that, The computer device includes a processor and a memory. The memory stores a computer program, and the processor is configured to execute the steps in the method according to any one of claims 1-7 by calling the computer program stored in the memory.