Information acquisition method, device, medium, product and system

By synchronizing the meta information update information written to the data to the target node in the distributed file system, and allowing the reading node to directly obtain the required meta information from the target node, the problem of excessive burden on the meta information node is solved, improving the efficiency of data reading and reducing costs.

CN114722020BActive Publication Date: 2025-05-16ALIBABA (CHINA) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210162053.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-22
Publication Date
2025-05-16
Estimated Expiration
2042-02-22

AI Technical Summary

Technical Problem

In distributed file systems, frequent data writing operations cause excessive burden on meta-information nodes, thereby reducing the efficiency of data reading and increasing costs.

Method used

When the write node of the distributed file system completes the data writing, the meta information of the written data is synchronized to the target node, and the meta information on the target node is updated according to these update information. The read node can directly obtain the required meta information from the target node for data reading.

Benefits of technology

By reducing the interaction of meta-information nodes, the burden on meta-information nodes is reduced, the efficiency of data reading is improved, and the cost of data reading is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114722020B_ABST
    Figure CN114722020B_ABST
Patent Text Reader

Abstract

The disclosed embodiment discloses an information acquisition method, device, medium, product and system, which is applied to the target node of a distributed file system, including: when the write node completes writing data to the data node, synchronizing the meta information update information of the written data from the write node; updating the meta information on the target node according to the meta information update information of the written data; obtaining the data read request sent by the read node; in response to the updated meta information on the target node including the target read meta information, returning the target read meta information to the read node. This solution enables the read node to obtain the required target read meta information from the target node without interacting with the meta information node, so as to complete the data reading according to the target read meta information, and reduces the burden of the meta information node, improves the efficiency of data reading and reduces the cost of data reading without affecting the completion of data reading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of network technology, and in particular to information acquisition methods, devices, media, products and systems. Background Art

[0002] In recent years, with the advancement of science and technology, the amount of data generated in scientific research and daily life has exploded, and the demand for data storage has also increased dramatically. Since local storage systems are difficult to meet the growing demand for data storage, and considering that mobile computing and enterprise-level large-scale computing have put higher requirements on the underlying storage system, people are beginning to use distributed file systems more and more. Distributed file systems can store data in multiple independent devices. They can store data as multiple copies, no longer back up data, improve the data storage rate, and save storage time.

[0003] Normally, when a user reads data from a distributed file system, a data read request can be sent to a read node through a client. The read node can obtain the metadata of the data to be read corresponding to the data read request from the metadata node based on the data read request. The metadata is used to indicate the storage location of the data to be read. The read node can read the data to be read from the corresponding data node based on the metadata and return the read data to the client.

[0004] Although the above scheme allows users to read the data they want from the distributed file system, since the reading node will obtain the corresponding metadata of the read data from the metadata node each time the data is read, when the data is read more frequently, the metadata node is burdened with a heavy burden, which increases the cost of data reading and reduces the efficiency of data reading. Summary of the invention

[0005] In order to solve the problems in the related art, the embodiments of the present disclosure provide information acquisition methods, devices, media, products and systems.

[0006] In a first aspect, an embodiment of the present disclosure provides an information acquisition method, wherein the method is applied to a target node of a distributed file system, and the method includes:

[0007] When the write node of the distributed file system completes writing data to the data node of the distributed file system, the metadata update information of the written data is synchronized from the write node;

[0008] Update the meta-information on the target node according to the meta-information update information of the written data;

[0009] Obtain data read requests sent by the read nodes of the distributed file system;

[0010] In response to the updated meta-information on the target node including target read meta-information, the target read meta-information is returned to the read node, the meta-information is used to indicate the storage location of the corresponding data, and the target read meta-information corresponds to the data requested to be read by the data read request.

[0011] In combination with the first aspect, in a first implementation of the first aspect of the present disclosure, the method further includes:

[0012] In response to the meta information on the target node not including the target read meta information, acquiring the target read meta information from the meta information node of the distributed file system;

[0013] Returns target read meta information to the read node.

[0014] In combination with the first aspect and any one of the first implementation manners of the first aspect, in a second implementation manner of the first aspect of the present disclosure, the metadata update information of the written data includes a data block identifier of a written data block corresponding to the written data, a data block length of the written data block, a file identifier of a written file corresponding to the written data block, and a file length of the written file;

[0015] The meta information on the target node is updated according to the meta information update information of the written data, including:

[0016] In response to the data block identifiers of the data blocks in the meta-information on the target node including the data block identifiers of all written data blocks in the meta-information corresponding to the written data, updating the meta-information on the target node according to the meta-information corresponding to the written data;

[0017] Or, in response to the data block identifier of the data block in the metadata on the target node not including the data block identifier of the target written data block in the metadata corresponding to the written data, the metadata of the target written data block is synchronized from the metadata node according to the data block identifier of the target written data block, and the metadata on the target node is updated according to the metadata of the target written data block.

[0018] In combination with the first aspect and any one of the first implementation manners of the first aspect, in a third implementation manner of the first aspect of the present disclosure, when the write node writes data to the data node without data block switching, the metadata update information of the written data includes a data block identifier of the written data block corresponding to the written data, a data block length of the written data block, a file identifier of the written file corresponding to the written data block, and a file length of the written file;

[0019] When a data block switch occurs when the write node writes data to the data node, the meta-information update information of the written data includes the meta-information corresponding to the written data.

[0020] In a second aspect, an information acquisition method is provided in an embodiment of the present disclosure, wherein the method is applied to a write node of a distributed file system, and the method includes:

[0021] In response to completing the writing of data to the data node of the distributed file system, the meta information update information of the written data is synchronized to the target node of the distributed file system, and the target node is used to update the meta information on the target node according to the meta information update information of the written data.

[0022] In combination with the second aspect, in a first implementation of the second aspect of the present disclosure, when no data block switching occurs when writing data to the data node, the metadata update information of the written data includes a data block identifier of the written data block corresponding to the written data, a data block length of the written data block, a file identifier of the written file corresponding to the written data block, and a file length of the written file;

[0023] When data block switching occurs when writing data to a data node, the meta-information update information of the written data includes the meta-information corresponding to the written data.

[0024] In a third aspect, an information acquisition method is provided in an embodiment of the present disclosure, wherein the method is applied to a reading node of a distributed file system, and the method includes:

[0025] Obtain a data read request and send the data read request to the target node of the distributed file system;

[0026] Receive target read meta information returned by the target node, where the target read meta information is meta information corresponding to the data requested to be read by the data read request, and the meta information is used to indicate the storage location of the corresponding data;

[0027] Read data according to the target reading meta information.

[0028] In a fourth aspect, an information acquisition device is provided in an embodiment of the present disclosure, wherein the device includes:

[0029] The information synchronization module is configured to synchronize the metadata update information of the written data from the write node when the write node of the distributed file system completes writing the data to the data node of the distributed file system;

[0030] An information updating module, configured to update the meta-information on the target node according to the meta-information update information of the written data;

[0031] A first request acquisition module is configured to acquire a data read request sent by a read node of the distributed file system;

[0032] The information return module is configured to return the target read meta-information to the read node in response to the updated meta-information on the target node, including the target read meta-information, wherein the meta-information is used to indicate the storage location of the corresponding data, and the target read meta-information corresponds to the data requested to be read by the data read request.

[0033] In a fifth aspect, an information acquisition device is provided in an embodiment of the present disclosure, wherein the device includes:

[0034] The meta-information synchronization module is configured to synchronize the meta-information update information of the written data to the target node of the distributed file system in response to completing the writing of data to the data node of the distributed file system, and the target node is used to update the meta-information on the target node according to the meta-information update information of the written data.

[0035] In a sixth aspect, an information acquisition device is provided in an embodiment of the present disclosure, wherein the device includes:

[0036] A second request acquisition module is configured to acquire a data read request and send the data read request to a target node of the distributed file system;

[0037] A meta-information acquisition module is configured to receive target read meta-information returned by a target node, wherein the target read meta-information is meta-information corresponding to the data requested to be read by the data read request, and the meta-information is used to indicate a storage location of the corresponding data;

[0038] The data reading module is configured to read data according to the target reading meta information.

[0039] In the seventh aspect, an electronic device is provided in an embodiment of the present disclosure, comprising a memory and at least one processor; the memory is used to store one or more computer instructions, and the one or more computer instructions are executed by at least one processor to implement the method steps described in any one of the first aspect, the first implementation method of the first aspect to the third implementation method, the second aspect, the first implementation method of the second aspect, and the third aspect.

[0040] In an eighth aspect, a computer-readable storage medium is provided in an embodiment of the present disclosure, on which computer instructions are stored. When the computer instructions are executed by a processor, the method steps described in any one of the first aspect, the first implementation method of the first aspect to the third implementation method, the second aspect, the first implementation method of the second aspect, and the third aspect are implemented.

[0041] In the ninth aspect, a computer program product is provided in an embodiment of the present disclosure, comprising computer instructions, which, when executed by a processor, implement the method steps described in any one of the first aspect, the first implementation method of the first aspect to the third implementation method, the second aspect, the first implementation method of the second aspect, and the third aspect.

[0042] In a tenth aspect, an embodiment of the present disclosure provides a distributed file system, the distributed file system comprising a write node, at least one data node, a read node, a target node, and a meta information node;

[0043] The write node is configured to synchronize meta information update information of the written data to the target node in response to completing the writing of data to the data node;

[0044] The target node is configured to synchronize the metadata update information of the written data from the write node when the write node completes writing the data to the data node; update the metadata on the target node according to the metadata update information of the written data; obtain the data read request sent by the read node; in response to the updated metadata on the target node including the target read metadata, return the target read metadata to the read node, the metadata is used to indicate the storage location of the corresponding data, and the target read metadata corresponds to the data requested to be read by the data read request;

[0045] The reading node is configured to obtain a data reading request and send the data reading request to a target node; receive target reading meta information returned by the target node; and read data according to the target reading meta information.

[0046] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects:

[0047] According to the technical solution provided by the embodiment of the present disclosure, the solution is applied to the target node of the distributed file system. The solution synchronizes the metadata update information of the written data from the write node when the write node of the distributed file system completes the data writing to the data node of the distributed file system, thereby ensuring that each time the data writing is completed, the target node can obtain the change content caused by the data writing to the corresponding metadata, that is, the metadata update information, and update the metadata on the target node according to the metadata update information of the written data, ensuring that the updated metadata on the target node can accurately reflect the storage location of the corresponding data after each data writing is completed, reducing the probability that the metadata cannot accurately reflect the storage location of the corresponding data due to frequent data writing. By obtaining the data read request sent by the read node of the distributed file system, and responding to the updated metadata on the target node including the target read metadata corresponding to the data requested to be read by the data read request, the target read metadata is returned to the read node. Among them, when a reading node in a distributed file system needs to read the data requested by a data reading request, the reading node does not need to obtain the meta information corresponding to the data, that is, the target reading meta information, from the meta information node, but can send a data reading request to enable the target node to respond to the data reading request to determine whether it has stored the target reading meta information. When the updated meta information stored by itself includes the target reading meta information, in response to the updated meta information on the target node including the target reading meta information corresponding to the data requested to be read by the data reading request, the target reading meta information is returned to the reading node, so that the reading node can obtain the target reading meta information, so that the reading node can read the data requested by the data reading request from the corresponding data node according to the target reading meta information. In summary, the technical solution provided by the embodiment of the present disclosure can enable the reading node to obtain the required target reading meta information from the target node when it needs to read the data requested by the data reading request without interacting with the meta information node, so as to complete the data reading of the data requested by the data reading request according to the target reading meta information, thereby reducing the burden of the meta information node, improving the efficiency of data reading, and reducing the cost of data reading without affecting the completion of data reading.

[0048] According to the technical solution provided by the embodiment of the present disclosure, by responding to the fact that the meta-information on the target node does not include the target read meta-information, obtaining the target read meta-information from the meta-information node of the distributed file system, and returning the target read meta-information to the reading node, it can be ensured that the reading node can read data according to the target read meta-information, thereby improving the success rate of data reading and the user experience.

[0049] According to the technical solution provided by the embodiments of the present disclosure, by limiting the metadata update information of the written data to include the data block identifier of the written data block corresponding to the written data, the data block length of the written data block, the file identifier of the written file corresponding to the written data block, and the file length of the written file, the volume of the metadata update information can be made smaller than the volume of the metadata. In response to the data block identifier of the data block in the metadata on the target node including the data block identifiers of all the written data blocks in the metadata corresponding to the written data, the metadata on the target node is updated according to the metadata corresponding to the written data, so that when no data block switching occurs during the data writing process, the updated metadata on the target node can include the metadata of the data block corresponding to the written data; or, in response to the data block identifier of the data block in the metadata on the target node not including the data block identifier of the target written data block in the metadata corresponding to the written data, the metadata of the target written data block is synchronized from the metadata node according to the data block identifier of the target written data block, and the metadata on the target node is updated according to the metadata of the target written data block, so that when a data block switching occurs during the data writing process, the updated metadata on the target node can include the metadata of the data block corresponding to the written data, thereby ensuring that the metadata on the target node includes the metadata of the data block corresponding to the written data while minimizing the size of the metadata, thereby minimizing the network resources occupied by transmitting the metadata and the processing resources occupied when serializing or deserializing the metadata.

[0050] According to the technical solution provided by the embodiment of the present disclosure, when no data block switching occurs when the write node writes data to the data node, the metadata update information of the written data includes the data block identifier of the written data block corresponding to the written data, the data block length of the written data block, the file identifier of the written file corresponding to the written data, and the file length of the written file; and when data block switching occurs when the write node writes data to the data node, the metadata update information of the written data includes the metadata corresponding to the written data. Under the premise of minimizing the volume of the metadata, the metadata updated according to the metadata update information of the written data on the target node includes the metadata of the data block corresponding to the written data, thereby minimizing the network resources occupied by transmitting the metadata and the processing resources occupied when serializing or deserializing the metadata.

[0051] According to the technical solution provided by the embodiment of the present disclosure, by synchronizing the metadata update information of the written data to the target node of the distributed file system in response to the completion of data writing to the data node of the distributed file system, it can be ensured that each time the data writing is completed, the target node can obtain the change content of the corresponding metadata caused by the data writing, that is, the metadata update information, and update the metadata on the target node according to the metadata update information of the written data, to ensure that the updated metadata on the target node can accurately reflect the storage location of the corresponding data after each data writing is completed, thereby reducing the probability that the metadata cannot accurately reflect the storage location of the corresponding data due to frequent data writing. Among them, when a reading node in a distributed file system needs to read the data requested by a data reading request, the reading node does not need to obtain the meta information corresponding to the data, that is, the target reading meta information, from the meta information node, but can send a data reading request so that the target node can determine whether it has stored the target reading meta information in response to the data reading request. When the updated meta information stored by itself includes the target reading meta information, in response to the updated meta information on the target node including the target reading meta information corresponding to the data requested to be read by the data reading request, the target reading meta information is returned to the reading node, so that the reading node can obtain the target reading meta information, so that the reading node can read the data requested by the data reading request from the corresponding data node according to the target reading meta information. In summary, the technical solution provided by the embodiment of the present disclosure can enable the reading node to obtain the data requested by the data reading request without interacting with the meta information node when it needs to read the data requested by the data reading request, but obtain the required target reading meta information from the target node, so as to complete the data reading of the data requested by the data reading request according to the target reading meta information, thereby reducing the burden of the meta information node, improving the efficiency of data reading, and reducing the cost of data reading without affecting the completion of data reading.

[0052] According to the technical solution provided by the embodiment of the present disclosure, when data block switching does not occur when writing data to a data node, the metadata update information of the written data includes the data block identifier of the written data block corresponding to the written data, the data block length of the written data block, the file identifier of the written file corresponding to the written data block, and the file length of the written file; when data block switching occurs when writing data to the data node, the metadata update information of the written data includes the metadata corresponding to the written data. Under the premise of minimizing the volume of the metadata, the metadata updated according to the metadata update information of the written data on the target node includes the metadata of the data block corresponding to the written data, thereby minimizing the network resources occupied by transmitting the metadata and the processing resources occupied when serializing or deserializing the metadata.

[0053] According to the technical solution provided by the embodiment of the present disclosure, by obtaining a data read request, sending the data read request to a target node, receiving target read meta-information returned by the target node, and reading data according to the target read meta-information, the reading node does not need to interact with the meta-information node when it needs to read the data requested by the data read request, but instead obtains the required target read meta-information from the target node, so as to complete the data reading of the data requested by the data read request according to the target read meta-information, thereby reducing the burden on the meta-information node, improving the efficiency of data reading, and reducing the cost of data reading without affecting the completion of data reading.

[0054] According to the technical solution provided by the embodiment of the present disclosure, in a distributed file system, when the writing node completes writing data to the data node of the distributed file system, it synchronizes the metadata update information of the written data to the target node, thereby ensuring that each time the data writing is completed, the target node can obtain the change content caused by the data writing to the corresponding metadata, that is, the metadata update information, and the target node updates the metadata on the target node according to the metadata update information of the written data, ensuring that the updated metadata on the target node can accurately reflect the storage location of the corresponding data after each data writing is completed, reducing the probability that the metadata cannot accurately reflect the storage location of the corresponding data due to frequent data writing. The reading node obtains a data reading request and sends a data reading request to the target node. The target node obtains the data reading request and responds to the updated metadata on the target node including the target reading metadata corresponding to the data requested to be read by the data reading request, and returns the target reading metadata to the reading node. Among them, when the reading node needs to read the data requested by the data reading request, the reading node does not need to obtain the meta information corresponding to the data, that is, the target reading meta information, from the meta information node, but can send a data reading request to enable the target node to respond to the data reading request to determine whether it has stored the target reading meta information. When the updated meta information stored by itself includes the target reading meta information, in response to the updated meta information on the target node including the target reading meta information corresponding to the data requested to be read by the data reading request, the target reading meta information is returned to the reading node, so that the reading node can obtain the target reading meta information, so that the reading node can read the data requested to be read by the data reading request from the corresponding data node according to the target reading meta information. In summary, the technical solution provided by the embodiment of the present disclosure can enable the reading node in the distributed system to read the data requested by the data reading request without interacting with the meta information node, but to obtain the required target reading meta information from the target node, so as to complete the data reading of the data requested to be read by the data reading request according to the target reading meta information, thereby reducing the burden of the meta information node, improving the efficiency of data reading, and reducing the cost of data reading without affecting the completion of data reading.

[0055] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Other features, objectives and advantages of the present disclosure will become more apparent through the following detailed description of non-limiting embodiments in conjunction with the accompanying drawings. In the accompanying drawings:

[0057] Figure 1 The system architecture of a distributed file system according to an embodiment of the present disclosure is shown.

[0058] Figure 2 A flowchart of an information acquisition method according to an embodiment of the present disclosure is shown.

[0059] Figure 3 A schematic block diagram showing meta information according to an embodiment of the present disclosure.

[0060] Figure 4 A schematic block diagram showing the process of an information acquisition method according to an embodiment of the present disclosure.

[0061] Figure 5 A schematic diagram showing meta-information update information of written data according to an embodiment of the present disclosure.

[0062] Figure 6 A flow chart of an information acquisition method according to an embodiment of the present disclosure is shown.

[0063] Figure 7 A flow chart of an information acquisition method according to an embodiment of the present disclosure is shown.

[0064] Figure 8 A structural block diagram of an information acquisition device according to an embodiment of the present disclosure is shown.

[0065] Fig. 9 A structural block diagram of an information acquisition device according to an embodiment of the present disclosure is shown.

[0066] Fig.10 A structural block diagram of an information acquisition device according to an embodiment of the present disclosure is shown.

[0067] Fig.11 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0068] Fig.12 It is a structural diagram of a computer system suitable for implementing an information acquisition method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0069] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. In addition, for the sake of clarity, parts not related to the description of the exemplary embodiments are omitted in the accompanying drawings.

[0070] In the present disclosure, it should be understood that terms such as "include" or "have" are intended to indicate the presence of labels, numbers, steps, behaviors, components, parts, or combinations thereof disclosed in the present specification, and are not intended to exclude the possibility that one or more other labels, numbers, steps, behaviors, components, parts, or combinations thereof exist or are added.

[0071] It should also be noted that, in the absence of conflict, the embodiments and labels in the embodiments of the present disclosure can be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0072] In the related art, in a distributed file system, files stored thereon are usually divided into a series of data blocks (chunks) for storage, and multiple replicas (Replicas) are generated for each data block, and the multiple replicas are respectively stored on different data nodes. When a user needs to read data from a distributed file system, a corresponding data read request can be sent to an access node of the distributed file system through a client, and the access node can obtain the meta information of the data block in the read data indicated by the data read request from the meta information node, wherein the meta information can be used to indicate the specific storage location of the corresponding data block, and the reading node can read the data block of the read data from the corresponding data node according to the meta information, so as to return the read data to the client and complete the data reading.

[0073] Disadvantages of the above scheme: Although the above scheme allows users to read the data they want to read from the distributed file system, in the distributed file system, it is often encountered that multiple different clients read the same file at the same time in a short period of time. For example, when a user uses a live broadcast application, after the live broadcast data uploaded by the anchor user through the client is written to a data node, multiple users as viewers may request to read the live broadcast data from other data nodes through the client in a very short period of time. In the case of such frequent data reading, a large number of demands for obtaining meta-information from the meta-information node will be generated in a short period of time. In the process of the meta-information node returning a large amount of meta-information in response to the demand, the burden on the meta-information node will be increased, thereby increasing the cost of data reading. In some cases, the meta-information node cannot even return the meta-information in time, resulting in the inability to complete data reading, thereby reducing the efficiency of data reading.

[0074] Considering the shortcomings of the above schemes, the inventors of the present disclosure have proposed a new scheme: the scheme is applied to the target node of the distributed file system, and the scheme synchronizes the metadata update information of the written data from the write node when the write node of the distributed file system completes the data writing to the data node of the distributed file system, thereby ensuring that each time the data writing is completed, the target node can obtain the change content caused by the data writing to the corresponding metadata, that is, the metadata update information, and update the metadata on the target node according to the metadata update information of the written data, ensuring that the updated metadata on the target node can accurately reflect the storage location of the corresponding data after each data writing is completed, reducing the probability that the metadata cannot accurately reflect the storage location of the corresponding data due to frequent data writing. By obtaining the data read request sent by the read node of the distributed file system, and responding to the updated metadata on the target node including the target read metadata corresponding to the data requested to be read by the data read request, the target read metadata is returned to the read node. Among them, when a reading node in a distributed file system needs to read the data requested by a data reading request, the reading node does not need to obtain the meta information corresponding to the data, that is, the target reading meta information, from the meta information node, but can send a data reading request to enable the target node to respond to the data reading request to determine whether it has stored the target reading meta information. When the updated meta information stored by itself includes the target reading meta information, in response to the updated meta information on the target node including the target reading meta information corresponding to the data requested to be read by the data reading request, the target reading meta information is returned to the reading node, so that the reading node can obtain the target reading meta information, so that the reading node can read the data requested by the data reading request from the corresponding data node according to the target reading meta information. In summary, the technical solution provided by the embodiment of the present disclosure can enable the reading node to obtain the required target reading meta information from the target node when it needs to read the data requested by the data reading request without interacting with the meta information node, so as to complete the data reading of the data requested by the data reading request according to the target reading meta information, thereby reducing the burden of the meta information node, improving the efficiency of data reading, and reducing the cost of data reading without affecting the completion of data reading.

[0075] In order to solve the above problems, the present disclosure proposes an information acquisition method, device, equipment and medium.

[0076] Figure 1 The system architecture of a distributed file system according to an embodiment of the present disclosure is shown. Figure 1 The system architecture shown is only an example of a system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0077] like Figure 1 As shown, the distributed file system includes a write node 101 , at least one data node 102 , a read node 103 , a target node 104 and a meta-information node 105 .

[0078] Among them, the write node 101 is used to obtain a data write request, and write the corresponding data to the corresponding data node 102 according to the data write request, and synchronize the metadata of the written data to the metadata node 105 after each data write is completed, wherein the metadata can be understood as being used to indicate the specific storage location of the corresponding data.

[0079] The data node 102 is used to store written data.

[0080] The client 106 may send a data read request to request to read corresponding data from the distributed file system, and the read node 103 is used to receive the data read request.

[0081] The meta-information node 105 is used to store the meta-information of the data stored in the data node 102 .

[0082] Figure 2 A flowchart of an information acquisition method according to an embodiment of the present disclosure is shown, and the method is applied to a target node of a distributed file system, such as Figure 2 As shown, the information acquisition method includes steps S101, S102, S103, and S104.

[0083] In step S101, when the write node of the distributed file system completes writing data to the data node of the distributed file system, the metadata update information of the written data is synchronized from the write node.

[0084] In step S102, the meta-information on the target node is updated according to the meta-information update information of the written data.

[0085] In step S103, a data read request sent by a read node of the distributed file system is obtained.

[0086] In step S104, in response to the updated meta-information on the target node including the target read meta-information, the target read meta-information is returned to the read node.

[0087] The meta information is used to indicate the storage location of the corresponding data, and the target read meta information corresponds to the data requested to be read by the data read request.

[0088] In one embodiment of the present disclosure, the metadata update information of the written data can be understood to include the content that the metadata corresponding to the data on the data node is changed due to the writing of the written data to the data node. Among them, the metadata can be understood to be used to indicate the storage location of the corresponding data. Exemplarily, since the files stored on the distributed file system are usually divided into a series of data blocks, the metadata can be used to indicate the file identifier of the file to which the corresponding data block belongs, the file length of the file to which the corresponding data block belongs, the write position of the corresponding data block in the corresponding file, and on which data node the corresponding data block is stored.

[0089] For example, Figure 3 A schematic block diagram showing meta information according to an embodiment of the present disclosure is shown as follows: Figure 3 As shown, the meta information 200 includes FileMeta 201, offsetTable 202, and chunkinfos 203, wherein chunkinfos 203 includes multiple csInfo 213, wherein FileMeta 201 includes the file identifier of the file to which the corresponding data block belongs and the file length of the file to which the corresponding data block belongs, offsetTable 202 is a mapping table between data blocks and write positions, and the mapping table is used to indicate the correspondence between the write position in the file and the data block, and csInfo 213 is used to indicate the correspondence between the node and the data block, that is, on which data node the corresponding data block is stored.

[0090] In one embodiment of the present disclosure, synchronizing the metadata update information of the written data from the write node can be understood as receiving the metadata update information sent by the write node when the write node completes writing data to the data node, and can also be understood as obtaining the forwarded metadata update information from other nodes, devices or systems.

[0091] In one embodiment of the present disclosure, the target node can be understood as a node in the distributed file system that is independent of the metadata node. The target node can be used to store metadata of part or all of the data in the corresponding data node. The metadata on the target node can be understood as a subset of the metadata on the metadata node.

[0092] It should be noted that a distributed file system may include multiple target nodes. For example, target nodes corresponding to the locations may be set at different locations, and a reading node at a certain location may send a data reading request to the target node corresponding to the location. In addition, the target node and other nodes in the distributed file system may be carried by different data centers, or the target node and any node in the distributed file system may be carried by the same data center. A data center may be understood as an entity that realizes centralized processing, storage, transmission, exchange, and management of information in a physical space. A data center may include computer equipment, server equipment, network equipment, storage equipment, etc. For example, a target node may be carried by a data center that carries any reading node.

[0093] In one embodiment of the present disclosure, updating the meta-information on the target node according to the meta-information update information of the written data can be understood as replacing or modifying the portion of the meta-information on the target node corresponding to the written data according to the meta-information update information of the written data.

[0094] In one embodiment of the present disclosure, a data read request can be understood as indicating a user request to read the data to be read. Obtaining a data read request can be receiving a data read request sent by a read node, or obtaining the forwarded data read request from other nodes, devices or systems. For example, other nodes in a distributed file system can forward the data read request sent by a read node to a target node.

[0095] It should be noted that the embodiments of the present disclosure do not specifically limit the order of step S101 - step S102 and step S103. Exemplarily, step S101 - step S102 and step S103 may be independently executed by two different threads, or may be executed sequentially by the same thread.

[0096] In one embodiment of the present disclosure, to determine whether the updated meta-information on the target node includes the target read meta-information, the entire content of the target read meta-information can be retrieved from the updated meta-information on the target node, or the search can be performed based on identification information of the target read meta-information, and based on the search results, it can be determined whether the updated meta-information on the target node includes the target read meta-information.

[0097] In one embodiment of the present disclosure, returning the target read meta information to the read node may be sending the target read meta information to the read node, or forwarding the target read meta information to the read node through other nodes, devices or systems.

[0098] For example, Figure 4 A schematic block diagram showing the process of an information acquisition method according to an embodiment of the present disclosure is shown as follows: Figure 4 As shown, the distributed file system includes a write node 101 , at least one data node 102 , a read node 103 , a target node 104 and a meta-information node 105 .

[0099] Among them, the write node 101 can obtain the written data 301 and write the written data 301 into the corresponding data node 102. When the write node 101 completes writing data to the data node 102, the meta information update information 302 of the written data 301 can be synchronized from the write node 101 to the target node 104. The target node 104 can update the meta information on itself according to the meta information update information 302. The read node 103 can receive the data read request 303 sent by the user end 106, and send the data read request 303. The target node 104 can receive the data read request 303 sent by the read node 103, and when the updated meta information on the target node includes the target read meta information 304 corresponding to the data requested to be read by the data read request 303, the target node 104 returns the target read meta information 304 to the read node 103. The reading node 103 can determine the data node 102 where the data requested to be read by the data reading request 303 is located and the specific storage location of the data requested to be read by the data reading request 303 on the data node 102 based on the target reading meta-information 304. Therefore, the reading node 103 can successfully read the data 305 requested to be read by the data reading request 303 from the corresponding data node 102 based on the target reading meta-information 304, and return the read data 305 to the user terminal 106, completing the process of reading data from the distributed file system.

[0100] The inventor of the present disclosure has proposed a new solution: the solution is applied to the target node of the distributed file system. The solution synchronizes the metadata update information of the written data from the write node when the write node of the distributed file system completes the data writing to the data node of the distributed file system, thereby ensuring that each time the data writing is completed, the target node can obtain the change content caused by the data writing to the corresponding metadata, that is, the metadata update information, and update the metadata on the target node according to the metadata update information of the written data, ensuring that the updated metadata on the target node can accurately reflect the storage location of the corresponding data after each data writing is completed, reducing the probability that the metadata cannot accurately reflect the storage location of the corresponding data due to frequent data writing. By obtaining the data reading request sent by the read node of the distributed file system, and responding to the updated metadata on the target node including the target reading metadata corresponding to the data requested to be read by the data reading request, the target reading metadata is returned to the read node. Among them, when a reading node in a distributed file system needs to read the data requested by a data reading request, the reading node does not need to obtain the meta information corresponding to the data, that is, the target reading meta information, from the meta information node, but can send a data reading request to enable the target node to respond to the data reading request to determine whether it has stored the target reading meta information. When the updated meta information stored by itself includes the target reading meta information, in response to the updated meta information on the target node including the target reading meta information corresponding to the data requested to be read by the data reading request, the target reading meta information is returned to the reading node, so that the reading node can obtain the target reading meta information, so that the reading node can read the data requested by the data reading request from the corresponding data node according to the target reading meta information. In summary, the technical solution provided by the embodiment of the present disclosure can enable the reading node to obtain the required target reading meta information from the target node when it needs to read the data requested by the data reading request without interacting with the meta information node, so as to complete the data reading of the data requested by the data reading request according to the target reading meta information, thereby reducing the burden of the meta information node, improving the efficiency of data reading, and reducing the cost of data reading without affecting the completion of data reading.

[0101] In one embodiment of the present disclosure, the method further comprises the following steps:

[0102] In response to the meta information on the target node not including the target read meta information, acquiring the target read meta information from the meta information node of the distributed file system;

[0103] Returns target read meta information to the read node.

[0104] In one embodiment of the present disclosure, obtaining target read meta-information from a meta-information node may be accomplished by sending a meta-information retrieval request for indicating the target read meta-information to the meta-information node, so that the meta-information node returns the target read meta-information in response to the meta-information retrieval request; or may be accomplished by receiving the target read meta-information obtained and forwarded from the meta-information node by other nodes, devices or systems.

[0105] In some scenarios, due to some reasons, within a period of time after the write node completes writing data to the data node, the metadata update information of the written data cannot be successfully synchronized from the write node to the target node, resulting in the target node still not storing the metadata corresponding to the data requested to be read by the data read request when the target node has obtained the data read request sent by the read node, resulting in the inability to return the corresponding metadata to the read node, making the read node unable to complete the data reading.

[0106] According to the technical solution provided by the embodiments of the present disclosure, in response to the problems arising in the above-mentioned scenarios, by responding to the fact that the meta-information on the target node does not include the target read meta-information, the target read meta-information is obtained from the meta-information node of the distributed file system, and the target read meta-information is returned to the reading node, it can be ensured that the reading node can read data according to the target read meta-information, thereby improving the success rate of data reading and the user experience.

[0107] In one embodiment of the present disclosure, the metadata update information of the written data includes a data block identifier of the written data block corresponding to the written data, a data block length of the written data block, a file identifier of the written file corresponding to the written data block, and a file length of the written file;

[0108] The meta information on the target node is updated according to the meta information update information of the written data, which can be achieved through the following steps:

[0109] In response to the data block identifiers of the data blocks in the meta-information on the target node including the data block identifiers of all written data blocks in the meta-information corresponding to the written data, updating the meta-information on the target node according to the meta-information corresponding to the written data;

[0110] Or, in response to the data block identifier of the data block in the metadata on the target node not including the data block identifier of the target written data block in the metadata corresponding to the written data, the metadata of the target written data block is synchronized from the metadata node according to the data block identifier of the target written data block, and the metadata on the target node is updated according to the metadata of the target written data block.

[0111] In one embodiment of the present disclosure, the written data block corresponding to the written data can be understood as a data block in which the data contained in the written data is changed during the process of writing the written data. The data block length of the written data block can be understood as the length of the data contained in the written data block. The written file corresponding to the written data block can be understood as the file to which the data contained in the written data block belongs. The file length of the written file can be understood as the length of the data corresponding to the written file.

[0112] For example, Figure 5 A schematic diagram showing metadata update information of written data according to an embodiment of the present disclosure is shown as follows: Figure 5 As shown, the metadata update information 400 of the written data includes fileid 401, fileLength 402, chunkId 403, and chunkLength 404, wherein fileid 401 is the file identifier of the written file corresponding to the written data block, fileLength 402 is the file length of the written file, chunkId 403 is the data block identifier of the written data block, and chunkLength 404 is the data block length of the written data block.

[0113] In one embodiment of the present disclosure, the data block identifier of the data block in the metadata on the target node includes the data block identifiers of all the written data blocks in the metadata corresponding to the written data. It can be understood that when the writing node completes the data writing to the data node, no data block switching occurs, that is, no new data block is generated, and only the content is written to the existing data block. In this case, since the target node has stored the metadata of the existing data block before the data writing is completed, it is only necessary to update the metadata on the target node according to the metadata corresponding to the written data, that is, to update the data block length of the written data block and the file length of the written file (both of which are the contents of the meta-information on the target node that are changed due to the data writing). The updated metadata on the target node includes the metadata of the data block corresponding to the written data after the data writing to the data node is completed.

[0114] In one embodiment of the present disclosure, in response to the fact that the data block identifier of the data block in the metadata on the target node does not include the data block identifier of the target written data block in the metadata corresponding to the written data, it can be understood that the write node has generated a data block switch during the process of completing the data writing to the data node, that is, a new data block has been generated. In this case, since the target node does not store the metadata of the new data block before the data is written, it is necessary to synchronize the metadata of the target written data block from the metadata node according to the data block identifier of the target written data block, and update the metadata on the target node according to the metadata of the target written data block to ensure that the target node stores the metadata of the target written data block.

[0115] Due to the large size of metadata, in some scenarios, the size of metadata may be greater than or equal to 3KB. Frequent transmission of metadata may occupy more network resources, and serializing or deserializing metadata on different nodes or clients will also occupy more processing resources from the perspective of different nodes or clients.

[0116] According to the technical solution provided by the embodiments of the present disclosure, by limiting the metadata update information of the written data to include the data block identifier of the written data block corresponding to the written data, the data block length of the written data block, the file identifier of the written file corresponding to the written data block, and the file length of the written file, the volume of the metadata update information can be made smaller than the volume of the metadata. In response to the data block identifier of the data block in the metadata on the target node including the data block identifiers of all the written data blocks in the metadata corresponding to the written data, the metadata on the target node is updated according to the metadata corresponding to the written data, so that when no data block switching occurs during the data writing process, the updated metadata on the target node can include the metadata of the data block corresponding to the written data; or, in response to the data block identifier of the data block in the metadata on the target node not including the data block identifier of the target written data block in the metadata corresponding to the written data, the metadata of the target written data block is synchronized from the metadata node according to the data block identifier of the target written data block, and the metadata on the target node is updated according to the metadata of the target written data block, so that when a data block switching occurs during the data writing process, the updated metadata on the target node can include the metadata of the data block corresponding to the written data, thereby ensuring that the metadata on the target node includes the metadata of the data block corresponding to the written data while minimizing the size of the metadata, thereby minimizing the network resources occupied by transmitting the metadata and the processing resources occupied when serializing or deserializing the metadata.

[0117] In one embodiment of the present disclosure, when a write node writes data to a data node without data block switching, the metadata update information of the written data includes a data block identifier of a written data block corresponding to the written data, a data block length of the written data block, a file identifier of a written file corresponding to the written data, and a file length of the written file;

[0118] When a data block switch occurs when the write node writes data to the data node, the meta-information update information of the written data includes the meta-information corresponding to the written data.

[0119] In one embodiment of the present disclosure, the write node may detect the data writing process of the data node, and determine whether data block switching occurs in the data writing according to the detection result.

[0120] According to the technical solution provided by the embodiment of the present disclosure, when no data block switching occurs when the write node writes data to the data node, the metadata update information of the written data includes the data block identifier of the written data block corresponding to the written data, the data block length of the written data block, the file identifier of the written file corresponding to the written data, and the file length of the written file; and when data block switching occurs when the write node writes data to the data node, the metadata update information of the written data includes the metadata corresponding to the written data. Under the premise of minimizing the volume of the metadata, the metadata updated according to the metadata update information of the written data on the target node includes the metadata of the data block corresponding to the written data, thereby minimizing the network resources occupied by transmitting the metadata and the processing resources occupied when serializing or deserializing the metadata.

[0121] Figure 6 A flowchart of an information acquisition method according to an embodiment of the present disclosure is shown, and the method is applied to a write node of a distributed file system, such as Figure 6 As shown, the information acquisition method includes step S201.

[0122] In step S201, in response to the completion of writing data to a data node of the distributed file system, the meta information update information of the written data is synchronized to a target node of the distributed file system.

[0123] The target node is used to update the meta-information on the target node according to the meta-information update information of the written data.

[0124] In one embodiment of the present disclosure, the metadata update information of the written data can be understood to include the content that the metadata corresponding to the data on the data node is changed due to the writing of the written data to the data node. Among them, the metadata can be understood to be used to indicate the storage location of the corresponding data. Exemplarily, since the files stored on the distributed file system are usually divided into a series of data blocks, the metadata can be used to indicate the identification of the file to which the corresponding data block belongs, the length of the file to which the corresponding data block belongs, the write position of the corresponding data block in the corresponding file, and on which data node the corresponding data block is stored.

[0125] In one embodiment of the present disclosure, synchronizing the metadata update information of the data to be written to the target node can be understood as the writing node sending the metadata update information to the target node when completing the writing of data to the data node, and can also be understood as forwarding the metadata update information to the target node through other nodes, devices or systems.

[0126] In one embodiment of the present disclosure, the target node can be understood as a node in the distributed file system that is independent of the metadata node. The target node can be used to store metadata of part or all of the data in the corresponding data node. The metadata on the target node can be understood as a subset of the metadata on the metadata node.

[0127] It should be noted that a distributed file system may include multiple target nodes. For example, target nodes corresponding to the locations may be set at different locations, and a reading node at a certain location may send a data reading request to the target node corresponding to the location. In addition, the target node and other nodes in the distributed file system may be carried by different data centers, or the target node and any node in the distributed file system may be carried by the same data center. A data center may be understood as an entity that realizes centralized processing, storage, transmission, exchange, and management of information in a physical space. A data center may include computer equipment, server equipment, network equipment, storage equipment, etc. For example, a target node may be carried by a data center that carries any reading node.

[0128] In one embodiment of the present disclosure, updating the meta-information on the target node according to the meta-information update information of the written data can be understood as replacing or modifying the portion of the meta-information on the target node corresponding to the written data according to the meta-information update information of the written data.

[0129] According to the technical solution provided by the embodiment of the present disclosure, by synchronizing the metadata update information of the written data to the target node of the distributed file system in response to the completion of data writing to the data node of the distributed file system, it can be ensured that each time the data writing is completed, the target node can obtain the change content of the corresponding metadata caused by the data writing, that is, the metadata update information, and update the metadata on the target node according to the metadata update information of the written data, to ensure that the updated metadata on the target node can accurately reflect the storage location of the corresponding data after each data writing is completed, thereby reducing the probability that the metadata cannot accurately reflect the storage location of the corresponding data due to frequent data writing. Among them, when a reading node in a distributed file system needs to read the data requested by a data reading request, the reading node does not need to obtain the meta information corresponding to the data, that is, the target reading meta information, from the meta information node, but can send a data reading request so that the target node can determine whether it has stored the target reading meta information in response to the data reading request. When the updated meta information stored by itself includes the target reading meta information, in response to the updated meta information on the target node including the target reading meta information corresponding to the data requested to be read by the data reading request, the target reading meta information is returned to the reading node, so that the reading node can obtain the target reading meta information, so that the reading node can read the data requested by the data reading request from the corresponding data node according to the target reading meta information. In summary, the technical solution provided by the embodiment of the present disclosure can enable the reading node to obtain the data requested by the data reading request without interacting with the meta information node when it needs to read the data requested by the data reading request, but obtain the required target reading meta information from the target node, so as to complete the data reading of the data requested by the data reading request according to the target reading meta information, thereby reducing the burden of the meta information node, improving the efficiency of data reading, and reducing the cost of data reading without affecting the completion of data reading.

[0130] In one embodiment of the present disclosure, when data is written to a data node without data block switching, the metadata update information of the written data includes a data block identifier of the written data block corresponding to the written data, a data block length of the written data block, a file identifier of the written file corresponding to the written data block, and a file length of the written file;

[0131] When data block switching occurs when writing data to a data node, the meta-information update information of the written data includes the meta-information corresponding to the written data.

[0132] In one embodiment of the present disclosure, the written data block corresponding to the written data can be understood as a data block in which the data contained in the written data is changed during the process of writing the written data. The data block length of the written data block can be understood as the length of the data contained in the written data block. The written file corresponding to the written data block can be understood as the file to which the data contained in the written data block belongs. The file length of the written file can be understood as the length of the data corresponding to the written file.

[0133] In one embodiment of the present disclosure, when data is written to a data node without data block switching, that is, no new data block is generated during the data writing process, and only content is written to the existing data block. In this case, since the target node already stores the metadata of the existing data block before the data writing is completed, it is only necessary to update the metadata on the target node according to the metadata corresponding to the written data, that is, to update the data block length of the written data block and the file length of the written file (both are the metadata on the target node that are changed due to the data writing). The updated metadata on the target node includes the metadata of the data block corresponding to the written data after the data writing to the data node is completed.

[0134] In one embodiment of the present disclosure, when a data block switch occurs when writing data to a data node, that is, a new data block is generated. In this case, since the target node does not store the metadata of the new data block before the data is written, it is necessary to update the metadata on the target node based on the metadata corresponding to the written data, so that the metadata on the updated target node includes the metadata of the data block corresponding to the written data after the data writing to the data node is completed.

[0135] According to the technical solution provided by the embodiment of the present disclosure, when data block switching does not occur when writing data to a data node, the metadata update information of the written data includes the data block identifier of the written data block corresponding to the written data, the data block length of the written data block, the file identifier of the written file corresponding to the written data block, and the file length of the written file; when data block switching occurs when writing data to the data node, the metadata update information of the written data includes the metadata corresponding to the written data. Under the premise of minimizing the volume of the metadata, the metadata updated according to the metadata update information of the written data on the target node includes the metadata of the data block corresponding to the written data, thereby minimizing the network resources occupied by transmitting the metadata and the processing resources occupied when serializing or deserializing the metadata.

[0136] Figure 7A flowchart of an information acquisition method according to an embodiment of the present disclosure is shown, and the method is applied to a reading node of a distributed file system, such as Figure 7 As shown, the information acquisition method includes steps S301, S302, and S303.

[0137] In step S301, a data read request is obtained and sent to a target node of the distributed file system.

[0138] In step S302, the target read meta information returned by the target node is received.

[0139] The target read meta-information is the meta-information corresponding to the data requested to be read by the data read request, and the meta-information is used to indicate the storage location of the corresponding data.

[0140] In step S303, data is read according to the target reading meta-information.

[0141] In one embodiment of the present disclosure, a data read request can be understood as indicating a user request to read the data to be read. Obtaining a data read request can be receiving a data read request sent by a read node, or obtaining the forwarded data read request from other nodes, devices or systems. For example, other nodes in a distributed file system can forward the data read request sent by a read node to a target node.

[0142] In one embodiment of the present disclosure, the target node can be understood as a node in the distributed file system that is independent of the metadata node. The target node can be used to store metadata of part or all of the data in the corresponding data node. The metadata on the target node can be understood as a subset of the metadata on the metadata node.

[0143] It should be noted that a distributed file system may include multiple target nodes. For example, target nodes corresponding to the locations may be set at different locations, and a reading node at a certain location may send a data reading request to the target node corresponding to the location. In addition, the target node and other nodes in the distributed file system may be carried by different data centers, or the target node and any node in the distributed file system may be carried by the same data center. A data center may be understood as an entity that realizes centralized processing, storage, transmission, exchange, and management of information in a physical space. A data center may include computer equipment, server equipment, network equipment, storage equipment, etc. For example, a target node may be carried by a data center that carries any reading node.

[0144] In one embodiment of the present disclosure, the meta information can be understood as being used to indicate the storage location of the corresponding data. For example, since the files stored on the distributed file system are usually divided into a series of data blocks, the meta information can be used to indicate the identifier of the file to which the corresponding data block belongs, the length of the file to which the corresponding data block belongs, the write position of the corresponding data block in the corresponding file, and the data node on which the corresponding data block is stored.

[0145] According to the technical solution provided by the embodiment of the present disclosure, by obtaining a data read request, sending the data read request to a target node, receiving target read meta-information returned by the target node, and reading data according to the target read meta-information, the reading node does not need to interact with the meta-information node when it needs to read the data requested by the data read request, but instead obtains the required target read meta-information from the target node, so as to complete the data reading of the data requested by the data read request according to the target read meta-information, thereby reducing the burden on the meta-information node, improving the efficiency of data reading, and reducing the cost of data reading without affecting the completion of data reading.

[0146] The following reference Figure 8 An information acquisition device according to an embodiment of the present disclosure is described, where the information acquisition device is located at a target node of a distributed file system. Figure 8 A structural block diagram of an information acquisition device 500 according to an embodiment of the present disclosure is shown.

[0147] like Figure 8 As shown, the information acquisition device 500 includes:

[0148] The information synchronization module 501 is configured to synchronize the meta information update information of the written data from the writing node when the writing node of the distributed file system completes writing the data to the data node of the distributed file system;

[0149] The information updating module 502 is configured to update the meta-information on the target node according to the meta-information update information of the written data;

[0150] A first request acquisition module 503 is configured to acquire a data read request sent by a read node of the distributed file system;

[0151] The information return module 504 is configured to return the target read meta-information to the read node in response to the updated meta-information on the target node including the target read meta-information, where the meta-information is used to indicate the storage location of the corresponding data, and the target read meta-information corresponds to the data requested to be read by the data read request.

[0152] According to the technical solution provided by the embodiment of the present disclosure, when the write node of the distributed file system completes the data writing to the data node of the distributed file system, the meta information update information of the written data is synchronized from the write node, so as to ensure that each time the data writing is completed, the target node can obtain the change content caused by the data writing to the corresponding meta information, that is, the meta information update information, and update the meta information on the target node according to the meta information update information of the written data, so as to ensure that the updated meta information on the target node can accurately reflect the storage location of the corresponding data after each data writing is completed, and reduce the probability that the meta information cannot accurately reflect the storage location of the corresponding data due to frequent data writing. By obtaining the data reading request sent by the read node of the distributed file system, and responding to the updated meta information on the target node including the target reading meta information corresponding to the data requested to be read by the data reading request, the target reading meta information is returned to the read node. Among them, when a reading node in a distributed file system needs to read the data requested by a data reading request, the reading node does not need to obtain the meta information corresponding to the data, that is, the target reading meta information, from the meta information node, but can send a data reading request to enable the target node to respond to the data reading request to determine whether it has stored the target reading meta information. When the updated meta information stored by itself includes the target reading meta information, in response to the updated meta information on the target node including the target reading meta information corresponding to the data requested to be read by the data reading request, the target reading meta information is returned to the reading node, so that the reading node can obtain the target reading meta information, so that the reading node can read the data requested by the data reading request from the corresponding data node according to the target reading meta information. In summary, the technical solution provided by the embodiment of the present disclosure can enable the reading node to obtain the required target reading meta information from the target node when it needs to read the data requested by the data reading request without interacting with the meta information node, so as to complete the data reading of the data requested by the data reading request according to the target reading meta information, thereby reducing the burden of the meta information node, improving the efficiency of data reading, and reducing the cost of data reading without affecting the completion of data reading.

[0153] The following reference Fig. 9 An information acquisition device according to an embodiment of the present disclosure is described, where the information acquisition device is located at a write node of a distributed file system. Fig. 9 A structural block diagram of an information acquisition device 600 according to an embodiment of the present disclosure is shown.

[0154] like Fig. 9 As shown, the information acquisition device 600 includes:

[0155] The metadata synchronization module 601 is configured to synchronize the metadata update information of the written data to the target node of the distributed file system in response to completing the writing of data to the data node of the distributed file system, and the target node is used to update the metadata on the target node according to the metadata update information of the written data.

[0156] According to the technical solution provided by the embodiment of the present disclosure, by synchronizing the metadata update information of the written data to the target node of the distributed file system in response to the completion of data writing to the data node of the distributed file system, it can be ensured that each time the data writing is completed, the target node can obtain the change content of the corresponding metadata caused by the data writing, that is, the metadata update information, and update the metadata on the target node according to the metadata update information of the written data, to ensure that the updated metadata on the target node can accurately reflect the storage location of the corresponding data after each data writing is completed, thereby reducing the probability that the metadata cannot accurately reflect the storage location of the corresponding data due to frequent data writing. Among them, when a reading node in a distributed file system needs to read the data requested by a data reading request, the reading node does not need to obtain the meta information corresponding to the data, that is, the target reading meta information, from the meta information node, but can send a data reading request so that the target node can determine whether it has stored the target reading meta information in response to the data reading request. When the updated meta information stored by itself includes the target reading meta information, in response to the updated meta information on the target node including the target reading meta information corresponding to the data requested to be read by the data reading request, the target reading meta information is returned to the reading node, so that the reading node can obtain the target reading meta information, so that the reading node can read the data requested by the data reading request from the corresponding data node according to the target reading meta information. In summary, the technical solution provided by the embodiment of the present disclosure can enable the reading node to obtain the data requested by the data reading request without interacting with the meta information node when it needs to read the data requested by the data reading request, but obtain the required target reading meta information from the target node, so as to complete the data reading of the data requested by the data reading request according to the target reading meta information, thereby reducing the burden of the meta information node, improving the efficiency of data reading, and reducing the cost of data reading without affecting the completion of data reading.

[0157] The following reference Fig.10 An information acquisition device according to an embodiment of the present disclosure is described, where the information acquisition device is located at a reading node of a distributed file system. Fig.10 A structural block diagram of an information acquisition device 700 according to an embodiment of the present disclosure is shown.

[0158] like Fig.10 As shown, the information acquisition device 700 includes:

[0159] The second request acquisition module 701 is configured to acquire a data read request and send the data read request to a target node of the distributed file system;

[0160] The meta information acquisition module 702 is configured to receive target read meta information returned by the target node, where the target read meta information is meta information corresponding to the data requested to be read by the data read request, and the meta information is used to indicate the storage location of the corresponding data;

[0161] The data reading module 703 is configured to read data according to the target reading meta-information.

[0162] According to the technical solution provided by the embodiment of the present disclosure, by obtaining a data read request, sending the data read request to a target node, receiving target read meta-information returned by the target node, and reading data according to the target read meta-information, the reading node does not need to interact with the meta-information node when it needs to read the data requested by the data read request, but instead obtains the required target read meta-information from the target node, so as to complete the data reading of the data requested by the data read request according to the target read meta-information, thereby reducing the burden on the meta-information node, improving the efficiency of data reading, and reducing the cost of data reading without affecting the completion of data reading.

[0163] Fig.11 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0164] The present disclosure also provides an electronic device, such as Fig.11 As shown, the electronic device 800 includes at least one processor 801; and a memory 802 in communication with the at least one processor 801; characterized in that the memory 802 stores instructions that can be executed by the at least one processor 801, and the instructions are executed by the at least one processor 801 to implement the following steps:

[0165] In a first aspect, an embodiment of the present disclosure provides an information acquisition method, wherein the method is applied to a target node of a distributed file system, and the method includes:

[0166] When the write node of the distributed file system completes writing data to the data node of the distributed file system, the metadata update information of the written data is synchronized from the write node;

[0167] Update the meta-information on the target node according to the meta-information update information of the written data;

[0168] Obtain data read requests sent by the read nodes of the distributed file system;

[0169] In response to the updated meta-information on the target node including target read meta-information, the target read meta-information is returned to the read node, the meta-information is used to indicate the storage location of the corresponding data, and the target read meta-information corresponds to the data requested to be read by the data read request.

[0170] In combination with the first aspect, in a first implementation of the first aspect of the present disclosure, the method further includes:

[0171] In response to the meta information on the target node not including the target read meta information, acquiring the target read meta information from the meta information node of the distributed file system;

[0172] Returns target read meta information to the read node.

[0173] In combination with the first aspect and any one of the first implementation manners of the first aspect, in a second implementation manner of the first aspect of the present disclosure, the metadata update information of the written data includes a data block identifier of a written data block corresponding to the written data, a data block length of the written data block, a file identifier of a written file corresponding to the written data block, and a file length of the written file;

[0174] The meta information on the target node is updated according to the meta information update information of the written data, including:

[0175] In response to the data block identifiers of the data blocks in the meta-information on the target node including the data block identifiers of all written data blocks in the meta-information corresponding to the written data, updating the meta-information on the target node according to the meta-information corresponding to the written data;

[0176] Or, in response to the data block identifier of the data block in the metadata on the target node not including the data block identifier of the target written data block in the metadata corresponding to the written data, the metadata of the target written data block is synchronized from the metadata node according to the data block identifier of the target written data block, and the metadata on the target node is updated according to the metadata of the target written data block.

[0177] In combination with the first aspect and any one of the first implementation manners of the first aspect, in a third implementation manner of the first aspect of the present disclosure, when the write node writes data to the data node without data block switching, the metadata update information of the written data includes a data block identifier of the written data block corresponding to the written data, a data block length of the written data block, a file identifier of the written file corresponding to the written data block, and a file length of the written file;

[0178] When a data block switch occurs when the write node writes data to the data node, the meta-information update information of the written data includes the meta-information corresponding to the written data.

[0179] In a second aspect, an information acquisition method is provided in an embodiment of the present disclosure, wherein the method is applied to a write node of a distributed file system, and the method includes:

[0180] In response to completing the writing of data to the data node of the distributed file system, the meta information update information of the written data is synchronized to the target node of the distributed file system, and the target node is used to update the meta information on the target node according to the meta information update information of the written data.

[0181] In combination with the second aspect, in a first implementation of the second aspect of the present disclosure, when no data block switching occurs when writing data to the data node, the metadata update information of the written data includes a data block identifier of the written data block corresponding to the written data, a data block length of the written data block, a file identifier of the written file corresponding to the written data block, and a file length of the written file;

[0182] When data block switching occurs when writing data to a data node, the meta-information update information of the written data includes the meta-information corresponding to the written data.

[0183] In a third aspect, an information acquisition method is provided in an embodiment of the present disclosure, wherein the method is applied to a reading node of a distributed file system, and the method includes:

[0184] Obtain a data read request and send the data read request to the target node of the distributed file system;

[0185] Receive target read meta information returned by the target node, where the target read meta information is meta information corresponding to the data requested to be read by the data read request, and the meta information is used to indicate the storage location of the corresponding data;

[0186] Read data according to the target reading meta information.

[0187] Fig.12 Schematic diagram of the structure of a computer system suitable for implementing the information acquisition method according to an embodiment of the present disclosure. Fig.12 As shown, the computer system 900 includes a processing unit 901, which can perform various processes in the embodiments shown in the above figures according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage part 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the system 900 are also stored. The CPU 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0188] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed, so that a computer program read therefrom is installed into the storage section 908 as needed. It is characterized in that the processing unit 901 can be implemented as a processing unit such as a CPU, a GPU, a TPU, an FPGA, an NPU, etc.

[0189] In particular, according to an embodiment of the present disclosure, the method described above with reference to the accompanying drawings can be implemented as a computer software program. Exemplarily, an embodiment of the present disclosure includes a computer program product, which includes a computer program tangibly contained on a readable medium thereof, and the computer program includes a program code for executing the method in the accompanying drawings. In such an embodiment, the computer program can be downloaded and installed from a network through a communication portion 909, and / or installed from a removable medium 911. Exemplarily, an embodiment of the present disclosure includes a readable storage medium on which computer instructions are stored, and when the computer instructions are executed by a processor, the program code for executing the method in the accompanying drawings is implemented.

[0190] like Figure 1 As shown, the distributed file system includes a write node, at least one data node, a read node, a target node, and a meta-information node;

[0191] The write node is configured to synchronize meta information update information of the written data to the target node in response to completing the writing of data to the data node;

[0192] The target node is configured to synchronize the metadata update information of the written data from the write node when the write node completes writing the data to the data node; update the metadata on the target node according to the metadata update information of the written data; obtain the data read request sent by the read node; in response to the updated metadata on the target node including the target read metadata, return the target read metadata to the read node, the metadata is used to indicate the storage location of the corresponding data, and the target read metadata corresponds to the data requested to be read by the data read request;

[0193] The reading node is configured to obtain a data reading request and send the data reading request to a target node; receive target reading meta information returned by the target node; and read data according to the target reading meta information.

[0194] According to the technical solution provided by the embodiment of the present disclosure, in a distributed file system, when the writing node completes writing data to the data node of the distributed file system, it synchronizes the metadata update information of the written data to the target node, thereby ensuring that each time the data writing is completed, the target node can obtain the change content caused by the data writing to the corresponding metadata, that is, the metadata update information, and the target node updates the metadata on the target node according to the metadata update information of the written data, ensuring that the updated metadata on the target node can accurately reflect the storage location of the corresponding data after each data writing is completed, reducing the probability that the metadata cannot accurately reflect the storage location of the corresponding data due to frequent data writing. The reading node obtains a data reading request and sends a data reading request to the target node. The target node obtains the data reading request and responds to the updated metadata on the target node including the target reading metadata corresponding to the data requested to be read by the data reading request, and returns the target reading metadata to the reading node. Among them, when the reading node needs to read the data requested by the data reading request, the reading node does not need to obtain the meta information corresponding to the data, that is, the target reading meta information, from the meta information node, but can send a data reading request to enable the target node to respond to the data reading request to determine whether it has stored the target reading meta information. When the updated meta information stored by itself includes the target reading meta information, in response to the updated meta information on the target node including the target reading meta information corresponding to the data requested to be read by the data reading request, the target reading meta information is returned to the reading node, so that the reading node can obtain the target reading meta information, so that the reading node can read the data requested to be read by the data reading request from the corresponding data node according to the target reading meta information. In summary, the technical solution provided by the embodiment of the present disclosure can enable the reading node in the distributed system to read the data requested by the data reading request without interacting with the meta information node, but to obtain the required target reading meta information from the target node, so as to complete the data reading of the data requested to be read by the data reading request according to the target reading meta information, thereby reducing the burden of the meta information node, improving the efficiency of data reading, and reducing the cost of data reading without affecting the completion of data reading.

[0195] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each box in the road map or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. Exemplary, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0196] The units or modules involved in the embodiments described in the present disclosure may be implemented by software or hardware. The units or modules described may also be set in a processor, and the names of these units or modules do not constitute limitations on the units or modules themselves in some cases.

[0197] As another aspect, the present disclosure further provides a computer-readable storage medium, which may be a computer-readable storage medium included in the node described in the above embodiment; or a computer-readable storage medium that exists independently and is not installed in a device. The computer-readable storage medium stores one or more programs, and the programs are used by one or more processors to execute the method described in the present disclosure.

[0198] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present disclosure is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. The above exemplary features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) to form a technical solution.

Claims

1. A method for obtaining information, wherein: The method is applied to a target node of a distributed file system, and the method comprises: When the write node of the distributed file system completes writing data to the data node of the distributed file system, synchronizing meta information update information of the written data from the write node; Updating the meta information on the target node according to the meta information update information of the written data; Obtain data read requests sent by the read nodes of the distributed file system; In response to the updated meta-information on the target node including target read meta-information, returning the target read meta-information to the read node, the meta-information being used to indicate a storage location of corresponding data, the target read meta-information corresponding to the data requested to be read by the data read request; Among them, when the write node writes data to the data node without data block switching, the metadata update information of the written data includes the data block identifier of the written data block corresponding to the written data, the data block length of the written data block, the file identifier of the written file corresponding to the written data block, and the file length of the written file; when the write node writes data to the data node and data block switching occurs, the metadata update information of the written data includes the metadata corresponding to the written data.

2. The information acquisition method according to claim 1, wherein: The method further comprises: In response to the meta information on the target node not including the target read meta information, acquiring the target read meta information from the meta information node of the distributed file system; The target read meta information is returned to the read node.

3. The information acquisition method according to claim 1 or 2, wherein: The metadata update information of the written data includes a data block identifier of the written data block corresponding to the written data, a data block length of the written data block, a file identifier of the written file corresponding to the written data block, and a file length of the written file; The updating of the meta-information on the target node according to the meta-information update information of the written data includes: In response to the data block identifiers of the data blocks in the meta-information on the target node including the data block identifiers of all written data blocks in the meta-information corresponding to the written data, updating the meta-information on the target node according to the meta-information corresponding to the written data; Or, in response to the data block identifier of the data block in the metadata on the target node not including the data block identifier of the target written data block in the metadata corresponding to the written data, the metadata of the target written data block is synchronized from the metadata node according to the data block identifier of the target written data block, and the metadata on the target node is updated according to the metadata of the target written data block.

4. A method for obtaining information, wherein: The method is applied to a write node of a distributed file system, and the method comprises: In response to completing the writing of data to the data node of the distributed file system, synchronizing the meta information update information of the written data to the target node of the distributed file system, and the target node is used to update the meta information on the target node according to the meta information update information of the written data; Among them, when no data block switching occurs when data is written to the data node, the metadata update information of the written data includes the data block identifier of the written data block corresponding to the written data, the data block length of the written data block, the file identifier of the written file corresponding to the written data block, and the file length of the written file; when data block switching occurs when data is written to the data node, the metadata update information of the written data includes the metadata corresponding to the written data.

5. A method for obtaining information, wherein: The method is applied to a reading node of a distributed file system, and the method comprises: Obtaining a data read request, and sending the data read request to a target node of the distributed file system; Receive target read meta-information returned by the target node, the target read meta-information being meta-information corresponding to the data requested to be read by the data read request, and the meta-information being used to indicate the storage location of the corresponding data; the target read meta-information is obtained after the target node updates the meta-information on the target node according to the meta-information update information of the written data; the meta-information update information of the written data is synchronously obtained by the target node from the write node of the distributed file system when the write node of the distributed file system completes writing the data to the data node of the distributed file system; Reading data according to the target reading meta information; Among them, when data block switching does not occur when data is written to the data node of the distributed file system, the metadata update information of the written data includes the data block mark of the written data block corresponding to the written data, the data block length of the written data block, the file mark of the written file corresponding to the written data block, and the file length of the written file; when data block switching occurs when data is written to the data node of the distributed file system, the metadata update information of the written data includes the metadata corresponding to the written data.

6. An electronic device, wherein: The method comprises a memory and at least one processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the at least one processor to implement the method steps described in any one of claims 1 to 5.

7. A computer-readable storage medium having computer instructions stored thereon, wherein: When the computer instructions are executed by a processor, the method steps described in any one of claims 1 to 5 are implemented.

8. A computer program product comprising a computer program / instructions, wherein: When the computer program / instructions are executed by a processor, the method steps described in any one of claims 1 to 5 are implemented.

9. A distributed file system, wherein: The distributed file system includes a write node, at least one data node, a read node, a target node and a meta information node; The write node is configured to synchronize meta information update information of the written data to the target node in response to completing the writing of data to the data node; The target node is configured to synchronize the meta-information update information of the written data from the write node when the write node completes writing the data to the data node; and update the meta-information on the target node according to the meta-information update information of the written data; Obtaining a data read request sent by a read node; in response to the updated meta-information on the target node including target read meta-information, returning the target read meta-information to the read node, wherein the meta-information is used to indicate a storage location of corresponding data, and the target read meta-information corresponds to the data requested to be read by the data read request; The reading node is configured to obtain a data reading request and send the data reading request to the target node; and receive target reading meta information returned by the target node; Reading data according to the target reading meta information; Among them, when the write node writes data to the data node without data block switching, the metadata update information of the written data includes the data block identifier of the written data block corresponding to the written data, the data block length of the written data block, the file identifier of the written file corresponding to the written data block, and the file length of the written file; when the write node writes data to the data node and data block switching occurs, the metadata update information of the written data includes the metadata corresponding to the written data.

Citation Information

Patent Citations

  • Efficient replication of distributed storage changes for read-only nodes of a distributed database

    US9507843B1