Optimization Method, System and Device for Transparent Fault Switching of Distributed File System
The method enhances HDFS fault handling by enabling automatic metadata switching and service node backup through DNS access and IP recovery, ensuring continuous data write operations in HDFS systems.
Patent Information
- Application Number
- CN202111215931.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-19
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-10-19
AI Technical Summary
During failover, existing HDFS systems have problems such as complex configuration, limited backup nodes, and inability to provide simultaneous backup of multiple service nodes, resulting in high risk of business interruption.
By starting the cluster database and service monitoring self-start function of the HDFS client service node, the domain name system provides domain name access, realizing automatic switching of metadata failure, and resuming the writing of data block cache information by switching the data service node IP when the service node fails, designing the service node monitoring function to monitor all HDFS service nodes and automatically switch in the event of a failure.
It realizes transparent failover during HDFS protocol access, ensures normal data writing and backup of service nodes, simplifies configuration, and improves system availability and reliability.
Smart Images

Figure CN114020503B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and more specifically, to an optimization method, system and device for transparent failure switching of a distributed file system. Background Art
[0002] The Hadoop Distributed File System (HDFS) is a distributed file system designed to run on commodity hardware. It has many common points with existing distributed file systems. At the same time, its differences from other distributed file systems are also obvious. HDFS is a highly fault-tolerant system suitable for deployment on inexpensive machines. HDFS can provide high-throughput data access and is very suitable for applications on large-scale data sets. HDFS relaxes some POSIX constraints to achieve the purpose of streaming reading of file system data. HDFS was initially developed as the infrastructure for the Apache Nutch search engine project. HDFS is part of the Apache Hadoop Core project.
[0003] During the read and write process of the distributed file system using the HDFS protocol, the support for transparent failures directly affects the availability of the system. In actual on-site services, how to ensure that the failed services do not interrupt has always been a major challenge that the storage industry must solve and face. Currently, in general HDFS scenarios, the primary-backup switching mode is usually used, but it has defects such as complex configuration, limited backup nodes, and the inability to provide the function of simultaneous backup of multiple service nodes, making it extremely inconvenient to use. Summary of the Invention
[0004] Aiming at the above problems, the purpose of the present invention is to provide an optimization method, system and device for transparent failure switching of a distributed file system, which can enable the distributed file system to achieve service failure switching and backup of service nodes during the read and write process using the HDFS protocol.
[0005] To achieve the above object, the present invention is realized through the following technical solutions: An optimization method for transparent failure switching of a distributed file system includes:
[0006] Starting the cluster database and service monitoring self-start function of the HDFS client service node;
[0007] The HDFS client service node provides domain name access to the distributed file system for the client through the domain name system;
[0008] When a node failure occurs in the metadata of the distributed file system, perform automatic metadata failure switching;
[0009] When the HDFS client service starts, it automatically registers node messages with the distributed file system monitor and monitors the HDFS service.
[0010] When it is detected that an HDFS service node fails, the HDFS client restores the write of data block cache information by switching the IP of the data service node.
[0011] Further, the HDFS client service node providing domain name access to the distributed file system through the domain name system to the client further includes:
[0012] The client obtains a virtual IP through the domain name round-robin service.
[0013] The client uses the virtual IP to access the HDFS client service and then accesses the distributed file storage service.
[0014] Further, the automatic metadata failure switchover specifically includes:
[0015] Automatically exclude the faulty node through the virtual IP of the cluster database and complete the virtual IP node switchover.
[0016] When the process of the node itself is abnormal, perform automatic abnormal restart through the self-monitoring function.
[0017] Further, when monitoring the HDFS service, the distributed file system monitor returns the storage node IP information to the HDFS client service node through the client service node list during the data read and write process.
[0018] Further, when the HDFS client service starts, automatically registering node messages with the distributed file system monitor and monitoring the HDFS service specifically includes:
[0019] When the HDFS client service starts, the HDFS client registers its own service node information with the distributed file monitor and regularly sends service heartbeats during the running phase.
[0020] If the HDFS client service exits or the heartbeat connection time exceeds 60 seconds, the corresponding HDFS client service node will be removed from the client service node list.
[0021] Further, the HDFS client restoring the write of data block cache information by switching the IP of the data service node includes:
[0022] When the HDFS client service receives a data block recovery message, obtain the written position information of the corresponding file; calculate the written position information of the current data block according to the written position information of the file and the size of the data block.
[0023] Construct a write cache object based on the position information where the current data block has been written and the data block identifier;
[0024] Write the received data into the constructed cache object.
[0025] Further, the writing the received data into the constructed cache object includes:
[0026] The constructed cache object receives the data and determines whether the data write position information is consistent with the position information where the current data block has been written. If so, write the received data into the constructed cache object; if not, return a write exception message.
[0027] Correspondingly, the present invention also discloses an optimization system for transparent failover of a distributed file system, including:
[0028] A startup unit for starting the cluster database and the service monitoring self-start function of the HDFS client service node; an access unit for providing domain name access to the distributed file system for clients through the domain name system by the HDFS client service node;
[0029] A failover unit for automatically switching metadata failures when node failures occur in the metadata of the distributed file system;
[0030] A monitoring unit for automatically registering node messages to the distributed file system monitor and performing HDFS service monitoring when the HDFS client service starts;
[0031] A recovery unit for the HDFS client to restore the write of the data block cache information by switching the IP of the data service node when it is monitored that a failure occurs in the HDFS service node.
[0032] Further, the recovery unit includes:
[0033] A first position acquisition module for acquiring the position information where the corresponding file has been written when the HDFS client service receives a data block recovery message;
[0034] A second position acquisition module for calculating the position information where the current data block has been written according to the position information where the file has been written and the size of the data block;
[0035] A cache object construction module for constructing a write cache object according to the position information where the current data block has been written and the data block identifier;
[0036] A writing module for writing the received data into the constructed cache object.
[0037] Correspondingly, the present invention discloses an optimization device for transparent failover of a distributed file system, including:
[0038] A memory for storing an optimization program for transparent failover of a distributed file system;
[0039] A processor for implementing the steps of the optimization method for transparent failover of the distributed file system as described in any one of the above when executing the optimization program for transparent failover of the distributed file system.
[0040] Correspondingly, the present invention discloses a readable storage medium, on which an optimization program for transparent failover of a distributed file system is stored, and when the optimization program for transparent failover of the distributed file system is executed by a processor, the steps of the optimization method for transparent failover of the distributed file system as described in any one of the above are implemented.
[0041] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention provides an optimization method, system and device for transparent failover of a distributed file system. For metadata failures, all service nodes can be accessed using a single domain name, enabling HDFS clients to access the HDFS storage service through the domain name, and realizing service switching by means of request resending in case of metadata failures; for transparent failover during the data writing process, a service node monitoring function is designed to monitor all HDFS service nodes. At the same time, multiple service IPs are returned to the HDFS client during data reading and writing, and the service node is automatically switched in case of failures. Meanwhile, the function of restoring the write data block cache information is realized, ensuring the normal writing of data.
[0042] The present invention realizes the failover of the metadata service during HDFS protocol access, the DFS client access is simple, and all service nodes are backup to each other.
[0043] It can be seen that compared with the prior art, the present invention has prominent substantial features and remarkable progress, and the beneficial effects of its implementation are also obvious. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for description in the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0045] Attached Figure 1 is the flowchart of the method of the present invention.
[0046] Attached Figure 2 is the schematic diagram of the data block recovery process of the present invention.
[0047] Attached Figure 3 is the system structure diagram of the present invention.
[0048] In the figure, 1 is the startup unit; 2 is the access unit; 3 is the fault switching unit; 4 is the monitoring unit; 5 is the recovery unit, where 51 is the first location acquisition module, 52 is the second location acquisition module, 53 is the cache object construction module, and 54 is the writing module. Specific implementation mode
[0049] The core of the present invention is to provide an optimized method for transparent fault switching of a distributed file system. In the prior art, in the general scenario of HDFS, the primary and standby switching mode is usually used, but there are defects such as complex configuration, limited backup nodes, and the inability to provide the function of simultaneous backup of multiple service nodes, which is extremely inconvenient to use.
[0050] For the optimized method for transparent fault switching of the distributed file system provided by the present invention, first, start the cluster database and the service monitoring self-start function of the HDFS client service node, and provide domain name access to the distributed file system to the client through the domain name system. When a node failure occurs in the metadata of the distributed file system, automatic switching of the metadata failure is performed; when the HDFS client service is started, node messages are automatically registered to the distributed file system monitor and HDFS service monitoring is performed; when it is monitored that a failure occurs in the HDFS service node, the HDFS client restores the writing of the data block cache information by switching the IP of the data service node.
[0051] It can be seen from this that the present invention can realize the fault switching of services and the backup of service nodes during the read and write process of the distributed file system using the HDFS protocol.
[0052] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation modes. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0053] Embodiment 1:
[0054] As Figure 1 shown, this embodiment provides an optimized method for transparent fault switching of a distributed file system, including the following steps:
[0055] S1: Start the cluster database and the service monitoring self-start function of the HDFS client service node.
[0056] S2: The HDFS client service node provides domain name access to the distributed file system to the client through the domain name system.
[0057] Among them, the customer obtains the virtual IP through the domain name round-robin service; the customer uses the virtual IP to access the HDFS client service and then accesses the distributed file storage service.
[0058] S3: When a node failure occurs in the metadata of the distributed file system, automatic metadata failure switchover is performed.
[0059] When performing automatic metadata failure switchover, first, the faulty node is automatically excluded through the virtual IP of the cluster database to complete the virtual IP node switchover; when the process of the node itself is abnormal, the self-monitoring function is set to automatically pull up the exception.
[0060] S4: When the HDFS client service starts, the node message is automatically registered with the distributed file system monitor and the HDFS service is monitored.
[0061] When the HDFS client service starts, the HDFS client first registers its own service node information with the distributed file monitor and sends service heartbeats regularly during the running phase. If the HDFS client service exits or the heartbeat connection time exceeds 60 seconds, the corresponding HDFS client service node is removed from the client service node list.
[0062] When performing HDFS service monitoring, the distributed file system monitor returns the storage node IP information to the HDFS client service node through the client service node list during the data reading and writing process.
[0063] S5: When it is detected that a failure occurs in the HDFS service node, the HDFS client restores the write of the data block cache information by switching the IP of the data service node.
[0064] Among them, the process of restoring the write of the data block cache information is as follows:
[0065] As Figure 2 shown, when the HDFS client service receives the data block recovery message, it obtains the position information of the corresponding file that has been written. First, the position information of the current data block that has been written is calculated based on the position information of the file that has been written and the size of the data block; then, a write cache object is constructed based on the position information of the current data block that has been written and the data block identifier; finally, the received data is written into the constructed cache object.
[0066] Specifically, when the constructed cache object receives data, it simultaneously judges whether the data write position information is consistent with the position information of the current data block that has been written. If so, the received data is written into the constructed cache object; if not, a write exception message is returned.
[0067] This embodiment provides an optimization method for transparent fault switching of a distributed file system, which can enable all service nodes to access using a single domain name, allowing the HDFS client to access the HDFS storage service through the domain name, and realizing service switching by means of request resending in case of metadata failure. This embodiment designs a service node monitoring function to monitor all HDFS service nodes. Meanwhile, when reading and writing data, it returns multiple service IPs to the HDFS client and automatically switches service nodes in case of failure. At the same time, it realizes the function of restoring the write data block cache information to ensure normal data writing.
[0068] Embodiment 2:
[0069] Based on Embodiment 1, this embodiment also discloses an optimization method for transparent fault switching of a distributed file system, which specifically includes the following steps:
[0070] Step 1: Start the CTDB and the self-starting function of service monitoring for the HDFS Client service node.
[0071] Step 2: The multi-node service of the HDFS Client provides a domain name for customers to access through PDNS. Customers obtain the virtual IP through the domain name (DNS round-robin), and customers use the virtual IP to access the HDFS Client service and then access the distributed file storage service.
[0072] Step 3: When a node failure occurs in the metadata, the faulty node is automatically excluded through the virtual IP of the CTDB to achieve virtual IP node switching; when the process of the node itself is abnormal, set the self-monitoring function, and the service will be automatically restarted when the service is abnormal to achieve automatic switching of metadata failure.
[0073] Step 4: When the HDFS Client service starts, automatically register the node message to the distributed file system monitor. During the reading and writing process, the IP information of multiple storage nodes is returned to the client through the ClientMap table. The client realizes multi-link reading and writing.
[0074] Among them, this embodiment uses the distributed file storage system monitor to implement HDFS service monitoring. The specific process is as follows:
[0075] When the HDFS service starts, it registers its own service node information, mainly including information such as IP, to the distributed file system monitor. Regularly send service heartbeats during the operation stage of the HDFS service. When the HDFS service exits or the heartbeat exceeds 60s, the node is removed from the ClientMap table.
[0076] Step 5: When a data service node fails, the client switches the IP of the data service node, and restores the written block cache information through the HDFSClient, so as to continue receiving written data.
[0077] This step realizes the restoration of the written block cache information through the HDFS Client. The specific process is as follows:
[0078] As Figure 2 shown, when the HDFS Client service receives a block restoration message, it first obtains the position where the file has been written. Then, it calculates the position where the current block has been written based on the written position of the file and the block_size of the file, and constructs a write cache object based on the written position of the block and the blockid. Finally, it starts to receive the written data and determines whether the written position is consistent with the restored position. If not, an exception is thrown; if so, the received data is written normally.
[0079] This embodiment provides an optimization method for transparent failover of a distributed file system, which realizes the HA transparent failover function of the HDFS service during the HDFS protocol access process. Specifically, for metadata failures, all service nodes are accessed using a domain name, and the client accesses the HDFS storage service through the domain name. When a failure occurs, service switching is achieved by resending requests; for transparent failover during the data writing process, a service node monitoring function is designed to monitor all HDFS service nodes. At the same time, multiple service IPs are returned to the client during data reading and writing. When a client failure occurs, the service node is automatically switched, and the function of restoring the write block cache information is also realized, ensuring the normal writing of data.
[0080] Embodiment Three:
[0081] Based on Embodiment One, as Figure 3 shown, the present invention also discloses an optimized system for transparent failover of a distributed file system, including: a startup unit 1, an access unit 2, a failover unit 3, a monitoring unit 4, and a restoration unit 5.
[0082] The startup unit 1 is used to start the cluster database and the self-start function of service monitoring of the HDFS client service node.
[0083] The access unit 2 is used to provide domain name access to the distributed file system for the client through the domain name system at the HDFS client service node.
[0084] The failover unit 3 is used to perform automatic failover of metadata failures when node failures occur in the metadata of the distributed file system.
[0085] The monitoring unit 4 is used to automatically register node messages to the distributed file system monitor and perform HDFS service monitoring when the HDFS client service is started.
[0086] Recovery unit 5 is used to, when it is monitored that a failure occurs in the HDFS service node, the HDFS client recovers the write of the data block cache information by switching the IP of the data service node.
[0087] Among them, the recovery unit 5 specifically includes:
[0088] The first position acquisition module 51 is used to, when the HDFS client service receives a data block recovery message, acquire the position information of the corresponding file that has been written.
[0089] The second position acquisition module 52 is used to calculate the position information of the currently written data block according to the position information of the written file and the size of the data block.
[0090] The cache object construction module 53 constructs a write cache object according to the position information of the currently written data block and the data block identifier.
[0091] The write module 54 is used to write the received data into the constructed cache object.
[0092] This embodiment provides an optimized system for transparent failure switching of a distributed file system, which can access all service nodes using a domain name, enabling the HDFS client to access the HDFS storage service through the domain name, and realizing service switching by using request retransmission in case of metadata failure. This embodiment designs a service node monitoring function to monitor all HDFS service nodes, and at the same time returns multiple service IPs to the HDFS client during data reading and writing, and automatically switches service nodes in case of failure, and at the same time realizes the function of recovering the write data block cache information, ensuring the normal writing of data.
[0093] Embodiment 4:
[0094] This embodiment discloses an optimized device for transparent failure switching of a distributed file system, including a processor and a memory; among them, when the processor executes the optimized program for transparent failure switching of the distributed file system saved in the memory, the following steps are implemented:
[0095] 1. Start the cluster database and service monitoring self-start function of the HDFS client service node.
[0096] 2. The HDFS client service node provides domain name access to the distributed file system for the client through the domain name system.
[0097] 3. When a node failure occurs in the metadata of the distributed file system, perform automatic metadata failure switching.
[0098] 4. When the HDFS client service starts, automatically register node messages to the distributed file system monitor and perform HDFS service monitoring.
[0099] 5. When it is monitored that a failure occurs in the HDFS service node, the HDFS client restores the writing of the data block cache information by switching the IP of the data service node.
[0100] Furthermore, the optimization device for transparent failover of the distributed file system in this embodiment may further include:
[0101] An input interface, which is used to obtain the optimization program for transparent failover of the distributed file system imported from the outside world, and save the obtained optimization program for transparent failover of the distributed file system into the memory. It can also be used to obtain various instructions and parameters transmitted from an external terminal device and transmit them to the processor, so that the processor can perform corresponding processing using the above various instructions and parameters. In this embodiment, the input interface may specifically include, but is not limited to, a USB interface, a serial interface, a voice input interface, a fingerprint input interface, a hard disk reading interface, etc.
[0102] An output interface, which is used to output various data generated by the processor to the terminal device connected thereto, so that other terminal devices connected to the output interface can obtain various data generated by the processor. In this embodiment, the output interface may specifically include, but is not limited to, a USB interface, a serial interface, etc.
[0103] A communication unit, which is used to establish a remote communication connection between the optimization device for transparent failover of the distributed file system and an external server, so that the optimization device for transparent failover of the distributed file system can mount the mirror file into the external server. In this embodiment, the communication unit may specifically include, but is not limited to, a remote communication unit based on wireless communication technology or wired communication technology.
[0104] A keyboard, which is used to obtain various parameter data or instructions input by the user by pressing the key caps in real time.
[0105] A display, which is used to display in real time the relevant information of the short-circuit location process of the server power supply line.
[0106] A mouse, which can be used to assist the user in inputting data and simplify the user's operation.
[0107] Embodiment Five:
[0108] This embodiment also discloses a readable storage medium. The readable storage medium mentioned here includes a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable hard disk, a CD-ROM, or any other form of storage medium well-known in the technical field. The optimization program for transparent failover of the distributed file system is stored in the readable storage medium. When the optimization program for transparent failover of the distributed file system is executed by the processor, the following steps are implemented:
[0109] 1. Start the cluster database and the self-starting function of service monitoring for the HDFS client service node.
[0110] 2. The HDFS client service node provides domain name access to the distributed file system for the client through the domain name system.
[0111] 3. When a node failure occurs in the metadata of the distributed file system, perform automatic metadata failure switching.
[0112] 4. When the HDFS client service starts, automatically register node messages to the distributed file system monitor and perform HDFS service monitoring.
[0113] 5. When it is detected that a failure occurs in the HDFS service node, the HDFS client restores the writing of data block cache information by switching the IP of the data service node.
[0114] In summary, the present invention enables the distributed file system to achieve service failure switching and service node backup during the read and write process using the HDFS protocol.
[0115] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts between the various embodiments, reference can be made to each other. For the methods disclosed in the embodiments, since they correspond to the systems disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0116] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0117] In several embodiments provided by the present invention, it should be understood that the disclosed systems, systems, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the systems or units can be in electrical, mechanical, or other forms.
[0118] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0119] In addition, the functional modules in each embodiment of the present invention can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in one unit.
[0120] Similarly, the processing units in each embodiment of the present invention can be integrated in a functional module, or each processing unit can exist physically, or two or more processing units can be integrated in a functional module.
[0121] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0122] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0123] The above has introduced in detail the optimization method, system, device and readable storage medium for transparent failure switching of the distributed file system provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. An optimization method for transparent failure switching of a distributed file system, characterized in that Including: Starting the cluster database of the HDFS client service node and the self-starting function of service monitoring; The HDFS client service node provides domain name access to the distributed file system for customers through the domain name system; When a node failure occurs in the metadata of the distributed file system, automatic metadata failure switching is performed; When the HDFS client service starts, automatically register node messages to the distributed file system monitor and perform HDFS service monitoring; When it is monitored that a failure occurs in the HDFS service node, the HDFS client restores the writing of data block cache information by switching the IP of the data service node; The HDFS client service node providing domain name access to the distributed file system for customers through the domain name system further includes: Customers obtain virtual IPs through the domain name round-robin service; Customers use the virtual IP to access the HDFS client service and then access the distributed file storage service; The performing of automatic metadata failure switching specifically includes: Automatically excluding the failed node through the virtual IP of the cluster database and completing the virtual IP node switching; When the process of the node itself is abnormal, perform automatic abnormal restart through setting the self-monitoring function; When performing HDFS service monitoring, the distributed file system monitor returns the storage node IP information to the HDFS client service node through the client service node list during the data reading and writing process; The HDFS client restores the writing of data block cache information by switching the IP of the data service node, including: when the HDFS client service receives a data block recovery message, obtaining the written position information of the corresponding file; calculating the written position information of the current data block according to the written position information of the file and the size of the data block; Constructing a write cache object according to the written position information of the current data block and the data block identifier; Writing the received data into the constructed cache object.
2. The optimization method for transparent failure switching of the distributed file system according to claim 1, characterized in that The automatically registering node messages to the distributed file system monitor and performing HDFS service monitoring when the HDFS client service starts specifically includes: When the HDFS client service starts, the HDFS client registers its own service node information to the distributed file monitor and sends service heartbeats regularly during the running phase; If the HDFS client service exits or the heartbeat connection time exceeds 60 seconds, the corresponding HDFS client service node is removed from the client service node list.
3. The optimization method for transparent failure switching of the distributed file system according to claim 2, wherein The writing the received data into the constructed cache object includes: The constructed cache object receives the data and judges whether the data writing position information is consistent with the written position information of the current data block. If so, write the received data into the constructed cache object; if not, return a write exception message.
4. An optimized system for transparent failover of a distributed file system, characterized in that, Including: A starting unit for starting the cluster database of the HDFS client service node and the self-starting function of service monitoring; An access unit for the HDFS client service node to provide domain name access to the distributed file system for customers through the domain name system; A failure switching unit for performing automatic metadata failure switching when a node failure occurs in the metadata of the distributed file system; A monitoring unit, which is used to automatically register node messages to the distributed file system monitor and monitor the HDFS service when the HDFS client service is started; A recovery unit, which is used to enable the HDFS client to restore the write of data block cache information by switching the IP of the data service node when a failure occurs in the HDFS service node; The recovery unit includes: A first position acquisition module, which is used to acquire the position information of the written corresponding file when the HDFS client service receives a data block recovery message; A second position acquisition module, which is used to calculate the position information of the currently written data block according to the position information of the written file and the size of the data block; A cache object construction module, which constructs a write cache object according to the position information of the currently written data block and the data block identifier; A write module, which is used to write the received data into the constructed cache object.
5. An optimized device for transparent failover of a distributed file system, characterized in that, It includes: A memory, which is used to store the optimization program for transparent failover of the distributed file system; A processor, which is used to implement the steps of the optimization method for transparent failover of the distributed file system as described in any one of claims 1 to 3 when executing the optimization program for transparent failover of the distributed file system.
Citation Information
Patent Citations
Distributed file system and failure processing method thereof
CN103019889A
Method and device for data recovery and cluster storage system
CN103064765A
Distributed storage system fault switching method and system, terminal and storage medium
CN112492011A