Calling Method, Device, Equipment and Storage Medium of HDFS API

Through multi-threaded reception and memory shared buffer processing, the efficiency of HDFS API in high concurrency environment is solved, and efficient data transmission and media file transmission is realized, suitable for distributed and high concurrency scenarios.

CN112765119BActive Publication Date: 2025-07-04SHENZHEN IPANEL TECH LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN201911001130.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-10-21
Publication Date
2025-07-04
Estimated Expiration
2039-10-21

AI Technical Summary

Technical Problem

The existing HDFS API cannot be called efficiently in high-concurrency environments, resulting in low data transmission efficiency. Especially in high-concurrency scenarios, HDFS bandwidth and DISK IO become bottlenecks, and the WebHDFS interface limits HDFS bandwidth usage and concurrent request processing performance.

Method used

Multi-threaded client requests are used, interface docking modules are used to connect with the HDFS API, obtain target file information, and send file data in sequence through the memory sharing buffer allocation and queueing mechanism, supporting efficient calls in high-concurrency scenarios.

Benefits of technology

In a high concurrency environment, efficient calls of HDFS API are realized, which improves data transmission efficiency, especially the asynchronous transmission efficiency of large files, and improves the transmission efficiency and playback quality of media files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112765119B_ABST
    Figure CN112765119B_ABST
Patent Text Reader

Abstract

The method, device, equipment and storage medium for calling the HDFS API of the present invention. The server uses multi-threading to receive file reading requests sent by the client, utilizes the interface docking module to dock with the HDFS API to obtain the target file identifier and send it to the client, providing support for distributed and high-concurrency scenarios to achieve efficient call of the HDFS API in a high-concurrency environment; receive the data reading request sent by the client carrying the preset offset value, preset data block size and target file identifier, obtain partial data of the target file accordingly, allocate a corresponding memory shared buffer for storage, send the stored partial data to the client in the queued order, and receive the data reading request sent by the client again until the file reading completion indication sent by the client is received, thereby optimizing the asynchronous transmission of large files on the basis of efficient call of the HDFS API and improving the data transmission efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data communication technology, and more specifically, to a method, apparatus, device and storage medium for calling HDFS API. Background Art

[0002] The Hadoop Distributed File System (HDFS) is an implementation of the Hadoop abstract file system. The Hadoop abstract file system can be integrated with local systems, Amazon S3, etc., and can even be operated through a Web protocol (webhsfs). The files of HDFS are distributed on cluster machines, and replicas are provided at the same time to ensure fault tolerance and reliability.

[0003] Currently, the Java / C API provided by HDFS is a return handle interface, that is, a blocking interface. For example, if 10M of data needs to be read, it must wait until all the data is read before returning. However, this will cause the calling thread to block, and nginx has a single-threaded working model within a process and cannot complete high-concurrency calls. Currently, a solution is to use Node.js and WebHDFS REST API to access Hadoop HDFS data, aiming to enable applications outside the HDFS cluster to not only not install Hadoop and Java libraries, but also access the HDFS cluster through a popular REST-style interface. However, the interfaces provided by WebHDFS are very limited, and HTTP GET / PUT needs to be used for uploading and downloading. Moreover, the reading and writing of HDFS files by WebHDFS will be redirected to the DataNode where the file is located and will completely occupy the bandwidth of HDFS. It can be seen that through WebHDFS for transit, the use of HDFS bandwidth will be severely restricted, resulting in poor processing performance and efficiency of concurrent requests. Moreover, since HDFS adopts a cluster deployment solution, bandwidth and DISK IO will also become bottlenecks, making the customization development difficult.

[0004] Therefore, there is an urgent need for an efficient calling scheme for HDFS API to achieve efficient calling of HDFS API in a high-concurrency environment and improve data transmission efficiency. Summary of the Invention

[0005] In view of this, the present invention provides a method, apparatus, device and storage medium for calling HDFS API to solve the technical problem that the current HDFS API cannot be efficiently called in a high-concurrency environment.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A method for calling the HDFS API, which is applied to the server side. The calling method includes:

[0008] Receiving file reading requests sent by the client using multiple threads. The file reading requests carry the target file path;

[0009] According to the target file path, using the interface docking module to obtain target file information from HDFS. The interface docking module is used to dock with the HDFS API;

[0010] Sending the target file information to the client. The target file information includes the target file identifier and the target file size;

[0011] Receiving data reading requests sent by the client using multiple threads. The data reading requests carry a preset offset value, a preset data block size, and the target file identifier. The data reading requests are sent by the client after receiving the target file information;

[0012] According to the preset offset value, the preset data block size, and the target file identifier, obtaining partial data of the target file and allocating a corresponding memory shared buffer;

[0013] Writing the partial data of the target file into the memory shared buffer for queuing;

[0014] Sending the partial data of the target file to the client in the queued order, and re-executing the step of receiving the data reading requests sent by the client until a file reading completion indication sent by the client is received.

[0015] Preferably, the receiving the file reading requests sent by the client using multiple threads includes:

[0016] When it is monitored that a client sends a request, performing multi-threaded scheduling through the operating system and allocating corresponding processing threads to receive and process the file reading requests sent by the client.

[0017] Preferably, the receiving the file reading requests sent by the client using multiple threads includes:

[0018] When there is an idle TCP long connection between the server and the client, based on the idle TCP long connection, receiving the file reading requests sent by the client using multiple threads.

[0019] Preferably, data transmission between the server and the client is based on sockets. The calling method further includes:

[0020] Monitor the heartbeat packets sent by the client based on the socket;

[0021] When the heartbeat packet is not monitored within a preset time period, close the socket.

[0022] Preferably, the calling method further includes:

[0023] Configure the ratio between the request processing speed of the processing thread and the data reading and writing speed of the HDFS file reading and writing interface to 1:8.

[0024] Preferably, the obtaining of the target file information from HDFS by using the interface docking module according to the target file path includes:

[0025] Obtain the mapping relationship between the preset file path and the file identifier;

[0026] According to the target file path and the mapping relationship, use the interface docking module to obtain the target file identifier from HDFS;

[0027] Obtain the target file size according to the target file identifier.

[0028] Preferably, the client is built based on OpenResty;

[0029] The server is built using a C program;

[0030] The interface docking module is built based on the SDK API.

[0031] A calling device for HDFS API, which is applied to the server, and the calling device includes:

[0032] A client request receiving unit, which is used to receive the file reading request sent by the client by using multi-threading, and the file reading request carries the target file path;

[0033] A file information obtaining unit, which is used to obtain the target file information from HDFS by using the interface docking module according to the target file path; the interface docking module is used to dock with the HDFS API;

[0034] A file information sending unit, which is used to send the target file information to the client, and the target file information includes the target file identifier and the target file size;

[0035] The client request receiving unit is further used to receive the data reading request sent by the client by using multi-threading, and the data reading request carries a preset offset value, a preset data block size and the target file identifier; the data reading request is sent by the client after receiving the target file information according to the target file information;

[0036] A file data processing unit, configured to obtain partial data of a target file according to the preset offset value, the preset data block size, and the target file identifier, and allocate a corresponding memory shared buffer;

[0037] A file data buffering unit, configured to write the partial data of the target file into the memory shared buffer for queuing;

[0038] A file data sending unit, configured to send the partial data of the target file to the client in the queued order, and trigger the client request receiving unit to re-execute the step of receiving the data reading request sent by the client until a file reading completion indication sent by the client is received.

[0039] A device for calling an HDFS API, including a processor and a memory;

[0040] Wherein, the memory is used to store an interface calling program;

[0041] The processor is used to call the interface calling program stored in the memory to execute the foregoing method for calling the HDFS API.

[0042] A computer-readable storage medium stores an interface calling program, and the interface calling program implements the foregoing method for calling the HDFS API when called by a computer device.

[0043] As can be seen from the above technical solutions, for the method, device, equipment, and storage medium for calling the HDFS API provided by the present invention, the server uses multi-threading to receive file reading requests sent by the client, uses an interface docking module to dock with the HDFS API to obtain a target file identifier and send it to the client, provides full support for distributed and high-concurrency scenarios, and can efficiently call the HDFS API in a high-concurrency environment; then receives a data reading request sent by the client carrying a preset offset value, a preset data block size, and a target file identifier, obtains partial data of the target file accordingly, and allocates a corresponding memory shared buffer for storage, sends the stored partial data to the client in the queued order, and re-receives the data reading request sent by the client until a file reading completion indication sent by the client is received, thereby optimizing the asynchronous transmission of large files on the basis of the efficient call of the HDFS API and improving the data transmission efficiency. Description of the Drawings

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0045] Figure 1 It is a schematic architecture diagram of the calling scheme of the HDFS API provided by the embodiment of the present invention;

[0046] Figure 2 It is a flowchart of the calling method of the HDFS API provided by the embodiment of the present invention;

[0047] Figure 3 It is a schematic diagram of the thread allocation method provided by the embodiment of the present invention;

[0048] Figure 4 It is an information interaction flowchart of the calling method of the HDFS API provided by the embodiment of the present invention;

[0049] Figure 5 It is a schematic structural diagram of the calling device of the HDFS API provided by the embodiment of the present invention;

[0050] Figure 6 It is a schematic structural diagram of the calling device of the HDFS API provided by the embodiment of the present invention. Detailed implementation manners

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0052] OpenResty is a high-performance Web platform based on Nginx and Lua, which integrates a large number of excellent Lua libraries, third-party modules and most dependencies internally, and can be used to easily build dynamic Web applications, Web services and dynamic gateways that can handle extremely high concurrency and have extremely high scalability.

[0053] OpenResty makes full use of Nginx's non-blocking I / O model, while Nginx uses a master-worker model. A master process manages multiple worker processes. Basic event processing is placed in the worker. The master is responsible for some global initialization and management of the worker. In OpenResty, each worker uses a LuaVM. When a request is assigned to a worker, a coroutine will be created in this LuaVM. The data between coroutines is isolated, and each coroutine has an independent global variable _G. Therefore, with the help of Nginx's event-driven model and non-blocking I / O, high-performance Web applications can be implemented.

[0054] For asynchronous non-blocking, when a thread calls a situation such as waiting for I / O, it will first handle other tasks instead of blocking here, waiting for the I / O to be ready, and then executing the current task.

[0055] epoll is an improved poll in the Linux kernel for processing large batches of file descriptors. It is an enhanced version of the multiplexed IO interface select / poll under Linux. It can significantly improve the system CPU utilization when only a small number of programs are active among a large number of concurrent connections. Another reason is that when obtaining events, it does not need to traverse the entire set of monitored descriptors, but only needs to traverse the set of descriptors that are asynchronously awakened by kernel IO events and added to the Ready queue.

[0056] The C10K problem refers to the inability of many servers to serve about 10,000 concurrent connections in their initial state. This is certainly related to the resource consumption of each service and the hardware configuration of the server, but in many cases it is limited by the default configuration of Linux and the selection of software stack. If there is no problem with the hardware configuration, the C10K problem occurs on a high-performance server. In many cases, it is related to the configuration and software stack, such as the maximum number of open files, the number of socket ports, and the IO basic stack.

[0057] The existing Hadoop file system (HDFS) user-facing interface is a blocking interface. For example, to read 1024 bytes, you must wait until all bytes are completed before returning, which will cause the calling thread to be blocked. Nginx is a working model with only a single thread within a process, and thus cannot complete high-concurrency calls.

[0058] The method for calling the HDFS API provided by the present invention is actually also a method for implementing file operations by calling the HDFS API interface in a high-concurrency environment of OpenResty. Through the HDFS proxy interface protocol server (i.e., the protocol server) provided by the present invention, it is used to receive high-concurrency requests, and internally adopts a multi-threaded request receiving and epoll multiplexing I / O development model, providing a unified process handling logic for different file operation interfaces, and using decoupled modular development with the HDFS API interface.

[0059] The technical implementation of the present invention mainly includes three parts, such as a protocol client (abbreviation: "client"), a protocol server (abbreviation: "server"), and an SDK interface docking module (abbreviation: "interface docking module"). Moreover, as Figure 1 shown, the client is built based on OpenResty; the server is built using a C program; the interface docking module is built based on the SDK API and is mainly responsible for docking with the HDFS API and the File API.

[0060] To improve the security and efficiency of communication and meet the need to customize fields independently, the communication protocol between the client and the server can adopt a TCP protocol with a "private data header + data body" because TCP is a streaming protocol and has a retransmission mechanism, which can achieve high availability of data communication. Among them, the two communication parties are the client and the server, belonging to a typical C / S architecture model. Since the TCP protocol is a streaming protocol, that is, it is transmitted to the receiver in the form of a byte stream without the inherent concept of "message" or "message boundary", when receiving data, first read the fixed data header, extract the length of its data body, and then read the data with the length of the data body, reading the data in two times.

[0061] Please refer to Figure 2 , Figure 2 which is a flowchart of a method for calling the HDFS API provided by an embodiment of the present invention.

[0062] The method for calling the HDFS API provided in this embodiment is applied to the server, that is, the method for calling the HDFS API executed from the perspective of the server.

[0063] As Figure 2 shown, the calling method includes:

[0064] S101: Receive the file reading request sent by the client using multiple threads.

[0065] The file reading request carries the target file path.

[0066] Before the client sends a file reading request to the server, it also needs to initiate a data connection establishment request first. After a data connection request is established between the client and the server, the file reading request is then initiated.

[0067] When it is monitored that a client sends a request, multi-threaded scheduling can be performed through the operating system to allocate corresponding processing threads to receive and process the file reading request sent by the client.

[0068] Since the number of concurrent request connections is large, and the waiting queue processed by each thread is limited, if a connection request waits too long to be processed, a timeout will occur. Therefore, after the client initiates a connection request, the time for the server to receive and process the connection request cannot be too long. Generally, if it exceeds about 2 - 3s, it is determined as a timeout. Therefore, in order to match the processing speed of the HDFS interface request and the network I / O speed, it is necessary to enable multi-threading on the server to receive and process requests, so as to reduce the waiting time during high concurrency of requests.

[0069] When the server starts the program, it first listens in the main thread, and then passes the corresponding file handle fd (file description) to all child threads while starting a new program. When a client initiates a request, multi-threaded scheduling is then performed by the operating system (as Figure 3 shown), and the upper-layer application does not participate. Among them, the main reason for using the operating system for multi-threaded scheduling is that the processing time, data bandwidth, and number of load links for each link are all different, and it is difficult to evenly distribute processing threads for load balancing.

[0070] The connection thread CID (connect id, connection identifier) is a globally unique ID value used to identify the session value. Before the client sends a request, it needs to perform a three-way TCP handshake with the server to establish a TCP connection. Generally, each time it takes a total of three handshakes to send and receive more than a dozen TCP packets to establish a connection, which is a significant resource consumption for thousands of connection requests. Therefore, the present invention also adopts the TCP multiplexing technology. When the client requests a connection, it first detects whether there is an idle long connection between the client and the server. If not, a new connection is established. If so, this long connection is directly reused, thereby avoiding the delay caused by establishing a new TCP connection and the consumption of server resources.

[0071] That is to say, when there is an idle TCP long connection between the server and the client, based on the idle TCP long connection, multi-threads are used to receive the file reading request sent by the client; when there is no idle TCP long connection between the server and the client, a new TCP connection is established, and based on the new TCP connection, multi-threads are used to receive the file reading request sent by the client.

[0072] During the usage process, the socket fds may be the same, but the CID values do not change and represent unique connection identifiers. The socket itself is an integer value. Generally, the number of sockets that can be allocated by a system process is limited. However, when the maximum value is reached, it will cycle through the minimum value and the unused values for allocation. The CID value is a value that identifies a session and is also a global incrementing value within the range of a 64-bit integer. It increments by 1 each time a new link is created and generally cannot be used up.

[0073] S102: Obtain the target file information from HDFS using the interface docking module according to the said target file path.

[0074] The said interface docking module is used to dock with the HDFS API.

[0075] The said target file information includes the target file identifier and the target file size. Step S102 may specifically include:

[0076] A1. Obtain the mapping relationship between the preset file path and the file identifier;

[0077] A2. According to the said target file path and the mapping relationship, use the interface docking module to obtain the target file identifier from HDFS;

[0078] A3. Obtain the target file size according to the said target file identifier.

[0079] Among them, the target file identifier can be the virtual file identifier VFID (Virtual File Identify). The mapping relationship between the file path and the VFID is established in advance. The file path is a string with a maximum length of 1024 bytes. Each time information interaction is required, header information needs to be transmitted to identify the unique link. Therefore, transmitting this long string each time will waste a large amount of bandwidth. The smaller the proportion of this bandwidth relative to the total bandwidth required for file transmission, the better. So, it is necessary to minimize the necessary transmission information content as much as possible. Therefore, in the present invention, the mapping relationship between the file path and the VFID is established. The VFID is a 64-bit integer value. For the sake of search efficiency, the present invention adopts the HASH algorithm and uses the VFID as the Hash_Key value among them.

[0080] The VFID is constructed by <machine serial number + process PID + current time + total accumulated value> and is a globally unique value. A new VFID is generated each time a file path is opened and deleted when the file is closed. The purpose is to ensure that only the VFID value needs to be transmitted during the file transmission process without repeating the transmission of the file path.

[0081] S103: Send the target file information to the client, where the target file information includes a target file identifier and a target file size.

[0082] S104: Receive the data reading request sent by the client using multiple threads. The data reading request carries a preset offset value, a preset data block size, and the target file identifier.

[0083] The data reading request is sent by the client after receiving the target file information based on the target file information.

[0084] S105: Obtain partial data of the target file according to the preset offset value, the preset data block size, and the target file identifier, and allocate a corresponding memory shared buffer.

[0085] S106: Write the partial data of the target file into the memory shared buffer for queuing.

[0086] The asynchronous non-blocking model for processing requests is a passive notification model. Add the corresponding fd to the epoll queue, and wake up the status of the fd through a centralized function trigger mechanism, such as read / write, exception, etc. Generally, a small amount of data requests and responses are relatively simple, but a special mechanism is required for large data volume read / write.

[0087] First, due to memory resource limitations, it is impossible to read an entire file of several gigabytes into memory and then send it slowly. Even if there is enough memory, data blocks are read into memory on demand and then sent, and new data blocks are requested after sending is completed.

[0088] In the video playback scenario, since the files read are basically large files for video playback, and users may fast forward or drag the progress bar at any time, it is more appropriate to read in chunks. For example, if the preset data block size is 4M, then data is transmitted in data blocks of a maximum of 4M. The client requests blocks with the VFID, preset offset value, and preset data block size. After receiving the data reading request, the server reads the corresponding data block into the memory shared buffer for the client to read. The client can send a new data reading request next time until all data blocks are read. On the one hand, this can achieve the repeated use of block memory, reduce memory fragmentation and efficiency losses caused by multiple allocations and releases. On the other hand, by using the method of dividing and conquering, fragmenting large files for processing can improve data transmission efficiency and data transmission flexibility.

[0089] The following introduces the process of data reading and data writing on the server side:

[0090] I. Process of data reading:

[0091] ① Allocate a shared buffer with a minimum block size of 4M. Inside the buffer is a doubly linked list structure.

[0092] ② According to the buffer size allocated by the doubly linked list, read data from the back-end file system to fill all blocks of the circular linked list and fill the file data.

[0093] ③ Update the file offset address and the length of the read data and return them to the client.

[0094] ④ Mount the data in the epoll send queue of this connection socket, and release the resources of the shared buffer after sending is completed.

[0095] II. Process of data writing:

[0096] ① Allocate a shared buffer with a minimum block size of 4M. Inside the buffer is a doubly linked list structure.

[0097] ② Mount the socket to the epoll receive queue, and receive the completion notification for the WR (write-read) thread according to the body length in the protocol header.

[0098] ③ All blocks of the circular linked list, and write the file to the corresponding file system.

[0099] ④ Update the current offset address of the received file and the length of the received data, and return them to the client.

[0100] ⑤ Recycle the resources of the shared buffer.

[0101] S107: Send part of the data of the target file to the client in the queued order.

[0102] S108: Determine whether a file read completion indication sent by the client is received. If not, continue to execute step S104; if so, execute step S109.

[0103] S109: Close the file data transfer.

[0104] Send part of the data of the target file to the client in the queued order, and re-execute the step of receiving the data read request sent by the client until a file read completion indication sent by the client is received.

[0105] The method for calling the HDFS API provided in this embodiment enables the server to receive file reading requests sent by the client through multiple threads, and uses the interface docking module to dock with the HDFS API to obtain the target file identifier and send it to the client, providing full support for distributed and high-concurrency scenarios and enabling efficient call of the HDFS API in a high-concurrency environment. Then, it receives the data reading request sent by the client carrying the preset offset value, preset data block size, and target file identifier, obtains partial data of the target file accordingly, allocates a corresponding memory shared buffer for storage, sends the stored partial data to the client in the queued order, and receives the data reading request sent by the client again until it receives the file reading completion indication sent by the client. Thus, on the basis of the efficient call of the HDFS API, the asynchronous transmission of large files is optimized, improving the data transmission efficiency, especially the transmission efficiency of large media files and the quality of media playback.

[0106] Moreover, through the protocolization of the interface, it can not only dock with the HDFS API, but also dock with the disk file operation API, enabling the technical solution of the present invention to be adapted and run on different platforms, providing support for distributed and high-concurrency scenarios, and being applicable to various application scenarios that need to solve the C10K problem.

[0107] Please refer to Figure 4 , Figure 4 which is the information interaction flowchart of the method for calling the HDFS API provided in the embodiment of the present invention.

[0108] The method for calling the HDFS API provided in this embodiment describes the information interaction process between the client, the server, and the HDFS system from the perspective of the entire interaction system.

[0109] As Figure 4 shown, the calling method includes:

[0110] S201: Open the file and pass in the target file path.

[0111] The client first sends the target file path to be read to the client protocol interface. After encapsulation, the data and length to be sent are returned. The client sends the encapsulated data (i.e., the file reading request) to the server.

[0112] S202: Query the target file information in HDFS.

[0113] The server docks with the HDFS API through the interface docking module to query the target file information in HDFS.

[0114] S203: Reply with the query result.

[0115] The HDFS system returns the query result to the server.

[0116] S204: Determine different feedback information according to different query results.

[0117] The server determines different feedback information according to different query results. The query results include found and not found.

[0118] The server receives the request and establishes a connection, analyzes the data header of the request to confirm the target file to be read, confirms the status and attribute information of the target file. If the target file does not exist, it returns the feedback information that the target file does not exist; if the target file exists, it returns information such as the VFID and file size of the target file (i.e., the target file information).

[0119] S205: Send the feedback information.

[0120] The server sends the feedback information to the client. In the case of finding the target file, the feedback information sent by the server to the client is the target file information, which may include the VFID and file size.

[0121] S206: If it is determined that the target file exists, continue; otherwise, directly close the connection.

[0122] S207: Request to read the data VFID, offset value, and data block size.

[0123] The client receives the response header information from the server, parses to obtain the VFID value and total length of the file. The client encapsulates according to the preset offset value plus the preset data block size and sends the data read request again. This data block size is generally 4M, that is, the target file is divided into 4M blocks (i.e., partial data of the target file), and only the last block may be transmitted with a length less than the complete block.

[0124] S208: Fill the shared buffer.

[0125] The server allocates a shared buffer for the requested data block.

[0126] S209: Read the data according to the offset value.

[0127] The server reads the requested data block from the HDFS system according to the offset value.

[0128] S210: Reply with data that meets the preset data block size.

[0129] The server receives the data read request, reads the data from the file service source end to the internal shared buffer according to the preset data block size and offset value, and finally hangs the data on the epoll waiting transmission queue. When the data can be sent, it is sent to the client.

[0130] The client waits for data reception and reads the received data and sends it to the required backend, such as a player.

[0131] S211: Repeat steps S207 - S210.

[0132] S212: Send an indication that the file reading is completed or the required size is reached.

[0133] After the client obtains the complete data of the target file, it will send an indication that the file reading is completed to the server, or an indication that the required size is reached.

[0134] S213: Close the file.

[0135] After the server receives the indication that the file reading is completed sent by the client, or the indication that the required size is reached, it means that all the data of the target file has been transmitted. Finally, it closes the data connection and the target file to end the process.

[0136] In an example, data transmission can be performed between the server and the client based on a socket. Correspondingly, the present invention can also monitor the heartbeat packet sent by the client based on the socket. When the heartbeat packet is not monitored within a preset time period, the socket is closed.

[0137] This example can solve the problem of abnormal socket closure. In any network communication process, there is no mechanism to monitor the closure of the peer socket. Therefore, only through an active heartbeat mechanism, continuously send its own heartbeat data information to tell the other party that it is still alive. If the heartbeat packet is not monitored within a period of time, it is considered that the peer has closed. In this way, on the one hand, the active queue that needs to be maintained can be reduced, and on the other hand, the number of unused or invalid connections can be closed, thereby improving the processing efficiency of the program and being beneficial to querying when the program has problems.

[0138] Each transmission channel has its own timeout time. As long as communication occurs on these channels, such as reading or writing data, its heartbeat time is updated, and this heartbeat time is used as the start time. If no heartbeat data information from the other party is received after a period of time, this channel is processed, such as closing this link and all resources on this link.

[0139] In another example, the calling method may further include: configuring the ratio between the request processing speed of the processing thread and the data reading and writing speed of the HDFS file reading and writing interface to 1:8.

[0140] This example mainly addresses the matching problem between the I / O speed of the HDFS file read / write interface and the request processing speed of the processing threads. According to the maximum output bandwidth capacity of a single HDFS server, which is 4000MB - 5000MB, through testing, the comparison between a processing thread and the WR thread for data shows that a performance-optimal configuration is achieved with a ratio of 1:8, that is, the best effects in terms of throughput, CPU, memory, and bandwidth output can be ensured, and the full utilization of existing resources can be guaranteed.

[0141] In the present invention, the protocol header and protocol body structures can be as follows:

[0142] Table 1 Protocol Header

[0143]

[0144] The protocol body has a variable length, which is identified by pack_length, and its maximum length is 4MB bytes.

[0145] The method for calling the HDFS API provided in this embodiment describes the information interaction process among the three parties from the perspectives of the server, the client, and the HDFS system. The server uses multi-threading to receive file read requests sent by the client, utilizes the interface docking module to dock with the HDFS API to obtain the target file identifier and send it to the client, providing full support for distributed and high-concurrency scenarios, and being able to efficiently call the HDFS API in a high-concurrency environment; and on the basis of the efficient call of the HDFS API, it optimizes the asynchronous transmission of large files, improving the data transmission efficiency; moreover, it also establishes a complete heartbeat mechanism for the abnormal closure of the socket, and determines which connection sockets need to be closed and eliminates the influence of invalid sockets through the heartbeat mechanism.

[0146] The embodiment of the present invention also provides a device for calling the HDFS API. The device for calling the HDFS API is used to implement the method for calling the HDFS API provided in the embodiment of the present invention. The technical content of the device for calling the HDFS API described below can be correspondingly referred to the technical content of the method for calling the HDFS API described above.

[0147] Please refer to Figure 5 , Figure 5 which is the structural schematic diagram of the device for calling the HDFS API provided in the embodiment of the present invention.

[0148] The device for calling the HDFS API provided in this embodiment is applied to the server.

[0149] As Figure 5As shown in the figure, the calling device includes: a client request receiving unit 10, a file information obtaining unit 20, a file information sending unit 30, a file data processing unit 40, a file data buffering unit 50, and a file data sending unit 60.

[0150] The client request receiving unit 10 is configured to receive a file reading request sent by a client using multi-threading, where the file reading request carries a target file path.

[0151] The file information obtaining unit 20 is configured to obtain target file information from HDFS using an interface docking module according to the target file path; the interface docking module is used to dock with the HDFS API.

[0152] The file information sending unit 30 is configured to send the target file information to the client, where the target file information includes a target file identifier and a target file size.

[0153] The client request receiving unit 10 is further configured to receive a data reading request sent by the client using multi-threading, where the data reading request carries a preset offset value, a preset data block size, and the target file identifier; the data reading request is sent by the client after receiving the target file information.

[0154] The file data processing unit 40 is configured to obtain partial data of the target file according to the preset offset value, the preset data block size, and the target file identifier, and allocate a corresponding memory shared buffer.

[0155] The file data buffering unit 50 is configured to write the partial data of the target file into the memory shared buffer for queuing.

[0156] The file data sending unit 60 is configured to send the partial data of the target file to the client in the queued order, and trigger the client request receiving unit to re-execute the step of receiving the data reading request sent by the client until a file reading completion indication sent by the client is received.

[0157] The calling device of the HDFS API provided in this embodiment uses multiple threads on the server side to receive file reading requests sent by the client, docks with the HDFS API through the interface docking module to obtain the target file identifier and send it to the client, providing full support for distributed and high-concurrency scenarios, and being able to efficiently call the HDFS API in a high-concurrency environment; then it receives the data reading request sent by the client carrying the preset offset value, preset data block size, and target file identifier, obtains a part of the data of the target file accordingly, and allocates a corresponding memory shared buffer for storage, sends the stored part of the data to the client in the queued order, and receives the data reading request sent by the client again until it receives the file reading completion indication sent by the client, thereby optimizing the asynchronous transmission of large files on the basis of the efficient call of the HDFS API, improving the data transmission efficiency, especially improving the transmission efficiency of large media files and the quality of media playback.

[0158] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of the calling device of the HDFS API provided by an embodiment of the present invention.

[0159] As Figure 6 shown, an embodiment of the present invention also provides a calling device for the HDFS API. The calling device may include: a processor 1, a communication interface 2, a memory 3, and a communication bus 4; wherein, the processor 1, the communication interface 2, and the memory 3 complete mutual communication through the communication bus 4.

[0160] The communication interface 2 may be an interface of a communication module, such as an interface of a GSM module;

[0161] The memory 3 is used to store an interface calling program;

[0162] The processor 1 is used to call the interface calling program stored in the memory 3 to execute the foregoing calling method of the HDFS API;

[0163] The program may include program code, and the program code includes operation instructions of the processor.

[0164] The processor may be a central processing unit CPU, or a specific integrated circuit ASIC, or one or more integrated circuits configured to implement the embodiments of the present application.

[0165] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0166] An embodiment of the present invention provides a computer-readable storage medium, in which an interface call program is stored, and when the interface call program is called by a computer device, the foregoing method for calling the HDFS API is implemented.

[0167] The device for calling the HDFS API in this article can be a server, a PC, etc.

[0168] This application also provides a computer program product, which, when executed on a computer device, is adapted to execute a program for initializing the steps of the foregoing method for calling the HDFS API.

[0169] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0170] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Based on such an understanding, all or part of the technical solution of this application that contributes to the background technology can be embodied in the form of a software product, and this computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0171] The various embodiments in this specification are described in a progressive manner, and the key point of each embodiment is the difference from other embodiments. The various embodiments can be combined with each other, and the same or similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0172] In this article, specific examples are used to elaborate on the principles and implementation manners of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application. At the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. To sum up, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for calling the HDFS API, characterized in that, Applied to the server side, the calling method includes: Receiving a file reading request sent by a client using multiple threads, where the file reading request carries a target file path; Obtaining the mapping relationship between the preset file path and the file identifier; According to the target file path and the mapping relationship, using the interface docking module to obtain the target file identifier from HDFS; According to the target file identifier, obtaining the target file size; the interface docking module is used to dock with the HDFS API; Sending the target file information to the client, where the target file information includes the target file identifier and the target file size; Receiving a data reading request sent by the client using multiple threads, where the data reading request carries a preset offset value, a preset data block size, and the target file identifier; the data reading request is encapsulated and sent by the client after receiving the target file information according to the preset offset value plus the preset data block size, and is used to split the target file into partial data according to the preset data block size; the target file identifier is a virtual file identifier VFID, and the VFID is used as the Hash_key value in the Hash algorithm to find the target file; According to the preset offset value, the preset data block size, and the target file identifier, obtaining partial data of the target file and allocating a corresponding memory shared buffer; Writing the partial data of the target file into the memory shared buffer for queuing; Sending the partial data of the target file to the client in the queued order, and re-executing the step of receiving the data reading request sent by the client until a file reading completion indication sent by the client is received; the file reading completion indication indicates that all the data of the target file has been transmitted.

2. The calling method according to claim 1, wherein The step of receiving a file reading request sent by a client using multiple threads includes: When it is monitored that a client sends a request, performing multi-threaded scheduling through the operating system, and allocating corresponding processing threads to receive and process the file reading request sent by the client.

3. The calling method according to claim 1, characterized in that, The step of receiving a file reading request sent by a client using multiple threads includes: When there is an idle TCP long connection between the server and the client, based on the idle TCP long connection, receiving the file reading request sent by the client using multiple threads.

4. The invocation method according to claim 1, wherein Data transmission is performed between the server and the client based on sockets; the calling method further includes: Monitoring the heartbeat message sent by the client based on the socket; When the heartbeat message is not monitored within a preset time period, closing the socket.

5. The invocation method according to claim 1, wherein The calling method further includes: Configuring the ratio between the request processing speed of the processing thread and the data reading and writing speed of the HDFS file reading and writing interface to 1:

8.

6. The calling method according to claim 1, wherein: The client is built based on OpenResty; The server is built using a C program; The interface docking module is built based on the SDK API.

7. A calling device for HDFS API, characterized in that Applied to the server side, the calling device includes: A client request receiving unit, configured to receive a file reading request sent by a client by using multi-threading, where the file reading request carries a target file path; A file information obtaining unit, configured to obtain a mapping relationship between a preset file path and a file identifier; according to the target file path and the mapping relationship, use an interface docking module to obtain a target file identifier from HDFS; according to the target file identifier, obtain a target file size; the interface docking module is used to dock with the HDFS API; A file information sending unit, configured to send target file information to the client, where the target file information includes the target file identifier and the target file size; The client request receiving unit is further configured to receive a data reading request sent by the client by using multi-threading, where the data reading request carries a preset offset value, a preset data block size, and the target file identifier; the data reading request is encapsulated and sent by the client after receiving the target file information according to the preset offset value plus the preset data block size, and is used to divide the target file into partial data according to the preset data block size; the target file identifier is a virtual file identifier VFID, and the VFID is used as a Hash_key value in the Hash algorithm to search for the target file; A file data processing unit, configured to obtain partial data of a target file according to the preset offset value, the preset data block size, and the target file identifier, and allocate a corresponding memory shared buffer; A file data buffering unit, configured to write the partial data of the target file into the memory shared buffer for queuing; A file data sending unit, configured to send the partial data of the target file to the client in the queued order, and trigger the client request receiving unit to re-execute the step of receiving the data reading request sent by the client until a file reading completion indication sent by the client is received; the file reading completion indication indicates that all data of the target file has been transmitted.

8. A device for calling the HDFS API, characterized in that, It includes a processor and a memory; Wherein, the memory is used to store an interface calling program; The processor is used to call the interface calling program stored in the memory to execute the method for calling the HDFS API according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, An interface calling program is stored in the computer-readable storage medium, and the interface calling program realizes the method for calling the HDFS API according to any one of claims 1 to 6 when called by a computer device.

Citation Information

Patent Citations

  • Distributed cache method and system

    CN104052824A

  • File reading method and device based on distributed system

    CN105426483A

  • Method for reading file in distributed storage system and server

    CN106161503A

  • File processing method and device

    CN106776720A