Distributed File System Access Method, Device, Host, and Medium

Through the parent-child process mode, the child process executes user operation code and jumps to the parent process to perform distributed file system access, solving the problems of inconvenience, insecurity and inubiquity in the existing technology, and achieving secure, convenient and cross-language distributed file system access.

CN113810434BActive Publication Date: 2025-07-22ALIBABA GROUP HOLDING LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010529578.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-11
Publication Date
2025-07-22
Estimated Expiration
2040-06-11

AI Technical Summary

Technical Problem

The methods of accessing distributed file systems in the prior art are inconvenient, unsafe and lack universality, especially when calling third-party file systems is expensive and has high thresholds, and cannot support multiple languages and syntaxes.

Method used

Adopt parent-child process mode. The child process executes user operation code and jumps to the security framework process (parent process) within the platform when encountering distributed file system access instructions. The parent process is responsible for the final access operation, isolates proxy access, and directly uses the underlying distributed file system.

Benefits of technology

It improves the security and ease of use of distributed file systems, reduces costs, supports multiple languages and syntaxes, achieves universality, and avoids dependence on third-party systems and additional expenses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113810434B_ABST
    Figure CN113810434B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, host, and medium for accessing a distributed file system. The method includes: receiving a user operation code; enabling a parent process and a child process, where the child process executes the user operation code and jumps to the parent process when an access instruction for the distributed file system is executed. The parent process executes the access instruction and returns to the child process after the access instruction is executed. The parent process is a security framework process within the platform to which the distributed file system belongs. Embodiments of the present disclosure improve the security, convenience, and universality of accessing the distributed file system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data, and more particularly, to a method, device, host, and medium for accessing a distributed file system. Background Art

[0002] Distributed big data platforms default to prohibit reading and writing to the local file system for security and other aspects. However, in many business scenarios, it is necessary to call the operating file system during the process of reading and writing files. For example, in scenarios of user personalized requirements, it is necessary to perform distributed parallel reading and writing on unstructured data, load, edit, and read and write special format resource files, etc., all of which require calling the distributed operating file system.

[0003] In the prior art, in order to be able to access the distributed file system, one solution is to implement it by calling a third-party file system. For example, MaxCompute performs distributed parallel reading and writing of unstructured data and supports reading and writing files on file systems such as Oss through the external table method. This solution has a high threshold and requires applying for a third-party storage, bringing a large cost, and requires many support measures, increasing user dependence. Another temporary solution is to encapsulate the sandbox method to allow users to run the code for reading and writing the file system. However, this method is still not absolutely secure and brings a lot of inconvenience in use. This method generally only supports individual user code with MapReduce as the interface, and does not support languages and grammars such as SQL that users are most familiar with, so it is not universal. The prior art lacks a method for accessing a distributed file system that is convenient to use and universal. Summary of the Invention

[0004] In view of this, the present disclosure aims to provide a technology for accessing a distributed file system that is convenient to use, secure, and universal.

[0005] To achieve this purpose, according to one aspect of the present disclosure, there is provided a method for accessing a distributed file system, including:

[0006] Receiving user operation code;

[0007] Enabling a parent process and a child process, wherein the child process executes the user operation code, and when it executes an access instruction to the distributed file system, it jumps to the parent process, and the parent process executes the access instruction, and after executing the access instruction, returns to the child process. The parent process is a security framework process within the platform to which the distributed file system belongs.

[0008] Optionally, the method is executed by a host in the platform, and the enabling of the parent process and the child process includes: allocating the parent process to a first machine other than the host in the platform for execution, and allocating the child process to a second machine other than the host in the platform for execution, where the second machine is different from the first machine.

[0009] Optionally, in addition to the host, the platform further includes worker machines and slave machines, and the first machine and the second machine are each selected from any one of the worker machines and the slave machines.

[0010] Optionally, the proxy between the parent process and the child process is isolated.

[0011] Optionally, the distributed file system is divided into a persistent file system type and a single-point file system type, the user operation code indicates the type of the distributed file system, and the parent process performs a first access operation on the persistent file system or a second access operation on the single-point file system according to the indicated type.

[0012] Optionally, before enabling the parent process and the child process, the method further includes: creating a child process by the parent process; after enabling the parent process and the child process, the method further includes: destroying the child process by the parent process.

[0013] Optionally, the receiving of the user operation code includes:

[0014] providing an execution layer context, where the execution layer context includes a handle of the distributed file system, and the handle points to a predefined function or interface;

[0015] responding to a user's request for obtaining the handle, and returning the predefined function or interface pointed to by the handle;

[0016] receiving user operation code written by the user using the predefined function or interface.

[0017] Optionally, after receiving the user operation code, the method further includes: performing adaptation on the file system targeted by the user operation code, and the adaptation at least includes: mapping the targeted file system to a file system prefix, providing permission authentication, and supporting file interfaces.

[0018] Optionally, the providing of the permission authentication includes: performing permission authentication according to the permission information of the file unit to be accessed according to the access instruction to the distributed file system, the identity of the user, the access content of the access instruction, and the access time.

[0019] Optionally, the user operation code includes parameter setting statement code, and the parameter setting statement code specifies the file to be accessed in the distributed file system.

[0020] Optionally, the user operation code includes a tool class, which is a program segment, and a file to be accessed in the distributed file system is specified based on the execution result of the tool class.

[0021] Optionally, the parent process has multiple child processes.

[0022] According to one aspect of the present disclosure, a distributed file system access device is provided, including:

[0023] A user interface unit for receiving a user operation code;

[0024] A parent-child process enabling unit for enabling a parent process and a child process, wherein the child process executes the user operation code, jumps to the parent process when an access instruction to the distributed file system is executed, the parent process executes the access instruction, and returns to the child process after the access instruction is executed, and the parent process is a security framework process within the platform to which the distributed file system belongs.

[0025] Optionally, the device is located in a host in the platform, and the parent-child process enabling unit is further configured to: allocate the parent process to a first machine other than the host in the platform for execution, and allocate the child process to a second machine other than the host in the platform for execution, and the second machine is different from the first machine.

[0026] Optionally, in addition to the host, the platform further includes a working machine and a slave machine, and the first machine and the second machine are each selected from any one of the working machine and the slave machine.

[0027] Optionally, the proxy between the parent process and the child process is isolated.

[0028] Optionally, the distributed file system is divided into a persistent file system type and a single-point file system type, the user operation code indicates the type of the distributed file system, and the parent process performs a first access operation on the persistent file system or a second access operation on the single-point file system according to the indicated type.

[0029] Optionally, the device further includes a parent-child process lifecycle management unit for creating the child process by the parent process before enabling the parent process and the child process, and destroying the child process by the parent process after enabling the parent process and the child process.

[0030] Optionally, the user interface unit is further configured to:

[0031] Provide an execution layer context, where the execution layer context includes a handle of the distributed file system, and the handle points to a predefined function or interface;

[0032] In response to a user's request for obtaining the handle, return a predefined function or interface pointed to by the handle;

[0033] Receive user operation code written by the user using the predefined function or interface.

[0034] Optionally, the device further includes: a file interface unit for adapting the file system targeted by the user operation code, and the adaptation at least includes: mapping the targeted file system to a file system prefix, providing permission authentication, and supporting file interfaces.

[0035] Optionally, the providing of permission authentication includes: performing permission authentication according to the permission information of the file unit to be accessed according to the access instruction to the distributed file system, the identity of the user, the access content of the access instruction, and the access time.

[0036] Optionally, the user operation code includes parameter setting statement code, and the parameter setting statement code specifies the file to be accessed in the distributed file system.

[0037] Optionally, the user operation code includes a tool class, and the tool class is a program segment, and the file to be accessed in the distributed file system is specified based on the execution result of the tool class.

[0038] Optionally, the parent process has multiple child processes.

[0039] According to one aspect of the present disclosure, there is provided a host, including: a memory for storing computer-executable code; a processor for executing the computer-executable code to implement the method as described above.

[0040] According to one aspect of the present disclosure, there is provided a computer-readable medium including computer-executable code, and when the computer-executable code is executed by a processor, the method as described above is implemented.

[0041] The embodiments of the present disclosure use parent and child processes to perform access to a distributed file system. The child process executes user operation code, and when it executes an access instruction to the distributed file system, it jumps to the parent process. The parent process executes the access instruction and returns to the child process after executing the access instruction. Since the parent process is a security framework process within the platform to which the distributed file system belongs, the security of accessing the distributed file system is guaranteed. This solution proposes to directly use the underlying distributed file system without introducing a third-party file system, improving the user's convenience and reducing additional expenses, and does not require various permission approvals, which is simple and efficient. In addition, this way of using parent and child processes is independent of language and syntax, supports various languages and syntax, and improves universality. Description of the Drawings

[0042] The above and other objects, features, and advantages of the present invention will become more apparent from the following description of embodiments of the present invention with reference to the accompanying drawings, in which:

[0043] Figure 1 A structural diagram of a big data platform to which the distributed file system to which the embodiments of the present disclosure are applied belongs is shown;

[0044] Figure 2 A flowchart of a distributed file system access method according to an embodiment of the present disclosure is shown.

[0045] Figure 3 A hierarchical structure diagram of a distributed file system access mechanism according to an embodiment of the present disclosure is shown.

[0046] Figure 4 A hierarchical structure diagram of a common file system interface showing the adaptation of a file interface unit and a distributed file system in a distributed file system access mechanism according to an embodiment of the present disclosure is shown.

[0047] Figure 5 A block diagram of a distributed file system access device according to an embodiment of the present disclosure is shown.

[0048] Figure 6 A structural diagram of a host according to an embodiment of the present disclosure is shown. Detailed implementation manners

[0049] The present invention is described below based on embodiments, but the present invention is not limited to these embodiments. In the following detailed description of the present invention, some specific details are described in detail. Those skilled in the art can fully understand the present invention without the description of these details. In order to avoid obscuring the essence of the present invention, well-known methods, processes, and procedures are not described in detail. Additionally, the accompanying drawings are not necessarily drawn to scale.

[0050] Figure 1 A structural diagram of a big data platform to which the distributed file system to which the embodiments of the present disclosure are applied belongs is shown.

[0051] A big data platform is a platform for computing the increasingly large amount of data generated in today's society for the purpose of storage, operation, and display. Big data technology refers to the technology of quickly obtaining valuable information from various types of data, including massively parallel processing (MPP) databases, data mining power grids, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems. A big data platform is a platform integrating data access, data processing, data storage, query retrieval, analysis and mining, application interfaces, etc.

[0052] A distributed file system refers to a system in which the physical storage resources managed by the file system are not necessarily directly connected to the local node, but are connected to the node through a computer network. The design of a distributed file system is based on the client / server model. A typical network may include multiple servers for multi-user access. In addition, the peer feature allows some systems to play the dual roles of client and server. For example, a user can "publish" a directory that allows other clients to access. Once accessed, this directory is just like using a local drive for the client.

[0053] As Figure 1 shown, the big data platform includes a host 110, worker machines 120, and slave machines 130. There is only one host 110 in the big data platform, which is used to issue tasks to the worker machines 120 and control the execution of tasks by the worker machines 120. There are multiple worker machines 120, which are used to execute the tasks assigned by the host 110. The worker machines 120 are equipped with slave machines 130, which are used to accept the control of the worker machines 120 and complete the tasks assigned by the worker machines 120.

[0054] The distributed file system is distributed in Figure 1 each of the machines shown. It is a computer program that manages the hardware and software resources of each of the Figure 1 machines, and is also the kernel and cornerstone of the big data platform. It is scattered in each host 110, worker machine 120, and slave machine 130, and is used to control the software resources and hardware resources of each machine.

[0055] A process is an execution activity of a program in a computer with respect to a certain data set. It is the basic unit for the system to allocate and schedule resources and is the basis of the operating system structure. A process is the basic execution entity of a program, the container of a thread, and the entity of a program. A process is both the basic allocation unit and the basic execution unit. A child process 20 refers to a process created by another process (correspondingly called the parent process 10); the child process 20 inherits most of the attributes of the corresponding parent process 10, such as file descriptors. A process may have multiple subordinate child processes 20, but can have at most 1 parent process 10. In the embodiments of the present disclosure, the child process 20 is used to undertake general user operations, and the parent process 10 undertakes the access to the distributed file system. The parent process 10 in the embodiments of the present disclosure is the framework process of the big data platform itself, and all business logics are developed and verified by the big data platform itself. In this way, the security of the platform and the security of user code are guaranteed. The embodiments of the present disclosure cleverly utilize the parent and child processes, moving the access to the distributed file system that is likely to cause insecure hidden dangers to the parent process 10 for execution, and the parent process 10 can be developed under the block diagram of the big data platform and has security. Therefore, in this way, the security of the access to the distributed file system is directly guaranteed by using the parent and child processes, and it is convenient and universal, does not require a third-party system, and there are no various additional operation and maintenance costs.

[0056] In the embodiments of the present disclosure, the user operation code defines all operations to be executed, including ordinary operations (operations that do not access the distributed file system) and operations that access the distributed file system. All of them are handed over to the child process 20 for execution. However, when the child process 20 executes an access instruction to the distributed file system, it jumps to the parent process 10 (for example, through Remote Procedure Call (RPC), which is a protocol for requesting services from a remote computer program over a network without the need to understand the underlying network technology), and the parent process 10 executes the access instruction and returns to the child process 20 (for example, through RPC) after executing the access instruction.

[0057] In one embodiment, the parent process 10 is assigned to a first machine (worker machine or slave machine) other than the host 110 in the big data platform for execution, and the child process is assigned to a second machine (worker machine or slave machine) other than the host 110 in the platform for execution, and the second machine is different from the first machine. Figure 1 For example, the parent process is in the worker machine 120 and the child process is in its slave machine 130, but it is also possible vice versa, that is, the parent process is in the slave machine 130 and the child process is in the worker machine 120. Or, the parent process and the child process are respectively in two worker machines 120, or respectively in two slave machines 130.

[0058] As Figure 1 shown, the parent process is in the worker machine 120 and the child process is in its slave machine 130. After receiving the user operation code, the host 110 first assigns it to the parent process 10 of the worker machine 120, and the parent process 10 indiscriminately hands it over to the child process 20 of the slave machine 130 for execution. When the child process 20 executes an access instruction to the distributed file system, it calls the parent process 10 through RPC or the like, and the parent process 10 executes the access instruction and returns to the child process 20 through RPC or the like after executing the access instruction.

[0059] The worker machine 120 has a calling server 121, and the slave machine 130 has a calling client 131. The main functions of the calling server 121 and the calling client 131 are to communicate (such as RPC calls), so as to realize the calls between the parent and child processes (including the above-mentioned parent process 10 calling the child process 20 and the child process 20 calling the parent process 10).

[0060] The calls between the parent process 10 and the child process 20 are divided into synchronous calls and asynchronous calls.

[0061] When calling synchronously, the calling server 121 sends a handshake message to the calling client 131, and the calling client 131 responds to the handshake message in real time. After the calling server 121 receives the real-time response, the two parties establish a handshake. The parent process 10 starts to call the child process 20, or the child process 20 calls the parent process 10. That is, the caller sends the program that needs to be executed by the callee to the callee.

[0062] When asynchronous calling is performed, a pipe is opened between the calling server 121 and the calling client 131. The calling client 131 monitors the pipe by reading the port. The calling server 121 writes a response request to the pipe, and the calling client 131 does not need to respond to the response request in real time, and the response request remains in the pipe. When the calling client 131 reads the last word of the response request, a response is returned. Then, the parent process 10 starts to call the child process 20, or the child process 20 calls the parent process 10. That is, the caller sends the program that needs to be executed by the callee to the callee. Since the embodiments of the present disclosure allow asynchronous calling, the flexibility of the calling operation is improved.

[0063] like Figure 1 As shown, when executing a task, the host 110 delegates the task plan (through the user operation code) to the parent process 10 of the working machine 120, and the parent process 10 delegates it to the child process 20 of the subordinate slave 130. The child process 20 returns the execution progress to the parent process 10, and the parent process 10 returns the execution progress to the host 110. When the host 110 wants to access data, it delegates the data access request (through the user operation code) to the parent process 10 of the working machine 120, and the parent process 10 delegates it to the child process 20 of the subordinate slave 130. If it is not an access to the distributed file system, the child process 20 executes the access, and if it involves modification, the modification result is returned to the parent process 10 of the working machine 120, and the parent process 10 returns the modification result to the host 110. If it is an access to a distributed file system, the child process 20 returns to the parent process 10 for execution (through a call between the calling server 121 and the calling client 131). After the execution is completed, it returns to the child process 20 (through a call between the calling server 121 and the calling client 131).

[0064] In the disclosed embodiment, the parent process 10 and the child process 20 are isolated by proxy, that is, the proxy between the parent process 10 and the child process 20 is isolated. In the traditional parent-child process access, the parent process 10 accesses the child process 20 through a proxy, which undoubtedly increases the possibility of leakage. The disclosed embodiment changes the traditional access method of the parent process 10 and the child process 20 through a proxy to direct access, isolates the proxy, and greatly reduces the potential safety hazards caused by proxy access.

[0065] likeFigure 2 As shown, a method for accessing a distributed file system is provided, and the method is executed by a host 110. The method includes:

[0066] Step 210, receiving a user operation code;

[0067] Step 220, enabling a parent process and a child process, where the child process executes the user operation code, and when an access instruction to the distributed file system is executed, it jumps to the parent process, and the parent process executes the access instruction, and after the access instruction is executed, it returns to the child process. The parent process is a security framework process within the platform to which the distributed file system belongs.

[0068] The above steps will be described in detail below.

[0069] In step 210, a user operation code is received.

[0070] The user operation code refers to the code written by the user for performing operations on the big data platform, which includes possible access codes to the distributed file system and also includes codes for other operations. For the codes for other operations, they are executed by the child process 20; for the access codes to the distributed file system, they are executed by the child process 20 jumping to the parent process 10.

[0071] In one embodiment, step 210 includes:

[0072] Providing an execution layer context, where the execution layer context includes a handle of the distributed file system, and the handle points to a predefined function or interface;

[0073] In response to a user's request for obtaining the handle, returning the predefined function or interface pointed to by the handle;

[0074] Receiving the user operation code written by the user using the predefined function or interface.

[0075] The execution layer context is the environment used for the execution of program code. Some information used when executing the program code can be placed in the context to facilitate the execution of the program code.

[0076] The user operation code is written by the user. Generally, a user interface (i.e., the user application layer, which is Figure 3It is programmed by the user in the user interface (executed by the user interface unit 310). However, in the embodiments of the present disclosure, the user does not mechanically type the written code into the user interface. Instead, some predefined functions or interfaces can be called to simplify the operation. For example, the general syntax of the distributed SQL language, such as user-defined functions Udf, user-defined aggregate functions Udaf, user-defined table functions Udtf, general Mapreduce usage, custom interfaces for graph computing Graph, etc., do not need to be developed separately. The handles pointing to these predefined functions or interfaces are placed in the execution layer context.

[0077] A handle refers to a unique integer value used, that is, a value 4 bytes long (8 bytes in a 64-bit program), to identify different objects in an application and different instances in the same category, such as a window, button, icon, scroll bar, output device, control, or file, etc. In the embodiments of the present disclosure, it identifies predefined functions or interfaces, such as user-defined functions Udf, user-defined aggregate functions Udaf, user-defined table functions Udtf, general Mapreduce usage, custom interfaces for graph computing Graph, etc. The application can access the information of the corresponding object through the handle, but the handle is not a pointer, and the program cannot use the handle to directly read the information in the file.

[0078] Therefore, an execution layer context is provided to the user, and the execution layer context includes handles pointing to predefined functions or interfaces. When the user writes user operation code, if a predefined function or interface is used, a request for obtaining the handle is made on the interface, and the interface returns the predefined function or interface pointed to by the handle to the user, and the user can use the predefined function or interface on the interface to write user operation code.

[0079] The advantage of this embodiment is that it unifies the use of the user application layer, making the general syntax of the distributed SQL, such as user-defined functions Udf, user-defined aggregate functions Udaf, etc., do not need to be developed separately, and can be used by uniformly obtaining the file system handle from the execution layer context. The user can use the original programming mode and general file I / O interface, without the need to incur new costs to learn, reducing the user's learning cost and improving the generality.

[0080] In addition, the embodiments of the present disclosure also provide a simplified and unified parameter setting mode, by inputting simple statements on the interface to specify the names, labels, etc. of the input and output files. That is, the user operation code includes parameter setting statement code, and the parameter setting statement code specifies the files to be accessed in the distributed file system.

[0081] In one embodiment, the file to be accessed is specified by a parameter setting statement code in the form of set K = V. K is an identifier specifying the file to be accessed, and V is the location or label where the file to be accessed is stored. The meaning of the label is that some storage locations are labeled in order to simplify the specification operation of these storage locations, so that when specifying the files stored at these locations later, only the label needs to be used for reference. The label is particularly effective when specifying files at multiple storage locations simultaneously. Multiple storage locations can be labeled with the same label at the same time, so that this label can be used to represent simultaneous access to the files at all these storage locations. This method can support various algorithms such as SQL, Map Reduce, and Graph, without the need for secondary specification and development. An exemplary parameter setting statement is as follows:

[0082] set odps.sql.volume.input[ / output].desc= <project> .] <partition> [: <label>;

[0083] Among them, the parameters in square brackets are optional. odps.sql.volume.input[ / output].desc represents the file to be accessed, where input or output represents inputting or outputting the file, project represents the group where the file to be accessed is located, table represents the table where the file to be accessed is located, label represents the partition where the file to be accessed is located, a group includes tables, and a table includes partitions. label represents the label of the file to be accessed.

[0084] In one embodiment, instead of specifying the file to be accessed in the distributed file system by using the above parameter setting statement, a tool class is used. The user operation code includes the tool class. A tool class is such a program fragment based on whose execution result the file to be accessed in the distributed file system can be specified.

[0085] The following way of the tool class has a certain degree of flexibility. This way is a variant of the parameter setting statement way. The tool class will convert the parameters set by the user in an objectified way (program fragment) into set parameters. The two ways only have different perceptions for the user, but are unified for the platform and will be submitted and recognized through the setting method. The following is an exemplary example of the tool class:

[0086] InputUtils.addVolume(new VolumeInfo([project,]inVolume,inPartition,"inLabel"),new JobConf());

[0087] OutputUtils.addVolume(new VolumeInfo([project,]outVolume,outPartition,"outLabel"),new JobConf());

[0088] The solution of the embodiment of the present disclosure unifies the use of the user application layer. The user can use the original programming mode and general file I / O interface without incurring new costs for learning.

[0089] In addition, in one embodiment, after step 210, the method further includes: performing adaptation of the file system targeted by the user operation code, and the adaptation at least includes: mapping from the targeted file system to a file system prefix, providing permission authentication, and supporting a file interface. This operation is executed by the file interface unit 320 as shown in Figure 3-4 Figure.

[0090] The user interface unit 310 in the user application layer is the interface unit of the big data platform facing the user, responsible for the input of user operation codes. The file interface unit 320 is the interface unit facing files, responsible for the adaptation of the user operation codes to the file system, that is, adapting which files the user operation codes need to access. Currently, this part can be implemented based on the current general file system interface Hadoopfilesystem system. However, the embodiments of the present disclosure have extended this general file system interface, mainly by adding the mapping from the file system to the file system prefix, providing permission authentication, and supporting file interfaces, including the support for interfaces related to symlink, snapshot, acl, and xattr, the support for the listCorruptFileBlocks interface, the support for the seekToNewSource interface, and implementing the concat and truncate interfaces.

[0091] The mapping from the file system to the file system prefix can quickly help users lock the storage area corresponding to the file system to be accessed by users. The file systems with the same file system prefix are stored in adjacent areas.

[0092] Permission authentication can separately set permission information for file units in the distributed file system. A file unit is the smallest unit for setting the same permission information. For example, in the persistent (volumn) file system, multiple backups of the same system file are set on multiple working machines or slave machines. The permission information of these multiple backups should be the same. Therefore, these multiple backups are combined as a file unit. Permission information is the information that stipulates the permissions for accessing file units. This permission includes user identity permissions, accessible file content permissions, access time permissions, etc. For example, if it is stipulated that user A can only access file units of type X from Monday to Friday, then it is not enough to only judge that it is user A accessing. It is also necessary to combine the access to the file content and the access time for judgment. Therefore, the judgment of this permission is a multi-faceted judgment. The provision of permission authentication includes: performing permission authentication according to the permission information of the file unit to be accessed by the access instruction to the distributed file system, the identity of the user, the access content of the access instruction, and the access time. If the identity of the user, the access content of the access instruction, and the access time conform to the permission information, the authentication is successful. In this way, the security of access is improved.

[0093] The support for various file interfaces can expand the range of files that users can access and improve the user experience.

[0094] Such as Figure 4 As shown in the figure, a common file system interface 321 is also provided between the distributed file system 330 and the file interface unit 320. Its function is to perform semantic adaptation between the distributed file system 330 and the file interface unit 320, achieve semantic compatibility, and provide support for the open source ecosystem. Due to the expansion of the file interface unit 320, such as adding mappings from the file system to the file system prefix, providing permission authentication, and supporting file interfaces, etc., mainly in an open source form, while the distributed file system 330 is generally in the cloud, and the common file system interface 321 is required for semantic adaptation. The common file system interface 321 is a Java interface for accessing the distributed file system 330, and open source HDFS compatibility is implemented based on it. The common file system interface 321 has two access paths to the distributed file system 330. One is the way of a single process 15 according to the prior art, and the other is the way of mutual calls between the subprocess 20 and the parent process 10 according to the embodiments of the present disclosure.

[0095] In step 220, the parent process and the subprocess are enabled.

[0096] The user operation code received in step 210 is executed by the subprocess 20. When an access instruction to the distributed file system is executed, it jumps to the parent process 10, and the parent process 10 executes the access instruction, and returns to the subprocess 20 after the access instruction is executed. The above process has been described in detail in the introduction of the Figure 1 framework, so it will not be elaborated here.

[0097] In one embodiment, as Figure 3 shown, the distributed file system is divided into a persistent (Volume) file system 331 type and a single-point (TempFile) file system 332 type, which respectively support operations on persistent files and single-point files. Persistent files are for persistent storage requirements and store multiple backups on different working machines 120 and slave machines 130 in the big data platform. Single-point files are used for temporary reading and writing and single-node storage (stored on one working machine 120 or slave machine 130), which are valid during the life cycle of the current job and will be automatically cleared after the job execution ends. Different file systems can be selected according to different demand scenarios.

[0098] The user operation code received in step 210 indicates the type of the distributed file system, and the parent process 10 executes the first access operation on the persistent file system 331 or the second access operation on the single-point file system 332 according to the indicated type. That is, if the type is the persistent file system 331 type, the parent process 10 executes the persistent file system access 101; if the type is the single-point file system 332, the parent process 10 executes the single-point file system access 102.

[0099] In addition, in one embodiment, before enabling the parent process and the child process, the parent process creates the child process. After enabling the parent process and the child process, the parent process destroys the child process. By specifying a complete way for the parent process to create, monitor, and destroy the child process, the lifecycle management of the child process in the parent-child process structure is improved. Moreover, a parent process can create multiple child processes, which respectively execute the above steps 210-220, and destroy the child processes respectively after their respective data access is completed, realizing the efficient management of the lifecycles of multiple child processes.

[0100] In the embodiments of the present disclosure, the parent-child process model can bring the following advantages: the framework of the distributed file system is separated from the user program, and the system structure is clearer. The parent process specifically solves some inherent problems of the distributed file system and major problems that affect security, and the child process solves some problems that do not affect the security of the distributed file system. The parent process and the child process can interact through RPC calls, making it easier to implement cross-language. The user execution logic is all executed in the child process, which is convenient for security control. The upper layer can be abstracted into a more general framework, and the common services can be extracted and executed by the parent process, while the child process focuses on specific execution. In addition, the parent-child process supports a one-to-many mode, that is, one parent process can correspond to multiple child processes. This is necessary in some scenarios. For example, some scripting languages (such as python, shell) do not support multithreading, and if concurrency is to be achieved, only the multi-process method can be used.

[0101] The embodiments of the present disclosure directly use the underlying distributed file system without introducing a third-party file system, improving the user's convenience and reducing additional expenses. Compared with the prior art method of accessing the distributed file system by calling a third-party file system, the threshold is relatively low. There is no need to apply for a third-party storage, and there is no need for secondary development during distributed computing for the characteristics of the third-party file system, such as file splitting, data serialization and deserialization, etc. There is no need to provide a framework method and syntax interface similar to the appearance to support and identify the user's input and output, configuration parameters, etc., without increasing the user's dependence. In addition, the embodiments of the present disclosure support calling the Java interface in a custom program to access and read and write files stored in the underlying distributed storage. The custom program that calls the file system supports running in an isolated environment, and the call to the file system operation is accessed through the parent-child process proxy, without the need to apply for and approve various complex permissions, and can be used by setting parameters in the internal environment of the group, which is simple, easy to use, and secure and controllable. The embodiments of the present disclosure unify the user interfaces of the upper-layer user application layer, and are compatible with and reuse the existing syntax SQL, MapReduce, etc. of the big data platform. These functions or interfaces do not need to be developed separately, and the file system handle can be uniformly obtained from the upper and lower layers of the execution layer for calling. Users can use the original programming mode and the general file I / O interface, reducing the learning cost of users.

[0102] As Figure 5 shown, according to an embodiment of the present disclosure, a distributed file system access device 300 is provided, including:

[0103] A user interface unit 310, configured to receive a user operation code;

[0104] A parent-child process enabling unit 320, configured to enable a parent process and a child process, wherein the child process executes the user operation code, and jumps to the parent process when an access instruction to the distributed file system is executed. The parent process executes the access instruction and returns to the child process after the execution of the access instruction. The parent process is a security framework process within the platform to which the distributed file system belongs.

[0105] Optionally, the device 300 is located in a host in the platform, and the parent-child process enabling unit 320 is further configured to: allocate the parent process to a first machine other than the host in the platform for execution, and allocate the child process to a second machine other than the host in the platform for execution. The second machine is different from the first machine.

[0106] Optionally, in addition to the host, the platform further includes worker machines and slave machines, and the first machine and the second machine are each selected from any one of the worker machines and the slave machines.

[0107] Optionally, the proxy between the parent process and the child process is isolated.

[0108] Optionally, the distributed file system is divided into a persistent file system type and a single-point file system type. The user operation code indicates the type of the distributed file system, and the parent process executes a first access operation on the persistent file system or a second access operation on the single-point file system according to the indicated type.

[0109] Optionally, the device further includes a parent-child process lifecycle management unit (not shown), configured to create the child process by the parent process before enabling the parent process and the child process, and destroy the child process by the parent process after enabling the parent process and the child process.

[0110] Optionally, the user interface unit 310 is further configured to:

[0111] Provide an execution layer context, where the execution layer context includes a handle of the distributed file system, and the handle points to a predefined function or interface;

[0112] In response to a user's request for obtaining the handle, return the predefined function or interface pointed to by the handle;

[0113] Receive user operation code written by the user using the predefined function or interface.

[0114] Optionally, the apparatus 300 further includes: a file interface unit 320, configured to perform adaptation of the file system targeted by the user operation code, where the adaptation includes at least: mapping of the targeted file system to a file system prefix, providing permission authentication, and support for file interfaces.

[0115] Optionally, the providing of permission authentication includes: performing permission authentication according to the permission information of the file unit to be accessed according to the access instruction to the distributed file system, the identity of the user, the access content of the access instruction, and the access time.

[0116] Optionally, the user operation code includes parameter setting statement code, and the parameter setting statement code specifies the file to be accessed in the distributed file system.

[0117] Optionally, the user operation code includes a tool class, and the tool class is a program segment, and the file to be accessed in the distributed file system is specified based on the execution result of the tool class.

[0118] Optionally, the parent process has multiple child processes.

[0119] Since the above has been described in detail in connection with Figure 2 the distributed file system access method of the present disclosure, the implementation details of the distributed file system access apparatus 300 are basically the same as those of the distributed file system access method, and thus will not be elaborated.

[0120] The method for automatically generating a picture according to an embodiment of the present disclosure can be implemented by Figure 6 the host 110. The host 110 according to an embodiment of the present disclosure will be described below with reference to Figure 6 The host 110 shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure. Figure 6 As

[0121] shown, the host 110 is presented in the form of a general computing device. The components of the host 110 may include, but are not limited to: at least one of the above-mentioned processing units 810, at least one of the above-mentioned storage units 820, and a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810). Figure 6

[0122] Figure 2 Among them, the storage unit stores program code, and the program code can be executed by the processing unit 810, so that the processing unit 810 executes the steps of various exemplary embodiments of the present invention described in the description part of the above exemplary method of this specification. For example, the processing unit 810 can execute the respective steps as Figure 2 shown in

[0123] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 8201 and / or a cache storage unit 8202, and may further include a read-only storage unit (ROM) 8203.

[0124] The storage unit 820 may also include a program / utilities 8204 having a set (at least one) of program modules 8205. Such program modules 8205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0125] The bus 830 may represent one or more of several types of bus structures, including a storage unit bus or storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0126] The host 110 may also communicate with one or more external devices 700 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the host 110, and / or may communicate with any device that enables the host 110 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be through an input / output (I / O) interface 850. Also, the host 110 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 860. As shown in the figure, the network adapter 860 communicates with other modules of the host 110 through the bus 830. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the host 110, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0127] It should be appreciated that the above description is only a preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, there are many variations in the embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should be included within the protection scope of the present invention.

[0128] It should be understood that the various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments.

[0129] It should be understood that the above describes specific embodiments of the present specification. Other embodiments are within the scope of the claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0130] It should be understood that an element described herein in the singular or shown only in one of the figures is not meant to limit the quantity of that element to one. Additionally, a module or element described or shown herein as separate may be combined into a single module or element, and a module or element described or shown herein as a single one may be split into multiple modules or elements.

[0131] It should also be understood that the terms and expressions used herein are for the purpose of description only, and one or more embodiments of the present specification should not be limited to these terms and expressions. Using these terms and expressions does not mean excluding any equivalent features of the illustration and description (or parts thereof), and it should be recognized that various modifications that may exist should also be included within the scope of the claims. Other modifications, variations, and substitutions may also exist. Accordingly, the claims should be regarded as covering all such equivalents.< / label> < / partition> . < / project>

Claims

1. A distributed file system access method, comprising: Receiving user operation code; Performing adaptation of the file system targeted by the user operation code, the adaptation at least including: mapping of the targeted file system to a file system prefix, providing permission authentication, and support for file interfaces; wherein, the providing of permission authentication includes: Performing permission authentication according to the permission information of the file unit to be accessed by the access instruction to the distributed file system, the identity of the user, the access content of the access instruction, and the access time; Enabling a parent process and a child process, wherein the child process executes the user operation code, and when it executes an access instruction to the distributed file system, it jumps to the parent process, and the parent process executes the access instruction and returns to the child process after executing the access instruction, and the parent process is a security framework process within the platform to which the distributed file system belongs.

2. The method according to claim 1, wherein The method is executed by a host in the platform, and the enabling of the parent process and the child process includes: Allocating the parent process to a first machine other than the host in the platform for execution, and allocating the child process to a second machine other than the host in the platform for execution, and the second machine is different from the first machine.

3. The method according to claim 2, wherein In addition to the host, the platform further includes worker machines and slave machines, and the first machine and the second machine are each selected from any one of the worker machines and the slave machines.

4. The method according to claim 1, wherein The proxy between the parent process and the child process is isolated.

5. The method according to claim 1, wherein The distributed file system is divided into a persistent file system type and a single-point file system type, the user operation code indicates the type of the distributed file system, and the parent process executes a first access operation on the persistent file system or a second access operation on the single-point file system according to the indicated type.

6. The method according to claim 1, wherein Before enabling the parent process and the child process, the method further includes: creating a child process by the parent process; After enabling the parent process and the child process, the method further includes: destroying the child process by the parent process.

7. The method according to claim 1, wherein The receiving of the user operation code includes: Providing an execution layer context, the execution layer context including a handle of the distributed file system, and the handle pointing to a predefined function or interface; In response to a user's request for obtaining the handle, returning the predefined function or interface pointed to by the handle; Receiving user operation code written by the user using the predefined function or interface.

8. The method according to claim 1, wherein The user operation code includes parameter setting statement code, and the parameter setting statement code specifies the file to be accessed in the distributed file system.

9. The method according to claim 1, wherein, The user operation code includes a tool class, and the tool class is a program fragment, and the file to be accessed in the distributed file system is specified based on the execution result of the tool class.

10. The method according to claim 1, wherein, The parent process has multiple child processes.

11. A distributed file system access device, comprising: A user interface unit for receiving user operation code; A file interface unit for performing adaptation of the file system targeted by the user operation code, the adaptation at least including: mapping of the targeted file system to a file system prefix, providing permission authentication, and support for file interfaces; wherein, the providing of permission authentication includes: Perform permission authentication based on the permission information of the file unit to be accessed according to the access instruction to the distributed file system, the identity of the user, the access content of the access instruction, and the access time; A parent-child process enabling unit, configured to enable a parent process and a child process, wherein the child process executes the user operation code, and when it reaches the access instruction to the distributed file system, it jumps to the parent process, and the parent process executes the access instruction, and returns to the child process after executing the access instruction, and the parent process is a security framework process within the platform to which the distributed file system belongs.

12. The apparatus according to claim 11, wherein, The apparatus is located in a host in the platform, and the parent-child process enabling unit is further configured to: Allocate the parent process to a first machine other than the host in the platform for execution, and allocate the child process to a second machine other than the host in the platform for execution, and the second machine is different from the first machine.

13. The apparatus according to claim 12, wherein, In addition to the host, the platform further includes a working machine and a slave machine, and the first machine and the second machine are each selected from any one of the working machine and the slave machine.

14. The apparatus according to claim 11, wherein, The proxy between the parent process and the child process is isolated.

15. The apparatus according to claim 11, wherein, The distributed file system is divided into a persistent file system type and a single-point file system type, the user operation code indicates the type of the distributed file system, and the parent process executes a first access operation on the persistent file system or a second access operation on the single-point file system according to the indicated type.

16. The apparatus according to claim 11, wherein, The apparatus further includes a parent-child process lifecycle management unit, configured to create a child process by the parent process before enabling the parent process and the child process, and destroy the child process by the parent process after enabling the parent process and the child process.

17. The device according to claim 11, wherein The user interface unit is further configured to: Provide an execution layer context, where the execution layer context includes a handle of the distributed file system, and the handle points to a predefined function or interface; In response to a user's request for obtaining the handle, return the predefined function or interface pointed to by the handle; Receive user operation code written by the user using the predefined function or interface.

18. The apparatus according to claim 11, wherein, The user operation code includes parameter setting statement code, and the parameter setting statement code specifies a file to be accessed in the distributed file system.

19. The apparatus according to claim 11, wherein, The user operation code includes a tool class, and the tool class is a program fragment, and the file to be accessed in the distributed file system is specified based on the execution result of the tool class.

20. The apparatus according to claim 11, wherein The parent process has multiple child processes.

21. A host, comprising: A memory, configured to store computer-executable code; A processor, configured to execute the computer-executable code to implement the method according to any one of claims 1-11.

22. A computer-readable medium, characterized in that, Including computer-executable code, which when executed by a processor implements the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Method, device and system for accessing file system

    CN108289080A