Computable storage system and service method for high-energy physics IO-intensive
By using spare CPU resources for data processing on storage nodes in the high-energy physics field, the problems of network bandwidth pressure and waste of computing resources are solved, efficient data processing and system stability are achieved, computing and analysis processes are optimized, and costs are reduced.
Patent Information
- Application Number
- CN202111136923.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-09-27
AI Technical Summary
During the IO-intensive data analysis and processing in the field of high-energy physics, the network transmission pressure is high, the performance of the storage system and computing system is limited by the network bandwidth, and the CPU computing power of the storage device is excessive, resulting in wasted computing resources.
Use the spare CPU computing power of the storage node to process experimental data, reduce the network bandwidth pressure of the storage system and computing system, and realize the computing service of the storage system. Through the metadata analysis module, the storage service analysis module, the storage computing function module and the metadata registration module, the data processing function and the registration result are realized.
It effectively improves the efficiency of the use of storage nodes, reduces network traffic and equipment pressure, improves the stability and operation efficiency of the system, optimizes the computing and analysis process of users, and reduces the cost of system construction.
Smart Images

Figure CN113901016B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of mass data storage and computable storage, and relates to a high-energy physics IO-intensive computable storage system and a service method. Background Art
[0002] With the development of experimental equipment and detector technology, the energy resolution of detectors is getting higher and higher, and the amount of experimental data has increased dramatically from TB to PB or even EB, which has put forward great requirements on computing power and network transmission capacity, and promoted technological innovation and development in the fields of high-performance computing, massive storage systems and network transmission. On the other hand, the development of chip design and manufacturing, storage and network technologies has also allowed high-energy physics to develop in the direction of higher energy and higher detector resolution.
[0003] At present, high-energy physics computational analysis can be divided into different computing modes according to the CPU / GPU resource usage. For example, accelerator and detector simulation design and physical simulation require a lot of CPU resources and are computationally intensive; while the analysis and processing of massive scientific data observed by particle detectors requires few CPU resources for a single task, but will process considerable data, which is a typical data-intensive type; high-intensity scientific computing in high-energy physics theoretical research such as lattice quantum chromodynamics and computational cosmology is also a typical computationally intensive type. According to the difference in computing power and data storage location, high-energy physics computational analysis is generally divided into two modes. One is a mode in which computing power and data storage are separated, represented by HTCondor and Slurm, which are commonly used computing task scheduling systems in computing fields such as high-energy physics; the other is a mode in which computing and storage are unified, and computing and data are stored on the same node, represented by Hadoop and Spark, which are commonly used data processing software in the field of big data. In the unified computing and storage model, data processing is performed on the storage node. The advantage is that there is no need to move a large amount of data through the network, but the disadvantages are also obvious: the density of computing devices is low, computing and storage must be expanded simultaneously, and the application needs to be modified to adapt to its workflow. This model is rarely used in the field of high-energy physics. In contrast, in the separate computing and storage model, computing nodes wait to read data for analysis through network file systems such as EOS and Lustre. The advantage is that computing and storage can be expanded independently, and the deployment is flexible and easy to use; the disadvantage is also obvious, requiring a large amount of network communication and data transmission. High-energy physics analysis is very suitable for and widely adopted this model.
[0004] In high energy physics computing, the general storage and computing cluster structure is as follows: Figure 1As shown in Figure 1. Users submit tasks from the login node, and job scheduling systems such as HTcondor / Slurm schedule the jobs to the computing nodes. The computing nodes transmit data from network file systems such as Lustre / EOS through the network for analysis. This structure is convenient for horizontal expansion, but it also has bottlenecks. The obvious point is that user login, job scheduling, data transmission, etc. all require a high-speed and stable Internet network. With the increasing density of computing nodes, the increasing number of CPU core data on a single node, and the increasing number of data on storage disk arrays, even the 100Gb / s upstream network is often occupied, resulting in various problems such as accumulation and loss of computing node server data packets, slow response of users accessing storage servers, and even inability to access files. In order to ensure the quality of service, one idea to solve the bottleneck is to upgrade the cluster network to a higher bandwidth, but this method requires the cluster to be transformed, the overall change is large, and the cost of upgrading to more than 100Gb / s is high. Compared with adding computing nodes and storage arrays, upgrading the network is not the most suitable way to solve the network bottleneck.
[0005] In general experimental physics data analysis, some data processing processes such as decode (converting experimental data from binary files to files of other formats), case reconstruction (reconstructing particle reaction processes from experimental data), simulation reconstruction (using simulation data to reconstruct cases), etc., occupy a considerable part of the data analysis process, and these processes only consume very little CPU resources, but will read and write files from the network storage system in large quantities, occupying a large amount of network bandwidth, causing network bottlenecks. On the other hand, the CPU utilization rate of the file system storage process on numerous storage nodes is not high, the CPU performance is excessive, and a large amount of CPU resources are idle. If the data processing methods such as Hadoop / Spark are used for reference, the file system storage process uses the idle CPU computing power on the storage node to directly perform data processing processes such as decode and case reconstruction, then a large amount of network transmission can be effectively avoided, reducing the network burden. The present invention has realized such a computable storage method for high-energy physics IO-intensive data processing.
[0006] The current method to solve network bottlenecks is mainly to upgrade the computing cluster network. According to computing needs, the cluster network environment is upgraded to a higher bandwidth, and the computing storage nodes are upgraded to a higher speed network interface.
[0007] The current method to solve the problem of CPU resource waste in storage nodes is to add storage nodes to the resource scheduling system to participate in calculations.
[0008] On the one hand, upgrading the network will affect the entire cluster architecture. On the other hand, after the network has been upgraded to a certain level, the cost of upgrading the network further will increase too much and will not be appropriate.
[0009] Adding storage nodes to the task scheduling system to participate in computing will increase the overall system complexity. The storage system and computing tasks will affect each other. Since computing tasks are uncontrollable, they will more easily cause the system to crash, restart, and other problems, which will lead to unstable storage system services. Summary of the invention
[0010] In view of the high pressure of network transmission during IO-intensive data analysis and processing in the field of high-energy physics, the performance and stability of the storage system and computing system are limited by the network bandwidth, and the storage device CPU computing power is excessive, resulting in a waste of computing resources. The purpose of the present invention is to provide a computable storage system and service method for high-energy physics IO-intensive tasks. The present invention uses the spare CPU computing power of the storage node to process experimental data, reduce the pressure of the network bandwidth of the storage system and computing system, and realize the computable service of the storage system.
[0011] The present invention provides a basis for realizing IO-intensive computing functions in a general distributed file storage system. By using the present invention, a storage system with computing capabilities not only retains all the functions of the original storage system, but also can utilize the idle computing power of the storage device to execute the established and self-defined data processing flow. This provides strong support for the widespread application of computing storage technology.
[0012] The technical solution of the present invention is:
[0013] A method for high-energy physics IO-intensive computable storage service, the steps of which include:
[0014] 1) Use the metadata parsing module to parse and filter the file access request sent by the client, delete the parameters in the logical address of the requested file except the parameters required for calling the computable storage function, and forward the modified file access request to the metadata service process MGM in the distributed storage system so that it can correctly obtain the file information, and pass the file's physical address and file size information back to the client, so that the client can redirect the file access request to the storage service process FST of the server storing the requested file in the distributed storage system according to the returned information;
[0015] 2) The storage service parsing module calls the corresponding method in the storage computing function module to process the data according to the file access request transmitted from the client after redirection;
[0016] 3) After the storage computing function module completes the data processing, it generates a file and sends it to the storage service analysis module; then the storage service analysis module calls the metadata registration module to register the file back to the distributed storage system, realizing the synchronization between the FST server local database and the distributed storage system metadata database.
[0017] Furthermore, in step 2), the storage service parsing module parses the file access request passed after redirection to obtain the file logical address, physical address, file size, object information of the request, and parameters required for the computable storage function; then, based on the parsed information, the corresponding method in the storage computing function module is called to process the requested object information.
[0018] Furthermore, the parameters required for the computable storage function are used to set a data processing method; the data processing method includes data compression, decompression, segmentation, format conversion, case screening, and noise filtering.
[0019] Furthermore, the metadata registration module first registers an empty file in the distributed storage system and assigns a relative storage path; then stores the file generated by the storage computing function module in the local data disk, and the file name is the relative storage path plus the absolute path prefix; then writes the file information into the local database of the FST server, and then synchronizes the records in the local database with the distributed storage system metadata database to complete the file registration process.
[0020] Furthermore, the file access request includes information about the object issuing the request, a logical address of the requested file, a requested operation type, and a request initiation time.
[0021] A high-energy physics IO-intensive computable storage system, characterized by comprising a metadata parsing module, a storage service parsing module, a storage computing function module and a metadata registration module; wherein,
[0022] The metadata parsing module is used to parse and filter the file access request sent by the client, delete the parameters in the logical address of the requested file except the parameters required for calling the computable storage function, and forward the modified file access request to the metadata service process MGM in the distributed storage system so that it can correctly obtain the file information, and pass the file's physical address and file size information back to the client, so that the client can redirect the file access request to the storage service process FST of the FST server in the distributed storage system where the requested file is stored according to the returned information;
[0023] The storage service parsing module runs on the FST server and is used to call the corresponding method in the storage computing function module to process the data according to the file access request passed from the client after redirection. After the data processing is completed, a file is generated and sent to the storage service parsing module; the storage service parsing module calls the metadata registration module to register the file back to the distributed storage system to obtain the logical address in the system, so as to achieve synchronization between the local database of the FST server and the metadata database of the distributed storage system;
[0024] A storage computing function module includes multiple data processing methods, and different data processing methods correspond to different parameters in the file access request;
[0025] The metadata registration module registers the files generated by the storage computing function module back into the metadata database of the distributed storage system to complete the storage computing process.
[0026] Furthermore, the storage service parsing module, the storage computing function module and the metadata registration module are compiled to generate a dynamic link library, which is added to the configuration file of the distributed storage system and loaded when the MGM and FST processes are started.
[0027] Compared with the prior art, the present invention has the following advantages:
[0028] 1) Make full use of idle CPU resources of storage devices
[0029] Modern server chips generally have many processing cores, and their single-core computing power and total computing power are very high. As a storage node in a distributed storage system, the CPU processing power that can be used by its storage process, such as the FST process of EOS, is limited, and the computing power of most cores is not fully utilized, resulting in a waste of resources. The present invention is based on the design concept of computable storage for high-energy physics massive storage systems, and utilizes the idle CPU resources of storage nodes to achieve certain data processing functions, effectively improving the utilization efficiency of storage nodes. In addition, compared to directly adding storage nodes with idle computing power to the computing resource scheduling system to participate in computing, the computable storage method isolates computing from storage services, making the system more stable.
[0030] 2) Effectively reduce the network load pressure of the storage system
[0031] High-energy physics data analysis generally requires processing a large amount of data, and its separation of computing and storage will result in network traffic of up to tens of GB / s between storage and computing nodes, which puts great pressure on both the storage system and the network. In order to cope with the growing computing needs, resources need to be continuously invested to increase the network speed. If some computing tasks that only perform simple data processing and do not require a lot of CPU computing power are converted into computing function modules of storage processes on storage nodes, and data files are processed locally ("local" processing) and subsequently stored, a large part of the network traffic can be reduced, the pressure on network devices and storage devices will be greatly reduced, and network services and file systems will be more stable. In this way, greater benefits can be achieved by utilizing limited network, storage and computing resources.
[0032] 3) Greatly improve the operating efficiency of scheduling and storage systems
[0033] Processing massive amounts of high-energy physics data requires tens of thousands of CPU computing resources. One of the characteristics of high-energy physics data processing jobs is data parallelism. A single job requires relatively few CPU resources, generally a few CPU cores instead of thousands or tens of thousands of CPU cores, but the number of jobs can reach tens of thousands or even hundreds of thousands. In order to improve resource utilization, a unified job scheduling system is generally required to coordinate the use of computing resources. A large number of jobs (tens of thousands or even hundreds of thousands) will put pressure on the job system, easily causing service failures and reducing the utilization of the entire system. If some of the tens of thousands to hundreds of thousands of jobs that only perform simple data processing and do not require a lot of CPU computing power are executed on the storage nodes of the storage system, the pressure on the scheduling system can be effectively reduced, and the service operation will be more stable; relatively speaking, the number of other analysis jobs can also be increased, and the operating efficiency of the analysis jobs can be improved. The reduction of job data will also reduce the pressure on the storage system, and the stability of the storage system will be improved accordingly.
[0034] 4) Effectively optimize the user's physical calculation analysis process
[0035] Given the limited computing resources, a large number of computing jobs will queue up in the job scheduling system waiting to be executed, and the scheduling system will schedule these jobs according to certain rules. Generally speaking, the more computing jobs there are, the lower the scheduling priority will be. A large number of jobs that only perform simple data processing not only "waste" computing resources, but also increase a lot of unnecessary waiting time, greatly extending the time of the entire analysis process. After moving these tasks to the storage node and being processed by the storage process FST, researchers can focus more time on real physical analysis, which will effectively optimize their computing process, effectively improve their analysis efficiency, and save time costs for sorting.
[0036] 5) Effectively reduce the construction cost of storage systems and computing clusters
[0037] The use of the present invention optimizes network transmission, reduces network pressure, and improves cluster computing efficiency. While meeting the needs of physical analysis computing, it can reduce the construction costs of storage systems, computing systems, and network systems, or provide higher computing power, better networks, and more storage.
[0038] 6) The design method is universal and scalable
[0039] The computable storage system method implemented by the present invention is to conduct secondary development of modules such as metadata services and storage services on top of the existing storage system, which not only retains the existing storage system characteristics and is backward compatible with existing programs, but also expands the system functions and realizes data processing functions. Although the entire design and development process is based on the EOS system, the design and development ideas are not limited to the EOS system, but are also applicable to other distributed file systems such as Lustre, and are universal. In addition, the data processing function of the storage system can be customized according to user needs, and is not limited to file compression and decompression, format conversion, etc. Different computing modules can also be called according to user needs, and it has very good scalability. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 Analytical computing cluster architecture for traditional high energy physics.
[0041] Figure 2 This is a diagram of the computable storage model based on EOS. DETAILED DESCRIPTION
[0042] The present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0043] In the following specific implementation examples, Figure 2 The present invention is further described in detail. These implementation examples are described in sufficient detail so that those skilled in the art and related fields can practice the present invention. Without departing from the spirit and scope of the present invention, logical, implementation and other changes can be made to the implementation. Therefore, the following detailed description should not be understood as limiting, and the scope of the present invention is limited only by the claims.
[0044] The present invention provides a method for solving the problems that the performance of storage system and computing system is limited by network bandwidth and the computing power of storage device is wasted. The method is mainly based on EOS file system, but can be extended to other file systems such as Lustre by analogy. Its main components or functional parts include: 1) metadata parsing module, 2) storage service parsing module, 3) storage computing function module and 4) metadata registration module. Figure 2 Each component is introduced in detail.
[0045] 1) Metadata parsing module
[0046] This is the starting point for implementing computational storage. The metadata service process analyzes and processes file access requests from storage clients. In the distributed storage system EOS, the metadata service process MGM is responsible for this part. The metadata parsing module is located before the metadata service process and filters user requests from the client.
[0047]
[0048] In this module, ThrottleMGM will parse the file access request sent by the client, which usually contains the object information of the request, the logical address of the requested file, the type of operation requested, the time when the request was initiated, etc. The parameters required to call the computable storage function are also in the file access request. The additional parameters are deleted and the modified file access request is forwarded to the metadata service process so that it can correctly obtain the file information and pass the file's physical address, file size and other information back to the client. The client then redirects the file access request to the storage service process (i.e., FST in EOS) based on the information.
[0049] For example: suppose a user wants to access a file in the storage system and rebuild it. The traditional computing method is that the user needs to first pull the file from the storage end, then run the reconstruction program locally, and finally generate the reconstructed file. In the computable storage method proposed in this patent, the user directly issues a file request to the storage system, that is, the reconstruction operation of the file can be completed directly on the storage end. For example, assuming that the file is root: / / castest2.ihep.ac.cn / / eos / data / 1.root, then in the computable storage system, the user can issue the following request to the storage system:
[0050] TFile*f1=TFile::Open("root: / / castest2.ihep.ac.cn / / eos / data / 1.root&CSS")
[0051] Here, the suffix "&CSS" is added after the file. After the storage system receives the request, the metadata parsing module first filters the request and converts the request field into TFile*f1=TFile::Open("root: / / castest2.ihep.ac.cn / / eos / data / 1.root"), so that the storage system can know the information of the FST server where the file 1.root is actually stored. After that, the FST server information is returned to the client.
[0052] 2) Storage service parsing module
[0053] This module is the main execution module for implementing computable storage. It parses the file access request sent by the client redirection, obtains the file logical address, physical address, file size, object information of the request, and additional parameters required for the computable storage function, and parses the file logical address and physical address and processes additional parameters here. This module is located before the storage service process (i.e. FST) in the message queue MQ and is deployed on the server where the storage service process is located. According to the options corresponding to the parameters, it calls different methods in the storage computing function module written by the user later (different additional parameters correspond to different data processing methods, and the object information of the request corresponds to the data object to be processed).
[0054]
[0055] After the data processing is completed, the metadata registration module needs to be called to register the files generated by the storage and computing function module back to the distributed storage system. This process does not require file transmission and is essentially to achieve synchronization between the local database and the metadata database.
[0056] 3) Storage and computing function module
[0057] This is one of the core parts of the computable storage system and the main part of the storage system to implement data processing functions. This module is provided in the form of a dynamic link library commonly used in Linux systems. It mainly extracts data compression, decompression, segmentation, format conversion, and simple data processing functions such as case screening and noise filtering in the data analysis process. The additional parameters added in the user request define the function to be called here. The storage service parsing module will also provide the required parameters to the storage computing function module according to the functions declared by the user. The following C++ code is a simple file decompression example:
[0058]
[0059] Then compile the storage service parsing module, storage computing function module and file metadata registration module to generate a dynamic link library such as libEosFstCompress.so, and configure the dynamic link library in the configuration files of MGM and FST. There is no need to recompile the original file system code, just restart the service to make the configuration effective. Different experiments have different data processing processes, so this module needs to be customized and developed according to different computing scenarios of different experiments.
[0060] 4) Metadata registration module
[0061] This module mainly registers the files generated in the storage and calculation function module back to the EOS storage system metadata to complete the final storage and calculation process. In this module, you first need to register an empty file such as / eos / site / decode / 1.dat.decode.root in EOS before starting the storage and calculation function module. EOS will not actually allocate a storage location and storage space for the empty file, but will allocate a relative storage path in the data disk such as 0000a716 / 197ee767. Then, after the storage and calculation function module completes the data processing, the generated results are stored in the local data disk in the form of files. The file name is the relative path in front plus the absolute path prefix: / data01 / eos / 0000a716 / 197ee767. Finally, the file metadata registration module writes the automatically generated file information such as timestamp, size, checksum, storage node, storage data disk, etc. to the local database levelDB where FST is located, and then synchronizes the records in the local database with the metadata database to complete the file registration process.
[0062]
[0063] The present application is not limited to the embodiments described in detail in the present invention. Those skilled in the art may make various modifications thereto. However, as long as these modifications do not deviate from the spirit and intent of the present invention, they are still within the protection scope of the present invention.
Claims
1. A method for high-energy physics IO-intensive computational storage service, comprising: 1) Use the metadata parsing module to parse and filter the file access request sent by the client, delete the parameters in the logical address of the requested file except the parameters required for calling the computable storage function, and forward the modified file access request to the metadata service process MGM in the distributed storage system so that it can correctly obtain the file information, and pass the file's physical address and file size information back to the client, so that the client can redirect and send the file access request to the storage service process FST of the FST server where the requested file is stored in the distributed storage system according to the returned information; 2) The storage service parsing module calls the corresponding method in the storage computing function module to process the data according to the file access request transmitted from the client after redirection; 3) After the storage computing function module completes the data processing, it generates a file and sends it to the storage service parsing module; then the storage service parsing module calls the metadata registration module to register the file back into the distributed storage system, thereby realizing synchronization between the FST server local database and the distributed storage system metadata database; wherein the metadata registration module first registers an empty file in the distributed storage system and assigns a relative storage path; then the file generated by the storage computing function module is stored in the local data disk, and the file name is the relative storage path plus the absolute path prefix; then the file information is written into the local database of the FST server, and then the records in the local database are synchronized with the distributed storage system metadata database to complete the file registration process.
2. The method according to claim 1, characterized in that In step 2), the storage service parsing module parses the file access request passed after redirection, obtains the file logical address, physical address, file size, object information of the request, and parameters required for the computable storage function; then, based on the parsed information, the corresponding method in the storage computing function module is called to process the requested object information.
3. The method according to claim 2, characterized in that The parameters required for the computable storage function are used to set the data processing method; the data processing method includes data compression, decompression, segmentation, format conversion, case screening, and noise filtering.
4. The method according to claim 1, characterized in that The file access request includes information about the object issuing the request, the logical address of the requested file, the type of operation requested, and the time when the request was initiated.
5. A high-energy physics IO-intensive computable storage system, characterized in that: It includes metadata parsing module, storage service parsing module, storage computing function module and metadata registration module; among which, The metadata parsing module is used to parse and filter the file access request sent by the client, delete the parameters in the logical address of the requested file except the parameters required for calling the computable storage function, and forward the modified file access request to the metadata service process MGM in the distributed storage system so that it can correctly obtain the file information, and pass the file's physical address and file size information back to the client, so that the client can redirect the file access request to the storage service process FST of the FST server in the distributed storage system where the requested file is stored according to the returned information; The storage service parsing module runs on the FST server and is used to call the corresponding method in the storage computing function module to process the data according to the file access request passed from the client after redirection. After the data processing is completed, a file is generated and sent to the storage service parsing module; the storage service parsing module calls the metadata registration module to register the file back to the distributed storage system to obtain the logical address in the system, so as to achieve synchronization between the local database of the FST server and the metadata database of the distributed storage system; A storage computing function module includes multiple data processing methods, and different data processing methods correspond to different parameters in the file access request; The metadata registration module registers the file generated by the storage and calculation function module back into the metadata repository of the distributed storage system to complete the storage and calculation process; wherein the metadata registration module first registers an empty file in the distributed storage system and assigns a relative storage path; then the file generated by the storage and calculation function module is stored in the local data disk, and the file name is the relative storage path plus the absolute path prefix; then the file information is written into the local database of the FST server, and then the records in the local database are synchronized with the metadata repository of the distributed storage system to complete the file registration process.
6. The computable storage system according to claim 5, characterized in that: The storage service parsing module parses the file access request passed after redirection to obtain the file logical address, physical address, file size, object information of the request, and parameters required for the computable storage function; then calls the corresponding method in the storage computing function module according to the parsed information to process the requested object information; the file access request includes the object information of the request, the logical address of the requested file, the requested operation type, and the request initiation time.
7. The computable storage system according to claim 6, characterized in that: The parameters required for the computable storage function are used to set the data processing method; the data processing method includes data compression, decompression, segmentation, format conversion, case screening, and noise filtering.
8. The computable storage system according to claim 5, wherein: The storage service parsing module, storage computing function module and metadata registration module are compiled to generate a dynamic link library, which is added to the configuration file of the distributed storage system and loaded when the MGM and FST processes are started.
Citation Information
Patent Citations
Data storage method and device
CN101539950A
Graph data function response method and system and electronic equipment
CN110489986A