High-concurrency Read / Write Optimization System, Medium and Device for Distributed File System
By adopting fine-grained read-write lock and neural network latency prediction models in distributed file systems, combined with concurrent execution of local and remote thread pools, the performance bottleneck problem of distributed file systems under high concurrency conditions is solved, and high performance and high concurrency read-writes are achieved.
Patent Information
- Application Number
- CN202310755386.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-06-25
AI Technical Summary
When a distributed file system faces high-concurrency file read and write requests, the performance bottleneck is serious and it is difficult to meet the needs of high performance and high concurrency.
A fine-grained read and write lock mechanism is adopted, combined with a concurrency control module based on binary tree and hash table, to achieve high concurrency control of file data. At the same time, a neural network-based read and write delay prediction model is introduced to predict the delay of read and write requests, and read and write tasks are performed concurrently through local and remote thread pools.
It realizes high-performance read and write of distributed file systems under high concurrency conditions, breaks through the performance bottleneck of traditional file read and write locks, and significantly improves the overall performance of the system.
Smart Images

Figure CN116737685B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer system architectures, and more particularly, to a high-concurrency read / write optimization system, medium, and device for a distributed file system. Background Art
[0002] Distributed storage systems are generally used to solve the problem of insufficient storage capacity of a single computer. With the advent of the network and big data era, many scientific research fields and commercial applications can no longer be satisfied by a single-node computer system. Generally speaking, a distributed storage system runs on a distributed cluster composed of multiple nodes, where some nodes store data and others perform computations, and here the computations can include data I / O requests and the logical functions of application programs. To judge the quality of a distributed storage system, it is mainly to judge whether it meets some basic specifications of the distributed system, such as consistency, availability, and partition tolerance, etc., and whether it provides high-performance data services for the computing nodes in the cluster.
[0003] In recent years, the emergence of some emerging hardware technologies is posing new challenges to the design of distributed storage systems. For example, the emergence of non-volatile memory and RDMA technology enables the design of storage systems and distributed systems to break through the original limitations and make new optimization schemes for new hardware.
[0004] PM: Persistent Memory, non-volatile memory or persistent memory, a new type of memory hardware technology. Non-volatile memory has the following three hardware characteristics: First, it can be byte-addressable and supports the direct load / store instructions of the processor; second, it has a latency and bandwidth close to DRAM; third, it is non-volatile, and the data stored is not lost when the power is off. These excellent hardware characteristics enable non-volatile memory to be used as a new storage hierarchy between memory and disk, or directly replace these two to form a single-level storage system.
[0005] RDMA: Remote Direct Memory Access, remote direct memory access technology, which enables direct memory access between different computers in a cluster. The meaning of "direct" here is to read and modify the memory of the remote end without notifying the remote CPU. From the perspective of memory computing, RDMA technology has a similar starting point to DMA (direct memory access) in a single-node computer. They both reduce the load of memory-related requests on the CPU, thereby improving the overall performance of the system. At the same time, from the perspective of network connection, RDMA technology is a new generation of network communication technology, which has lower communication latency and higher cross-node data throughput rate than the original server cluster connected by Ethernet.
[0006] Therefore, a new technical solution needs to be proposed to improve the above technical problems. Summary of the Invention
[0007] Aiming at the defects in the prior art, the purpose of the present invention is to provide a high-concurrency read-write optimization system, medium and device for a distributed file system.
[0008] According to a high-concurrency read-write optimization system for a distributed file system provided by the present invention, the system includes the following modules:
[0009] Module M1: The data read-write concurrency control module uses a fine-grained read-write lock to obtain consistent and highly concurrent data;
[0010] Module M2: The file data cache module controls the data cache of the client system for the server-side file system;
[0011] Module M3: The read-write request latency prediction module predicts the runtime latency of file read-write requests in the file system client;
[0012] Module M4: The read-write task execution module obtains the read-write performance under concurrent conditions by simultaneously executing local and remote data read-write operations;
[0013] The data read-write concurrency control module processes multi-threaded parallel file data read-write requests in multiple clients to obtain consistent file data and concurrent read-write requests;
[0014] The file data cache module constructs a file data cache in the client system, and the application saves data cache locally for the recent read-write requests of the file;
[0015] The read-write latency prediction module records the latency of data read-write tasks in the system and uses a prediction model to predict the runtime read-write request latency;
[0016] The read-write task execution module maintains two thread pools to execute data read-write tasks on the remote server and data read-write tasks on the local cache respectively, and obtains the read-write performance under concurrent conditions.
[0017] Preferably, the module M1 includes:
[0018] Module M1.1: A fine-grained file data read-write lock module based on a binary tree;
[0019] Module M1.2: A client data permission lease module based on a hash table;
[0020] The fine-grained file data read-write lock module contains a binary tree list, where each binary tree in the list corresponds to a file data read-write lock management unit in an open state; the binary tree list uses the unique identifier of the file in the distributed file system as an index;
[0021] In each binary tree, it contains the data range occupied by the read-write requests for the current file, and this range is represented by a pair of numbers consisting of the starting offset and length of the read-write request; the root node of the binary tree corresponds to the range of all data of the file, and each child node corresponds to the data range after bisecting the parent node's range; the data segment corresponding to each node contains a read-write lock, and the requester obtains the permission of one write and multiple reads simultaneously;
[0022] When a new read-write request is submitted to the file server, the binary tree of the data read-write lock corresponding to the file number is indexed, and the read-write permission record on the corresponding node is searched, added, or modified; when the read-write permission is recovered, the read-write permission on the corresponding node is modified or deleted;
[0023] The client data permission lease module contains a hash table composed of lease records, and this hash table maintains all the read-write lease permissions currently reserved by the client system;
[0024] When the client application initiates a file read-write request, it submits the read-write request to the file system server, and at the same time submits a data cache request to the module M2;
[0025] When the read-write permission request and the cache request are processed, a read-write description item is added to the hash table, and at the same time the read-write request that needs to be executed is passed to the module M3.
[0026] Preferably, the module M2 includes:
[0027] A data cache management module running in the client system of the distributed file system;
[0028] The data cache management module includes a metadata cache unit, a data cache unit, and a metadata management unit for the local data cache;
[0029] When the client system initiates a request related to metadata operations, the metadata cache unit requests metadata permission from the file system server and caches the relevant metadata area to the client local;
[0030] When the client system initiates a request related to data read-write, the data cache unit reads or writes back data to the file system server; at the same time, the metadata management unit for the local data cache modifies the metadata information in the local cache area.
[0031] Preferably, the module M3 includes:
[0032] Module M3.1: Data Read / Write Task Latency Recording Module;
[0033] Module M3.2: Read / Write Latency Prediction Model Training Module;
[0034] Module M3.3: Real-time Read / Write Task Latency Prediction Module;
[0035] When the module M4 executes a read / write task, the read / write task latency recording module is awakened and tracks the description parameters, start and end times of the read / write task; meanwhile, this module sequentially records the data structure composed of the description parameters and start / end times of each read / write task into a persistent file storage.
[0036] When the change amount of the size of the read / write task latency recording file output by the module M3.1 exceeds a certain threshold, the latency prediction model training module reads the newly added task latency records in the file and converts them into the empirical knowledge of the prediction model through the model training method.
[0037] When receiving the read / write task passed by the module M1, the real-time read / write task latency prediction module passes the received real-time read / write task into the trained latency prediction model, and simultaneously predicts the latency of the read / write task executed via the remote server node and the latency of bypassing the remote and executing via the local cache respectively; compare the two latencies to obtain the options of the read / write path scheme.
[0038] Preferably, the module M3.1 includes:
[0039] Module M3.1.1: File System Read / Write Operation Tracking Module;
[0040] Module M3.1.2: Read / Write Latency Recording Persistence and Summarization Module;
[0041] When the module M4 executes a remote or local read / write task, the file system read / write operation tracking module records the task parameters of the read / write task, and the parameters include the read / write file identification number, read / write identification bit, remote / local identification bit, read / write offset, and read / write length; meanwhile, the file system read / write operation tracking module records the start time of the read / write task.
[0042] When the module M4 completes a remote or local read / write task, the file system read / write operation tracking module records the end time of the read / write task.
[0043] When the file system read / write operation tracking module completes a task record, the read / write latency recording persistence and summarization module writes the record into a persistent local file; the file is opened in append mode and new content is sequentially written into the file each time.
[0044] When the length of the persistent local file record exceeds a threshold, the newly appended content in the file is submitted to the read-write latency prediction model training module;
[0045] When the data in the persistent local file record has been processed by the read-write latency prediction model training module, the read-write latency record persistence and summarization module truncates the old data in the local file record so that the record file size does not exceed a threshold.
[0046] Preferably, the module M3.2 includes:
[0047] Module M3.2.1: Read-write latency prediction model construction module;
[0048] Module M3.2.2: Latency prediction model training and updating module;
[0049] When the high-concurrency read-write optimization system for a distributed file system is constructed, the read-write latency prediction model construction module constructs a read-write latency prediction model based on a neural network;
[0050] The read-write latency prediction model based on a neural network includes a two-layer fully connected neural network; the input data of the neural network is the read-write task description parameters, and the output data is the real-time latency prediction result for the read-write task, with the unit being the real-world standard time unit;
[0051] Before the construction of the read-write latency prediction model is completed, there are two initialization schemes for the neural network model parameters included in the prediction model: in Scheme 1, the model parameters are initialized to all 0; in Scheme 2, the model parameters are initialized to the latency prediction model parameters that have been trained in other systems;
[0052] When the construction of the read-write latency prediction model is completed, the model is persistently stored as a file that is readable and writable by the read-write latency prediction model training module and read-only for other modules in the system;
[0053] When the latency prediction model training and updating module receives the read-write task record from the read-write latency record persistence and summarization module, it performs the latency prediction model training and updating operation;
[0054] The latency prediction model training and updating module uses the input read-write task record and the corresponding real-time latency result as the input and output training data of the neural network model respectively, and trains the latency prediction model using the gradient descent method until the error between the predicted result of the model for the latency and the actual latency data converges;
[0055] When the error obtained from the convergence of the model is less than the average error at the start of training, the delay prediction model training and update module performs model update, that is, persists the trained model as a new storage file and overwrites the original model storage file.
[0056] Preferably, the module M3.3 includes:
[0057] Module M3.3.1: Real-time device read / write pressure collection module;
[0058] Module M3.3.2: Real-time read / write task delay prediction module;
[0059] When the real-time read / write task delay prediction module receives a real-time read / write delay prediction request, the real-time device read / write pressure collection module collects the summary information of the read / write tasks being executed in the read / write task execution module in the current system, and passes this information as part of the parameters to the read / write delay prediction model based on neural network;
[0060] When the real-time device read / write pressure collection module finishes, the real-time read / write task delay prediction module calls the neural network prediction model to perform real-time read / write delay prediction;
[0061] After the real-time read / write delay prediction is completed, the read / write task delay prediction module submits the prediction result to the read / write task execution module.
[0062] Preferably, the module M4 includes:
[0063] Module M4.1: RDMA-based remote data read / write thread pool module;
[0064] Module M4.2: Non-volatile memory-based cache data read / write thread pool module;
[0065] When the high-concurrency read / write optimization system for a distributed file system is built, the RDMA-based remote data read / write thread pool module initializes a certain number of RDMA connections and completes the information exchange test from the client system to the server file system;
[0066] After the RDMA connection is established, the RDMA-based remote data read / write thread pool module creates a remote data read / write thread pool based on RDMA primitives and sets all threads to the idle state;
[0067] When a new remote read / write task is submitted to the read / write task execution module, the RDMA-based remote data read / write thread pool module schedules an idle thread to execute the task;
[0068] When the high-concurrency read-write optimization system for a distributed file system is constructed, the cache data read-write thread pool module based on non-volatile memory completes the read-write operation test of the client system on the local non-volatile memory;
[0069] After the non-volatile memory read-write operation test is completed, the cache data read-write thread pool module based on non-volatile memory creates a non-volatile memory data read-write thread pool based on load / store primitives and sets all threads to the idle state;
[0070] When a new local read-write task is submitted to the read-write task execution module, the cache data read-write thread pool module based on non-volatile memory schedules an idle thread to execute the task.
[0071] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the functions of the above-mentioned modules are implemented.
[0072] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, the functions of the above-mentioned modules are implemented.
[0073] Compared with the prior art, the present invention has the following beneficial effects:
[0074] 1. The present invention proposes optimization methods in three directions: concurrency control, latency optimization, and parallel execution, aiming at the performance bottleneck problem of the distributed file system in the face of high-concurrency file read-write requests;
[0075] 2. In terms of concurrency control, the fine-grained concurrency control mechanism used in the present invention enables concurrent read-write requests on the same file to be executed simultaneously, breaking through the performance bottleneck of the original file read-write lock and bringing a significant performance improvement;
[0076] 3. In terms of latency optimization, the present invention adopts a multi-path latency prediction and read-write path selection scheme based on neural networks, which essentially explores more hardware performance potential and makes real-time optimization decisions during system operation;
[0077] 4. The present invention improves the overall performance of the distributed file system, adapts to mainstream operating systems and file system interfaces, and has good market prospects and application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present invention will become more apparent:
[0079] Figure 1Schematic diagram of the overall device module in the embodiments of the present invention;
[0080] Figure 2 Schematic diagram of the device data read / write concurrency control module in the embodiments of the present invention;
[0081] Figure 3 Schematic diagram of the device file data cache module in the embodiments of the present invention;
[0082] Figure 4 Schematic diagram of the device read / write request delay prediction module in the embodiments of the present invention;
[0083] Figure 5 Schematic diagram of the device read / write task execution module in the embodiments of the present invention. Detailed implementation manners
[0084] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all fall within the protection scope of the present invention.
[0085] Example 1:
[0086] According to a high-concurrency read / write optimization system for a distributed file system provided by the present invention, the system includes the following modules:
[0087] Module M1: The data read / write concurrency control module uses fine-grained read / write locks to obtain consistent and highly concurrent data;
[0088] Module M1.1: Fine-grained file data read / write lock module based on a binary tree;
[0089] Module M1.2: Client data permission lease module based on a hash table;
[0090] The fine-grained file data read / write lock module includes a binary tree list, and each binary tree in the list corresponds to a file data read / write lock management unit in an open state; the binary tree list uses the unique identification number of the file in the distributed file system as an index;
[0091] In each binary tree, it includes the data interval occupied by the read / write request of the current file, and this interval is represented by a pair of numbers composed of the starting offset and length of the read / write request; the root node of the binary tree corresponds to the interval of all data of the file, and each child node corresponds to the data interval after the division of the parent node interval; the data segment corresponding to each node includes a read / write lock, and the requester obtains the permissions of one write and multiple reads at the same time;
[0092] When a new read / write request is submitted to the file server, the data read / write lock binary tree corresponding to the file number is indexed, and the read / write permission records on the corresponding nodes are looked up, added, or modified; when the read / write permissions are recycled, the read / write permissions on the corresponding nodes are modified or deleted.
[0093] The client data permission lease module contains a hash table composed of lease records, which maintains all the read / write lease permissions currently reserved by the client system.
[0094] When the client application initiates a file read / write request, it submits a read / write request to the file system server and simultaneously submits a data caching request to the module M2.
[0095] When the read / write permission request and the caching request are processed, a read / write description item is added to the hash table, and the read / write request that needs to be executed is passed to the module M3.
[0096] Module M2: The file data caching module controls the data caching of the client system for the server-side file system.
[0097] The module M2 includes: a data caching management module running in the client system of the distributed file system.
[0098] The data caching management module includes a metadata caching unit, a data caching unit, and a metadata management unit for local data caching.
[0099] When the client system initiates a request related to metadata operations, the metadata caching unit requests metadata permissions from the file system server and caches the relevant metadata area to the client local.
[0100] When the client system initiates a request related to data read / write, the data caching unit reads or writes back data from / to the file system server; meanwhile, the metadata management unit for local data caching modifies the metadata information in the local cache area.
[0101] Module M3: The read / write request latency prediction module predicts the runtime latency of file read / write requests in the file system client.
[0102] Module M3.1: The data read / write task latency recording module.
[0103] Module M3.1.1: The file system read / write operation tracking module.
[0104] Module M3.1.2: The read / write latency recording persistence and summarization module.
[0105] When the module M4 executes a remote or local read / write task, the file system read / write operation tracking module records the task parameters of this read / write task. The parameters include the read / write file identification number, the read / write identification bit, the remote / local identification bit, the read / write offset, and the read / write length. At the same time, the file system read / write operation tracking module records the start time of this read / write task;
[0106] When the module M4 completes a remote or local read / write task, the file system read / write operation tracking module records the end time of this read / write task;
[0107] When the file system read / write operation tracking module completes a task record, the read / write latency record persistence and summarization module writes this record into a persistent local file. This file is opened in append mode and new content is written to the file sequentially each time;
[0108] When the record length of the persistent local file exceeds a threshold, the newly appended content in the file is submitted to the read / write latency prediction model training module;
[0109] When the data in the record of the persistent local file has been processed by the read / write latency prediction model training module, the read / write latency record persistence and summarization module truncates the old data in the local file record, and the record file size does not exceed a threshold.
[0110] Module M3.2: Read / write latency prediction model training module;
[0111] Module M3.2.1: Read / write latency prediction model construction module;
[0112] Module M3.2.2: Latency prediction model training and update module;
[0113] When the high-concurrency read / write optimization system for a distributed file system is constructed, the read / write latency prediction model construction module constructs a read / write latency prediction model based on a neural network;
[0114] The read / write latency prediction model based on a neural network includes a two-layer fully connected neural network. The input data of the neural network is the read / write task description parameters, and the output data is the real-time latency prediction result of this read / write task, with the unit being the real-world standard time unit;
[0115] Before the construction of the read / write latency prediction model is completed, there are two initialization schemes for the neural network model parameters included in the prediction model: In Scheme 1, the model parameters are initialized to all 0; in Scheme 2, the model parameters are initialized to the latency prediction model parameters that have been trained in other systems.
[0116] After the read / write latency prediction model is constructed, the model is persistently stored as a file that is readable and writable by the read / write latency prediction model training module and read-only by other modules in the system;
[0117] When the latency prediction model training and update module receives read / write task records from the read / write latency record persistence and summary module, it performs latency prediction model training and update operations;
[0118] The latency prediction model training and update module uses the input read / write task records and the corresponding real-time latency results as the input and output training data of the neural network model respectively, and trains the latency prediction model using the gradient descent method until the error between the model's latency prediction result and the actual latency data converges;
[0119] When the error obtained by model convergence is less than the average error at the start of training, the latency prediction model training and update module performs model update, that is, persists the trained model as a new storage file and overwrites the original model storage file.
[0120] Module M3.3: Real-time read / write task latency prediction module;
[0121] Module M3.3.1: Real-time device read / write pressure collection module;
[0122] Module M3.3.2: Real-time read / write task latency prediction module;
[0123] When the real-time read / write task latency prediction module receives a real-time read / write latency prediction request, the real-time device read / write pressure collection module collects the summary information of the read / write tasks being executed in the read / write task execution module of the current system, and passes this information as part of the parameters to the neural network-based read / write latency prediction model;
[0124] When the real-time device read / write pressure collection module is completed, the real-time read / write task latency prediction module calls the neural network prediction model to perform real-time read / write latency prediction;
[0125] When the real-time read / write latency prediction is completed, the read / write task latency prediction module submits the prediction result to the read / write task execution module.
[0126] When module M4 executes a read / write task, the read / write task latency record module is awakened and tracks the description parameters and start / end times of the read / write task; at the same time, this module sequentially records the data structure composed of the description parameters and start / end times of each read / write task into the persistent file storage;
[0127] When the change amount of the read / write task delay record file size output by the module M3.1 exceeds a certain threshold, the delay prediction model training module reads the newly added task delay records recorded in the file and converts them into the empirical knowledge of the prediction model through the model training method;
[0128] When receiving the read / write task transmitted by the module M1, the real-time read / write task delay prediction module passes the received real-time read / write task into the trained delay prediction model, and simultaneously predicts the delay of the read / write task executed via the remote server node and the delay of bypassing the remote end and executing via the local cache respectively; compare the two delays to obtain the options of the read / write path scheme.
[0129] Module M4: The read / write task execution module obtains the read / write performance under concurrent conditions by simultaneously performing local and remote data read / write operations;
[0130] Module M4.1: The remote data read / write thread pool module based on RDMA;
[0131] Module M4.2: The cache data read / write thread pool module based on non-volatile memory;
[0132] When the high-concurrency read / write optimization system for a distributed file system is constructed, the remote data read / write thread pool module based on RDMA initializes a certain number of RDMA connections and completes the information exchange test from the client system to the server file system;
[0133] After the RDMA connection is established, the remote data read / write thread pool module based on RDMA creates a remote data read / write thread pool based on RDMA primitives and sets all threads to the idle state;
[0134] When a new remote read / write task is submitted to the read / write task execution module, the remote data read / write thread pool module based on RDMA schedules an idle thread to execute the task;
[0135] When the high-concurrency read / write optimization system for a distributed file system is constructed, the cache data read / write thread pool module based on non-volatile memory completes the read / write operation test of the client system on the local non-volatile memory;
[0136] After the non-volatile memory read / write operation test is completed, the cache data read / write thread pool module based on non-volatile memory creates a non-volatile memory data read / write thread pool based on load / store primitives and sets all threads to the idle state;
[0137] When a new local read / write task is submitted to the read / write task execution module, the cache data read / write thread pool module based on non-volatile memory schedules an idle thread to execute the task.
[0138] The data read / write concurrency control module processes multi-threaded parallel file data read / write requests in multiple clients to obtain consistent file data and concurrent read / write requests;
[0139] The file data caching module constructs a file data cache in the client system, and the application saves data cache locally for the most recent read / write requests for the file;
[0140] The read / write latency prediction module records the latency of data read / write tasks in the system and uses a prediction model to predict the read / write request latency during runtime;
[0141] The read / write task execution module maintains two thread pools to execute data read / write on the remote server and data read / write on the local cache respectively, so as to obtain the read / write performance under concurrent conditions.
[0142] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the functions of the above-mentioned modules are implemented.
[0143] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, the functions of the above-mentioned modules are implemented.
[0144] Example 2:
[0145] Aiming at the defects in the prior art, the purpose of the present invention is to provide a high-concurrency read / write optimization method and system for a distributed file system.
[0146] According to the high-concurrency read / write optimization method and system for a distributed file system provided by the present invention, it includes:
[0147] Module M1: Data read / write concurrency control module, which uses fine-grained read / write locks to ensure data consistency and high concurrency;
[0148] Module M2: File data caching module, which controls the data caching of the client system for the server-side file system;
[0149] Module M3: Read / write request latency prediction module, which predicts the runtime latency of file read / write requests in the file system client;
[0150] Module M4: Read / write task execution module, which ensures the optimal read / write performance under concurrent conditions by simultaneously executing local and remote data read / write operations;
[0151] The data read / write concurrency control module: processes multi-threaded parallel file data read / write requests in multiple clients, ensuring file data consistency and the concurrency of read / write requests.
[0152] The file data cache module: constructs a file data cache in the client system, ensuring that the application has data caches locally saved for the most recent read / write requests to the file.
[0153] The read / write latency prediction module: records the latency of data read / write tasks within the system and uses a prediction model to predict the read / write request latency during runtime.
[0154] The read / write task execution module: maintains two thread pools to execute data read / write tasks on the remote server and data read / write tasks on the local cache respectively, ensuring optimal read / write performance under concurrent conditions.
[0155] Preferably, the module M1 includes:
[0156] Module M1.1: a fine-grained file data read / write lock module based on a binary tree;
[0157] The fine-grained file data read / write lock module contains a binary tree list, and each binary tree in the list corresponds to a file data read / write lock management unit in an open state. The binary tree list uses the unique identification number of the file in the distributed file system as an index.
[0158] In each binary tree, it contains the data interval occupied by the read / write request for the current file, which is represented by a pair of numbers consisting of the starting offset and length of the read / write request. The root node of the binary tree corresponds to the interval of all data of the file, and each child node corresponds to the data interval after the parent node interval is bisected. Each data segment corresponding to a node contains a read / write lock, allowing the requester to obtain the permission of one write and multiple reads simultaneously.
[0159] When a new read / write request is submitted to the file server, the binary tree of the data read / write lock corresponding to the file number is indexed, and the read / write permission record on the corresponding node is searched, added, or modified. When the read / write permission is recycled, the read / write permission on the corresponding node is modified or deleted.
[0160] Module M1.2: a client data permission lease module based on a hash table;
[0161] The client data permission lease module contains a hash table composed of lease records, which maintains all the read / write lease permissions currently reserved by the client system.
[0162] When the client application initiates a file read / write request, it submits the read / write request to the file system server and simultaneously submits a data cache request to the module M2 described above.
[0163] When the read / write permission request and the cache request are processed, add a read / write description item to the hash table, and at the same time pass the read / write request that needs to be executed to the module M3 described above.
[0164] Preferably, the module M2 includes:
[0165] A data cache management module running in the distributed file system client system.
[0166] The data cache management module includes: a metadata cache unit, a data cache unit, and a metadata management unit for the local data cache.
[0167] When the client system initiates a request related to metadata operations, the metadata cache unit requests metadata permissions from the file system server and caches the relevant metadata area locally on the client;
[0168] When the client system initiates a request related to data read / write, the data cache unit reads or writes back data to / from the file system server; at the same time, the metadata management unit for the local data cache modifies the metadata information in the local cache area to ensure the consistency, correctness, and integrity of the data cache area.
[0169] Preferably, the module M3 includes:
[0170] Module M3.1: Data read / write task delay recording module;
[0171] When the module M4 described above executes a read / write task, the read / write task delay recording module is awakened and tracks the description parameters, start and end times of the read / write task; at the same time, this module sequentially records the data structure composed of the description parameters and start and end times of each read / write task into a persistent file storage.
[0172] Module M3.2: Read / write delay prediction model training module;
[0173] When the change amount of the size of the read / write task delay record file output by the module M3.1 described above exceeds a certain threshold, the delay prediction model training module reads the newly added task delay records recorded in the file and converts them into the empirical knowledge of the prediction model through the model training method.
[0174] Module M3.3: Real-time read / write task delay prediction module;
[0175] When receiving the read / write task passed by the module M1 described above, the real-time read / write task delay prediction module passes the received real-time read / write task into the trained delay prediction model, and at the same time predicts the delay of the read / write task executed via the remote server node and the delay of bypassing the remote and executing via the local cache respectively; the comparison of the two delays provides options for the read / write path optimization plan.
[0176] Preferably, the module M3.1 includes:
[0177] Module M3.1.1: File system read / write operation tracking module;
[0178] When the module M4 performs a remote or local read / write task, the file system read / write operation tracking module records the task parameters of the read / write task, and the parameters include the read / write file identification number, read / write identification bit, remote / local identification bit, read / write offset, and read / write length; at the same time, the file system read / write operation tracking module records the start time of the read / write task;
[0179] When the module M4 completes a remote or local read / write task, the file system read / write operation tracking module records the end time of the read / write task;
[0180] Module M3.1.2: Read / write latency record persistence and summary module;
[0181] When the file system read / write operation tracking module completes a task record, the read / write latency record persistence and summary module writes the record into a persistent local file; the file is opened in append mode and new content is written sequentially each time new content is appended;
[0182] When the record length of the persistent local file exceeds a threshold, the newly appended content in the file is submitted to the read / write latency prediction model training module described above;
[0183] When the data in the record of the persistent local file has been processed by the read / write latency prediction model training module described above, the read / write latency record persistence and summary module truncates the old data in the local file record to ensure that the record file size does not exceed a threshold.
[0184] Preferably, the module M3.2 includes:
[0185] Module M3.2.1: Read / write latency prediction model construction module;
[0186] When the high-concurrency read / write optimization system for a distributed file system is constructed, the read / write latency prediction model construction module constructs a read / write latency prediction model based on a neural network;
[0187] The read / write latency prediction model based on a neural network includes: a two-layer fully connected neural network; the input data of the neural network is the read / write task description parameters, and the output data is the real-time latency prediction result of the read / write task, with the unit of real-world standard time unit;
[0188] Before the construction of the read-write latency prediction model is completed, there are two initialization schemes for the neural network model parameters included in the prediction model: in Scheme 1, the model parameters are initialized to all 0; in Scheme 2, the model parameters are initialized to the latency prediction model parameters after training in other systems.
[0189] After the construction of the read-write latency prediction model is completed, the model is persistently stored as a file, which is readable and writable by the read-write latency prediction model training module and read-only for other modules in the system.
[0190] Module M3.2.2: Latency Prediction Model Training and Update Module
[0191] When the latency prediction model training and update module receives the read-write task records from the read-write latency record persistence and summary module, it performs the latency prediction model training and update operation.
[0192] The latency prediction model training and update module uses the input read-write task records and the corresponding real-time latency results as the input and output training data of the neural network model respectively, and trains the latency prediction model using the gradient descent method until the error between the model's predicted latency result and the actual latency data converges.
[0193] When the error obtained by the model convergence is less than the average error at the start of training, the latency prediction model training and update module performs model update, that is, persists the trained model as a new storage file and overwrites the original model storage file.
[0194] Preferably, the module M3.3 includes:
[0195] Module M3.3.1: Real-time Device Read-Write Pressure Collection Module
[0196] When the real-time read-write task latency prediction module receives a real-time read-write latency prediction request, the real-time device read-write pressure collection module collects the summary information of the read-write tasks being executed in the read-write task execution module in the current system, and passes this information as part of the parameters to the neural network-based read-write latency prediction model.
[0197] Module M3.3.2: Real-time Read-Write Task Latency Prediction Module
[0198] When the real-time device read-write pressure collection module is completed, the real-time read-write task latency prediction module calls the neural network prediction model to perform real-time read-write latency prediction.
[0199] After the real-time read-write latency prediction is completed, the read-write task latency prediction module submits the prediction result to the read-write task execution module.
[0200] Preferably, the module M4 includes:
[0201] Module M4.1: RDMA-based Remote Data Read / Write Thread Pool Module;
[0202] When the high-concurrency read / write optimization system for a distributed file system is constructed, the RDMA-based remote data read / write thread pool module initializes a certain number of RDMA connections and completes the information exchange test from the client system to the server file system;
[0203] After the RDMA connection is established, the RDMA-based remote data read / write thread pool module creates a remote data read / write thread pool based on RDMA primitives and sets all threads to the idle state;
[0204] When a new remote read / write task is submitted to the read / write task execution module, the RDMA-based remote data read / write thread pool module schedules an idle thread to execute the task;
[0205] Module M4.2: Non-Volatile Memory-Based Cache Data Read / Write Thread Pool Module;
[0206] When the high-concurrency read / write optimization system for a distributed file system is constructed, the non-volatile memory-based cache data read / write thread pool module completes the read / write operation test of the client system on the local non-volatile memory;
[0207] After the non-volatile memory read / write operation test is completed, the non-volatile memory-based cache data read / write thread pool module creates a non-volatile memory data read / write thread pool based on load / store primitives and sets all threads to the idle state;
[0208] When a new local read / write task is submitted to the read / write task execution module, the non-volatile memory-based cache data read / write thread pool module schedules an idle thread to execute the task.
[0209] A computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it realizes the functions of the above modules.
[0210] As Figure 1 shown, a high-concurrency read / write optimization system for a distributed file system provided by the present invention includes:
[0211] A data read / write concurrency control module, which constructs fine-grained read / write locks at the file system server side to ensure data consistency and high concurrency, and constructs a client data permission lease record table based on a hash table at the file system client side to reduce the time overhead caused by concurrent lock requests;
[0212] A file data caching module that controls the client system to cache data from the server-side file system;
[0213] A read / write request latency prediction module that constructs a read / write task latency prediction model based on a neural network structure to predict the real-time file read / write request latency in the file system client;
[0214] A read / write task execution module that schedules local and remote read / write tasks by constructing two read / write task thread pools, and simultaneously performs local and remote data read / write operations to improve read / write performance under concurrent conditions.
[0215] When a file read / write request is initiated by the file system client application, the request is first submitted to the concurrency control module. After parsing the read / write request, the module submits a data caching request to the data caching module, and simultaneously passes the read / write task description to the latency prediction module to obtain the read / write latency prediction result;
[0216] When both the data caching operation and the read / write latency prediction are completed, the read / write task is submitted by the latency prediction module to the task execution module and is concurrently executed by the read / write thread pool therein.
[0217] As Figure 2 shown, the read / write concurrency control module provided by the present invention includes:
[0218] A fine-grained file data read / write lock module based on a binary tree, which contains a binary tree list. Each binary tree in the list corresponds to a file data read / write lock management unit in an open state. The binary tree list uses the unique identification number of the file in the distributed file system as an index.
[0219] In each binary tree, it contains the data range occupied by the read / write request for the current file, and this range is represented by a pair of numbers consisting of the start offset and length of the read / write request. The root node of the binary tree corresponds to the range of all data of the file, and each child node corresponds to the data range after the parent node range is bisected. Each node corresponds to a data segment that contains a read / write lock, allowing the requester to obtain the permission of one write and multiple reads simultaneously.
[0220] When a new read / write request is submitted to the file server, the binary tree of the data read / write lock corresponding to the file number is indexed, and the read / write permission record on the corresponding node is searched, added, or modified. When the read / write permission is recycled, the read / write permission on the corresponding node is modified or deleted.
[0221] A client data permission lease module based on a hash table, which contains a hash table composed of lease records. This hash table maintains all the read / write lease permissions currently reserved by the client system.
[0222] When the client application initiates a file read / write request, it submits the read / write request to the file system server, and at the same time submits a data caching request to the file data caching module of the present invention.
[0223] When the read / write permission request and the caching request are processed, a read / write description item is added to the hash table, and at the same time, the read / write request to be executed is passed to the read / write latency prediction module of the present invention.
[0224] In the distributed file system, the file data read / write lock module and the data permission lease module perform information transmission through the RDMA communication unit between the client and the server.
[0225] As Figure 3 shown, the file data caching module provided by the present invention includes:
[0226] A metadata caching unit, a file data caching unit, and a metadata management unit for local data caching.
[0227] The metadata caching unit caches the metadata information in the file system inode, and provides caching services for the metadata-related operations of the client to reduce the time overhead;
[0228] The file data caching unit only caches the actual data in the file data block, and provides caching services for the data read / write-related operations of the client to reduce the time overhead. This unit also uses the MESI cache coherence protocol;
[0229] The metadata management unit for local data caching is responsible for the space allocation of the local cache area, allocating non-volatile memory when new cache space is needed, sorting and merging the cache content when the cache area is full, or replacing outdated cache data blocks and metadata data blocks;
[0230] When the client system initiates a request related to metadata operations, the metadata caching unit requests metadata permissions from the file system server and caches the relevant metadata area locally on the client;
[0231] When the client system initiates a request related to data read / write, the data caching unit reads or writes back data to / from the file system server; at the same time, the metadata management unit for local data caching modifies the metadata information in the local cache area to ensure the consistency and correctness of the data cache area.
[0232] As Figure 4 shown, the read / write latency prediction module provided by the present invention includes three sub-modules: a data read / write task latency recording module, a read / write latency prediction model training module, and a real-time read / write task latency prediction module;
[0233] Among them, the data read / write task delay recording module further includes two sub-modules: a file system read / write operation tracking module and a read / write delay recording persistence and summarization module;
[0234] During system operation, the file system read / write operation tracking module performs the following operations:
[0235] When the read / write task execution module of the present invention executes a remote or local read / write task, the file system read / write operation tracking module records the task parameters of the read / write task, and the parameters include the read / write file identification number, the read / write identification bit, the remote / local identification bit, the read / write offset, and the read / write length; at the same time, the file system read / write operation tracking module records the start time of the read / write task;
[0236] When the read / write task execution module of the present invention completes a remote or local read / write task, the file system read / write operation tracking module records the end time of the read / write task;
[0237] During system operation, the read / write delay recording persistence and summarization module performs the following operations:
[0238] When the file system read / write operation tracking module described above completes a task record, the read / write delay recording persistence and summarization module writes the record into a persistent local file; the file is opened in append mode, and new content is written to the file sequentially each time;
[0239] When the record length of the persistent local file exceeds a threshold, the newly appended content in the file is submitted to the read / write delay prediction model training module;
[0240] When the data in the record of the persistent local file has been processed by the read / write delay prediction model training module, the read / write delay recording persistence and summarization module truncates the old data in the local file record to ensure that the record file size does not exceed a threshold. The setting of this threshold can ensure that the file system storage space occupied by the read / write delay prediction model is controllable, and at the same time does not affect the storage of the latest data to support model training.
[0241] In addition, the read / write delay prediction model training module includes two sub-modules: a read / write delay prediction model construction module and a delay prediction model training and update module;
[0242] During the initialization of the distributed file system, the read / write delay prediction model construction module completes the task of constructing a read / write delay prediction model based on a neural network;
[0243] The neural network-based read / write latency prediction model includes: a two-layer fully connected neural network; the input data of the neural network is the read / write task description parameters, and the output data is the real-time latency prediction result for the read / write task, with the unit being the real-world standard time unit, such as microseconds;
[0244] Before the construction of the read / write latency prediction model is completed, there are two initialization schemes for the neural network model parameters included in the prediction model: in Scheme 1, the model parameters are initialized to all 0; in Scheme 2, the model parameters are initialized to the latency prediction model parameters that have been trained in other systems.
[0245] When the read / write latency prediction model is constructed and completed, the model is persistently stored as a file, which is readable and writable by the read / write latency prediction model training module and read-only by other modules in the system of the present invention.
[0246] During system operation, when the latency prediction model training and update module receives read / write task records from the read / write latency record persistence and summary module, it performs latency prediction model training and update operations.
[0247] The training data is sent to the latency prediction model training and update module in batches. We set the single-batch data volume of the training data to a power of 2, such as 2048 task records, to optimize the memory overhead of the neural network.
[0248] At the start of training, the latency prediction model training and update module uses the input read / write task records and the corresponding real-time latency results as the input and output training data of the neural network model, and trains the latency prediction model using the gradient descent method until the error between the model's predicted latency result and the actual latency data converges, that is, there is no more than a 1% error decrease in the training error within one training cycle.
[0249] When the error obtained by the model convergence is less than the average error at the start of training, the latency prediction model training and update module performs model update, that is, persists the trained model as a new storage file and overwrites the original model storage file.
[0250] In addition, the real-time read / write task latency prediction module includes two sub-modules: a real-time device read / write pressure collection module and a real-time read / write task latency prediction module;
[0251] During system operation, when the task latency prediction module receives a real-time read / write latency prediction request, the first sub-module: the real-time device read / write pressure collection module collects the summary information of the read / write tasks being executed in the read / write task execution module of the current system, and passes this information as part of the parameters to the neural network-based read / write latency prediction model to assist in real-time latency prediction.
[0252] When the system is running, when the real-time device read / write pressure collection module is completed, the real-time read / write task delay prediction module calls the neural network prediction model to perform real-time read / write delay prediction;
[0253] After the real-time read / write delay prediction is completed, the read / write task delay prediction module submits the prediction result to the read / write task execution module to assist this module in selecting an appropriate read / write task thread pool and completing efficient data read / write.
[0254] As Figure 5 shown, the read / write task execution module provided by the present invention includes two sub-modules: a remote data read / write thread pool module based on RDMA and a cache data read / write thread pool module based on non-volatile memory;
[0255] The remote data read / write thread pool module based on RDMA is responsible for scheduling and executing remote read / write tasks assigned to the distributed file system server;
[0256] When the high-concurrency read / write optimization system for the distributed file system of the present invention is constructed, the remote data read / write thread pool module based on RDMA initializes a certain number of RDMA connections and completes the information exchange test from the client system to the server file system, including send / recv operation tests and read / write tests;
[0257] After the RDMA connection is established, the remote data read / write thread pool module based on RDMA creates a remote data read / write thread pool based on RDMA primitives and sets all threads to the idle state;
[0258] When a new remote read / write task is submitted to the read / write task execution module, the remote data read / write thread pool module based on RDMA schedules an idle thread to execute the task;
[0259] The cache data read / write thread pool module based on non-volatile memory is responsible for scheduling and executing read / write tasks assigned to the local non-volatile memory buffer;
[0260] When the high-concurrency read / write optimization system for the distributed file system of the present invention is constructed, the cache data read / write thread pool module based on non-volatile memory completes the read / write operation test of the client system on the local non-volatile memory, including non-volatile memory read / write tests using load / store;
[0261] After the non-volatile memory read / write operation test is completed, the non-volatile memory cache data read / write thread pool module creates a non-volatile memory data read / write thread pool based on memory read / write primitives and sets all threads to the idle state; when a new local read / write task is submitted to the read / write task execution module, the non-volatile memory cache data read / write thread pool module schedules an idle thread to execute the task.
[0262] Those skilled in the art can understand this embodiment as a more specific description of Embodiment 1.
[0263] Those skilled in the art know that in addition to implementing the system and its various devices, modules, and units provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system and its various devices, modules, and units provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc. to achieve the same functions. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be regarded as the structure within the hardware component; the devices, modules, and units for implementing various functions can also be regarded as either software modules for implementing the method or the structure within the hardware component.
[0264] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined arbitrarily with each other.
Claims
1. A high-concurrency read-write optimization system for a distributed file system, characterized in that The system includes the following modules: Module M1: The data read-write concurrency control module adopts fine-grained read-write locks to obtain consistent and highly concurrent data; Module M2: The file data cache module controls the data cache of the client system for the server-side file system; Module M3: The read-write request latency prediction module predicts the runtime latency of file read-write requests in the file system client; Module M4: The read-write task execution module obtains read-write performance under concurrent conditions by simultaneously executing local and remote data read-write operations; The data read-write concurrency control module processes multi-threaded parallel file data read-write requests in multiple clients to obtain consistent file data and concurrent read-write requests; The file data cache module constructs the file data cache in the client system, and the application saves data cache locally for the recent read-write requests of the file; The read-write request latency prediction module records the latency of data read-write tasks in the system and uses a prediction model to predict the runtime read-write request latency; The read-write task execution module maintains two thread pools to execute data read-write tasks on the remote server and local cache respectively to obtain read-write performance under concurrent conditions; The module M1 includes: Module M1.1: The fine-grained file data read-write lock module based on a binary tree; Module M1.2: The client data permission lease module based on a hash table; The fine-grained file data read-write lock module contains a binary tree list, and each binary tree in the list corresponds to a file data read-write lock management unit in the open state; the binary tree list uses the unique identification number of the file in the distributed file system as an index; Each binary tree contains the data range occupied by the read-write request for the current file, and this range is represented by a pair of numbers composed of the starting offset and length of the read-write request; the root node of the binary tree corresponds to the range of all data of the file, and each child node corresponds to the data range after the parent node range is bisected; the data segment corresponding to each node contains a read-write lock, and the requester obtains the permission of one write and multiple reads at the same time; When a new read-write request is submitted to the file server, the binary tree of the data read-write lock corresponding to the file number is indexed, and the read-write permission record on the corresponding node is searched, added, or modified; when the read-write permission is recycled, the read-write permission on the corresponding node is modified or deleted; The client data permission lease module contains a hash table composed of lease records, and this hash table maintains all read-write lease permissions currently reserved by the client system; When the client application initiates a file read-write request, it submits the read-write request to the file system server and simultaneously submits a data cache request to the module M2; When the read-write permission request and cache request are processed, a read-write description item is added to the hash table, and at the same time, the read-write request to be executed is passed to the module M3; 2. The high-concurrency read-write optimization system for a distributed file system according to claim 1, wherein The module M2 includes: A data cache management module running in the client system of the distributed file system; The data cache management module includes a metadata cache unit, a data cache unit, and a metadata management unit for the local data cache; When the client system initiates a request related to metadata operations, the metadata cache unit requests metadata permissions from the file system server and caches the relevant metadata area locally on the client; When the client system initiates a request related to data reading and writing, the data cache unit reads or writes back data from / to the file system server; meanwhile, the metadata management unit of the local data cache modifies the metadata information in the local cache area.
3. The high-concurrency read-write optimization system for a distributed file system according to claim 1, characterized in that The module M3 includes: Module M3.1: Data Read / Write Task Delay Recording Module; Module M3.2: Read / Write Delay Prediction Model Training Module; Module M3.3: Real-Time Read / Write Task Delay Prediction Module; When the module M4 executes a read / write task, the read / write task delay recording module is awakened and tracks the description parameters, start and end times of the read / write task; meanwhile, this module sequentially records the data structure composed of the description parameters and start and end times of each read / write task into a persistent file storage. When the change amount of the size of the read / write task delay recording file output by the module M3.1 exceeds a certain threshold, the delay prediction model training module reads the newly added task delay records recorded in the file and converts them into the empirical knowledge of the prediction model through the model training method. When receiving the read / write task passed by the module M1, the real-time read / write task delay prediction module passes the received real-time read / write task into the trained delay prediction model, and simultaneously predicts the delay of the read / write task executed via the remote server node and the delay of bypassing the remote and executing via the local cache respectively; compare the two delays to obtain the options of the read / write path scheme.
4. The high-concurrency read-write optimization system for a distributed file system according to claim 3, characterized in that The module M3.1 includes: Module M3.1.1: File System Read / Write Operation Tracking Module; Module M3.1.2: Read / Write Delay Recording Persistence and Summarization Module; When the module M4 executes a remote or local read / write task, the file system read / write operation tracking module records the task parameters of the read / write task, and the parameters include the read / write file identification number, read / write identification bit, remote / local identification bit, read / write offset, and read / write length; meanwhile, the file system read / write operation tracking module records the start time of the read / write task. When the module M4 completes a remote or local read / write task, the file system read / write operation tracking module records the end time of the read / write task. When the file system read / write operation tracking module completes a task record, the read / write delay recording persistence and summarization module writes the record into a persistent local file; the file is opened in append mode and new content is sequentially written into the file each time. When the length of the persistent local file record exceeds a threshold, the newly appended content in the file is submitted to the read / write delay prediction model training module; When the data in the persistent local file record has been processed by the read / write delay prediction model training module, the read / write delay recording persistence and summarization module truncates the old data in the local file record, and the record file size does not exceed a threshold.
5. The high-concurrency read-write optimization system for a distributed file system according to claim 3, characterized in that The module M3.2 includes: Module M3.2.1: Read / Write Delay Prediction Model Construction Module; Module M3.2.2: Latency Prediction Model Training and Update Module; When the high-concurrency read-write optimization system for a distributed file system is constructed, the read-write latency prediction model construction module constructs a read-write latency prediction model based on a neural network; The read-write latency prediction model based on a neural network includes a two-layer fully-connected neural network; the input data of the neural network is the read-write task description parameters, and the output data is the real-time latency prediction result for the read-write task, with the unit being the real-world standard time unit; Before the construction of the read-write latency prediction model is completed, there are two initialization schemes for the neural network model parameters included in the prediction model: in Scheme 1, the model parameters are initialized to all 0; in Scheme 2, the model parameters are initialized to the parameters of the latency prediction model that has been trained in other systems; When the read-write latency prediction model is constructed, the model is persistently stored as a file that is readable and writable by the read-write latency prediction model training module and read-only by other modules in the system; When the latency prediction model training and update module receives the read-write task records from the read-write latency record persistence and summary module, it performs the latency prediction model training and update operation; The latency prediction model training and update module uses the input read-write task records and the corresponding real-time latency results as the input and output training data of the neural network model, and trains the latency prediction model using the gradient descent method until the error between the predicted result of the model for latency and the actual latency data converges; When the error obtained by the model convergence is less than the average error at the start of training, the latency prediction model training and update module performs model update, that is, persists the trained model as a new storage file and overwrites the original model storage file.
6. The high-concurrency read-write optimization system for a distributed file system according to claim 3, characterized in that The module M3.3 includes: Module M3.3.1: Real-time Device Read-Write Pressure Collection Module; Module M3.3.2: Real-time Read-Write Task Latency Prediction Module; When the real-time read-write task latency prediction module receives a real-time read-write latency prediction request, the real-time device read-write pressure collection module collects the summary information of the read-write tasks being executed in the read-write task execution module in the current system, and passes this information as part of the parameters to the read-write latency prediction model based on a neural network; When the real-time device read-write pressure collection module is completed, the real-time read-write task latency prediction module calls the neural network prediction model to perform real-time read-write latency prediction; When the real-time read-write latency prediction is completed, the read-write task latency prediction module submits the prediction result to the read-write task execution module.
7. The high-concurrency read-write optimization system for a distributed file system according to claim 1, characterized in that The module M4 includes: Module M4.1: RDMA-based Remote Data Read-Write Thread Pool Module; Module M4.2: Non-Volatile Memory-based Cache Data Read-Write Thread Pool Module; When the high-concurrency read-write optimization system for a distributed file system is constructed, the RDMA-based remote data read-write thread pool module initializes a certain number of RDMA connections and completes the information exchange test from the client system to the server file system; After the RDMA connection is established, the RDMA-based remote data read / write thread pool module creates a remote data read / write thread pool based on RDMA primitives and sets all threads to the idle state; When a new remote read / write task is submitted to the read / write task execution module, the RDMA-based remote data read / write thread pool module schedules an idle thread to execute the task; When the high-concurrency read / write optimization system for a distributed file system is constructed, the non-volatile memory-based cache data read / write thread pool module completes the read / write operation test of the client system on the local non-volatile memory; After the non-volatile memory read / write operation test is completed, the non-volatile memory-based cache data read / write thread pool module creates a non-volatile memory data read / write thread pool based on load / store primitives and sets all threads to the idle state; When a new local read / write task is submitted to the read / write task execution module, the non-volatile memory-based cache data read / write thread pool module schedules an idle thread to execute the task.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the functions of the module described in any one of claims 1 to 7.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the functions of the module described in any one of claims 1 to 7.
Citation Information
Patent Citations
Small file access method accelerated based on solid state disk for distributed file system
CN106775446A
Distributed file data block read-write method and system based on RDMA and nonvolatile memory
CN111125049A