Data processing system, method and equipment

By configuring the JuiceFS module and Redis service module in the terminal device, it directly communicates with the ceph storage system, solving the performance loss problem caused by the complex architecture of the Ceph storage system, and achieving efficient data read and write performance.

CN120162009APending Publication Date: 2025-06-17BEIJING TIANDI CHAOYUN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510323175.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Due to the complex architecture of Ceph storage systems, the storage performance loss is large, especially when the object gateway RGW running in the user state communicates with the underlying RADOS through an additional abstraction layer, unnecessary performance overhead will be introduced.

Method used

The JuiceFS module and the Redis service module are configured in the terminal device. The JuiceFS module communicates with the switching unit through the network card, and directly communicates with the storage unit composed of multiple node devices of the ceph storage system, avoiding additional abstraction layers, thereby improving data read and write performance.

Benefits of technology

Through distributed design and high-performance storage characteristics, the storage performance loss is reduced, the read and write performance is improved, and scenarios with high concurrency and high random I/O requirements are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162009A_ABST
    Figure CN120162009A_ABST
Patent Text Reader

Abstract

The invention provides a data processing system, method and equipment, terminal equipment is provided with a JuiceFS module and a Redis service module, after the JuiceFS module receives a data writing request, to-be-written data is written into a preset memory buffer area through the Redis service module, then first target data in the preset memory buffer area is subjected to data splitting, and the first target data is subjected to data processing; the obtained first data blocks are written into a storage unit through a network card and a switching unit in sequence; the JuiceFS module receives a data reading request, reads the multiple second data blocks from the storage unit through the network card and the exchange unit in sequence, combines the multiple second data blocks and outputs complete second target data, and the system is based on the storage architecture of the JuiceFS module and utilizes the high-performance storage characteristic of the Redis service module at the same time. The distributed design of the data processing system is realized, and an additional abstraction layer does not need to be introduced, so that the read-write performance is improved, and the loss of the storage performance is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a data processing system, method and device. Background Art

[0002] As a highly scalable storage system, Ceph provides services such as block storage, object storage, and file storage, and different storage types can be selected according to different requirements. Its underlying uses RADOS (RADically Scalable Object Store) for data storage and management, ensuring high availability, reliability, and scalability. Generally speaking, Ceph storage refers to the file systems such as Samba / nfs, objects such as S3, and protocols such as blocks exposed by the Ceph cluster after building a cluster on multiple nodes. Customers perform operations such as reading, writing, and deleting data stored in the cluster through the protocols exposed by the Ceph cluster. However, the related technology has the problem of performance loss due to the complex system architecture. For example, the object gateway RGW (RADOS Gateway) of Ceph runs in the user state and communicates with the underlying RADOS through an additional abstraction layer, which may introduce unnecessary storage performance overhead and cause relatively large storage performance loss. Summary of the Invention

[0003] The purpose of the present invention is to provide a data processing system, method and device to reduce storage performance loss.

[0004] A data processing system provided by the present invention includes: a terminal device, a switching unit, and a storage unit that are communicatively connected in sequence; wherein, a JuiceFS module and a Redis service module that are communicatively connected to each other are configured in the terminal device, and the JuiceFS module is communicatively connected to the switching unit through a network card; the storage unit is a Ceph storage system composed of multiple node devices; the JuiceFS module is used to receive a data writing request for data to be written, write the data to be written into a preset memory buffer of the terminal device through the Redis service module, perform data splitting processing on the first target data in the preset memory buffer to obtain multiple split first data blocks, and write the multiple first data blocks into the storage unit through the network card and the switching unit in sequence; wherein, the first target data includes at least the data to be written; the JuiceFS module is used to receive a data reading request for second target data, read multiple second data blocks corresponding to the second target data from the storage unit through the network card and the switching unit in sequence, and perform merging processing on the multiple second data blocks to obtain the second target data.

[0005] Further, each node device is configured with a gigabit network interface and a ten-gigabit network interface.

[0006] Further, the switching unit includes: multiple switches; the multiple switches are connected in a stacking mode.

[0007] Further, the JuiceFS module is used for: dividing the first target data in the preset memory buffer according to a preset data volume to obtain multiple logical blocks; splitting each logical block to obtain multiple data segments respectively corresponding to each logical block; determining multiple first data blocks based on each data segment.

[0008] Further, the JuiceFS module is used for: saving the read second target data to a preset directory through the Redis service module.

[0009] Further, the storage unit is configured with multiple storage pools; wherein, at least a part of the storage pools are configured according to a preset erasure code policy, and the other storage pools except at least a part of the storage pools are configured according to a preset replication configuration policy.

[0010] Further, the other storage pools use solid state drives as the storage medium.

[0011] Further, the data stored in the storage unit is stored in a pre-created index-free container.

[0012] A data processing method provided by the present invention is applied to the data processing system of any one of the above. The system includes: a terminal device, a switching unit, and a storage unit that are communicatively connected in sequence; wherein, the terminal device is configured with a JuiceFS module and a Redis service module that are communicatively connected to each other, and the JuiceFS module is communicatively connected to the switching unit through a network card; the storage unit is a ceph storage system composed of multiple node devices. The method includes: the JuiceFS module receives a data write request for the data to be written, writes the data to be written into the preset memory buffer of the terminal device through the Redis service module, performs data splitting processing on the first target data in the preset memory buffer to obtain multiple split first data blocks, and writes the multiple first data blocks into the storage unit through the network card and the switching unit in sequence; wherein, the first target data includes at least the data to be written; the JuiceFS module receives a data read request for the second target data, reads multiple second data blocks corresponding to the second target data from the storage unit through the network card and the switching unit in sequence, and performs merging processing on the multiple second data blocks to obtain the second target data.

[0013] A data processing device provided by the present invention includes the data processing system of any one of the above.

[0014] The data processing system, method, and device provided by the present invention configure a JuiceFS module and a Redis service module in a terminal device. After receiving a data writing request for data to be written, the JuiceFS module can write the data to be written into a preset memory buffer of the terminal device through the Redis service module, and then perform data splitting processing on the first target data in the preset memory buffer, and write the obtained multiple first data blocks into a storage unit in sequence through a network card and a switching unit; after receiving a data reading request for second target data, the JuiceFS module can sequentially read multiple second data blocks corresponding to the second target data from the storage unit through the network card and the switching unit, and then obtain the second target data. Based on the storage architecture of the JuiceFS module itself and utilizing the high-performance storage characteristics of the Redis service module, the system realizes the distributed design of the data processing system, does not require the introduction of an additional abstraction layer, thereby helping to improve the reading and writing performance and further reducing the loss of storage performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0016] Figure 1 The structural schematic diagram of a data processing system provided by an embodiment of the present invention;

[0017] Figure 2 The application schematic diagram of a JuiceFS module provided by an embodiment of the present invention;

[0018] Figure 3 The schematic diagram of a data writing process provided by an embodiment of the present invention;

[0019] Figure 4 The schematic diagram of a data reading process provided by an embodiment of the present invention;

[0020] Figure 5 The architecture schematic diagram of a data processing system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0022] Currently, there are performance problems in the file system, object, and block services directly provided by Ceph storage, which are as follows:

[0023] 1. Metadata Server (MDS) bottleneck: Metadata operations (such as creating, deleting, and renaming files) rely on MDS. When the MDS load is high, it will become a performance bottleneck; the scalability of the MDS cluster is limited. Although the pressure can be shared through the multi-active mode, the complex distributed lock management may affect performance. Poor performance of small files: Ceph storage has poor read and write performance for small files. Frequent metadata operations and scattered data result in high latency.

[0024] 2. The multi-layer architecture affects performance: The RGW service runs in the user space and communicates with the underlying RADOS through librados. The additional abstraction layer may introduce performance overhead. Limited high-concurrency processing ability: For high-concurrency object storage access scenarios (such as massive picture uploads and downloads), the processing ability of RGW may not be comparable to that of dedicated object storage services (such as Amazon S3). S3 interface compatibility: Although Ceph's object storage supports the S3 API, some advanced functions (such as AWS Lambda triggers and lifecycle management rules) may not be fully compatible. Insufficient management tools: Compared with dedicated object storage solutions, the monitoring and management tools for RGW are relatively few, and the customization functions are insufficient.

[0025] 3. High latency issue: The I / O operations of RBD (RADOS Block Device, an important component in the Ceph storage system) involve multiple network hops (from the client to the OSD cluster, where OSD stands for Object Storage Device, which is the basic storage unit in the Ceph distributed storage system). The latency is higher than that of traditional SAN (Storage Area Network) storage, especially in application scenarios with low latency requirements. Weak random I / O performance: For scenarios with high random I / O requirements (such as database loads), the performance of RBD may be inferior to that of local SSD (Solid State Drive) or NVMe (Non-Volatile Memory Express, a storage protocol and interface specification specifically designed for SSDs). Network dependence: The performance of RBD highly depends on network bandwidth and stability, and network bottlenecks directly affect performance. Host operating system support: Although RBD supports mainstream operating systems, in some environments (such as Windows), additional drivers or tools may be required, increasing the deployment difficulty. Difficulty in integrating with traditional storage: When integrating with traditional SAN and NAS (Network Attached Storage) systems, additional gateway devices or conversion tools may be needed. Based on this, the embodiments of the present invention provide a data processing system, method, and device. This technology can be applied to application scenarios that require improving the data read and write performance of storage units, etc.

[0026] To facilitate the understanding of this embodiment, first, a data processing system disclosed in the embodiments of the present invention will be introduced, as Figure 1 shown, the system includes: a terminal device 10, a switching unit 11, and a storage unit 12 that are communicatively connected in sequence; among them, the terminal device 10 is configured with a JuiceFS module and a Redis service module that are communicatively connected to each other. The JuiceFS module is communicatively connected to the switching unit through a network card; the storage unit 12 is a ceph storage system composed of multiple node devices; the above terminal device 10 can be a testing machine, a computer, or other devices. The terminal device 10 is configured with a JuiceFS module and a Redis service module communicatively connected thereto. The JuiceFS module is an open-source cloud-native distributed file system designed to provide an efficient, scalable, and easy-to-manage solution for large-scale data storage and processing. The advantages of the JuiceFS module mainly include the following aspects:

[0027] 1. Cloud-native design: The JuiceFS module is a distributed file system designed for cloud-native applications. It is built on top of a Redis service module and storage units. It stores file metadata through the Redis service module, enabling massive cloud storage to be directly deployed into production environments and suitable for various application platforms such as big data, machine learning, and artificial intelligence.

[0028] 2. High compatibility: The JuiceFS module supports mainstream public cloud object storage such as Amazon S3 and can use multiple open-source software as metadata engines, such as the Redis service module, TiKV (Ti Distributed Key-Value Store, a distributed key-value storage system), MySQL (a relational database management system), etc. This design allows the JuiceFS module to adapt to different storage requirements and environments.

[0029] 3. High performance: The JuiceFS module adopts an architecture that separates the storage of "data" and "metadata". Leveraging the high-performance key-value data storage characteristics of the Redis service module, it realizes the distributed design of the file system and improves read and write performance.

[0030] 4. Flexibility and scalability: The architecture design of the JuiceFS module enables users to choose a suitable database as the metadata engine according to their own scenarios. This flexibility accelerates product iteration and can better meet specific requirements.

[0031] 5. Strong consistency and stability: The JuiceFS module is a distributed file system that achieves strong consistency guarantees by sharing the same database and object storage. It performs well in large model training under a multi-cloud architecture, ensuring storage stability.

[0032] 6. Low cost: The JuiceFS module has achieved good results in reducing operation and maintenance costs and storage costs, making it suitable for large-scale deployment and use.

[0033] The following is provided as Figure 2The application schematic diagram of a JuiceFS module is shown. At the client layer (Client), the application running environment supports multi-platform clients of Linux / Windows / macOS; the file system protocols supported by the user-space file system FUSE (Filesystem in Userspace) include: POSIX (Portable Operating System Interface, POSIX is a series of standards aimed at ensuring the portability of application programs across different operating systems. These standards define the functions and behaviors that an operating system should support, including the file system interface), SMB / CIFS (Server Message Block / Common Internet File System, SMB and CIFS are two file sharing protocols), and NFS (Network File System, a distributed file system protocol). The client can support Java development interfaces and CSI (Container Storage Interface), and can also provide Web services through the HTTP protocol. The JuiceFS module can support Local Metadata Cache and Local Data Cache. The Metadata Engine includes Redis, MySQL, etc., and Data Storage (object storage) includes: S3, etc. For details, reference can be made to relevant technologies and will not be elaborated here.

[0034] In this embodiment, the Redis service module is an open-source, memory-based data structure storage system that can be used as a database, cache, and message middleware; the above network card is a hardware device used to connect the JuiceFS module and the switching unit; the above switching unit can be used for data forwarding between the terminal device and the storage unit; the above storage unit includes multiple node devices, and a ceph cluster can be built on the multiple node devices.

[0035] The JuiceFS module is used to receive a data write request for the data to be written, write the data to be written into the preset memory buffer of the terminal device 10 through the Redis service module, perform data splitting processing on the first target data in the preset memory buffer to obtain multiple split first data blocks, and write the multiple first data blocks into the storage unit 12 through the network card and the switching unit in sequence; where the first target data at least includes the data to be written.

[0036] The above data to be written can be understood as the file or dataset to be written this time, etc.; the above first target data usually includes at least the data to be written, that is, the first target data can only include the data to be written this time, and can also include the data to be written that has not been written into the storage unit before this time, etc.; in actual implementation, when the user needs to write the data to be written into the storage unit, a data write request for the data to be written can be sent through an operation, and the JuiceFS module can receive the data write request. After receiving the data write request, the JuiceFS module can first temporarily store the first target data in a preset memory buffer of the terminal device through the Redis service module, and perform splitting processing on the first target data in the preset memory buffer to split it into multiple small data blocks, and these small data blocks correspond to the above-mentioned multiple first data blocks. The JuiceFS module can write the split multiple first data blocks into the storage unit through the network card and the switching unit in sequence.

[0037] The JuiceFS module is used to receive a data read request for the second target data, and sequentially read multiple second data blocks corresponding to the second target data from the storage unit 12 through the network card and the switching unit, and perform merging processing on the multiple second data blocks to obtain the second target data.

[0038] The above second target data can be a file or dataset to be read, etc.; in actual implementation, when the user needs to read the second target data in the storage unit, a data read request for the second target data can be sent through an operation, and the JuiceFS module can receive the data read request, and through the network card and the switching unit, read multiple second data blocks related to the second target data from the storage unit 12 according to the data read request. These second data blocks are the multiple small data blocks split when the second target data was previously stored in the storage unit. After reading the multiple second data blocks, the multiple second data blocks can be merged and integrated in sequence to restore the original second target data.

[0039] In the above data processing system, a JuiceFS module and a Redis service module are configured in the terminal device. After receiving a data writing request for the data to be written, the JuiceFS module can write the data to be written into a preset memory buffer of the terminal device through the Redis service module, and then perform data splitting processing on the first target data in the preset memory buffer. The multiple first data blocks obtained by splitting are sequentially written into the storage unit through the network card and the switching unit; after receiving a data reading request for the second target data, the JuiceFS module can sequentially read the multiple second data blocks corresponding to the second target data from the storage unit through the network card and the switching unit, and then obtain the second target data. Based on the storage architecture of the JuiceFS module itself and by utilizing the high-performance storage characteristics of the Redis service module, the system realizes the distributed design of the data processing system without introducing an additional abstraction layer, which helps to improve the reading and writing performance and further reduces the loss of storage performance.

[0040] Furthermore, each node device is configured with a gigabit network interface and a 10-gigabit network interface.

[0041] The above-mentioned gigabit network interface usually refers to a network interface with a speed of 1 Gbps (gigabits per second). The above-mentioned 10-gigabit network interface usually refers to a network interface with a speed of 10 Gbps (ten gigabits per second); such high-speed networks are usually used for data transmission that requires high bandwidth, such as storage networks. In actual implementation, each node device in the storage unit is usually set with a gigabit network interface and a 10-gigabit network interface, and the appropriate network interface can be selected according to actual needs. For example, in scenarios that require high bandwidth and low latency, a 10-gigabit network interface can be selected.

[0042] Furthermore, the switching unit includes: multiple switches; the multiple switches are connected in a stacking mode.

[0043] The number of the above-mentioned multiple switches can be two or more, etc.; the above-mentioned stacking mode can be understood as a technology of connecting multiple switches to form a logically single high-performance switch; in actual implementation, multiple switches can be connected together through dedicated stacking ports or general Ethernet ports to form a logically single high-performance switch, and this configuration is usually used in network environments that require higher bandwidth, larger port density, and higher reliability. After multiple switches are connected in the stacking mode, the following beneficial effects can be brought: 1. High bandwidth: Through stacking, the bandwidths of multiple switches can be combined to provide a higher total bandwidth. 2. High reliability: The stacking mode usually has a redundant design. If a certain switch fails, other switches can take over its functions, thereby improving the reliability of the network. 3. Simplified management: Administrators can centrally manage and configure the entire stacking group instead of managing each switch separately, which simplifies network management. 4. Load balancing: Data streams can be load-balanced among the stacked switches to improve the overall network performance.

[0044] Further, the JuiceFS module is used for: dividing the first target data in the preset memory buffer according to a preset data volume to obtain multiple logical blocks; splitting each logical block to obtain multiple data segments respectively corresponding to each logical block; and determining multiple first data blocks based on each data segment.

[0045] The size of the above-mentioned preset data volume can be set according to actual needs. For example, it can be 64MiB, etc.; in actual implementation, after the data to be written is written into the preset memory buffer, the JuiceFS module can divide the first target data in the preset memory buffer according to the offset of the first target data and the size of the preset data volume to obtain multiple logical blocks. These logical blocks are only logical divisions, not physical divisions, and each logical block represents a logical data segment; the JuiceFS module can further split each logical block into smaller data segments according to the actual situation. Finally, the JuiceFS module can split each data segment into one or more consecutive first data blocks according to the default data volume size, and these first data blocks are written into the storage unit through the network card and the switching unit in sequence as the smallest units.

[0046] For easy understanding, see Figure 3 the schematic diagram of a data writing process shown below, specifically as follows:

[0047] a) The JuiceFS module writes the data to be written into the client memory buffer;

[0048] b) According to the file offset in the client memory buffer, continuous logical chunk units (corresponding to the above-mentioned logical blocks) split by a size of 64MiB;

[0049] c) Each chunk unit will be further split into Slices (corresponding to the above data segments) according to the actual situation of the application write request;

[0050] d) When the newly written data in the client memory buffer is continuous or overlaps with the existing Slices, it will be directly updated on that Slice; otherwise, a new Slice will be created. Usually, the client of the JuiceFS module will maintain a data structure in memory (such as a hash table or a red-black tree, etc.) to efficiently manage the Slice information in each chunk. Through this data structure, the Slice related to the current write request can be quickly located and the continuity can be judged.

[0051] e) During flush, each data segment will be split into one or more continuous Blocks according to the default size of 4MiB and uploaded to the storage unit as the smallest unit; such as Figure 3 In it, the Application (application program) sends each 1M-sized data to be written to the Kernel FUSE (a mechanism that enables user-space programs to implement a file system), and the Kernel FUSE splits it into 128K-sized chunks; the Kernel FUSE sends multiple 128K-sized data to the JuiceFS Client (client software for interacting with the JuiceFS file system), and the JuiceFS Client processes the multiple 128K-sized data to obtain multiple continuous 4M-sized Blocks and sends them to the Object Storage as the smallest unit. Among them, each data segment can include multiple 4M-sized Blocks;

[0052] f) Update the metadata and write the new Slice information.

[0053] In the internal implementation logic of the above JuiceFS module, writing the data to the client memory buffer first can avoid frequent interactions with the remote storage, thereby reducing network latency and increasing throughput; dividing the file into chunks of a fixed size (such as 64MiB) is convenient for management, indexing, and distributed storage; the existence of Slices enables the JuiceFS module to handle write requests of different sizes and patterns more flexibly; for discontinuous write requests, creating new Slices can maintain the independence and integrity of the data and avoid interference with the existing data.

[0054] Furthermore, the JuiceFS module is used to: save the read second target data to a preset directory through the Redis service module.

[0055] The above preset directory can be understood as a directory on the terminal device for temporarily storing the second target data read from the storage unit. This directory helps improve the efficiency of data reading and reduce the overhead of network transmission. For example, in subsequent applications, when the user needs to read the data of a certain file, they can first check the preset directory. If the required data already exists in the preset directory, the data can be directly read from the preset directory, avoiding the network latency of reading data from the remote storage unit. In actual implementation, after obtaining the second target data, the JuiceFS module can save the second target data to the preset directory through the Redis service module.

[0056] For ease of understanding, see Figure 4 the schematic diagram of a data reading process shown below:

[0057] a) Completely read the object corresponding to the Block through the GetObject interface of Object Storage (object storage, corresponding to the above storage unit);

[0058] b) Generally read from the object storage in a way aligned with 4MiB Blocks; in the design of the JuiceFS module, the default size of a Block is 4MiB, which is a reasonable value selected after balancing performance and cost, adapting to the characteristics of object storage, taking into account both random and sequential reads and writes, reducing the metadata scale, industry practices and experience (many distributed storage systems such as HDFS, Ceph, etc. also use a similar Block size).

[0059] c) Write the read data to the local Cache directory (corresponding to the above preset directory); as Figure 4 shown, the JuiceFSClient sends multiple 4M-sized Blocks read to the Kernel FUSE. Taking one 4M-sized Block as an example, the Kernel FUSE processes the received 4M-sized Block to obtain multiple 128K data. The KernelFUSE merges the multiple 128K data to obtain the merged data and sends it to the Application.

[0060] d) Disable caching; considering that caching may cause data inconsistency or the complexity of cache management, no form of caching (whether it is memory caching or other forms of caching) is used during the reading process.

[0061] e) Disable read-ahead; that is, prohibit the behavior of reading ahead in advance, which can avoid unnecessary data reading.

[0062] Further, the storage unit is configured with multiple storage pools; among them, at least a part of the storage pools are configured according to a preset erasure code policy, and the other storage pools except at least a part of the storage pools are configured according to a preset replica configuration policy. The other storage pools use solid-state drives as the storage medium.

[0063] The above erasure code policy can be used for data protection and can provide data redundancy by splitting data into multiple data blocks and parity blocks; the above replica configuration policy can provide data redundancy by replicating data to multiple locations; in actual implementation, multiple storage pools can be configured for the storage unit, and the number of multiple storage pools can be set according to actual needs and is not limited here. For example, taking the number of storage pools as 4, 3 of the storage pools can be configured according to the preset erasure code policy. For example, taking a group of data as an example, the preset erasure code policy is 4+2, which means splitting this group of data into 4 data blocks and 2 parity blocks, and these data blocks and parity blocks are distributed on different node devices. Even if some node devices fail, this group of data can still be restored through the remaining data blocks and parity blocks. In this way, even if some data blocks are lost, this group of data can still be restored through the remaining data blocks and parity blocks; the advantage of this policy is that it can provide higher storage efficiency because there is no need to fully replicate the data. The other 1 storage pool is configured according to the preset replica configuration policy. For example, taking a group of data as an example, the preset replica configuration policy is 2 replicas, which means this group of data will be replicated into two copies and stored in two different locations respectively. This policy is simple and easy to implement and can ensure high availability and fast access of the data. By mixing the erasure code policy and the replica policy, the storage unit can find a balance between storage efficiency and data redundancy, so as to meet the different needs of different types of data.

[0064] Further, the data stored in the storage unit is stored in a pre-created indexless container. The indexless container is a way of storing data. In the indexless container, no additional index structure is created or maintained when storing data. Using the indexless container can reduce the additional index management overhead and improve the data writing and reading speed. This storage method is suitable for the application scenario of high-performance storage, can reduce the complexity and overhead of the storage unit, and improve the storage efficiency and storage performance. When the user's terminal device has high requirements for read-write deletion performance, the data can be placed in the indexless container. Since there is no index, the read-write deletion efficiency is relatively fast.

[0065] For easy understanding, refer to Figure 5 the schematic architecture diagram of a data processing system shown in the figure. The system includes a storage layer (corresponding to the above storage unit), a switching unit, a Windows test machine that requires high performance (corresponding to the above terminal device), and other client machines that do not require read-write performance; the specific system construction process is as follows:

[0066] 1) Install 9 CS13000 nodes of the distributed storage system with Ceph as the core according to the plan, and set up the gigabit service network bond0 (corresponding to the above gigabit network interface) and the 10-gigabit storage network bond1 (corresponding to the above 10-gigabit network interface). Set the VIP (Virtual IP Address) for the cluster, and specifically set a virtual IP for each node device. Figure 5 In it, the public IP in the storage layer corresponds to bond0; the storageIP corresponds to bond1.

[0067] 2) Use two switches, and both switches adopt the stacking mode;

[0068] 3) For the test machines with performance requirements from users, use 40G network cards and connect to the switches. Install the self-compiled JuiceFS module and Redis service module on this Windows test machine;

[0069] 4) Create 4 pools (corresponding to the above storage pools) in the cluster. Among them, 3 storage pools are configured as EC 4+2 (Erasure Coding), where 4 is the number of data blocks and 2 is the number of parity blocks; 1 storage pool is configured with 2 replicas, and use the SSD indexpool;

[0070] 5) Manually create a placement using zonegroup;

[0071] 6) Manually add the just-created placement to the zone;

[0072] 7) Set the default placement of the zonegroup to the pre-created bucket without index (corresponding to the above container without index);

[0073] 8) Modify other relevant optimization configuration parameters of Ceph; for example, read-write locks, etc. By optimizing the configuration parameters, ensure the optimal read-write and deletion performance of users.

[0074] 9) Restart Ceph;

[0075] 10) Create an S3 user;

[0076] 11) The JuiceFS module uses the local Redis service module and caches to mount the bucket in the cluster mode;

[0077] 12) The JuiceFS module creates a drive letter and tests the read-write performance.

[0078] Each of the above node devices can be configured according to the device configuration table in Table 1 below:

[0079] Table 1

[0080]

[0081] Based on the above Figure 5 For the high-performance implementation method based on distributed storage implemented above, configure 9 4U physical servers, with each server equipped with 32 SATA HDDs and 2 m.2 NVMe; install the latest version of CS13000 storage software and build a cluster; create a pool and bucket mapped to the number of clusters; install the Redis service module and JuiceFS module on the client; use JuiceFS on the client to mount the storage bucket and map the drive letter; the client uses the mapped dedicated drive letter for business read and write; with the above devices, software, and hardware configurations, this solution can achieve high performance when reading and writing the cluster in single-client or multi-client (Windows) mode.

[0082] The above data processing system solves the problem in the prior art that the performance does not meet customer requirements when users use the file system scenario in a multi-node cluster mode. The storage end of this system adopts EC 4+2:2. Since the cluster mode is 9 nodes, when a total of 2 disks are lost or 2 nodes go down, the data availability can be guaranteed, improving the reliability.

[0083] It is targeted at the storage layer. Since the JuiceFS module is used, the performance of the storage layer is improved. After verification, for the write performance of the storage layer, it was 523 MB / s before optimization and 3.4 GB / s after optimization, with a 565% increase; for the deletion performance, it was 380 MB / s before optimization and 4 GB / s after optimization, with a 977% increase; for writing while deleting, it was 3.4 GB / s for writing and 1.5 GB / s for deleting after optimization.

[0084] This system uses the JuiceFS module as the upper-layer access to ceph. The main technical improvement points are concentrated in the following aspects:

[0085] 1) Hierarchical storage architecture optimization: Improve performance through the combination of JuiceFS caching and ceph distributed storage.

[0086] 2) Data sharding and parallel processing: Optimize the reading and writing of large files by using Slice / Block splitting and parallel uploading.

[0087] 3) Metadata management optimization: Reduce the burden on ceph through independent metadata storage.

[0088] 4) Separation of hot and cold data: Implement data life cycle management and optimize storage costs.

[0089] 5) Concurrency and consistency optimization: Ensure performance and consistency during concurrent access by multiple clients.

[0090] 6) Cost - performance balance: Flexibly adjust the storage strategy to meet different requirements.

[0091] 7) Ecological compatibility and ease of use: Provide standard interfaces and rich tools to simplify integration and operation and maintenance.

[0092] These improvement points not only enhance the performance of the data processing system, but also strengthen the flexibility and scalability of the system, making it more suitable for modern distributed storage scenarios.

[0093] An embodiment of the present invention discloses a data processing method. The method is applied to the data processing system in the above - mentioned embodiment. The system includes: a terminal device, a switching unit, and a storage unit that are communicatively connected in sequence; wherein, a JuiceFS module and a Redis service module that are communicatively connected to each other are configured in the terminal device, and the JuiceFS module is communicatively connected to the switching unit through a network card; the storage unit is a cept storage system composed of multiple node devices. The method includes:

[0094] Step 1, the JuiceFS module receives a data writing request for the data to be written, writes the data to be written into a preset memory buffer of the terminal device through the Redis service module, performs data splitting processing on the first target data in the preset memory buffer to obtain multiple split first data blocks, and writes the multiple first data blocks into the storage unit through the network card and the switching unit in sequence; wherein, the first target data includes at least the data to be written.

[0095] Step 2, the JuiceFS module receives a data reading request for the second target data, reads multiple second data blocks corresponding to the second target data from the storage unit through the network card and the switching unit in sequence, and performs merging processing on the multiple second data blocks to obtain the second target data.

[0096] The above - mentioned data processing method, based on the storage architecture of the JuiceFS module itself and making use of the high - performance storage characteristics of the Redis service module, realizes the distributed design of the data processing system, does not require the introduction of an additional abstraction layer, thereby helping to improve the read - write performance and further reducing the loss of storage performance.

[0097] An embodiment of the present invention discloses a data processing device, and the device includes the data processing system in any one of the above.

[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A data processing system, characterized in that: The system comprises: a terminal device, a switching unit and a storage unit which are sequentially connected in communication; wherein the terminal device is configured with a JuiceFS module and a Redis service module which are mutually connected in communication, and the JuiceFS module is connected in communication with the switching unit via a network card; the storage unit is a ceph storage system composed of a plurality of node devices; The JuiceFS module is used to receive a data write request for data to be written, write the data to be written into a preset memory buffer of the terminal device through the Redis service module, perform data splitting processing on the first target data in the preset memory buffer to obtain a plurality of split first data blocks, and write the plurality of first data blocks into the storage unit in sequence through the network card and the switching unit; wherein the first target data at least includes the data to be written; The JuiceFS module is used to receive a data read request for second target data, read multiple second data blocks corresponding to the second target data from the storage unit through the network card and the switching unit in sequence, and merge the multiple second data blocks to obtain the second target data.

2. The system according to claim 1, characterized in that Each of the node devices is configured with a Gigabit network interface and a 10 Gigabit network interface.

3. The system according to claim 1, characterized in that The switching unit includes: a plurality of switches; the plurality of switches are connected in a stacking mode.

4. The system according to claim 1, characterized in that The JuiceFS module is used to: Dividing the first target data in the preset memory buffer according to a preset data amount to obtain a plurality of logic blocks; Splitting each of the logic blocks to obtain a plurality of data segments corresponding to each of the logic blocks; Based on each of the data segments, a plurality of first data blocks are determined.

5. The system according to claim 1, characterized in that The JuiceFS module is used to: The read second target data is saved in a preset directory through the Redis service module.

6. The system according to claim 1, characterized in that The storage unit is configured with a plurality of storage pools; wherein at least a portion of the storage pools are configured according to a preset erasure code strategy, and other storage pools except the at least a portion of the storage pools are configured according to a preset replica configuration strategy.

7. The system according to claim 6, characterized in that The other storage pool uses a solid state drive as a storage medium.

8. The system according to claim 1, characterized in that The data stored in the storage unit is stored in a pre-created index-free container.

9. A data processing method, characterized in that: The method is applied to the data processing system according to any one of claims 1 to 8, the system comprising: a terminal device, a switching unit and a storage unit which are sequentially connected in communication; wherein the terminal device is configured with a JuiceFS module and a Redis service module which are mutually connected in communication, and the JuiceFS module is connected in communication with the switching unit through a network card; the storage unit is a ceph storage system composed of a plurality of node devices; the method comprises: The JuiceFS module receives a data write request for data to be written, writes the data to be written into a preset memory buffer of the terminal device through the Redis service module, performs data splitting processing on the first target data in the preset memory buffer to obtain a plurality of split first data blocks, and writes the plurality of first data blocks into the storage unit in sequence through the network card and the switching unit; wherein the first target data at least includes the data to be written; The JuiceFS module receives a data read request for second target data, reads multiple second data blocks corresponding to the second target data from the storage unit through the network card and the switching unit in sequence, and merges the multiple second data blocks to obtain the second target data.

10. A data processing device, characterized in that: The device comprises the data processing system according to any one of claims 1-8.