Mass file data stream storage method and system and big data processing system

By introducing a rotating state transition model and B+ tree index structure in the distributed file system, the performance bottlenecks in massive small file storage are solved and efficient data flow storage and access are achieved.

CN120295570AActive Publication Date: 2025-07-11CHINESE PEOPLES LIBERATION ARMY 92493 UNIT INFORMATION TECH CENT
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510357998.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

When the existing distributed file system processes massive small files, the metadata management pressure is high, resulting in a decline in write performance and low access efficiency. Traditional storage calling technology cannot meet storage requirements and access speed requirements.

Method used

The rotating state transition model and B+ tree index structure are adopted to distribute storage pressure through multi-node cache, optimize the storage method of data flow, distribute small file data flow to multiple disks of multiple nodes, and generate index information to improve access performance.

Benefits of technology

It effectively alleviates the pressure of writing high-speed data streams on a single node, improves the storage performance of the cluster and the access efficiency of small files, reduces the memory consumption of the master node, and solves the performance bottlenecks of traditional systems when facing massive files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295570A_ABST
    Figure CN120295570A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer application, and discloses a mass file data stream storage method, a mass file data stream storage system and a big data processing system, a rotation state conversion model is deployed between a data stream and a storage node, and the storage pressure is dispersed through multi-node cache. The storage method comprises the following steps: presetting a file size demarcation point; inputting massive file data, and judging whether the size of the input file is smaller than or equal to the demarcation point one by one; file data smaller than or equal to the demarcation point is stored in a selected storage node through a rotation state conversion model; generating index information of all the small files, combining the index information to index files, and organizing data in a B + tree form for each index file to obtain an index data bucket; a plurality of index data buckets are stored in a storage node in a data block form, and all the index data buckets are stored as an index data bucket queue in sequence. The method can effectively solve the problems that the writing performance is reduced and the efficiency is low when a traditional distributed storage system faces massive files.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer application technologies, and in particular, to a storage method for a massive file data stream, a storage system, and a big data processing system. Background Art

[0002] Against the backdrop of the rapid development of technologies such as the Internet of Things, cloud computing, artificial intelligence, and big data, the amount of data has shown an explosive growth. These data not only require a huge amount of storage space but also have prominent characteristics such as a wide variety of data types, large variations in data size, and fast data flow. Generally, data files can be classified into structured files, semi-structured files, and unstructured files, with different file sizes. That is to say, these file data often contain a large number of large files and hundreds of millions of massive small files. There are bottlenecks in aspects such as metadata management, storage efficiency, and access performance for these file data. In particular, the storage problem of massive small files (LSOF, lots of small files) is a recognized difficult problem in the current industrial and academic circles. Traditional storage call systems are difficult to meet the storage requirements and access speed requirements of the growing file data. On the one hand, the read and write speed of ordinary hard disk storage call technology is slow and not suitable for the storage of large-scale data. Although the read and write speed of solid-state drive and flash storage call technology is fast, the cost is relatively high. On the other hand, traditional relational databases have certain limitations in processing massive data. Although distributed storage systems provide some solutions, they still face huge challenges in aspects such as performance optimization, cost control, and scalability.

[0003] Hadoop is the most popular big data processing architecture at present, which provides a processing framework for managing and analyzing massive data. The Hadoop Distributed File System (HDFS) is the core component for storing data in the Hadoop system. It performs well in storing and managing large files. However, its performance is extremely poor when dealing with a large number of small files. This is because existing distributed file systems represented by HDFS generally adopt a master-slave node architecture. The client node issues a write request, and the slave node DataNode is responsible for storing the data, while the master node NameNode (the management node of the HDFS file system, which maintains the file directory tree and the data block index, as well as the correspondence between data blocks and data nodes) is responsible for maintaining the entire namespace and recording all changes to the namespace. Each file stored in HDFS needs to store metadata in the NameNode. Regardless of the size of the file itself, the metadata occupies 150 bytes. Therefore, when faced with a large number of small files, the master node needs to maintain too much metadata, consuming memory resources rapidly, causing great pressure on the master node, resulting in a serious decline in the write performance of HDFS and low efficiency in accessing small files.

[0004] In view of this, there is an urgent need to propose a method for efficiently storing a large number of small files. Summary of the Invention

[0005] The object of the present invention is to provide a storage method for a large number of file data streams, a storage system and a big data processing system. By providing a rotation state conversion model and a policy architecture design of ORFS, an optimized rotation intermediate layer is built between the small file data stream and the underlying storage nodes, and the data stream is dispersed to multiple disks on multiple nodes, making full use of the performance of the storage device and improving the utilization rate of the cluster. At the same time, a B+ tree is used to generate an index structure and integrated with ORFS. By additional capabilities, the access performance of a single small file is improved, data operations are simplified, the memory consumption of the NameNode can be effectively reduced, and it provides convenience for efficiently storing and analyzing a large number of small files.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a storage method for a large number of file data streams, deploying a rotation state conversion model between the data stream and the storage nodes to disperse the storage pressure through multi-node caching; the storage method includes the following steps:

[0008] S1. Preset the demarcation point of the file size;

[0009] S2. Input a large number of file data, and judge one by one whether the size of the input file is less than or equal to the demarcation point. If so, execute S3-S4;

[0010] S3. The file data less than or equal to the demarcation point is stored in the selected storage node through the rotation state conversion model;

[0011] S4. Generate index information for all files less than or equal to the demarcation point, and merge the index information into an index file. Each index file organizes data in the form of a B+ tree to obtain an index data bucket; multiple index data buckets are stored in the storage node in the form of data blocks, and all index data buckets are saved as an index data bucket queue in sequence.

[0012] As a possible implementation, the rotation state conversion model has m data buckets, and the m data buckets are incorporated into a queue; the first k data buckets in the queue are data filling buckets, k < m, and the other data buckets are idle waiting buckets;

[0013] Storing the file data less than or equal to the demarcation point in the selected storage node through the rotation state conversion model specifically includes the following steps:

[0014] S30. Write the file data in parallel into multiple data filling buckets;

[0015] When the data filling bucket is full, remove the full data filling bucket from the data filling bucket queue, and at the same time move the full data filling bucket to the write waiting bucket queue, that is, change the status of the full data filling bucket to the write waiting bucket; remove a new bucket from the idle waiting bucket to the data filling bucket queue for receiving subsequent file data writing;

[0016] The rotation state conversion model is configured with a data bucket scheduling module for polling storage nodes and selecting a storage node based on a load balancing policy;

[0017] Use hash technology to route the write waiting bucket to the selected storage node for data dumping preparation. At this time, change the status of the write waiting bucket to the data dumping bucket;

[0018] After the data in the data dumping bucket is written to the selected storage node, change the status of the data dumping bucket to the idle waiting bucket, waiting for task distribution.

[0019] As a possible implementation, each selected storage node corresponds to a waiting queue composed of multiple write waiting buckets. Refresh the write waiting buckets one by one. The write waiting bucket to be refreshed is simultaneously rewritten as the data dumping bucket, and the refreshed data dumping bucket is rewritten as the idle waiting bucket.

[0020] As a possible implementation, the storage method of the storage node includes:

[0021] Receive the data dumping bucket, judge whether the remaining memory is enough to hold the file data in the data dumping bucket. If not, send a resend request message and wait for resending. If so, add the data dumping bucket to the write queue;

[0022] After the write thread completes the previous write task, take out the first data dumping bucket in the write queue and hand it over to the write thread to complete the task of writing the file data in the data dumping bucket to the storage node.

[0023] As a possible implementation, judge one by one whether the size of the input file is less than or equal to the demarcation point. If not, execute the following steps:

[0024] S5. Judge whether the file data needs to be segmented. If so, execute S6 - S7; if not, execute S8;

[0025] S6. Segment the file data;

[0026] S7. Judge whether the segmented file data is less than or equal to the demarcation point. If so, execute S3 - S4; if not, execute S5 - S7; until the judgment result of S5 is that no segmentation is required and then execute S8;

[0027] S8. Store the file data in HDFS.

[0028] As a possible implementation, the demarcation point is 4.35 MB.

[0029] In a second aspect, the present invention provides a storage system for a massive file data stream, including:

[0030] An input port for receiving a massive file data stream;

[0031] A first judgment module, configured with a demarcation point, for attentively judging whether the size of the input file is less than or equal to the demarcation point;

[0032] A second judgment module for judging whether the file data greater than the demarcation point needs to be segmented;

[0033] A segmentation module for segmenting file data that is greater than the demarcation point and needs to be segmented;

[0034] HDFS for storing file data that is greater than the demarcation point and does not need to be segmented;

[0035] A rotation state conversion module for storing file data less than or equal to the demarcation point in a selected storage node;

[0036] An index information generation module for generating index information of all files less than or equal to the demarcation point, and merging the index information into an index file. Each index file organizes data in the form of a B+ tree to obtain an index data bucket; multiple index data buckets are stored in the storage node in the form of data blocks, and all index data buckets are sequentially saved as an index data bucket queue.

[0037] A storage node for storing file data less than or equal to the demarcation point and multiple index data buckets.

[0038] As a possible implementation, the rotation state conversion module has m data buckets, and the m data buckets are incorporated into a queue; the first k data buckets in the queue are data filling buckets, k < m, and the other data buckets are idle waiting buckets; when storing file data less than the demarcation point in a selected storage node through the rotation state conversion model, steps S30 to S34 described in claim 2 are executed.

[0039] As a possible implementation, the storage system has a master node and a storage node. The master node controls the rotation state conversion model, and the storage node is responsible for data storage.

[0040] In a third aspect, the present invention provides a big data processing system, which implements the storage of a massive file data stream by applying the storage method provided in the first aspect.

[0041] Compared with the prior art, the present invention has the following effects:

[0042] 1. The storage method for massive file data streams provided by the present invention designs a rotation state conversion model, deploys this model between high-speed data streams and underlying storage nodes, and disperses the storage pressure through multi-node caching, which can effectively relieve the writing pressure of high-speed data streams on a single node, make full use of each node in the cluster, and maximize the development of the storage performance of the cluster.

[0043] 2. The storage method for massive file data streams provided by the present invention deploys the rotation state conversion model on the upper layer of the storage nodes, so it will not be affected by the iterative update of the database version, and also enables many optimization works for storage to still be applicable.

[0044] 3. The storage method for massive file data streams provided by the present invention organizes index information into a B+ tree for storage. The storage mode based on the B+ tree can quickly retrieve the index information of small files, thereby effectively improving the overall performance of processing small files.

[0045] 4. By providing the rotation state conversion model ORFS, using the B+ tree to generate an index structure at the same time, and integrating it with ORFS, the present invention can effectively solve the problems of serious decline in write performance and low efficiency in accessing small files when the traditional distributed storage system faces massive files. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0047] Figure 1 It is a rotation state conversion diagram of data buckets in four states: idle waiting, data filling, write waiting, and data dumping in the embodiment of the present invention;

[0048] Figure 2 It is a flowchart of the storage method for massive file data streams provided by the embodiment of the present invention;

[0049] Figure 3 It is a schematic diagram of the working principle of the rotation state conversion model provided by the embodiment of the present invention;

[0050] Figure 4 It is a schematic diagram of the storage system for massive file data streams provided by the embodiment of the present invention;

[0051] Figure 5 It is the rotation process of the main node of the storage system for massive file data streams provided by the embodiment of the present invention;

[0052] Figure 6 It is a flowchart of the working process of the storage node of the storage system for massive file data streams provided by the embodiment of the present invention.

[0053] Reference numerals:

[0054] 10 - Input port, 20 - First judgment module, 30 - Second judgment module, 40 - Splitting module, 50 - HDFS, 60 - Rotation state conversion module, 70 - Index information generation module, 80 - Storage node. Detailed implementation manners

[0055] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0056] It should be noted that when an element is referred to as being "fixedly disposed on" or "disposed on" another element, it can be directly on the other element or indirectly on the other element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or indirectly connected to the other element.

[0057] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, the meaning of "a plurality of" is two or more unless otherwise specifically defined.

[0058] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "upper" and "lower" is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention.

[0059] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the term "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a direct connection or an indirect connection through an intermediate medium, and can be the internal communication of two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meaning of the above terms in the present invention can be understood according to specific circumstances.

[0060] An embodiment of the present invention provides a storage method, a storage system, and a big data processing system for a massive file data stream. The method adopts an efficient storage and call technology for massive data based on an optimized rotation strategy, and is used for the efficient storage and call of a massive small file data stream. Different from traditional solutions, the present invention provides a rotation state conversion model ORFS and a policy architecture design. This model is deployed between a high-speed data stream and underlying storage nodes. By using multi-node caching to disperse the storage pressure, it alleviates the high-speed data stream writing pressure on a single node, makes full use of each node in the cluster, and maximally develops the storage performance of the cluster. Since it is deployed on the upper layer of the storage nodes, it will not be affected by the iterative update of the database version, and many optimization works for storage are still applicable. At the same time, ORFS organizes index information into a B+ tree for storage. Based on the B+ tree storage mode, the index information of small files can be quickly retrieved, thereby effectively improving the overall performance of processing small files.

[0061] In a first aspect, an embodiment of the present invention provides a storage method for a massive file data stream, which deploys a rotation state conversion model between the data stream and the storage nodes, and disperses the storage pressure through multi-node caching.

[0062] As a possible implementation, the rotation state conversion model has m data buckets, and the m data buckets are incorporated into a queue; the first k data buckets in the queue are data filling buckets, k < m, and the other data buckets are idle waiting buckets.

[0063] The rotation algorithm, also known as the rotation method, is an optimization strategy widely used in computer science. It rotates the elements in a data set according to certain rules to achieve the purpose of optimizing data access and improving data processing efficiency. The core idea of the rotation algorithm is to rotate the elements in the data set in a certain order so that the data can be accessed more efficiently. The basic principle of the rotation algorithm is as follows:

[0064] (1) Define the rotation direction: Determine the rotation direction of the data set, which can be rotating to the left or to the right.

[0065] (2) Determine the number of rotations: Determine the number of rotations according to actual needs, that is, how many positions to rotate.

[0066] (3) Perform the rotation operation: Rotate the data set according to the defined rotation direction and number of rotations.

[0067] According to the different states of the data buckets, it can be divided into four states: idle waiting, data filling, write waiting, and data dumping. See the rotation state conversion diagram in Figure 1, Idle waiting: When a data bucket is in the idle waiting state, the data bucket is empty and can accept the insertion and storage of new data streams at any time. In this embodiment, all data buckets will first be set to this idle waiting state. Data filling: When a data bucket is in the data filling state, all data insertions will be directed to the data bucket. That is to say, such a data bucket is the data bucket that is actually to be written when the system is running. In the implementation of this method, at least one data bucket is in the data filling state, otherwise there will be no available data bucket to accommodate the newly inserted data, thereby triggering the system stop mechanism. Another key point is that when the data filling causes the data bucket to be full, it will not accept the newly inserted data. In this case, the data bucket will become a write waiting state. At the same time, the system will automatically select one from the data buckets in the idle waiting state and use it as a new data filling bucket. Write waiting: The data buckets in the write waiting state are all converted from the data buckets in the data filling state, that is, they are all full of newly written temporary storage data and are waiting to be persistently stored on the disk or solid-state drive. The system will not immediately persist a full data bucket because the write bandwidth of the underlying disk or solid-state drive is limited. Therefore, all full bucket data must wait until the system arranges them to be persisted. Each storage node will have a waiting queue to store write-wait state data buckets that have not yet been persisted, and the system refreshes the data buckets in the waiting queue one by one. Data dumping: When the data in a bucket in the write-wait state is about to be persisted, its state will change to data dumping. This means that the data writes temporarily stored in the bucket will be written to persistent storage. When all data in the data dumping bucket is written to persistent storage, the system will change its state to idle waiting. The system can have multiple data dumping buckets storing data persistently at the same time.

[0068] This embodiment defines a large number of data buckets at the upper layer of the round-robin state transition model, and a data bucket represents the memory space of a node in the data center. Therefore, all data buckets constitute a huge distributed memory space. The key idea of ​​this patent design is to use the huge distributed memory space to accept the massive data streams that arrive quickly.

[0069] See also Figure 2 The method for storing a massive file data stream provided in this embodiment includes the following steps:

[0070] S1. Preset the file size cutoff point; illustratively, the cutoff point is 4.35MB.

[0071] S2. A large amount of file data is input, and the size of the input files is determined one by one to see whether it is less than or equal to the delimiting point. If so, execute S3 to S4;

[0072] In this embodiment, files are differentiated according to their sizes. Files with a size greater than 4.35 MB are defined as large files, and files with a size less than or equal to 4.35 MB are defined as small files. Small files will be stored in a non-clustered index, that is, their index information is stored in the rotation state transition model, and the data information is stored in the selected storage node through the rotation state transition model.

[0073] As an example, the data structure of the small file set is as follows:

[0074] {

[0075] "_id":ObjectId("IDENTIFIER"),

[0076] "filename":"FILENAME",

[0077] "format":"FORMAT",

[0078] "uploadDate":UPLOAD_TIMESTAMP,

[0079] "length":INTEGER,

[0080] "downloadCount":INTEGER,

[0081] "data":BINARY DATA

[0082] }。

[0083] Files with a size greater than 4.35 MB are defined as large files. If a large file is indivisible, it will be directly stored in HDFS without further processing; if a large file needs to be split, after the splitting process, it will be returned to the previous process to determine whether it is a small file, and then further processing will be carried out. The data structure of the large file is as follows:

[0084] {

[0085] "_id":ObjectId("IDENTIFIER"),

[0086] "fileId":ObjeetId("FILE_IDENTIFIER"),

[0087] "n":INTEGER,

[0088] "data":BINARY DATA

[0089] }。

[0090] S3. The file data less than or equal to the demarcation point is stored in the selected storage nodes through the rotation state conversion model;

[0091] Exemplarily, each selected storage node corresponds to a waiting queue composed of multiple write waiting buckets. The write waiting buckets are refreshed one by one, and the write waiting buckets to be refreshed are simultaneously rewritten as data dumping buckets. After refreshing, the data dumping buckets are rewritten as idle waiting buckets.

[0092] See Figure 3 , as a possible implementation, S3 includes the following steps:

[0093] S30. Write the file data into multiple data filling buckets in parallel;

[0094] When the data filling bucket is full, remove the full data filling bucket from the data filling bucket queue, and at the same time move the full data filling bucket into the write waiting bucket queue, that is, the state of the full data filling bucket is changed to a write waiting bucket; Move a new bucket from the idle waiting bucket to the data filling bucket queue for receiving subsequent file data writing;

[0095] It should be noted that the size of the data filling bucket is determined by the system configuration and depends on the specific data distribution strategy. Exemplarily, the data filling bucket size is set to 10GB, that is, when the data size in each data filling bucket reaches the range of 10000M to 10240M, it can be considered that the data filling bucket is full.

[0096] The rotation state conversion model is configured with a data bucket scheduling module for polling the storage nodes and selecting storage nodes based on the load balancing strategy;

[0097] Use the hash technology to route the write waiting bucket to the selected storage node for data dumping preparation. At this time, the state of the write waiting bucket is changed to a data dumping bucket;

[0098] After the data in the data dumping bucket is written into the selected storage node, the state of the data dumping bucket is changed to an idle waiting bucket, waiting for task distribution;

[0099] S4. Generate the index information of all files less than or equal to the demarcation point, and merge the index information into the index file. Each index file organizes data in the form of a B+ tree to obtain an index data bucket; Multiple index data buckets are stored in the storage node in the form of data blocks, and all index data buckets are saved in order as an index data bucket queue.

[0100] As a possible implementation, the storage method of the storage node includes:

[0101] Receive a data dumping bucket, determine whether the remaining memory is sufficient to hold the file data in the data dumping bucket. If not, send a retransmission request message and wait for retransmission. If so, add the data dumping bucket to the write queue;

[0102] After the write thread completes the previous write task, take out the first data dumping bucket in the write queue and hand it over to the write thread to complete the task of writing the file data in the data dumping bucket to the storage node.

[0103] As a possible implementation, if it is determined that the input file size is greater than the demarcation point, the following steps are executed:

[0104] S5. Determine whether the file data needs to be segmented. If so, execute S6 - S7; if not, execute S8;

[0105] Exemplarily, the standard for large file segmentation is that if a large file is segmented into several small files and stored in the storage node, each small file can be retrieved and quickly merged and called, then segment; otherwise, the large file is not segmented and directly stored in the HDFS system. The greater the regularity in the general data, the easier it is to segment.

[0106] S6. Segment the file data;

[0107] As an example, use an iterative segmentation algorithm to segment large file data, a method to improve data processing efficiency through parallel computing and distributed storage. The basic principle is as follows:

[0108] (1) Data segmentation: Segment the large data set into multiple small data sets, and each small data set contains a part of the data set;

[0109] (2) Parallel computing: Allocate the segmented data sets to multiple computing nodes and perform independent computing on each data set;

[0110] (3) Result merging: Merge the calculation results of each computing node to obtain the final result.

[0111] S7. Determine whether the segmented file data is less than or equal to the demarcation point. If so, execute S3 - S4; if not, execute S5 - S7; until the result of S5 determines that segmentation is not required and then execute S8;

[0112] S8. Store the file data in HDFS.

[0113] In a second aspect, an embodiment of the present invention provides a storage system for a massive file data stream. Refer to Figure 4 , the storage system includes:

[0114] An input port 10 for receiving a massive file data stream;

[0115] The first judgment module 20 is configured with a demarcation point for attentively judging whether the size of the input file is less than or equal to the demarcation point;

[0116] The second judgment module 30 is used for judging whether the file data greater than the demarcation point needs to be segmented;

[0117] The segmentation module 40 is used for segmenting the file data that is greater than the demarcation point and needs to be segmented;

[0118] HDFS 50 is used for storing the file data that is greater than the demarcation point and does not need to be segmented;

[0119] The round-robin state conversion module 60 is used for storing the file data less than or equal to the demarcation point into the selected storage node;

[0120] As an example, the round-robin state conversion module 60 has m data buckets, and the m data buckets are incorporated into a queue; the first k data buckets in the queue are data filling buckets, k < m, and the other data buckets are idle waiting buckets; when storing the file data less than or equal to the demarcation point in the selected storage node through the round-robin state conversion model, S30 to S34 of the first aspect are executed.

[0121] The index information generation module 70 is used for generating the index information of all files less than or equal to the demarcation point and merging the index information into the index file. Each index file organizes data in the form of a B+ tree to obtain an index data bucket; multiple index data buckets are stored in the storage node in the form of data blocks, and all index data buckets are saved in order as an index data bucket queue.

[0122] The storage node 80 is used for storing the file data less than or equal to the demarcation point and multiple index data buckets.

[0123] As a possible implementation manner, the storage system has a master node and storage nodes. The master node controls the round-robin state conversion model, and the storage nodes are responsible for data storage;

[0124] For the parameter description of the storage system, see Table 1:

[0125] Parameter Description c Data flow size t Moment k Number of storage nodes <![CDATA[B t > Bucket data bucket carrying data flow at moment t b Bucket size i Storage node label (i ∈ [0, k - 1]) <![CDATA[D i > The i-th storage node <![CDATA[WriteQueue i > Write queue on the i-th storage node <![CDATA[WriteThread i > Write thread on the i-th storage node

[0126] Among them, B t refers to the bucket data bucket carrying the data stream at time t; D i is the i-th storage node for persistent storage data; WriteQueue i is the write queue on the i-th storage node; WriteThread i is the write thread on the i-th storage node.

[0127] For the rotation process of the master node, see Figure 5 , including:

[0128] (1) The i is marked as a storage node and initialized to 0;

[0129] (2) The data stream continuously flows in. At time t, the master node allocates a predefined data bucket space B t . After successful allocation, the data will be temporarily stored in B t ; If the allocation fails and the master node cannot allocate a new bucket, the system fails and exits;

[0130] (3) When the data bucket B t is full, calculate the storage node number i to be sent, and send the data bucket to D i . At the same time, turn to phase (4); After the sending is completed, release the data bucket B t , and the data writing work is completed by the node itself; If the sending fails, resend after waiting for a period of time (the reason for failure may be that the write work queue of the lower layer node is full or the network transmission fails);

[0131] (4) i is incremented by 1 in sequence, and a new data bucket is opened to continuously receive the data stream.

[0132] For the storage strategy of the underlying storage nodes, see Figure 6 . In the round-robin state transition model, if the number of nodes does not reach the minimum requirement, the data will be temporarily resident in the memory. Therefore, the memory resources become very precious. This system chooses to use a single thread to complete the write throughput. The writing process is as follows:

[0133] (1) After receiving the data bucket B t from the upper layer, judge whether the remaining memory is sufficient to hold the data bucket. If not, send a message to the upper layer and wait for retransmission; If there is still remaining memory, add the data bucket to the write waiting queue WriteQueue i ;

[0134] (2) Wait for the write thread WriteThread i to complete the previous write task, then take out the first bucket in the write waiting queue and hand it over to the write thread to continue to complete.

[0135] Usually, the index information of each small file has its fixed fields and lengths. The information it contains is shown in Table 2:

[0136] Table 2 Small file index information

[0137] Field name Description Length File name hash File name hash of small file 32 bytes File size Size information of small file 32 bytes Merged file information Information about merged file 16 bytes Offset Offset position of small file 16 bytes

[0138] The B+ tree is a variant of the B tree. The main difference is that all actual data is stored in the leaf nodes, and the non-leaf nodes are only used for indexing. This design enables the B+ tree to significantly improve query efficiency when dealing with a large amount of data, especially performing excellently in range queries. The leaf nodes of the B+ tree are connected by a linked list, facilitating interval search and traversal. This patent uses the size information of small files as an index to construct an index system and ensures the orderliness of the data in the index system. At the same time, the hash value of the small file name is saved in the metadata to ensure that the small file can be uniquely identified. Through the index queue, ORFS can obtain the index information of small files and merge this index information into the index file. Each index file organizes data in the form of a B+ tree. Such an index file is called an index data bucket. Since the metadata of each small file occupies 96B, our index data buckets are also stored in ORFS in the form of data blocks. Therefore, in a system with a default data block size of 128MB, an index data bucket can store information about approximately 1 million small files. This may not be sufficient for tens of millions of small files, so this patent splits the buckets and saves the metadata of each bucket in a Bucket queue to ensure the order of the index data buckets stored in the queue. This storage method expands the number of small file indexes stored in the index system without affecting the access efficiency.

[0139] In a third aspect, the present invention provides a big data processing system that implements the storage of a massive file data stream by applying the storage method provided in the first aspect.

[0140] In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in a suitable manner in any one or more embodiments or examples.

[0141] As described above, these are only the specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claimed rights.

Claims

1. A storage method for a massive file data stream, characterized in that, Deploy a round-robin state transition model between the data stream and the storage nodes to disperse the storage pressure through multi-node caching; the storage method includes the following steps: S1. Preset the breakpoints of the file sizes. S2. Input massive file data, and judge one by one whether the size of the input file is less than or equal to the breakpoint. If so, execute S3-S4. S3. The file data smaller than or equal to the breakpoint is stored in the selected storage node through the round-robin state transition model. S4. Generate the index information of all files smaller than or equal to the breakpoint, and merge the index information into the index file. Each index file organizes data in the form of a B+ tree to obtain an index data bucket; multiple index data buckets are stored in the storage node in the form of data blocks, and all index data buckets are saved as an index data bucket queue in sequence.

2. The storage method of the massive file data stream according to claim 1, characterized in that, The round-robin state transition model has m data buckets, and the m data buckets are incorporated into a queue; the first k data buckets in the queue are data filling buckets, k < m, and the other data buckets are idle waiting buckets. Storing the file data smaller than or equal to the breakpoint in the selected storage node through the round-robin state transition model specifically includes the following steps: S30. Write the file data in parallel into multiple data filling buckets. S31. When the data filling bucket is full, remove the full data filling bucket from the data filling bucket queue, and at the same time move the full data filling bucket into the write waiting bucket queue, that is, change the state of the full data filling bucket to a write waiting bucket; move a new bucket from the idle waiting bucket to the data filling bucket queue for accepting subsequent file data writing. S32. The round-robin state transition model is configured with a data bucket scheduling module for polling the storage nodes and selecting a storage node based on the load balancing strategy. S33. Use the hash technology to route the write waiting bucket to the selected storage node for data dumping preparation. At this time, the state of the write waiting bucket is changed to a data dumping bucket. S34. After the data in the data dumping bucket is written into the selected storage node, the state of the data dumping bucket is changed to an idle waiting bucket, waiting for task distribution.

3. The storage method of the massive file data stream according to claim 2, characterized in that Each selected storage node corresponds to a waiting queue composed of multiple write waiting buckets. Refresh the write waiting buckets one by one, and the write waiting buckets to be refreshed are simultaneously rewritten as data dumping buckets, and the refreshed data dumping buckets are rewritten as idle waiting buckets.

4. The storage method of the massive file data stream according to claim 2, wherein The storage method of the storage node includes: Receive the data dumping bucket, judge whether the remaining memory is sufficient to accommodate the file data in the data dumping bucket. If not, send a resend request message and wait for resending. If so, add the data dumping bucket to the write queue. After the write thread completes the previous write task, take out the first data dumping bucket in the write queue and hand it over to the write thread to complete the task of writing the file data in the data dumping bucket into the storage node.

5. The storage method of the massive file data stream according to claim 1, characterized in that, Judge one by one whether the size of the input file is less than or equal to the breakpoint. If not, execute the following steps: S5. Judge whether the file data needs to be segmented. If so, execute S6-S7; if not, execute S8. S6. Segment the file data. S7. Determine whether the segmented file data is less than or equal to the demarcation point. If so, execute S3 - S4; if not, execute S5 - S7; until the result of S5 is that no segmentation is required, then execute S8; S8. Store the file data in HDFS.

6. The storage method of the massive file data stream according to any one of claims 1 to 5, characterized in that, The demarcation point is 4.35MB.

7. A storage system for a massive file data stream, characterized in that, It includes: An input port for receiving a massive file data stream; A first judgment module configured with a demarcation point for attentively judging whether the size of the input file is less than or equal to the demarcation point; A second judgment module for judging whether the file data greater than the demarcation point needs to be segmented; A segmentation module for segmenting the file data greater than the demarcation point and requiring segmentation; HDFS for storing the file data greater than the demarcation point and not requiring segmentation; A rotation state conversion module for storing the file data less than or equal to the demarcation point in the selected storage node; An index information generation module for generating index information of all files less than or equal to the demarcation point, and merging the index information into an index file. Each index file organizes data in the form of a B+ tree to obtain an index data bucket; multiple index data buckets are stored in the storage node in the form of data blocks, and all index data buckets are saved in sequence as an index data bucket queue; A storage node for storing the file data less than or equal to the demarcation point and multiple index data buckets.

8. The storage system for a massive file data stream according to claim 7, characterized in that, The rotation state conversion module has m data buckets, and the m data buckets are incorporated into a queue; the first k data buckets in the queue are data filling buckets, k < m, and the other data buckets are idle waiting buckets; when storing the file data less than the demarcation point in the selected storage node through the rotation state conversion model, execute S30 - S34 described in claim 2.

9. The storage system for a massive file data stream according to claim 7, characterized in that, The storage system has a master node and storage nodes. The master node controls the rotation state conversion model, and the storage nodes are responsible for data storage.

10. A big data processing system, characterized in that, Apply the storage method described in claims 1 to 6 to implement the storage of a massive file data stream.

Citation Information

Patent Citations

  • Method for storing mass of small files on basis of master-slave distributed file system

    CN103020315A

  • Method for accessing archived file based on B + tree index

    CN114116612A

  • Local file system management method, electronic equipment and readable storage medium

    CN117112496A