Distributed cloud native storage oriented small file merging optimization method
By analyzing access logs in the HDFS system and using the FP-Growth algorithm to discover association rules, combining the Huffman tree merging strategy and distributed processing, the memory bottleneck and merge efficiency problems of HDFS when processing massive small files are solved, and more efficient file storage and access are achieved.
Patent Information
- Application Number
- CN202510311863.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-17
AI Technical Summary
The existing HDFS system has memory bottlenecks when processing massive small files, the file correlation analysis is not in-depth enough, and the file merging method is too rigid and difficult to scale, resulting in insufficiency of storage and access efficiency.
By analyzing the HDFS access log, using the FP-Growth algorithm to discover association rules from the file access records, an association analysis algorithm based on user access mode was proposed, and small files were merged according to the generated rules, the small file merging strategy of the Huffman tree was adopted, and the distributed characteristics of HDFS were used for parallel processing.
Optimize the file storage space of HDFS, improve the access efficiency of massive small files, reduce the load pressure of NameNode data blocks, and improve the utilization rate of memory space.
Smart Images

Figure CN120162006A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a small file merging optimization method for distributed cloud-native storage, belonging to the field of computer communication technology. Background Art
[0002] According to the 54th Statistical Report on the Development of the Internet in China released by the China Internet Network Information Center (CNNIC), as of June 2024: The scale of Internet users in China reached 1.099 billion, and the Internet penetration rate reached 78.0%. Among them, the users of micro short dramas accounted for 52.4% of the overall Internet users, and the users of short videos accounted for 95.5% of the overall Internet users, indicating that watching online videos has become an indispensable part of the daily life of Chinese Internet users. Short videos, audio, and various document cache data are the main components of small files. According to IDC data, the number of Internet of Things connections in China exceeded 6.6 billion in 2023, and the compound annual growth rate in the next 5 years will be about 16.4%, maintaining rapid development.
[0003] Facing the trend of a large amount of data pouring in and the corresponding storage requirements of data files, the traditional file storage framework and database technology are obviously unable to cope. Therefore, an efficient and stable big data storage technology in the cloud environment has become an obvious research goal.
[0004] HDFS is a distributed cloud storage system composed of a NameNode, multiple DataNodes, and some auxiliary software. Among them, the NameNode is the main server of HDFS, mainly responsible for managing files and client access requests, while the DataNode is used to store and query file data. HDFS was initially born for storing large files. However, since small files are basically stored in the metadata block space provided by the NameNode, it greatly increases the pressure on block memory, cannot fully ensure the improvement of memory performance, and is prone to memory bottleneck problems when processing a large number of metadata requests for small files. The Hadoop platform itself provides several solution mechanisms, and later generations have continuously improved the problem of storing a large number of small files in HDFS, but there are still problems such as the lack of sufficient exploration of the correlation of small files and the difficulty in accessing the merged small files. In summary, there are some technical problems in the current storage method for a large number of small files in HDFS, including but not limited to the following aspects:
[0005] 1. Insufficient exploration of small file correlation: Traditional small file correlation analysis methods only perform grouped preprocessing based on basic characteristics such as the type and size of small files, while ignoring the user characteristics of accessing files, resulting in insufficient in-depth correlation analysis and affecting the robustness of the analysis results.
[0006] 2. The file merging method is too rigid: Most small file merging methods adopt a global processing method, failing to consider the merging requirements of small files with different characteristics. This may lead to poor data locality in the merged file. As the number of files increases, it is difficult for the merging method to be effectively extended, and the management complexity increases significantly.
[0007] 3. The distributed characteristics of HDFS are underutilized: Few methods utilize the distributed concept of the HDFS system to run the merging process of small files in parallel, which can greatly reduce the compression efficiency. Summary of the Invention
[0008] The object of the present invention is to provide an optimization method for merging small files for distributed cloud-native storage in view of the defects and deficiencies of the above-mentioned prior art. This method analyzes the HDFS access logs, converts the data into a structure suitable for analysis, and proposes a correlation analysis algorithm based on the user access pattern. The FP-Growth algorithm is used to discover association rules from the file access records: by constructing an FP tree to compress the transaction data and determine the frequent item sets, and generating association rules that meet the support and confidence from the frequent item sets. The small files with different degrees of correlation are processed separately according to the generated rules: find the small files with strong correlation, and use the small file merging strategy based on the Huffman tree to merge the files; for the small files with weak correlation, continuously add them to the waiting queue, and when the queue meets the storage size condition of the HDFS data node, perform merged storage. This method optimizes the file storage space of HDFS and improves the access efficiency of a large number of small files. It belongs to the field of correlation analysis and small file storage.
[0009] The technical solution adopted by the present invention to solve its technical problems is: An optimization method for merging small files for distributed cloud-native storage, which includes the following steps:
[0010] Step 1: Define the small file size threshold size. According to the HDFS file access logs, if the file f i has a size smaller than the small file size threshold size, then append the file f i to the massive small file set, and obtain a partial massive small file set Files = {f1, f2, f3,..., f n}, and the historical order of user access requests for the file system is Requests = {r1, r2, r3,..., r n}, where each request r i , 1 ≤ i ≤ n, corresponds to a small file f i, 1 ≤ i ≤ n. The small files contain user attributes u, file types t, and time attributes s. Any small file can be represented by the three basic attributes (u, t, s). The access request sequence containing the basic attributes can be expressed as: Requests = {(u1, t1, s1), (u2, t2, s2), (u3, t3, s3), …, (u n ,t n ,s n )}
[0011] Step 2: Scan the user access request sequence Requests and partition the basic attribute data set. Copy the data set to HDFS, and HDFS decomposes the data set into continuous Blocks and saves the corresponding replicas, and distributes each Block to N nodes for storage. This step is automatically completed by Hadoop
[0012] Step 3: Define respective transaction dictionaries F_dict j = {key1: value1, key2: value2, …}, 1 ≤ j ≤ N, where the key is the non-repeated attribute in the access request sequence, and the value is the frequency of the attribute occurrence. Each node traverses the user access request sequence Requests j , 1 ≤ j ≤ N, and determines whether the attribute exists in the transaction dictionary F_dict j . If it does not exist, add it to the transaction dictionary F_dict j , set the key to the attribute, and set the value to 1; if it exists, increment the value corresponding to the attribute by 1. Finally, each node obtains the updated transaction dictionary F_dict j . Count the local frequencies of each item in all nodes, sort them in descending order of frequency to obtain the global transaction set FList. The formula for calculating the frequency of the statistical data item occurrence is:
[0013]
[0014] where value_I i represents the frequency of data item I i appearance, and value_T j I i represents the frequency of data item I i appearance on the j-th node
[0015] Step 4: Set the minimum support minSup, and construct the local frequent one-item set header table on each node. The input is the global transaction set FList, and the output is the local frequent one-item set header table L(x, sup)
[0016] Step 5: Build local FP-trees in parallel on each node. The input is the local frequent 1-itemset item header table L(x, sup), and the output is the root of the local FP-tree j 。
[0017] Step 6: Build local conditional FP-trees in parallel. The inputs are the local FP-tree root j , the header pointer list header_table j and the sorted local dataset sorted_F_dict j , and the output is the local conditional database paths j 。
[0018] Step 7: Mine local frequent itemsets on each node. The input is the local conditional database paths j , and the output is the set of local frequent itemsets Frequents j 。
[0019] Step 8: Generate association rules according to confidence and lift thresholds. The input is the local frequent itemsets Frequents j , and the output is the association rules rules
[0020] Step 9: Divide massive small files into strong and weak relevance groups. The inputs are the association rules rules and the set of massive small files Files, and the outputs are the strongly relevant file group SR_Files and the weakly relevant file group WR_Files
[0021] Step 10: Merge files in the strongly relevant file group. The input is the strongly relevant file group SR_Files, and the output is the root node huffman_root of the Huffman tree
[0022] Regarding the encoding on the left branch as "0" and the encoding on the right branch as "1", the string composed of the encoding characters on the path from the root node huffman_root to each file node is regarded as the index of the file
[0023] Step 11: Create a file merge queue MergeFile, append the files in the weakly relevant file group WR_Files to the queue in turn, and calculate the size of the merge queue MergeFile. If it exceeds 3 / 4 of the HDFS data block size, directly merge the files in the merge queue MergeFile and store them in the HDFS file system; otherwise, return to Step 9 and wait for the generation of subsequent new weakly relevant file groups WR_Files
[0024] Step 12: The metadata server assigns a globally unique key r value to the root node huffman_root to identify different root nodes. This keyr The value is also saved in the attributes of each file in its corresponding strongly associated file group SR_Files. At the same time, record the size of the root node as Size. Sequentially allocate a globally unique key for the files in the strongly associated file group SR_Files f value, which is used to identify each file. At the same time, record the index of the file in step 10 as index, which is used for the data storage device to obtain the physical address of the file.
[0025] Step 13: The metadata server creates an object Object, allocates the object on the data storage device, and the data storage device stores the root node huffman_root according to the key r allocated by the metadata server and Size into the object Object, and sequentially stores the small files in the strongly associated file group into the object Object according to the allocated key f 、index and key r into the object Object.
[0026] Beneficial effects
[0027] 1. By pre-analyzing the association of small files based on the user access pattern and then performing different merging processes on the small files, the present invention improves the utilization rate of the HDFS memory space, reduces the file load pressure in the NameNode data block, and improves the efficiency of users reading files.
[0028] 2. By pre-analyzing the association of small files based on the user access pattern and then performing different merging processes on the small files, the present invention plays a huge role in improving the utilization rate of the HDFS memory space, reducing the file load pressure in the NameNode data block and further improving the efficiency of users reading files, and improves the access efficiency of a large number of small files.,. Description of the drawings
[0029] Figure 1 is a schematic flowchart of the small file merging method in the embodiment of the present invention.
[0030] Figure 2 is a schematic deployment diagram of the small file merging module in the embodiment of the present invention. Detailed implementation manners
[0031] The following further describes the present invention in detail with reference to the accompanying drawings of the specification.
[0032] As Figure 1 and Figure 2 shown, the present invention provides a small file merging and optimization method for distributed cloud native storage, and the method includes the following steps:
[0033] Step 1: Define the small file size threshold size. According to the file access logs of HDFS, if the size of file f i is less than the small file size threshold size, then append file f i to the massive small file set, and obtain a partial massive small file set Files = {f1, f2, f3, …, f n}. The historical order of user access requests in the file system is Requests = {r1, r2, r3, …, r n}, where each request r i , 1 ≤ i ≤ n, corresponds to a small file f i , 1 ≤ i ≤ n. The small file contains user attribute u, file type t, and time attribute s. Any small file can be represented by three basic attributes (u, t, s). The access request sequence containing the basic attributes can be expressed as: Requests = {(u1, t1, s1), (u2, t2, s2), (u3, t3, s3), …, (u n , t n , s n )}.
[0034] Step 2: Scan the user access request sequence Requests and perform partitioning processing on the basic attribute data set therein. Copy the data set to HDFS, and HDFS decomposes the data set into continuous Blocks and saves the corresponding replicas, and distributes each Block to N nodes for storage. This step is automatically completed by Hadoop.
[0035] Step 3: Define respective transaction dictionaries F_dict j = {key1: value1, key2: value2, …} on each node, 1 ≤ j ≤ N. The key is the non-repeated attribute in the access request sequence, and the value is the frequency of the attribute occurrence. Each node traverses the user access request sequence Requests j , 1 ≤ j ≤ N, and determines whether the attribute exists in the transaction dictionary F_dict j . If it does not exist, add it to the transaction dictionary F_dict j , set the key to this attribute, and set the value to 1; if it exists, increment the value corresponding to this attribute by 1. Finally, obtain the updated transaction dictionary F_dict j respectively. Statistically calculate the local frequency of each item in all nodes, and sort them in descending order of frequency to obtain the global transaction set FList. The formula for calculating the frequency of the statistical data item used is:
[0036]
[0037] Among them, value_I i represents the frequency of occurrence of data item I i , value_T j I i represents the frequency of occurrence of data item I i on the j-th node.
[0038] Step 4: Set the minimum support minSup, and construct a local frequent one-item set item header table on each node. The input is the global transaction set FList, and the output is the local frequent one-item set item header table L(x, sup). The following specific operations are carried out:
[0039] Step 4-1: On each node, input the globally statistically obtained transaction set FList. Step 4-2: According to the global transaction set FList, parallelly sort the local data sets F_dict j of each node to obtain the occurrence frequency value of the single attribute key, that is, the support sup, and put the combination (key, value) with value greater than or equal to the minimum support minSup into the local frequent one-item set item header table L, denoted as L(x, sup). The traversal results are sorted in descending order of support sup.
[0040] Step 4-3: Each node generates a local frequent one-item set item header table L(x, sup) according to the scanning results. This table records all access request attributes x that meet the support requirements and their corresponding supports sup.
[0041] Step 5: Parallelly construct a local FP tree on each node. The input is the local frequent one-item set item header table L(x, sup), and the output is the local FP tree root j , and the following specific operations are carried out:
[0042] Step 5-1: Each node parallelly inputs its own local frequent one-item item header table L(x, sup), rearranges the local data set F_dict j according to the order of the item header table in the table, eliminates the attributes with support less than the minimum support minSup, and records the sorted local data set sorted_F_dict j .
[0043] Step 5-2: Create a local FP tree root with the root node being null j , traverse the sorted local data set sorted_F_dict j , and insert each basic attribute x into the local FP tree root j one by one.
[0044] Step 5-3: For each basic attribute x, check whether there is a node node of this attribute in the current path x . If not, create a new node node x ; if the node exists, increment the item support sup of the node node x by 1.
[0045] Step 5-4: Assume that a certain path of the attribute x in the local FP-tree root j is non-frequent, and its corresponding node node x has prefix paths L1 and L2, and L1 is a sub-path of L2. Then merge the node x of path L2 with path L1, and at the same time delete the node node x .
[0046] Step 5-5: During the process of each node constructing the local FP-tree root j , use the linked list header_table j = {x: [node1, node2, …, node n} to record the local FP-tree nodes node i corresponding to each frequent item x.
[0047] Step 6: Parallelly construct the local conditional FP-tree, input the local FP-tree root j , the header pointer list header_table j and the sorted local data set sorted_F_dict j , and the output is the local conditional database paths j , and the specific steps are as follows:
[0048] Step 6-1: Each node parallelly selects a frequent item x from the local data set sorted_F_dict j , and records all the items (node, x1, x2, …, x j (excluding root j ) from this frequent item node node to the local FP-tree root node root n ), which is the path path containing this frequent item x.
[0049] Step 6-2: For each path path, delete this frequent item x and record the remaining items path x = (x1, x2, …, x n ), count the frequency of occurrence of each remaining item as cnt x , and perform pruning operations on the paths with frequencies less than the minimum support minSup.
[0050] Step 6-3: Define the conditional pattern base of the frequent item x as cpb x ={x:{path1:cnt1},…{path x :cnt x}…,{path n :cnt n}} where path x is a prefix path between the element item x and the local FP tree root node root j .
[0051] Step 6-4: Traverse the local dataset sorted_F_dict j in parallel. For each frequent item x, generate their respective local conditional databases paths j ={cpb1,cpb2,…,cpb n} according to Steps 6-1, 6-2, and 6-3
[0052] Step 7: Each node performs mining of local frequent item sets. The input is the local conditional database paths j and the output is the set of local frequent item sets Frequents j . The specific operations are as follows
[0053] Step 7-1: Each node traverses the local conditional database paths j in parallel. For each conditional pattern base cpb x , use the local conditional FP tree construction mechanism to construct the local conditional FP tree cfp_root j and obtain the local conditional FP tree header pointer list cfp_header_table j .
[0054] Step 7-2: Delete some nodes in the local conditional FP tree cfp_root j whose frequency cnt x is less than the minimum support minSup. If this tree is not a single path, recursively use the local conditional FP tree construction mechanism to mine the local conditional FP tree cfp_root j until there are no elements in the local conditional FP tree cfp_root j . Step 7-3: After forming a single path in the above Step 7-2, combine the element item x with each element in the local conditional FP tree and count the occurrences to obtain the frequent item set of x, frequents_x = {{new_path1:cnt1},{new_path2:cnt2},…,{new_path n :cnt n}}。
[0055] Step 7-4: Consider the remaining elements in the local condition database paths j and execute according to Steps 7-1, 7-2, and 7-3. Then, the local frequent itemset Frequents of each element item of each node can be obtained j ={item: frequents_item}.
[0056] Step 8: Generate association rules according to the confidence and lift thresholds. The input is the local frequent itemset Frequents j , and the output is the association rule rules. The specific operations are as follows:
[0057] Step 8-1: Integrate the local frequent itemset Frequents of each node j , remove the same local frequent itemset, and obtain the global frequent itemset Frequents.
[0058] Step 8-2: Define the rule set Rules = {A, B}, where A is the premise of the rule and B is the conclusion of the rule. It is composed of all possible non-empty subsets A and the corresponding remaining part B generated by the frequent itemset Frequents = {f1, f2, f3,..., f n}, that is, satisfying A ∪ B = Frequents.
[0059] Step 8-3: According to the calculation formula of confidence, calculate the confidence of each rule
[0060] where support(A ∪ B) is the frequency of the itemset A ∪ B in all transactions, and support(A) is the frequency of the itemset A in all transactions.
[0061] Step 8-4: According to the calculation formula of lift, calculate the lift of each rule
[0062] where support(B) is the frequency of the itemset B in all transactions.
[0063] Step 8-5: Set the minimum confidence minCon and the minimum lift minLift, traverse the rule set Rules, and filter out the rules whose confidence is greater than or equal to the minimum confidence minCon and the lift is greater than or equal to the minimum lift minLift. These are the association rules rules that meet the requirements.
[0064] Step 9: Divide a large number of small files into strongly and weakly related groups. The input is the association rules "rules" and the set of large numbers of small files "Files", and the output is the strongly related file group "SR_Files" and the weakly related file group "WR_Files". The following specific operations are carried out:
[0065] Step 9-1: Traverse the association rules "rules" to obtain the association rule "rule" in turn x , where
[0066] Step 9-2: Define a file association determination mechanism: Discuss "rule" x . If there are certain attributes "x" in "a" corresponding to the file "f" x , and there are certain attributes "y" in "b" corresponding to the file "f" y , then it is considered that there is a strong association between the file "f" x and the file "f" y . If all attributes in "a" correspond to the file "f" z , and all attributes in "b" also correspond to the file "f" z , then it is considered that the file "f" z has a weak association with other files.
[0067] Step 9-3: Traverse the set of large numbers of small files "Files". According to the determination mechanism in (2), append the small files with strong associations to the strongly related file group "SR_Files" in the form of key-value, and go to Step 10; append the small files with weak associations to the weakly related file group "WR_Files" in the form of key-value, and go to Step 11; where the key is the small file "f" x , and the value is its size.
[0068] Step 10: Merge the files in the strongly related file group. The input is the strongly related file group "SR_Files", and the output is the root node "huffman_root" of the Huffman tree. The following specific operations are carried out:
[0069] Step 10-1: Traverse the strongly related file group "SR_Files", regard each file as an independent node "node", and the weight of the node is the size "size" of the file.
[0070] Step 10-2: Insert all file nodes into the priority queue "heap", and the priority queue "heap" will be automatically sorted according to the file size.
[0071] Step 10-3: Take out the two nodes "node" l and "node" r(i.e., the two files with the smallest file sizes), merge these two nodes into a new node merged, and the weight of the new node merged is the sum of the two file sizes.
[0072] Step 10-4: Re-insert the new node merged into the priority queue heap, and repeatedly execute (3) until there is only one node huffman_root left in the queue heap.
[0073] Regarding the encoding on the left branch as "0" and the encoding on the right branch as "1", consider the string composed of the encoded characters on the path from the root node huffman_root to each file node as the index of the file.
[0074] Step 11: Create a file merge queue MergeFile, append the files in the weakly associated file group WR_Files to the queue in sequence, and calculate the size of the merge queue MergeFile. If it exceeds 3 / 4 of the HDFS data block size, directly merge the files in the merge queue MergeFile and store them in the HDFS file system; otherwise, return to Step 9 and wait for the generation of subsequent new weakly associated file groups WR_Files.
[0075] Step 12: The metadata server assigns a globally unique key r value to the root node huffman_root to identify different root nodes, and this key r value is also saved in the attributes of each file in its corresponding strongly associated file group SR_Files. At the same time, record the size of the root node as Size. Assign globally unique key f values to the files in the strongly associated file group SR_Files in sequence to identify each file, and at the same time record the index of this file in Step 10 as index, which is used for the data storage device to obtain the physical address of the file.
[0076] Step 13: The metadata server creates an object Object, allocates the object on the data storage device, and the data storage device stores the root node huffman_root into the object Object according to the key r and Size assigned by the metadata server, and then stores the small files in the strongly associated file group into the object Object in sequence according to the assigned key f , index and key r respectively.
[0077] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A small file merging optimization method for distributed cloud native storage, characterized in that: The method comprises the following steps: Step 1: Define the small file size threshold size. According to the HDFS file access log, if the file f i If the size is smaller than the small file size threshold, the file f i Append to the massive small file collection, and obtain a partial massive small file collection as Files = {f1,f2,f3,…,f n }, the historical order of user access requests to the file system is Requests = {r1, r2, r3, ..., r n }, where each request r i , 1≤i≤n, corresponding to a small file f i , 1≤i≤n, a small file contains user attribute u, file type t and time attribute s. Any small file can be represented by three basic attributes (u, t, s). The access request sequence containing the basic attributes can be expressed as: Requests = {(u1, t1, s1), (u2, t2, s2), (u3, t3, s3), …, (u n , t n ,s n )}; Step 2: Scan the user access request sequence Requests, divide the basic attribute data set, copy the data set to HDFS, and use HDFS to decompose the data set into continuous blocks and save the corresponding copies. Each block is distributed and stored on N nodes. This step is automatically completed by Hadoop. Step 3: Define the transaction dictionary F_dict on each node j ={key1:value1,key2:value2,…}, 1≤j≤N, key is a unique attribute in the access request sequence, value is the frequency of the attribute, and each node traverses the user access request sequence Requests in parallel j , 1≤j≤N, determine whether the attribute exists in the transaction dictionary F_dict j If it does not exist, add the transaction dictionary F_dict j , key is set to the attribute, value is set to 1; if it exists, the value corresponding to the attribute is increased by 1, and finally each obtains the updated transaction dictionary F_dict j , count the local frequency of each item in all nodes, sort them in descending order of frequency, and get the global transaction set FList. The frequency formula of the statistical data item used is: Among them, value_I i Represents data item I i Frequency of occurrence, value_T j I i Represents the data item I on the jth node i frequency of occurrence; Step 4: Set the minimum support minSup, build a local frequent item set header table on each node, input is the global transaction set FList, and output is the local frequent item set header table L(x,sup); Step 5: Construct a local FP tree in parallel on each node. The input is the local frequent item set header table L(x,sup), and the output is the local FP tree root. j ; Step 6: Construct the local conditional FP tree in parallel and input the local FP tree root j , header pointer list header_table j and the sorted local dataset sorted_F_dict j , the output is the local condition database paths j ; Step 7: Each node mines local frequent itemsets, and the input is the local condition database paths j , the output is the local frequent itemset set Frequents j ; Step 8: Generate association rules according to confidence and lift thresholds, with input as local frequent item sets Frequents j , the output is association rules; Step 9: Divide the massive small files into strong and weak associations. The input is association rules, the massive small file set Files, and the output is the strongly associated file group SR_Files and the weakly associated file group WR_Files. Step 10: Merge the strongly associated file groups. The input is the strongly associated file group SR_Files, and the output is the root node huffman_root of the Huffman tree. The code on the left branch is "0", and the code on the right branch is "1", and the string consisting of the coded characters on the path from the root node huffman_root to each file node is regarded as the index of the file; Step 11: Create a file merge queue MergeFile, append the files in the weakly associated file group WR_Files to the queue in turn, and calculate the size of the merge queue MergeFile. If it exceeds 3 / 4 of the HDFS data block size, directly merge the files in the merge queue MergeFile and store them in the HDFS file system; otherwise, return to step 9 to wait for the generation of a subsequent new weakly associated file group WR_Files; Step 12: The metadata server assigns a globally unique key to the root node huffman_root r Value, used to identify different root nodes, the key r The value is also saved in the attributes of each file in the corresponding strongly associated file group SR_Files, and the root node size is recorded as Size. The globally unique key is assigned to the files in the strongly associated file group SR_Files in turn. f A value is used to identify each file, and the index of the file in step 10 is recorded as index, which is used by the data storage device to obtain the physical address of the file; Step 13: The metadata server creates an object Object and allocates the object to the data storage device. The data storage device stores the root node huffman_root according to the key allocated by the metadata server. r The size and size are stored in the object Object, and the strongly related files are grouped into small files according to the assigned key. f , index and key r Stored in the object Object.
2. According to the small file merging optimization method for distributed cloud native storage according to claim 1, it is characterized in that: The step 4 comprises: Step 4-1: On each node, input the global transaction set FList obtained by statistics; Step 4-2: Based on the global transaction set FList, parallelize the local data set F_dict of each node j Sort and obtain the frequency of occurrence value of the single attribute key, that is, the support sup, and put the combination (key, value) whose value is greater than or equal to the minimum support minSup into the local frequent item set header table L, recorded as L(x,sup), and sort the traversal results in descending order of support sup; Step 4-3: Each node generates a local frequent item set header table L(x, sup) based on the scanning results. The table records all access request attributes x whose support meets the requirements and their corresponding support sup.
3. According to the small file merging optimization method for distributed cloud native storage according to claim 1, it is characterized in that: The step 5 comprises: Step 5-1: Each node inputs its own local frequent item header table L(x,sup) in parallel, and rearranges the local data set F_dict according to the order of the item header table in the table j , remove attributes whose support is less than the minimum support minSup, and record the sorted local data set sorted_F_dict j ; Step 5-2: Create a local FP tree root with null as the root node j , traverse the sorted local data set sorted_F_dict j , insert each basic attribute x one by one into the local FP tree root j middle; Step 5-3: For each basic attribute x, check whether there is a node node with this attribute in the current path. x If not, create a new node x ; If the node exists, the node node x The item support sup is increased by 1; Step 5-4: Assume that attribute x is in the local FP tree root j A path in is infrequent, and its corresponding node node x If there are prefix paths L1 and L2, and L1 is a subpath of L2, then the node x Merge with path L1 and add node node x delete; Step 5-5: Each node builds a local FP tree root j In the process, the linked list header_table is used j ={x:[node1,node2,…,node n ]}Record the local FP tree node node corresponding to each frequent item x i .
4. According to the small file merging optimization method for distributed cloud native storage according to claim 1, it is characterized in that: The step 6 comprises: Step 6-1: Each node in parallel extracts the local dataset sorted_F_dict j Select a frequent item x from the local FP tree root node and record the data from the frequent item node node to the local FP tree root node root j (excluding root j ) of all items (node,x1,x2,…,x n ), which is the path path containing the frequent item x; Step 6-2: For each path, delete the frequent item x and record the remaining item path x =(x1,x2,…,x n ), count the frequency of each remaining item as cnt x , prune the paths whose frequency is less than the minimum support minSup; Step 6-3: Define the conditional pattern base of the frequent item x as cpb x ={x:{path1:cnt1},…{path x :cnt x }…,{path n :cnt n }}, where path x is a line between the element item x and the root node of the local FP tree j The prefix path between ; Step 6-4: Parallel traversal of the local data set sorted_F_dict j For each frequent item x, generate the respective local conditional database paths according to steps 6-1, 6-2, and 6-3 j ={cpb1,cpb2,…,cpb n }.
5. According to the small file merging optimization method for distributed cloud native storage according to claim 1, it is characterized in that: The step 7 comprises: Step 7-1: Each node traverses the local condition database paths in parallel j , for each conditional pattern base cpb x , use the local conditional FP tree construction mechanism to construct the local conditional FP tree cfp_root j , get the local conditional FP tree head pointer list cfp_header_table j ; Step 7-2: Delete the local conditional FP tree cfp_root j Medium frequency cnt x If the tree is not a single path, the local conditional FP tree construction mechanism is recursively used to construct the local conditional FP tree cfp_root. j Dig until the local conditional FP tree cfp_root j Until there are no elements in the Step 7-3: After forming a single path in step 7-2 above, combine the element item x with each element in the local conditional FP tree and count them to obtain the frequent item set of x, frequents_x = {{new_path1:cnt1},{new_path2:cnt2},…,{new_path n :cnt n }}; Step 7-4: Consider local conditional database paths j For the remaining elements in, follow steps 7-1, 7-2, and 7-3 to get the local frequent item sets Frequents of each element item in each node. j ={item:frequents_item}.
6. According to the small file merging optimization method for distributed cloud native storage according to claim 1, it is characterized in that: The step 8 comprises: Step 8-1: Frequents of each node j Integration, eliminating the same local frequent item sets, to obtain the global frequent item sets Frequents; Step 8-2: Define the rule set Rules = {A, B}, where A is the premise of the rule and B is the conclusion of the rule, which is composed of the frequent item set Frequents = {f1, f2, f3, ..., f n All possible non-empty subsets A generated by} and the corresponding remaining parts B, that is, A∪B=Frequents; Step 8-3: Calculate each rule according to the confidence calculation formula Confidence Where support(A∪B) is the frequency of item set A∪B in all transactions, and support(A) is the frequency of item set A in all transactions; Step 8-4: Calculate each rule according to the calculation formula of lift The degree of improvement Where support(B) is the frequency of item set B appearing in all transactions; Step 8-5: Set the minimum confidence minCon and the minimum lift minLift, traverse the rule set Rules, and filter out the rules whose confidence is greater than or equal to the minimum confidence minCon and whose lift is greater than or equal to the minimum lift minLift, which are the association rules that meet the requirements.
7. According to the small file merging optimization method for distributed cloud native storage according to claim 1, it is characterized in that: The step 9 comprises: Step 9-1: Traverse the association rules and get the association rules rule in turn x ,in Step 9-2: Define the file association determination mechanism: rule x For discussion, if there is some attribute x in a corresponding to file f x , and there is some attribute y in b that corresponds to file f y , then the file f x and file f y There is a strong correlation between them. If all attributes in a correspond to file f z , and all attributes in b also correspond to file f z , then the file f z There is a weak correlation with other files; Step 9-3: Traverse the massive small file set Files, and according to the judgment mechanism of (2), add the small files with strong correlation to the strongly correlated file group SR_Files in the form of key-value, and go to step 10; add the small files with weak correlation to the weakly correlated file group WR_Files in the form of key-value, and go to step 11; where key is the small file f x , value is its size.
8. According to the small file merging optimization method for distributed cloud native storage according to claim 1, it is characterized in that: The step 10 comprises: Step 10-1: Traverse the strongly associated file group SR_Files, and regard each file as an independent node node. The weight of the node is the size of the file size; Step 10-2: Insert all file nodes into the priority queue heap, which will be automatically sorted according to file size; Step 10-3: Take out the two nodes with the smallest weight from the priority queue heap l and node r (i.e. the two files with the smallest file size), merge these two nodes into a new node merged, and the weight of the new node merged is the sum of the sizes of the two files; Step 10-4: Re-insert the new node merged into the priority queue heap, and repeat (3) until only one node huffman_root remains in the queue heap; The code on the left branch is "0", and the code on the right branch is "1", and the string consisting of the coded characters on the path from the root node huffman_root to each file node is regarded as the index of the file.
Citation Information
Cited By
Crop phenotype data storage and retrieval method and device, equipment and storage medium
CN120723776A
Crop phenotype data storage and retrieval method, device, equipment and storage medium
CN120723776B