Distributed processing method for large-scale data
By classifying and grouping files to be processed and matching the data processing algorithm of the computing nodes, the efficiency and accuracy of large-scale data processing under the distributed architecture are solved, and efficient data processing and adaptability of computing nodes are achieved.
Patent Information
- Application Number
- CN202510051281.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-23
AI Technical Summary
The existing large-scale data processing methods under distributed architectures have problems such as unclear division of labor, low processing efficiency and poor processing results.
By classifying and grouping the to-processed files according to the file type, matching the data processing algorithm of the computing node, generating the mapping relationship between the file group and the computing node, and realizing distributed computing and result merging.
Improve the efficiency and accuracy of data processing, and ensure the data uniformity and adaptability of computing nodes during processing.
Smart Images

Figure CN120029980A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of distributed computing and relates to a distributed processing method for large-scale data. Background Art
[0002] File data processing refers to the process of parsing, extracting, converting and analyzing various types of files. These files can be text files, spreadsheets, image files, audio files, etc. Through data processing technology, useful information can be obtained from them and further analyzed and applied. The first step in file data processing is to parse the file and extract the data content according to the format and structure of the file. Different parsing tools and techniques can be used for different types of files, such as text parsers, image processing algorithms, etc. Based on file parsing, data with specific meaning or value in the file is extracted according to needs. These data can be text information, digital data, image features, audio signals, etc., for subsequent analysis and processing.
[0003] However, the existing data processing under the distributed architecture mostly relies on unified data processing algorithms to process data, and the division of labor between multiple computing nodes is not clear. When processing simple data such as images, if the data processing algorithm used is too complex, it will not be able to process large-scale and simple data, and the efficiency of simple data processing will be greatly reduced. If a simple data processing algorithm is used to analyze some complex data, the data processing effect may not meet expectations, reducing the accuracy and speed of data processing. Summary of the invention
[0004] In order to solve the problems existing in the background technology, the present invention proposes a distributed processing method for large-scale data.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] Classify the files to be processed according to the file type, and store the files to be processed in different storage nodes according to the classification results. Storage nodes with the same file type constitute a storage cluster;
[0007] Grouping the to-be-processed files of the same file type in the storage cluster according to the data size and data structure of the to-be-processed files stored in the same storage cluster;
[0008] According to the data processing logic, data structure and real-time performance of the same file group, the data processing algorithm of the computing node is matched, and a mapping relationship is generated between each file group and different computing nodes;
[0009] Each computing node calculates the data in the corresponding file group it maps, and multiple computing nodes perform intermediate calculations at the same time, and merge the intermediate calculation results to obtain the processing results of large-scale data.
[0010] Furthermore, the specific method of classifying the files to be processed according to the file type is:
[0011] When uploading the files to be processed, the user adds tag information to the files to be processed;
[0012] Traverse the files to be processed and extract the file name, file extension and tag information;
[0013] Firstly, the files to be processed are classified into different file types according to their file extensions, thus completing a classification;
[0014] The files to be processed are secondary classified according to their tag information. The tag information of the files to be processed is extracted using the transformer model, and the files to be processed with similar semantics and the same file type are classified to complete the secondary classification.
[0015] Furthermore, the specific method for classifying the to-be-processed files of the same file type in the storage cluster is as follows:
[0016] The files to be processed in the same storage cluster are grouped according to data size and data structure, and all file types are traversed to obtain multiple file groups of files to be processed under the same file type.
[0017] Furthermore, the specific method for generating a mapping relationship between each file group and different computing nodes is:
[0018] According to the processing logic, data structure and real-time requirements of the files to be processed in the file group, a decision tree model is used to match the data processing algorithm for the current file group;
[0019] The algorithm database sends the data processing algorithm output by the decision tree and required for processing the current file group to any computing node without a mapping relationship, associates the computing node with the currently matched file group, and generates a mapping relationship.
[0020] Furthermore, the decision tree is constructed as follows:
[0021] When a file group M is input, a feature set X is created based on the file group, X = {S, D, R}, where S is the processing logic of the file group, S ∈ {serial, parallel, distributed}, D is the data structure of the file group, D ∈ {a 1 ,a 2 ,...,a n}, n is a natural number, R is the real-time requirement of the file group during processing, R∈{b 1 ,b 2 ,...,b n};
[0022] First, based on the entropy h(M) of the file group, the specific formula is:
[0023]
[0024] where p i is the probability of the i-th category. After determining the entropy value of the file group, add feature A to divide the file group and obtain the conditional entropy h(M|A) of the file group. The specific formula is:
[0025]
[0026] Where V(A) is the set of all values of feature A, P(v) is the probability when feature A takes value v, and h(M|A=v) is the entropy of the current file group M when feature A takes value v;
[0027] The information gain IG(M,A) of the file group is calculated by the entropy value h(M) and the conditional entropy value h(M|A). The specific formula is:
[0028] IG(M,A)=h(M)-h(M|A)
[0029] By inputting training data test M Train the decision tree model, test M is the training data set, where test M Each file group in includes a feature set {S, D, R}, which completes the training of the decision tree;
[0030] The processing logic, data structure and real-time requirements are as follows:
[0031] Processing logic S∈{serial,parallel,distributed}, where serial is serial computing, parallel is parallel computing, and distributed is distributed computing;
[0032] Data structure D∈{a 1 ,a 2 ,a 3 ,...}, where a 1 is a linear data structure, a 2 is a tree data structure, a 3 For graph data structures…;
[0033] Real-time R∈{b1 ,b 2 ,...,b n}, where b 1 ,b 2 ,...,b n The real-time level is determined by users themselves, which limits the data processing speed.
[0034] Furthermore, the method for merging the intermediate calculation results is:
[0035] The computing nodes include primary computing nodes and secondary computing nodes, the primary computing nodes are responsible for computing and processing the data in the corresponding file group, and the secondary computing nodes are responsible for processing the intermediate computing results of the corresponding primary computing nodes;
[0036] After the primary computing node completes the current task, the intermediate computing results are sent to the corresponding secondary computing node. The intermediate computing results are merged through data exchange between multiple secondary computing nodes to complete the merging of the intermediate computing results.
[0037] Furthermore, the processor of the required computing node includes a CPU and / or a GPU processor.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] The present invention achieves effective data classification by classifying the files to be processed according to the file type, and grouping the files to be processed of the same file type in the storage cluster according to the data size and data structure of the files to be processed stored in different storage nodes, thereby ensuring that the data types in the same type of data are the same and the data sizes are similar, so that the computing nodes can be fully adapted during the data processing process, and the computing nodes maintain the uniformity of the data during the calculation process.
[0040] According to the data processing logic, data structure and real-time performance of the file group, the data processing algorithm of the computing node is matched, and a mapping relationship is generated between each file group and different computing nodes. Each computing node calculates the data in the corresponding category storage node to which it is mapped, so that each computing node can use the data processing algorithm to adapt to the files to be processed, which greatly improves the efficiency of computing nodes in processing data. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a flow chart of the method of the present invention. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0043] like Figure 1 As shown, the technical solution adopted by the present invention is as follows: a distributed processing method for large-scale data, comprising:
[0044] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0045] The files to be processed are classified according to the file type, and based on the classification results of the files to be processed, the files to be processed are stored in different storage nodes. Storage nodes with the same file type constitute a storage cluster.
[0046] The to-be-processed files of the same file type in the storage cluster are grouped according to the data size and data structure of the to-be-processed files stored in the same storage cluster.
[0047] According to the data processing logic, data structure and real-time performance of the same file group, the data processing algorithm of the computing node is matched, and a mapping relationship is generated between each file group and different computing nodes.
[0048] Each computing node calculates the data in the corresponding file group it maps, and multiple computing nodes perform intermediate calculations at the same time, and merge the intermediate calculation results to obtain the processing results of large-scale data.
[0049] When uploading files to be processed, users can add tag information to the files to be processed.
[0050] Traverse the files to be processed and extract the file name, file extension and tag information. Use the file name as an index tag to extract the relevant information of all the files to be processed, and extract similar content from them as the basis for classification.
[0051] First, the files to be processed are classified into different file types according to their file extensions, completing a classification. First, the file type is determined according to the file extension, and a classification is completed according to the type. In the subsequent classification process, only the same file type needs to be classified, reducing the amount of data processing.
[0052] The files to be processed are secondary classified according to their tag information. The tag information of the files to be processed is extracted using the transformer model, and the files to be processed with similar semantics and the same file type are classified to complete the secondary classification. Transformer is a deep learning model architecture used for natural language processing and other sequence-to-sequence tasks. It has a self-attention mechanism and can effectively extract and classify the user's tag information in the files to be processed.
[0053] After the files to be processed are classified according to the storage nodes, the clarity of the data is greatly improved. However, in the same storage node, the data is still not clear enough, and the data may still be inconsistent during data processing. Therefore, the data needs to be further grouped.
[0054] The file type of the to-be-processed file stored in any storage node is read to obtain the file type stored in the storage node, and the file types stored in other storage nodes are searched based on the file type to obtain the storage cluster corresponding to the file type.
[0055] Storage nodes with the same file type form a storage cluster, and the to-be-processed files of the same file type in the storage cluster are grouped according to the data size and data structure of the to-be-processed files stored in the same storage cluster.
[0056] According to the data size and data structure of the to-be-processed files stored in the same storage cluster, the specific method for grouping the to-be-processed files of the same file type in the storage cluster is as follows:
[0057] The files to be processed in the storage cluster are judged based on the data size and data structure. The files to be processed with similar file sizes and data structures are extracted, and all file types are traversed to obtain multiple file groups of the same file type. The file groups divided in this way classify the data size and data structure, which greatly improves the uniformity of the files in subsequent processing.
[0058] According to the data processing logic, data structure and real-time performance of the same file group, the data processing algorithm of the computing node is matched, and a mapping relationship is generated between each file group and different computing nodes.
[0059] The processor of the computing node includes a CPU and a GPU processor. The processor of the computing node can be a CPU processor or a GPU processor alone, or a CPU and GPU combined processor. Different processors can execute different algorithms, and can process different data types, and can process different data sizes. Therefore, it is necessary to match the algorithm and processor for the file group to realize the use of computing power when processing different file groups.
[0060] According to the processing logic, data structure and real-time requirements of the files to be processed in the file group, a decision tree model is used to match the data processing algorithm for the current file group.
[0061] A decision tree is a model that uses a tree-shaped data structure to display decision rules and classification results. As an inductive learning algorithm, its focus is to transform seemingly disordered and messy known data into a tree-shaped model that can predict unknown data through some technical means. Each path from the root node to the leaf node represents a decision rule. The decision tree model can match different file groups with data processing algorithms with higher adaptability.
[0062] When a file group M is input, a feature set X is created based on the file group, X = {S, D, R}, where S is the processing logic of the file group, S ∈ {serial, parallel, distributed}, D is the data structure of the file group, D ∈ {a 1 ,a 2 ,...,a n}, n is a natural number, R is the real-time requirement of the file group during processing, R∈{b 1 ,b 2 ,...,b n}.
[0063] The processing logic, data structure and real-time requirements are as follows:
[0064] The processing logic S∈{serial,parallel,distributed}, where serial is serial computing, parallel is parallel computing, and distributed is distributed computing. Serial computing, parallel computing, and distributed computing have their own restrictions, and can match data processing algorithms based on different processing logics, ensuring that the matched data processing algorithms meet the corresponding requirements of the data group in terms of processing logic.
[0065] Data structure D∈{a 1 ,a 2 ,a 3 ,...}, where a 1is a linear data structure, a 2 is a tree data structure, a 3 It is a graph data structure, which also includes hash tables, sets, maps, etc.
[0066] Linear data structures include arrays, linked lists, stacks, queues, etc.; tree structures include binary trees, binary search trees, balanced trees, heaps, dictionary trees, etc.; image data structures include graphs, adjacency matrices, adjacency lists, etc. Using different data structures and different data processing algorithms can effectively save time and space complexity in the calculation process.
[0067] Real-time R∈{b 1 ,b 2 ,...,b n}, where b 1 ,b 2 ,...,b n The real-time level is determined by the user and limits the data processing speed. The real-time requirement can effectively control the operating efficiency of the data processing algorithm.
[0068] The decision tree is constructed as follows:
[0069] First, based on the entropy h(M) of the file group, the specific formula is:
[0070]
[0071] where p i is the probability of the i-th category. After determining the entropy value of the file group, add feature A to divide the file group and obtain the conditional entropy h(M|A) of the file group. The specific formula is:
[0072]
[0073] Where V(A) is the set of all values of feature A, P(v) is the probability when feature A takes value v, and h(M|A=v) is the entropy of the current file group M when feature A takes value v. The conditional entropy h(M|A) measures the uncertainty of the target variable Y when feature A is known. It reflects the classification ability of feature A for the target variable. The smaller the conditional entropy, the lower the uncertainty of predicting Y through feature A and the stronger the classification ability.
[0074] The information gain IG(M,A) of the file group is calculated by the entropy value h(M) and the conditional entropy value h(M|A). The specific formula is:
[0075] IG(M,A)=h(M)-h(M|A)
[0076] Information gain measures the reduction in information uncertainty after selecting feature A to classify dataset M.
[0077] By inputting training data test M Train the decision tree model, test M is the training data set, where test M Each file group in includes a feature set {S, D, R}, which completes the training of the decision tree.
[0078] The algorithm database sends the data processing algorithm required for processing the current file group output by the decision tree to any computing node without a mapping relationship, associates the computing node with the currently matched file group, and generates a mapping relationship. Sending to any computing node without a mapping relationship can effectively prevent the computing nodes that already have a mapping relationship from being overwritten by the new algorithm, and effectively improves the fault tolerance when assigning data processing algorithms to computing nodes. The processor of the computing node should meet the application requirements of the data processing algorithm.
[0079] Each computing node calculates the data in the corresponding category storage node to which it is mapped. Multiple computing nodes perform calculations simultaneously, and the final calculation results are merged to obtain the processing results of large-scale data.
[0080] The method for merging the final calculation results is:
[0081] The computing nodes include primary computing nodes and secondary computing nodes. The primary computing nodes are responsible for computing and processing the data in the corresponding file group, and the secondary computing nodes are responsible for processing the intermediate computing results of the corresponding primary computing nodes.
[0082] After the primary computing node completes the current task, the intermediate computing results are sent to the corresponding secondary computing node. The intermediate computing results are merged through data exchange between multiple secondary computing nodes to complete the merging of the intermediate computing results.
[0083] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A distributed processing method for large-scale data, characterized in that: Included are: Classify the files to be processed according to the file type, and store the files to be processed in different storage nodes according to the classification results. Storage nodes with the same file type constitute a storage cluster; Grouping the to-be-processed files of the same file type in the storage cluster according to the data size and data structure of the to-be-processed files stored in the same storage cluster; According to the data processing logic, data structure and real-time performance of the same file group, the data processing algorithm of the required computing nodes is matched, and a mapping relationship is generated between each file group and different computing nodes; Each computing node calculates the data in the corresponding file group it maps, and multiple computing nodes perform intermediate calculations at the same time, and merge the intermediate calculation results to obtain the processing results of large-scale data.
2. A distributed processing method for large-scale data according to claim 1, characterized in that: The specific method of classifying the files to be processed according to the file type is as follows: When uploading the files to be processed, the user adds tag information to the files to be processed; Traverse the files to be processed and extract the file name, file extension and tag information; Firstly, the files to be processed are classified into different file types according to their file extensions, thus completing a classification; The files to be processed are secondary classified according to their tag information. The tag information of the files to be processed is extracted using the transformer model, and the files to be processed with similar semantics and the same file type are classified to complete the secondary classification.
3. The distributed processing method for large-scale data according to claim 1, characterized in that: The specific method for classifying the to-be-processed files of the same file type in the storage cluster is as follows: The files to be processed in the same storage cluster are grouped according to data size and data structure, and all file types are traversed to obtain multiple file groups of files to be processed under the same file type.
4. A distributed processing method for large-scale data according to claim 3, characterized in that: The specific method for generating a mapping relationship between each file group and different computing nodes is: According to the processing logic, data structure and real-time requirements of the files to be processed in the file group, a decision tree model is used to match the data processing algorithm for the current file group; The algorithm database sends the data processing algorithm output by the decision tree and required for processing the current file group to any computing node without a mapping relationship, associates the computing node with the currently matched file group, and generates a mapping relationship.
5. A distributed processing method for large-scale data according to claim 4, characterized in that: The decision tree is constructed as follows: When a file group M is input, a feature set X is created based on the file group, X = {S, D, R}, where S is the processing logic of the file group. S∈{ serial, parallel, distributed}, D is the data structure of the file group, D∈{a1,a2,...,a n }, n is a natural number, R is the real-time requirement of the file group during processing, R∈{b1,b2,...,b n }; First, based on the entropy h(M) of the file group, the specific formula is: where p i is the probability of the i-th category. After determining the entropy value of the file group, add feature A to divide the file group and obtain the conditional entropy h(M|A) of the file group. The specific formula is: Where V(A) is the set of all values of feature A, P(v) is the probability when feature A takes value v, and h(M|A=v) is the entropy of the current file group M when feature A takes value v; The information gain IG(M,A) of the file group is calculated by the entropy value h(M) and the conditional entropy value h(M|A). The specific formula is: IG(M,A)=h(M)-h(M|A) By inputting training data test M Train the decision tree model, test M is the training data set, where test M Each file group in includes a feature set {S, D, R}, which completes the training of the decision tree; The processing logic, data structure and real-time requirements are as follows: Processing Logic S∈{ serial, parallel, distributed}, where serial refers to serial computing, parallel refers to parallel computing, and distributed refers to distributed computing; Data structure D∈{a1,a2,a3,...}, where a1 is a linear data structure, a2 is a tree data structure, a3 is a graph data structure, etc.; Real-time R∈{b1,b2,...,b n }, where b1, b2, ..., b n The real-time level is determined by users themselves, which limits the data processing speed.
6. A distributed processing method for large-scale data according to claim 1, characterized in that: The method for merging the intermediate calculation results is: The computing nodes include primary computing nodes and secondary computing nodes, the primary computing nodes are responsible for computing and processing the data in the corresponding file group, and the secondary computing nodes are responsible for processing the intermediate computing results of the corresponding primary computing nodes; After the primary computing node completes the current task, the intermediate computing results are sent to the corresponding secondary computing node. The intermediate computing results are merged through data exchange between multiple secondary computing nodes to complete the merging of the intermediate computing results.
7. The distributed processing method for large-scale data according to claim 1, characterized in that: The processor of the required computing node includes a CPU and / or a GPU processor.