File Fuzzy Copy Method and System Based on Big Data File Cluster

The method and system utilize deep learning for file content feature extraction and distributed parallel processing to address inefficiencies in existing file copy technologies, improving accuracy and reliability in file copying.

CN119739537BActive Publication Date: 2025-07-15北京科杰科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510243632.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-15
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The existing file copying technology based on the distributed computing framework cannot identify files with similar content but different file names, resulting in duplicate copying or missed copying, and lacks refined control and management, making it difficult to dynamically adjust the priority and resource allocation of replication tasks according to file importance and real-time.

Method used

The deep learning model is used to extract file content features, combine file names and metadata features to perform similarity calculations, and file copy operations are performed in parallel through a distributed computing framework, dynamically allocate task nodes, and perform parallel transmission and integrity verification at the data block level.

Benefits of technology

Improve file copying efficiency and accuracy, ensure data integrity and reliability, optimize resource utilization, and adapt to cluster environments of different sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119739537B_ABST
    Figure CN119739537B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for fuzzy file copying based on a big data file cluster, which relates to the technical field of file copying. The method includes extracting file content, file names, and metadata feature vectors from a set of files to be matched. Among them, the file content feature vector is obtained by encoding based on a deep learning model. Then, a distributed computing framework is used to parallelly calculate the similarity scores between the files to be matched and the files in the target file set. The score is obtained by weighted calculation of the similarities of the file content, file names, and metadata feature vectors, and a list of files to be copied is generated by screening according to a preset threshold. Finally, the distributed file system dynamically allocates copy tasks according to system resources, executes file copying based on a data block-level parallel transmission mechanism, and verifies data integrity to generate a copy task execution report. The present invention can efficiently and accurately perform fuzzy file copying in a big data file cluster, improve the efficiency and accuracy of file copying, and reduce system resource consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to file copy technology, and in particular to a file fuzzy copy method and system based on a big data file cluster. Background Art

[0002] File copy is a basic and important operation in computer systems, and is widely used in scenarios such as data backup, data migration, and data synchronization. With the explosive growth of data volume, the traditional single-machine file copy method has been difficult to meet the processing requirements of massive data, and the parallel file copy technology based on the distributed computing framework has emerged as the times require. Such technologies usually divide the files to be copied into multiple data blocks, and utilize the parallel processing ability of the distributed computing framework to distribute the data blocks to multiple computing nodes for concurrent replication, thereby significantly improving the file copy efficiency.

[0003] However, the existing file copy technologies based on the distributed computing framework still have some defects and deficiencies:

[0004] 1. Lack of semantic understanding of file content: The existing technologies usually perform file matching and copying only based on file names or file paths, and cannot identify files with similar content but different file names, resulting in duplicate copying or missed copying. Especially when dealing with a large amount of unstructured data, the efficiency is low and errors are prone to occur.

[0005] 2. Single feature extraction method: The existing technologies usually rely only on simple file attributes (such as file names, file sizes, modification times, etc.) for file comparison, ignoring the features of the file content itself, and it is difficult to distinguish files with similar content but different attributes, such as different versions of documents, slightly modified pictures, etc.

[0006] 3. Lack of refined control and management: The existing technologies usually lack a refined control and management mechanism for the file replication process. For example, it is difficult to dynamically adjust the priority and resource allocation of the replication task according to the importance, timeliness, etc. of the file, and it is also difficult to effectively process and monitor errors and exceptions during the replication process. Summary of the Invention

[0007] The embodiments of the present invention provide a file fuzzy copy method and system based on a big data file cluster, which can solve the problems in the existing technologies.

[0008] In the first aspect of the embodiments of the present invention,

[0009] A file fuzzy copy method based on a big data file cluster is provided, including:

[0010] Acquire a set of files to be matched, perform feature extraction on each file in the set of files to be matched, and construct a set of feature vectors of files to be matched, wherein the set of feature vectors of files to be matched includes a file content feature vector, a file name feature vector, and a file metadata feature vector of each file, wherein the file content feature vector is obtained by encoding the file content through a deep learning model, the file name feature vector is obtained by encoding the file name through a character-level encoding model, and the file metadata feature vector includes a numerical representation of the file size, creation time, and modification time;

[0011] Based on a distributed computing framework, the feature vector set of files to be matched is divided into multiple subtasks, feature vector similarity calculation is performed in parallel on multiple computing nodes of the distributed computing framework, and a similarity score between each file to be matched and a file in the target file set is calculated, wherein the similarity score is obtained by weighted calculation of the file content feature vector similarity, the file name feature vector similarity and the file metadata feature vector similarity, and the similarity score is filtered according to a preset similarity threshold to generate a list of files to be copied;

[0012] The list of files to be copied is submitted to the task scheduler of the distributed file system. The task scheduler dynamically allocates the execution node of the copy task according to the system resource status, and starts the file copy process on the execution node. The file copy process executes the file copy operation based on the parallel transmission mechanism at the data block level, verifies the data integrity during the copying process, stores the copied files to the target storage location, and generates a copy task execution report.

[0013] The file content feature vector is obtained by encoding the file content through a deep learning model, including:

[0014] Execute the WordPiece word segmentation algorithm on the input file content to perform subword segmentation to obtain a word unit sequence, unify the word unit sequence into a preset length by dynamic truncation or padding, perform word unit encoding and position encoding on each word unit in the word unit sequence, wherein the position encoding uses a sine function and a cosine function to encode the word unit position information, and superimpose the word unit encoding and the position encoding to obtain a word unit embedding sequence;

[0015] Inputting the word-element embedding sequence into a multi-layer BERT encoder, each layer of which includes a multi-head self-attention sublayer and a feed-forward neural network sublayer, and the output of each of the multi-head self-attention sublayer and the feed-forward neural network sublayer is residually connected with the input and subjected to layer normalization processing to obtain a hidden feature representation of the file content;

[0016] Perform feature fusion processing on the hidden layer feature representation, map the hidden layer feature representation to a query matrix, a key-value matrix, and a value matrix respectively, calculate the outputs of multiple attention heads based on the query matrix, the key-value matrix, and the value matrix, concatenate the outputs of the multiple attention heads and perform a linear transformation to obtain a file content feature vector, and the file content feature vector is used to represent the semantic features of the file content.

[0017] Based on a distributed computing framework, divide the set of feature vectors of the files to be matched into multiple subtasks, and perform parallel feature vector similarity calculation on multiple computing nodes of the distributed computing framework. Calculating the similarity score between each file to be matched and the files in the target file set includes:

[0018] Two-dimensionally partition the set of feature vectors of the files to be matched and the set of feature vectors of the target files according to a preset row block size and column block size to obtain multiple feature vector calculation blocks, construct a calculation dependency graph based on the feature vector calculation blocks, the vertices of the calculation dependency graph represent the feature vector calculation blocks, the edges of the calculation dependency graph represent the data dependency relationships between the feature vector calculation blocks, and obtain the block size information of the feature vector calculation blocks;

[0019] Calculate the inter-block load balancing metric value based on the block size information of the feature vector calculation blocks, the inter-block load balancing metric value is the ratio of the largest calculation block size to the average calculation block size in the feature vector calculation blocks, and obtain the data transmission volume between the computing nodes in the calculation dependency graph;

[0020] Construct a task scheduling optimization objective function based on the inter-block load balancing metric value and the data transmission volume; allocate the feature vector calculation blocks to multiple computing nodes of the distributed computing framework according to the optimization result of the task scheduling optimization objective function, and establish a feature vector calculation task in each computing node, and the feature vector calculation task includes calculating the similarity of the feature vectors of the corresponding calculation block;

[0021] Parallelly execute the feature vector calculation tasks on the multiple computing nodes. For the feature vector calculation task of each computing node, perform L2 normalization processing on the feature vectors in the corresponding calculation block to obtain normalized feature vectors, and calculate the cosine similarity between the normalized feature vectors using matrix multiplication to obtain the local similarity calculation result;

[0022] Construct a multi-layer tree-shaped aggregation structure based on the calculated dependency graph. The leaf nodes of the multi-layer tree-shaped aggregation structure correspond to the local similarity calculation results of the multiple computing nodes. At each layer, an aggregation operation is performed on the local similarity calculation results of multiple child nodes. The aggregation operation selects the maximum similarity value between the same feature vector pairs as the aggregation result until the aggregation of all levels is completed to obtain the final similarity calculation result.

[0023] Construct an optimization objective function for task scheduling based on the inter-block load balancing metric value and the data transfer volume; allocating the feature vector calculation blocks to multiple computing nodes of the distributed computing framework according to the optimization result of the task scheduling optimization objective function includes:

[0024] Construct a calculation dependency adjacency matrix, where the elements of the calculation dependency adjacency matrix represent the dependency relationships between calculation blocks; based on the calculation dependency adjacency matrix, the block size information of the feature vector calculation blocks, and the task allocation status, construct a data transfer volume matrix, where the elements of the data transfer volume matrix are the total data transfer volumes between calculation blocks with dependency relationships.

[0025] Construct a multi-objective optimization function by using the inter-block load balancing metric value, the sum of the elements of the data transfer volume matrix, and the maximum load of the computing nodes. The multi-objective optimization function includes a load balancing weight coefficient, a communication overhead weight coefficient, and a computing load weight coefficient.

[0026] Generate an initial task allocation plan based on the min-max criterion. Under the conditions of satisfying the task integrity constraint, the node capacity constraint, and the dependency satisfaction constraint, select the plan with the minimum objective function value from the neighborhood solutions of the current allocation plan as the allocation plan for the next iteration through iterative optimization.

[0027] When the optimization iteration converges or reaches the preset number of iterations, send the optimal allocation plan to each computing node through the distributed scheduler. The optimal allocation plan includes the mapping relationship between the calculation blocks and the computing nodes and the task execution time constraint.

[0028] Submit the list of files to be copied to the task scheduler of the distributed file system. The task scheduler dynamically allocates the execution nodes of the copy task according to the system resource status, starts a file copy process on the execution nodes. The file copy process performs file copy operations based on the data block-level parallel transmission mechanism, verifies the data integrity during the copy process, stores the copied files in the target storage location, and generates a copy task execution report including:

[0029] Collect the resource status information of the execution nodes in the distributed file system and construct a resource status vector; set a resource weight vector, where the weight values of each dimension in the resource weight vector correspond to the importance of each dimension of resources in the resource status vector, and perform weighted calculation on the resource status vector and the resource weight vector to obtain the comprehensive load index of the execution node;

[0030] Perform feature modeling on the file to be replicated to obtain a file feature vector including file size, file access frequency, and file priority; calculate the resource requirement vector of the replication task based on the file feature vector, and the resource requirement vector includes the resource requirements of each dimension corresponding to the resource status vector;

[0031] Multiply the reciprocal of the comprehensive load index of the execution node, the reciprocal of the resource status vector and the resource requirement vector, and the file priority in the file feature vector by a preset balance factor respectively and then sum to obtain the objective function value; select the optimal execution node based on the objective function value, and the optimal execution node is the execution node with the largest objective function value;

[0032] Calculate the initial data block size according to the file size in the file feature vector and the preset parallelism, and limit the initial data block size within the range of the preset minimum data block size and maximum data block size to obtain the actual data block size; start a file replication process on the optimal execution node, and the file replication process divides the file to be replicated into multiple data blocks according to the actual data block size;

[0033] Generate a data block-level checksum containing a random salt value and a timestamp for each data block, and generate a file-level checksum based on the data block-level checksum; perform integrity verification on the data transmission process based on the data block-level checksum and the file-level checksum;

[0034] During the file replication process, count the amount of data transmitted per unit time to obtain the instantaneous transmission rate, and calculate the cumulative average transmission rate of the entire transmission process; when the instantaneous transmission rate is lower than the performance threshold determined based on the cumulative average transmission rate, adjust the preset parallelism to obtain the updated parallelism, and recalculate the data block size based on the updated parallelism; at the same time, recalculate the objective function value based on the latest resource status information and select a new optimal execution node; record the performance statistics information, resource utilization status, integrity verification result, and execution time of each stage during the transmission process based on the new optimal execution node to generate a task execution report.

[0035] During the file copying process, the amount of data transferred per unit time is counted to obtain the instantaneous transfer rate, and the cumulative average transfer rate of the entire transfer process is calculated; when the instantaneous transfer rate is lower than the performance threshold determined based on the cumulative average transfer rate, the preset parallelism is adjusted to obtain the updated parallelism, and the data block size is recalculated based on the updated parallelism, including:

[0036] During the file copying process, count the number of data blocks transferred and their corresponding data amounts within each unit statistical time window, and divide the data amount by the duration of the unit statistical time window to obtain the instantaneous transfer rate; count the total amount of data transferred from the start time of copying to the current time, and divide the total amount of data transferred by the elapsed transfer time to obtain the cumulative average transfer rate;

[0037] Multiply the cumulative average transfer rate by a preset performance coefficient to obtain the ideal transfer rate, and correct the ideal transfer rate based on the current system resource occupancy rate to obtain the performance threshold; compare the instantaneous transfer rate with the performance threshold;

[0038] When it is detected that the instantaneous transfer rate is lower than the performance threshold, calculate the difference between the performance threshold and the instantaneous transfer rate, and divide the difference by the performance threshold to obtain the performance degradation ratio; calculate the parallelism adjustment coefficient based on the performance degradation ratio;

[0039] Multiply the preset parallelism by the parallelism adjustment coefficient to obtain the updated parallelism; obtain the current resource usage status of the system, and calculate the maximum parallelism that can be supported based on the remaining available resources; limit the updated parallelism within the range of the maximum parallelism to obtain the actual parallelism;

[0040] Recalculate the data block size based on the actual parallelism, divide the remaining size of the file to be copied by the actual parallelism to obtain the new data block size; limit the new data block size within the range of the preset minimum data block size and maximum data block size to obtain the actual data block size; re-partition the data content that has not been transferred yet according to the actual data block size.

[0041] The method further includes:

[0042] Collect the system load factor and historical error rate of the distributed system, calculate the load sensitivity coefficient based on the system load factor, and add the product of the initial retry interval, the load sensitivity coefficient, and the system load factor to obtain the base retry interval; calculate the exponential backoff retry interval based on the base retry interval, the number of retries, the error rate impact factor, and the historical error rate;

[0043] Statistically analyze the historical data of task status transitions, calculate the transition frequencies between various states, and calculate the state transition probabilities based on the transition frequencies; weight the historical state transition probabilities and the state transition probabilities according to the historical weight factor to obtain an updated state transition probability matrix;

[0044] Calculate the retry priority based on the current retry count, task importance factor, and task error rate of the task. The retry priority decreases as the retry count increases, increases as the task importance factor increases, and decreases as the task error rate increases; calculate the resource reservation ratio according to the ratio of the number of failed tasks to the total number of tasks;

[0045] Collect the data verification failure rate, and adjust the retry threshold based on the data verification failure rate. The retry threshold increases as the data verification failure rate increases; collect the data inconsistency rate, and adjust the consistency check frequency based on the data inconsistency rate. The consistency check frequency increases as the data inconsistency rate increases;

[0046] When a data block transfer fails, sort the failed tasks according to the retry priority, initiate retries at the exponential backoff retry intervals, and allocate computing resources within the range of the resource reservation ratio; when a data block verification fails, determine whether to continue retrying based on the retry threshold, and perform data consistency verification at the consistency check frequency.

[0047] In a second aspect of the embodiments of the present invention, a file fuzzy copy system based on a big data file cluster is provided, including:

[0048] A first unit for obtaining a set of files to be matched, extracting features from each file in the set of files to be matched, and constructing a set of feature vectors of files to be matched. The set of feature vectors of files to be matched includes a file content feature vector, a file name feature vector, and a file metadata feature vector for each file. The file content feature vector is obtained by encoding the file content through a deep learning model, the file name feature vector is obtained by encoding the file name through a character-level encoding model, and the file metadata feature vector includes numerical representations of the file size, creation time, and modification time;

[0049] A second unit for dividing the set of feature vectors of files to be matched into multiple subtasks based on a distributed computing framework, and parallelly executing feature vector similarity calculations on multiple computing nodes of the distributed computing framework to calculate the similarity scores between each file to be matched and the files in the target file set. The similarity scores are obtained by weighted calculation of the file content feature vector similarity, the file name feature vector similarity, and the file metadata feature vector similarity, and filtering the similarity scores according to a preset similarity threshold to generate a list of files to be copied;

[0050] A third unit is configured to submit the list of files to be copied to a task scheduler of a distributed file system. The task scheduler dynamically allocates an execution node for the copy task according to the system resource status, starts a file copy process on the execution node. The file copy process performs file copy operations based on a parallel transmission mechanism at the data block level, verifies the data integrity during the copying process, stores the copied files in a target storage location, and generates a copy task execution report.

[0051] In a third aspect of the embodiments of the present invention

[0052] There is provided an electronic device, including:

[0053] a processor;

[0054] a memory for storing instructions executable by the processor;

[0055] Wherein, the processor is configured to call the instructions stored in the memory to execute the foregoing method.

[0056] In a fourth aspect of the embodiments of the present invention,

[0057] There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the foregoing method is implemented.

[0058] The beneficial effects of the present application are as follows:

[0059] 1. Improve file copy efficiency: By using a distributed computing framework to perform parallel computing of file similarity and parallel execution of file copy operations, as well as a parallel transmission mechanism at the data block level, the copy speed of large-scale file clusters is significantly improved.

[0060] 2. Enhance file matching accuracy: Using a deep learning model to extract file content features, and combining file name features and file metadata features for similarity calculation, can more accurately identify files with fuzzy matches and avoid incorrect copying.

[0061] 3. Ensure data integrity and reliability: Data integrity verification is performed during the file copy process, and the task scheduling and resource management mechanisms provided by the distributed file system ensure the reliability of file copying and the integrity of data. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 is a schematic flowchart of a method for fuzzy file copying based on a big data file cluster according to an embodiment of the present invention;

[0063] Figure 2 is a schematic structural diagram of a system for fuzzy file copying based on a big data file cluster according to an embodiment of the present invention. Detailed Implementation Manner

[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0065] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0066] Figure 1 is a schematic flowchart of a file fuzzy copy method based on a big data file cluster according to an embodiment of the present invention. As Figure 1 shown, the method includes:

[0067] S101. Obtain a set of files to be matched, extract features from each file in the set of files to be matched, and construct a set of feature vectors of files to be matched. The set of feature vectors of files to be matched includes a file content feature vector, a file name feature vector, and a file metadata feature vector for each file. The file content feature vector is obtained by encoding the file content through a deep learning model, the file name feature vector is obtained by encoding the file name through a character-level encoding model, and the file metadata feature vector includes numerical representations of file size, creation time, and modification time;

[0068] S102. Based on a distributed computing framework, divide the set of feature vectors of files to be matched into multiple subtasks, perform parallel calculation of feature vector similarity on multiple computing nodes of the distributed computing framework, calculate the similarity scores between each file to be matched and the files in the target file set. The similarity scores are obtained by weighted calculation of the similarity of file content feature vectors, the similarity of file name feature vectors, and the similarity of file metadata feature vectors. Filter the similarity scores according to a preset similarity threshold to generate a list of files to be copied;

[0069] S103. Submit the list of files to be copied to the task scheduler of the distributed file system. The task scheduler dynamically allocates the execution nodes of the copy task according to the system resource status, starts a file copy process on the execution nodes. The file copy process performs a file copy operation based on a data block-level parallel transmission mechanism, verifies the data integrity during the copy process, stores the copied files in the target storage location, and generates a copy task execution report.

[0070] In an alternative embodiment, the file content feature vector is obtained by encoding the file content through a deep learning model, including:

[0071] Perform the WordPiece tokenization algorithm on the input file content for sub-word segmentation to obtain a token sequence, unify the token sequence to a preset length through dynamic truncation or padding, perform token encoding and position encoding on each token in the token sequence, where the position encoding uses sine and cosine functions to encode the token position information, and stack the token encoding and the position encoding to obtain a token embedding sequence;

[0072] Input the token embedding sequence into a multi-layer BERT encoder. Each layer of the multi-layer BERT encoder includes a multi-head self-attention sub-layer and a feed-forward neural network sub-layer. The output of each multi-head self-attention sub-layer and the feed-forward neural network sub-layer is connected to the input through a residual connection and undergoes layer normalization processing to obtain a hidden layer feature representation of the file content;

[0073] Perform feature fusion processing on the hidden layer feature representation, map the hidden layer feature representation to a query matrix, a key-value matrix, and a value matrix respectively, calculate the output of multiple attention heads based on the query matrix, the key-value matrix, and the value matrix, splice the output of the multiple attention heads and perform a linear transformation to obtain a file content feature vector, and the file content feature vector is used to characterize the semantic features of the file content.

[0074] A method for extracting a file content feature vector, which is used to characterize the semantic features of the file content, and its specific implementation is as follows:

[0075] First, perform the WordPiece tokenization algorithm on the input file content for sub-word segmentation. For example, if the input file content is “This is a good example”, the token sequence obtained after WordPiece tokenization may be “[CLS] Thisis a good example [SEP]”. Here, “[CLS]” and “[SEP]” are special tokens, representing the start and end of the sequence respectively.

[0076] Next, perform dynamic truncation or padding on the token sequence. Assume the preset length is 10. Since the current token sequence length is less than 10, padding is required. A common padding method is to add the “[PAD]” token at the end of the sequence until the preset length is reached. The padded token sequence is “[CLS] This is a good example [SEP] [PAD] [PAD][PAD]”. If the token sequence length exceeds the preset length, the extra part is truncated.

[0077] Then, token encoding and positional encoding are performed on each token in the token sequence. Token encoding is to convert each token into a corresponding vector representation. Assuming the vocabulary size is 10,000, each token can be represented by a 10,000-dimensional one-hot vector. Positional encoding is to encode the position information of the token into the vector. For example, the positional encoding of the first token can be represented as [sin(0), cos(0), sin(2*0), cos(2*0),...], the positional encoding of the second token can be represented as [sin(1), cos(1), sin(2*1), cos(2*1),...], and so on. The token encoding and positional encoding are superimposed to obtain the token embedding sequence. For example, the embedding vector of the first token is the sum of the token encoding vector and the positional encoding vector of the first token.

[0078] After that, the token embedding sequence is input into the multi-layer BERT encoder. Assume the BERT encoder has 12 layers. Each layer of the BERT encoder contains a multi-head self-attention sub-layer and a feed-forward neural network sub-layer. In the multi-head self-attention sub-layer, the token embedding sequence is respectively mapped into a query matrix, a key-value matrix, and a value matrix. For example, for the embedding vector of the first token, it is multiplied by the query matrix, the key-value matrix, and the value matrix respectively to obtain the corresponding query vector, key-value vector, and value vector. Based on the query matrix, the key-value matrix, and the value matrix, the outputs of multiple attention heads are calculated. Each attention head calculates the attention weights between all tokens and performs a weighted sum on the value matrix according to the weights. The outputs of multiple attention heads are concatenated and linearly transformed to obtain the output of the multi-head self-attention. The output of the multi-head self-attention sub-layer is connected with the input through a residual connection and undergoes layer normalization processing. The feed-forward neural network sub-layer performs a non-linear transformation on the output of the multi-head self-attention, and its output is also connected with the input through a residual connection and undergoes layer normalization processing. Finally, the hidden layer feature representation of the file content is obtained.

[0079] Finally, feature fusion processing is performed on the hidden layer feature representation. The hidden layer feature representation is respectively mapped into a query matrix, a key-value matrix, and a value matrix. Based on the query matrix, the key-value matrix, and the value matrix, the outputs of multiple attention heads are calculated. The outputs of multiple attention heads are concatenated and linearly transformed to obtain the file content feature vector. For example, taking the output of the last layer of the BERT encoder as the hidden layer feature representation, a 768-dimensional feature vector is obtained after feature fusion processing. This feature vector is used to characterize the semantic features of the file content.

[0080] The beneficial effects of this method can be summarized in the following three aspects:

[0081] 1. Improve the feature expression ability: Through WordPiece tokenization and the BERT encoder, semantic information in the document content can be better captured, thus improving the expression ability of feature vectors. For example, compared with the traditional bag-of-words model, this method can better handle polysemous words and word order information.

[0082] 2. Enhance the model generalization ability: The BERT encoder adopts a pre-training mechanism and is trained on large-scale text data, which can learn rich language knowledge, thereby enhancing the model's generalization ability. This means that even when there is less training data, the model can still achieve good results.

[0083] 3. Simplify feature engineering: This method automates the feature extraction process without the need for complex manual feature engineering, thus reducing development costs and time. For example, there is no need to manually design feature templates or rules.

[0084] In an alternative embodiment, based on a distributed computing framework, the set of feature vectors of the files to be matched is divided into multiple subtasks, and the calculation of feature vector similarity is performed in parallel on multiple computing nodes of the distributed computing framework. Calculating the similarity score between each file to be matched and the files in the target file set includes:

[0085] Two-dimensionally partition the set of feature vectors of the files to be matched and the set of feature vectors of the target files according to a preset row block size and column block size to obtain multiple feature vector calculation blocks. Based on the feature vector calculation blocks, construct a calculation dependency graph. The vertices of the calculation dependency graph represent the feature vector calculation blocks, and the edges of the calculation dependency graph represent the data dependency relationships between the feature vector calculation blocks, and obtain the block size information of the feature vector calculation blocks;

[0086] Calculate the inter-block load balancing metric value based on the block size information of the feature vector calculation blocks. The inter-block load balancing metric value is the ratio of the largest calculation block size to the average calculation block size in the feature vector calculation blocks, and obtain the data transmission volume between the calculation nodes in the calculation dependency graph;

[0087] Construct a task scheduling optimization objective function based on the inter-block load balancing metric value and the data transmission volume; according to the optimization result of the task scheduling optimization objective function, allocate the feature vector calculation blocks to multiple computing nodes of the distributed computing framework, and establish a feature vector calculation task in each computing node. The feature vector calculation task includes the calculation of the similarity of the feature vectors of the corresponding calculation blocks;

[0088] Execute the feature vector calculation task in parallel on the multiple computing nodes. For the feature vector calculation task of each computing node, perform L2 normalization on the feature vectors in the corresponding calculation block to obtain normalized feature vectors, and calculate the cosine similarity between the normalized feature vectors using matrix multiplication to obtain the local similarity calculation result;

[0089] Construct a multi-layer tree-shaped aggregation structure based on the calculation dependency graph. The leaf nodes of the multi-layer tree-shaped aggregation structure correspond to the local similarity calculation results of the multiple computing nodes. Perform an aggregation operation on the local similarity calculation results of multiple child nodes at each layer. The aggregation operation selects the maximum similarity value between the same feature vector pairs as the aggregation result until the aggregation of all levels is completed to obtain the final similarity calculation result.

[0090] A method for calculating the similarity of file feature vectors based on a distributed computing framework, which is used to efficiently calculate the similarity between large-scale file feature vectors. The core idea of this method is to partition the set of feature vectors of the files to be matched and the set of feature vectors of the target files, and calculate the similarity in parallel under the distributed computing framework, and finally aggregate the results through a tree structure.

[0091] First, prepare the set of feature vectors of the files to be matched and the set of feature vectors of the target files. Assume there are 10,000 files to be matched, and the dimension of the feature vector of each file is 512. There are 1,000 target files, and the dimension of the feature vector of each file is also 512. Represent these two sets of feature vectors as matrices A and B respectively. The dimension of matrix A is 10,000x512, and the dimension of matrix B is 1,000x512.

[0092] Next, partition matrices A and B two-dimensionally according to the preset row block size and column block size. For example, set the row block size to 1,000 and the column block size to 128. Then matrix A will be divided into 10 row blocks and 4 column blocks, a total of 40 feature vector calculation blocks; matrix B will be divided into 1 row block and 4 column blocks, a total of 4 feature vector calculation blocks.

[0093] Then, construct a calculation dependency graph based on these feature vector calculation blocks. The vertices of the graph represent the feature vector calculation blocks, and the edges represent the data dependency relationships between the calculation blocks. Since calculating the similarity between A and B requires each row block of A to be calculated with each row block of B, there is a dependency relationship between each feature vector calculation block of A and each feature vector calculation block of B. At the same time, record the size information of each feature vector calculation block. For example, the size of a certain block of A is 1,000x128, and the size of a certain block of B is 1,000x128.

[0094] Next, calculate the inter-block load balancing metric. Calculate the size of each eigenvector calculation block, find the largest calculation block size, and calculate the average size of all calculation blocks. The inter-block load balancing metric is equal to the largest calculation block size divided by the average calculation block size. For example, if the largest calculation block size is 1000x128 and the average calculation block size is 1000x128, then the inter-block load balancing metric is 1. At the same time, obtain the data transfer volume between calculation nodes in the calculation dependency graph, which depends on the size of each calculation block and the network topology of the calculation nodes. Assume the data transfer volume is an estimated value, such as 1GB.

[0095] Construct an optimization objective function for task scheduling based on the inter-block load balancing metric and the data transfer volume. The objective of the objective function is to minimize the calculation time, which can be expressed as a weighted sum of the inter-block load balancing metric and the data transfer volume. For example, the objective function can be defined as: Objective function = 0.8 * inter-block load balancing metric + 0.2 * data transfer volume. By optimizing this objective function, an allocation scheme of eigenvector calculation blocks to calculation nodes can be obtained.

[0096] Then, allocate the eigenvector calculation blocks to multiple calculation nodes of the distributed computing framework and establish eigenvector calculation tasks on each calculation node. Each task is responsible for calculating the eigenvector similarity of the corresponding calculation block.

[0097] On each calculation node, perform L2 normalization on the eigenvectors in the allocated calculation blocks. For example, perform L2 normalization on the eigenvectors in a certain block of A and a certain block of B respectively. Then, use matrix multiplication to calculate the cosine similarity between the normalized eigenvectors to obtain the local similarity calculation result. For example, calculate the matrix product of a certain normalized block of A and a certain normalized block of B to obtain a local similarity matrix.

[0098] Finally, construct a multi-layer tree-shaped aggregation structure based on the calculation dependency graph. The leaf nodes correspond to the local similarity calculation results of each calculation node. At each layer, perform an aggregation operation on the local similarity calculation results of multiple child nodes. The aggregation operation selects the maximum similarity value between the same eigenvector pairs as the aggregation result until all levels of aggregation are completed to obtain the final similarity calculation result. For example, assume that two calculation nodes have calculated the similarity between the first block of A and the first block of B, and the similarity between the second block of A and the first block of B respectively. The aggregation operation will compare the similarity values of the same eigenvector pairs in these two results and take the larger value as the final result.

[0099] The beneficial effects of this method can be summarized in the following three aspects:

[0100] 1. Improve computational efficiency: Through distributed computing and block processing, large-scale similarity calculation tasks are decomposed into multiple small tasks for parallel execution, significantly shortening the calculation time.

[0101] 2. Optimize resource utilization: Through task scheduling optimization, the load of each computing node is balanced, avoiding situations of idle or overloaded computing resources, and improving the utilization rate of computing resources.

[0102] 3. Improve computational accuracy: Through a tree-shaped aggregation structure, the local results of each computing node can be effectively integrated, and the optimal similarity value can be selected, thereby improving the accuracy of the final calculation result.

[0103] In an alternative embodiment, a task scheduling optimization objective function is constructed based on the inter-block load balancing metric value and the data transfer volume; allocating the feature vector calculation blocks to multiple computing nodes of the distributed computing framework according to the optimization result of the task scheduling optimization objective function includes:

[0104] Construct a computational dependency adjacency matrix, and the elements of the computational dependency adjacency matrix represent the dependency relationships between calculation blocks; based on the computational dependency adjacency matrix, the block size information of the feature vector calculation blocks, and the task allocation status, construct a data transfer volume matrix, and the elements of the data transfer volume matrix are the total data transfer volume between calculation blocks with dependency relationships;

[0105] Construct a multi-objective optimization function with the inter-block load balancing metric value, the sum of the elements of the data transfer volume matrix, and the maximum load of the computing node. The multi-objective optimization function includes a load balancing weight coefficient, a communication overhead weight coefficient, and a computational load weight coefficient;

[0106] Generate an initial task allocation scheme based on the min-max criterion. Under the conditions of satisfying task integrity constraints, node capacity constraints, and dependency satisfaction constraints, select a scheme with the minimum objective function value from the neighborhood solutions of the current allocation scheme as the allocation scheme for the next iteration through iterative optimization;

[0107] When the optimization iteration converges or reaches a preset number of iterations, send the optimal allocation scheme to each computing node through the distributed scheduler. The optimal allocation scheme includes the mapping relationship between calculation blocks and computing nodes and task execution time constraints.

[0108] A distributed task scheduling optimization method based on load balancing and data transfer volume aims to minimize the calculation time and improve resource utilization. The core idea of this method is to comprehensively consider the dependency relationships between calculation blocks, the data transfer volume, and the load conditions of computing nodes, and reasonably allocate computing tasks to each computing node.

[0109] First, construct a computational dependency adjacency matrix. Traverse all computational blocks. If the output of block A is the input of block B, then the element value at the intersection of row A and column B in the adjacency matrix is 1; otherwise, it is 0.

[0110] Next, construct a data transfer volume matrix based on the computational dependency adjacency matrix, the block size information of the eigenvector calculation blocks, and the task allocation status. Traverse the adjacency matrix. If the element value at the intersection of row A and column B is 1, indicating a dependency relationship between block A and block B, then calculate the output data volume of block A, which is the data volume that needs to be transferred between block A and block B. Fill this data volume into the intersection of row A and column B in the data transfer volume matrix.

[0111] Then, construct a multi-objective optimization function. The objectives of this function are to minimize the load difference between blocks, minimize the total data transfer volume, and ensure that the load of the computing nodes does not exceed the maximum value. The function contains three weight coefficients: the load balancing weight coefficient, the communication overhead weight coefficient, and the computing load weight coefficient. These coefficients are used to adjust the importance of the three objectives. For example, if more attention is paid to load balancing, the load balancing weight coefficient can be set larger.

[0112] Subsequently, generate an initial task allocation scheme based on the min-max criterion. This scheme needs to satisfy three constraint conditions: the task integrity constraint (all tasks must be allocated to computing nodes), the node capacity constraint (the load of each computing node cannot exceed its maximum capacity), and the dependency satisfaction constraint (if block A depends on block B, then block A must be executed after block B, and they need to be allocated to the same computing node or different computing nodes with network connections).

[0113] Next, improve the task allocation scheme through iterative optimization. In each iteration, select the scheme with the minimum objective function value from the neighborhood solutions of the current allocation scheme as the allocation scheme for the next iteration. Neighborhood solutions refer to the schemes obtained by making minor changes to the current scheme, such as moving a computational block from one computing node to another.

[0114] Finally, when the optimization iteration converges or reaches the preset number of iterations, send the optimal allocation scheme to each computing node through the distributed scheduler. The optimal allocation scheme includes the mapping relationship between computational blocks and computing nodes and the task execution time constraints. For example, the scheme may specify that block 1 is allocated to node 1 with an execution time of 10 seconds; blocks 2 and 3 are allocated to node 2 with execution times of 5 seconds and 5 seconds respectively; block 4 is allocated to node 1 with an execution time of 5 seconds.

[0115] The beneficial effects of this method can be summarized in the following three aspects:

[0116] 1. Improved cluster resource utilization: Through load balancing, the situation where some computing nodes are overloaded while others are idle is avoided, thus making more full use of the cluster resources.

[0117] 2. Reduced task completion time: By optimizing the data transfer volume, the data exchange time between computing nodes is reduced, thus shortening the completion time of the entire task.

[0118] 3. Enhanced system scalability: This method can adapt to clusters of different scales and can adjust parameters according to the actual situation to achieve the best scheduling effect.

[0119] In an optional implementation manner, the list of files to be replicated is submitted to the task scheduler of the distributed file system. The task scheduler dynamically allocates the execution nodes of the replication tasks according to the system resource status, starts a file replication process on the execution nodes. The file replication process performs file replication operations based on a parallel transmission mechanism at the data block level, verifies the data integrity during the replication process, stores the replicated files in the target storage location, and generates a replication task execution report including:

[0120] Collect the resource status information of the execution nodes in the distributed file system and construct a resource status vector; set a resource weight vector, where the weight values of each dimension in the resource weight vector correspond to the importance of the resources of each dimension in the resource status vector, and perform weighted calculation on the resource status vector and the resource weight vector to obtain the comprehensive load index of the execution nodes;

[0121] Perform feature modeling on the files to be replicated to obtain a file feature vector including file size, file access frequency, and file priority; calculate a resource requirement vector for the replication task based on the file feature vector, and the resource requirement vector includes the resource demand quantities of each dimension corresponding to the resource status vector;

[0122] Multiply the reciprocal of the comprehensive load index of the execution nodes, the reciprocal of the resource status vector and the resource requirement vector, and the file priority in the file feature vector by a preset balance factor respectively and then sum to obtain a target function value; select the optimal execution node based on the target function value, and the optimal execution node is the execution node with the largest target function value;

[0123] Calculate the initial data block size according to the file size in the file feature vector and the preset parallelism, and limit the initial data block size within the range of the preset minimum data block size and maximum data block size to obtain the actual data block size; start a file replication process on the optimal execution node, and the file replication process divides the file to be replicated into multiple data blocks according to the actual data block size;

[0124] Generate a block - level checksum for each of the said data blocks, which includes a random salt value and a timestamp, and generate a file - level checksum based on the block - level checksum; perform integrity verification on the data transmission process based on the block - level checksum and the file - level checksum;

[0125] During the file copying process, count the amount of data transferred per unit time to obtain the instantaneous transfer rate, and calculate the cumulative average transfer rate for the entire transfer process; when the instantaneous transfer rate is lower than the performance threshold determined based on the cumulative average transfer rate, adjust the preset parallelism to obtain the updated parallelism, recalculate the data block size based on the updated parallelism; at the same time, recalculate the objective function value based on the latest resource status information and select a new optimal execution node; record the performance statistics, resource utilization status, integrity verification results, and execution time of each stage during the transfer process based on the new optimal execution node to generate a task execution report.

[0126] First, collect the resource status information of each execution node in the distributed file system. This information includes CPU usage rate, memory occupancy rate, network bandwidth usage rate, disk I / O rate, etc. For example, the CPU usage rate of node A is 20%, the memory occupancy rate is 30%, the network bandwidth usage rate is 10%, and the disk I / O rate is 40MB / s. Construct these information into a resource status vector. For example, the resource status vector of node A is (20, 30, 10, 40).

[0127] Then, set a resource weight vector to represent the importance of different resource dimensions in the comprehensive load evaluation. For example, if it is considered that CPU and memory are more important than network bandwidth and disk I / O, the resource weight vector can be set as (0.4, 0.4, 0.1, 0.1). Multiply the corresponding dimensions of the resource status vector and the resource weight vector and sum them to obtain the comprehensive load index of the execution node. For example, the comprehensive load index of node A is 20 * 0.4 + 30 * 0.4 + 10 * 0.1 + 40 * 0.1 = 23.

[0128] Next, perform feature modeling on the file to be copied. Extract feature information such as file size, file access frequency, and file priority, and construct a file feature vector. For example, if the size of file F1 is 1GB, the access frequency is 10 times per day, and the priority is high, then the file feature vector of file F1 is (1GB, 10, high).

[0129] Calculate the resource requirement vector for the replication task based on the file feature vector. Each dimension in the resource requirement vector corresponds to a dimension in the resource status vector, representing the corresponding resource amount required to replicate the file. For example, to replicate file F1, it requires a CPU usage rate of 20%, a memory occupancy rate of 30%, a network bandwidth of 10 Mbps, and a disk I / O rate of 50 MB / s. Then the resource requirement vector is (20, 30, 10, 50).

[0130] To select the optimal execution node, it is necessary to calculate the objective function value of each execution node. The objective function value is obtained by summing the reciprocal of the comprehensive load index of the execution node, the product of the corresponding dimensions of the resource status vector and the reciprocal of the resource requirement vector, and the sum of the product of the file priority and the preset balance factor respectively. Assume the balance factors are 0.5, 0.3, 0.2 respectively, the reciprocal of the comprehensive load index of node A is 1 / 23, the sum of the product of the corresponding dimensions of the resource status vector and the reciprocal of the resource requirement vector is (1 / 20 * 20) + (1 / 30 * 30) + (1 / 10 * 10) + (1 / 40 * 50) = 3.25, and the file priority is high. Assume the value corresponding to "high" is 3. Then the objective function value of node A is 0.5 * (1 / 23) + 0.3 * 3.25 + 0.2 * 3 = 1.59. Select the execution node with the largest objective function value as the optimal execution node.

[0131] After determining the optimal execution node, calculate the initial data block size according to the file size and the preset parallelism. For example, if the size of file F1 is 1 GB and the preset parallelism is 4, then the initial data block size is 256 MB. Limit the initial data block size within the preset minimum and maximum data block sizes. For example, the minimum data block size is 64 MB and the maximum data block size is 512 MB, then the actual data block size is 256 MB.

[0132] Start the file replication process on the optimal execution node, and divide the file into multiple data blocks according to the actual data block size. Generate a data block - level checksum containing a random salt value and a timestamp for each data block, and generate a file - level checksum based on the data block - level checksums. For example, the checksum of data block 1 is "ABCD1234567890", and the file - level checksum is the combination of all data block checksums "ABCD1234567890EFGH9876543210". During the data transmission process, use the data block - level checksum and the file - level checksum for integrity verification.

[0133] During the file copying process, count the amount of data transferred per unit time to obtain the instantaneous transfer rate, and calculate the cumulative average transfer rate for the entire transfer process. If the instantaneous transfer rate is lower than the performance threshold determined based on the cumulative average transfer rate, adjust the preset parallelism, recalculate the data block size, and reselect the optimal execution node based on the latest resource status information. For example, if the instantaneous transfer rate is lower than 80% of the cumulative average transfer rate, adjust the parallelism from 4 to 8, recalculate the data block size, and reselect the optimal execution node.

[0134] Finally, generate a task execution report to record performance statistics, resource utilization, integrity check results, execution times of each stage, and other information during the transfer process.

[0135] The beneficial effects of this method can be summarized in the following three aspects:

[0136] 1. Improve file copying efficiency: Through the parallel transfer mechanism at the data block level and dynamic adjustment of parallelism, make full use of system resources to improve file copying efficiency.

[0137] 2. Ensure data integrity: Through the checksum mechanism at the data block level and file level, ensure the integrity during data transfer and avoid data corruption or loss.

[0138] 3. Optimize resource utilization: Through dynamic resource allocation and load balancing strategies, allocate the copying task to the most suitable execution node, optimize resource utilization, and avoid resource waste and bottlenecks.

[0139] In an alternative embodiment, during the file copying process, count the amount of data transferred per unit time to obtain the instantaneous transfer rate, and calculate the cumulative average transfer rate for the entire transfer process; when the instantaneous transfer rate is lower than the performance threshold determined based on the cumulative average transfer rate, adjust the preset parallelism to obtain the updated parallelism, and recalculating the data block size based on the updated parallelism includes:

[0140] During the file copying process, count the number of data blocks transferred and their corresponding data amounts within each unit statistical time window, divide the data amount by the duration of the unit statistical time window to obtain the instantaneous transfer rate; count the total amount of data transferred from the start time of copying to the current time, and divide the total amount of completed transfer by the elapsed transfer time to obtain the cumulative average transfer rate;

[0141] Multiply the cumulative average transfer rate by a preset performance coefficient to obtain the ideal transfer rate, correct the ideal transfer rate based on the current system resource occupancy rate to obtain the performance threshold; compare the instantaneous transfer rate with the performance threshold;

[0142] When it is detected that the instantaneous transmission rate is lower than the performance threshold, calculate the difference between the performance threshold and the instantaneous transmission rate, and divide the difference by the performance threshold to obtain the performance degradation ratio; calculate the parallelism adjustment coefficient based on the performance degradation ratio;

[0143] Multiply the preset parallelism by the parallelism adjustment coefficient to obtain the updated parallelism; obtain the current resource usage status of the system, and calculate the maximum parallelism that can be supported based on the remaining available resources; limit the updated parallelism within the range of the maximum parallelism to obtain the actual parallelism;

[0144] Recalculate the data block size based on the actual parallelism, and divide the remaining size of the file to be copied by the actual parallelism to obtain the new data block size; limit the new data block size within the range of the preset minimum data block size and maximum data block size to obtain the actual data block size; re-partition the data content that has not been completed in transmission according to the actual data block size.

[0145] First, initialize the file copy parameters. Set the preset parallelism, for example, initially set to 4, which means that initially the file is divided into 4 data blocks for simultaneous transmission. Set the performance coefficient, for example, 0.8, for calculating the ideal transmission rate. Set the minimum data block size and the maximum data block size, for example, 1MB and 16MB respectively, to limit the fluctuation range of the data block size. Set the unit statistical time window, for example, 1 second, for calculating the instantaneous transmission rate.

[0146] Next, start file copying and perform real-time performance monitoring. At the end of each unit statistical time window, for example, every 1 second, count the number of data blocks completed in transmission within this time window and their corresponding data volumes. For example, within the first 1 second, 2 data blocks are completed in transmission, and the data volume is 2MB. Divide this data volume by the duration of the unit statistical time window, that is, 2MB / 1 second = 2MB / s, to obtain the instantaneous transmission rate. At the same time, count the total amount of data completed in transmission from the start time of copying to the current time and the elapsed transmission duration. For example, assuming that 1 second has elapsed and the total amount of data completed in transmission is 2MB, the cumulative average transmission rate is 2MB / 1 second = 2MB / s.

[0147] Then, calculate the performance threshold. Multiply the cumulative average transmission rate by the preset performance coefficient to obtain the ideal transmission rate. For example, 2MB / s * 0.8 = 1.6MB / s. Consider the current system resource occupancy rate to correct the ideal transmission rate to obtain the performance threshold. For example, assuming that the current CPU usage rate is 50% and the memory usage rate is 30%, according to the preset rules for the impact of resource occupancy rate, reduce the ideal transmission rate by 20%, and obtain the performance threshold 1.6MB / s * (1 - 0.2) = 1.28MB / s.

[0148] Next, compare the instantaneous transmission rate with the performance threshold. If the instantaneous transmission rate is lower than the performance threshold, the parallelism and data block size need to be adjusted. For example, assume that in the second 1-second period, the instantaneous transmission rate drops to 1 MB / s, which is lower than the performance threshold of 1.28 MB / s.

[0149] If the instantaneous transmission rate is lower than the performance threshold, calculate the parallelism adjustment coefficient. Calculate the difference between the performance threshold and the instantaneous transmission rate, and divide this difference by the performance threshold to obtain the performance degradation ratio. For example, (1.28 MB / s - 1 MB / s) / 1.28 MB / s = 0.22, and the performance degradation ratio is 22%. According to the preset correspondence between the performance degradation ratio and the parallelism adjustment coefficient, for example, when the performance degradation is 20% - 30%, the parallelism adjustment coefficient is 1.2. Then multiply the preset parallelism by the parallelism adjustment coefficient to obtain the updated parallelism. For example, 4 * 1.2 = 4.8.

[0150] Obtain the current resource usage status of the system, such as CPU usage rate, memory usage rate, etc., and calculate the maximum parallelism that can be supported based on the remaining available resources. For example, assume that the current system resources allow a maximum parallelism of 6. Then limit the updated parallelism within the range of the maximum parallelism to obtain the actual parallelism, which is 4.8 here. Since the parallelism must be an integer, it is rounded down to 4.

[0151] Recalculate the data block size based on the actual parallelism. Divide the remaining size of the file to be copied by the actual parallelism to obtain the new data block size. For example, assume that the remaining size of the file to be copied is 100 MB and the actual parallelism is 4. Then the new data block size is 100 MB / 4 = 25 MB. Limit the new data block size within the range of the preset minimum data block size and maximum data block size to obtain the actual data block size. For example, since the maximum data block size is 16 MB, the actual data block size is 16 MB. Redivide the data content that has not been completed in the transmission according to the actual data block size. Subsequent data transmissions will be carried out according to the new data block size and parallelism. Repeat the above steps until the file copy is completed.

[0152] This application can achieve:

[0153] 1. Improve the copying efficiency: This method can dynamically adjust the parallelism and data block size according to the network conditions and system resources, avoiding the low transmission efficiency caused by fixed parameter settings, thereby improving the overall copying efficiency.

[0154] 2. Resource self-adaptation: This method can dynamically adjust the parallelism according to the system resource usage, avoiding system jams caused by excessive resource occupation and ensuring the normal operation of other applications.

[0155] 3. Stable and reliable: This method can adjust according to the real-time transmission rate, avoiding transmission interruptions or failures caused by network fluctuations, and improving the stability and reliability of file copying.

[0156] In an optional implementation, the method further includes:

[0157] Collect the system load factor and historical error rate of the distributed system, calculate the load sensitivity coefficient based on the system load factor, and add the product of the initial retry interval and the load sensitivity coefficient and the system load factor to obtain the base retry interval; calculate the exponential backoff retry interval based on the base retry interval, the number of retries, the error rate impact factor, and the historical error rate;

[0158] Statistically analyze the historical data of task state transitions, calculate the transition frequencies between each state, and calculate the state transition probability based on the transition frequencies; weight the historical state transition probability and the state transition probability according to the historical weight factor to obtain an updated state transition probability matrix;

[0159] Calculate the retry priority based on the current number of retries, task importance factor, and task error rate of the task. The retry priority decreases as the number of retries increases, increases as the task importance factor increases, and decreases as the task error rate increases; calculate the resource reservation ratio according to the ratio of the number of failed tasks to the total number of tasks;

[0160] Collect the data verification failure rate, and adjust the retry threshold based on the data verification failure rate. The retry threshold increases as the data verification failure rate increases; collect the data inconsistency rate, and adjust the consistency check frequency based on the data inconsistency rate. The consistency check frequency increases as the data inconsistency rate increases;

[0161] When a data block transmission fails, sort the failed tasks according to the retry priority, initiate a retry according to the exponential backoff retry interval, and allocate computing resources within the resource reservation ratio; when a data block verification fails, determine whether to continue the retry based on the retry threshold, and perform data consistency verification according to the consistency check frequency.

[0162] A method and device for processing data in a distributed system, aiming to improve data processing efficiency and reliability.

[0163] First, obtain the system load factor and historical error rate of the distributed system. The system load factor can be comprehensively calculated through indicators such as CPU utilization, memory usage, and network traffic. For example, if the current CPU utilization is 70%, the memory usage is 80%, and the network traffic is 50 Mbps, the system load factor is calculated to be 0.75 according to the preset weights. The historical error rate refers to the proportion of task execution failures in the past period. For example, if a total of 1000 tasks were executed in the past hour and 10 tasks failed, the historical error rate is 1%.

[0164] Next, calculate the load sensitivity coefficient. The load sensitivity coefficient is used to dynamically adjust the retry interval according to the system load. The calculation method of the load sensitivity coefficient can be adjusted according to the actual situation. For example, it can be calculated based on the difference between the system load factor and the preset threshold. Assuming the preset threshold is 0.8, the load sensitivity coefficient is 0.8 - 0.75 = 0.05. Add the product of the initial retry interval, the load sensitivity coefficient, and the system load factor to obtain the base retry interval. Assuming the initial retry interval is 1 second, the base retry interval is 1 + 0.05 * 0.75 = 1.0375 seconds.

[0165] Then, calculate the exponential backoff retry interval based on the base retry interval, the number of retries, the error rate impact factor, and the historical error rate. The error rate impact factor is used to adjust the retry interval according to the historical error rate. For example, it can be calculated based on the difference between the historical error rate and the preset threshold. Assuming the preset threshold is 2%, the error rate impact factor is 2% - 1% = 1%. Assuming the current number of retries is 2, the exponential backoff retry interval is 1.0375 * (1 + 1% * 2) = 1.058 seconds.

[0166] At the same time, count the historical data of task state transitions and calculate the transition frequencies between each state. For example, the task states include "to be executed", "executing", "success", and "failure", and count the number of transitions between each state in the past period. For example, the number of times the "to be executed" state transitions to the "executing" state is 1000 times, the number of times the "executing" state transitions to the "success" state is 900 times, and the number of times the "executing" state transitions to the "failure" state is 100 times.

[0167] Calculate the state transition probability based on the conversion frequency. For example, the probability of the "to be executed" state transitioning to the "executing" state is 1000 / (1000 + 0) = 100%, the probability of the "executing" state transitioning to the "successful" state is 900 / (900 + 100) = 90%, and the probability of the "executing" state transitioning to the "failed" state is 100 / (900 + 100) = 10%. Weight the historical state transition probability and the currently calculated state transition probability according to the historical weight factor to obtain the updated state transition probability matrix. Assume the historical weight factor is 0.5, then the probability of the updated "executing" state transitioning to the "successful" state is (0.5 * historical probability) + (0.5 * 90%).

[0168] In addition, calculate the retry priority based on the current retry count of the task, the task importance factor, and the task error rate. The task importance factor is used to distinguish the priorities of different tasks. For example, the weight of an important task can be set to 1.5, and the weight of an ordinary task can be set to 1. Assume the current retry count is 2, the task importance factor is 1, and the task error rate is 1%, then the retry priority can be calculated according to a preset rule. For example, retry priority = 1 / (2 * 1 * 1%) = 50. The retry priority decreases as the retry count increases, increases as the task importance factor increases, and decreases as the task error rate increases. Calculate the resource reservation ratio based on the ratio of the number of failed tasks to the total number of tasks. For example, if there are currently 10 failed tasks and the total number of tasks is 1000, then the resource reservation ratio is 1%.

[0169] Meanwhile, collect the data verification failure rate. For example, within the past hour, a total of 1000 data blocks were verified, and 10 of them failed the verification, so the data verification failure rate is 1%. Adjust the retry threshold based on the data verification failure rate. Assume the initial retry threshold is 3, then the adjusted retry threshold is 3 + 1% = 3.01. The retry threshold increases as the data verification failure rate increases. Collect the data inconsistency rate. For example, within the past hour, a total of 1000 data consistency checks were performed, and 10 of them found data inconsistencies, so the data inconsistency rate is 1%. Adjust the data consistency check frequency based on the data inconsistency rate. Assume the initial data consistency check frequency is once per minute, then the adjusted data consistency check frequency is once every 54 seconds. The data consistency check frequency increases as the data inconsistency rate increases.

[0170] When a data block transmission fails, sort the failed tasks according to the retry priority, initiate retries at exponential backoff retry intervals, and allocate computing resources within the resource reservation ratio range. When a data block verification fails, determine whether to continue retrying based on the retry threshold, and perform data consistency verification according to the data consistency check frequency.

[0171] This application can achieve the following:

[0172] 1. Improve data processing efficiency: Dynamically adjust the retry interval according to the system load to avoid frequent retries during high system load, reduce the system burden, and improve data processing efficiency. Reasonably allocate computing resources to avoid resource waste and further improve data processing efficiency.

[0173] 2. Enhance data processing reliability: Predict the task execution result based on the task state transition probability, take measures in advance to avoid task failures. Prioritize important tasks to ensure the stable operation of critical services. Data verification and consistency check mechanisms effectively guarantee data quality and enhance data processing reliability.

[0174] 3. Make the system more flexible: Dynamically adjust the retry threshold and consistency check frequency according to the data verification failure rate and data inconsistency rate, enabling the system to adapt to different data quality conditions, and improving the flexibility and robustness of the system.

[0175] Figure 2 FIG. [FIG. NUMBER] is a schematic structural diagram of a file fuzzy copy system based on a big data file cluster according to an embodiment of the present invention. As Figure 2 shown, the system includes:

[0176] A first unit for obtaining a set of files to be matched, extracting features from each file in the set of files to be matched, and constructing a set of feature vectors of files to be matched. The set of feature vectors of files to be matched includes a file content feature vector, a file name feature vector, and a file metadata feature vector for each file. The file content feature vector is obtained by encoding the file content through a deep learning model, the file name feature vector is obtained by encoding the file name through a character-level encoding model, and the file metadata feature vector includes numerical representations of file size, creation time, and modification time;

[0177] A second unit for dividing the set of feature vectors of files to be matched into multiple subtasks based on a distributed computing framework, parallelly executing feature vector similarity calculations on multiple computing nodes of the distributed computing framework, calculating the similarity scores between each file to be matched and the files in the target file set. The similarity scores are obtained by weighted calculation of the file content feature vector similarity, the file name feature vector similarity, and the file metadata feature vector similarity, and filtering the similarity scores according to a preset similarity threshold to generate a list of files to be copied;

[0178] It should be noted that the specific FIG. NUMBER in needs to be filled in according to the actual figure number in the original text.A third unit is used to submit the list of files to be copied to a task scheduler of a distributed file system. The task scheduler dynamically allocates an execution node for the copy task according to the system resource status, starts a file copying process on the execution node. The file copying process performs file copying operations based on a parallel transmission mechanism at the data block level, verifies the data integrity during the copying process, stores the copied files at a target storage location, and generates a copy task execution report.

[0179] In a third aspect of the embodiments of the present invention,

[0180] An electronic device is provided, including:

[0181] A processor;

[0182] A memory for storing instructions executable by the processor;

[0183] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0184] In a fourth aspect of the embodiments of the present invention,

[0185] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0186] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium on which computer-readable program instructions for executing various aspects of the present invention are uploaded.

[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A file fuzzy copy method based on a big data file cluster, characterized in that, Including: Obtain a set of files to be matched, extract features from each file in the set of files to be matched, and construct a set of feature vectors of files to be matched. The set of feature vectors of files to be matched includes a file content feature vector, a file name feature vector, and a file metadata feature vector for each file. The file content feature vector is obtained by encoding the file content through a deep learning model, the file name feature vector is obtained by encoding the file name through a character-level encoding model, and the file metadata feature vector includes numerical representations of file size, creation time, and modification time; Based on a distributed computing framework, divide the set of feature vectors of files to be matched into multiple subtasks, perform feature vector similarity calculations in parallel on multiple computing nodes of the distributed computing framework, calculate the similarity scores between each file to be matched and the files in the target file set. The similarity scores are obtained by performing weighted calculations on the similarities of file content feature vectors, file name feature vectors, and file metadata feature vectors, and filter the similarity scores according to a preset similarity threshold to generate a list of files to be copied; Submit the list of files to be copied to the task scheduler of the distributed file system. The task scheduler dynamically allocates the execution nodes of the copy task according to the system resource status, starts a file copy process on the execution nodes. The file copy process performs file copy operations based on a data block-level parallel transmission mechanism, verifies the data integrity during the copying process, stores the copied files in the target storage location, and generates a copy task execution report.

2. The method according to claim 1, wherein The file content feature vector is obtained by encoding the file content through a deep learning model, including: Perform WordPiece tokenization algorithm on the input file content to obtain a sequence of tokens, unify the sequence of tokens to a preset length by dynamic truncation or padding, perform token encoding and position encoding on each token in the sequence of tokens. The position encoding uses sine and cosine functions to encode the token position information, and superimpose the token encoding and the position encoding to obtain a sequence of token embeddings; Input the sequence of token embeddings into a multi-layer BERT encoder. Each layer of the multi-layer BERT encoder includes a multi-head self-attention sublayer and a feed-forward neural network sublayer. The outputs of each multi-head self-attention sublayer and the feed-forward neural network sublayer are connected with the input by residual connection and are processed by layer normalization to obtain a hidden layer feature representation of the file content; Perform feature fusion processing on the hidden layer feature representation, map the hidden layer feature representation into a query matrix, a key-value matrix, and a value matrix respectively, calculate the outputs of multiple attention heads based on the query matrix, the key-value matrix, and the value matrix, splice the outputs of the multiple attention heads and perform a linear transformation to obtain a file content feature vector, and the file content feature vector is used to represent the semantic features of the file content.

3. The method according to claim 1, wherein Based on a distributed computing framework, divide the set of feature vectors of the file to be matched into multiple subtasks, and perform parallel calculation of the similarity of feature vectors on multiple computing nodes of the distributed computing framework. Calculating the similarity score between each file to be matched and the files in the target file set includes: Two-dimensionally partition the set of feature vectors of the file to be matched and the set of feature vectors of the target file according to a preset row block size and column block size to obtain multiple feature vector calculation blocks. Based on the feature vector calculation blocks, construct a calculation dependency graph. The vertices of the calculation dependency graph represent the feature vector calculation blocks, and the edges of the calculation dependency graph represent the data dependency relationships between the feature vector calculation blocks, and obtain the block size information of the feature vector calculation blocks; Calculate the inter-block load balancing metric value based on the block size information of the feature vector calculation blocks. The inter-block load balancing metric value is the ratio of the maximum calculation block size to the average calculation block size in the feature vector calculation blocks, and obtain the data transmission volume between the computing nodes in the calculation dependency graph; Construct a task scheduling optimization objective function based on the inter-block load balancing metric value and the data transmission volume; according to the optimization result of the task scheduling optimization objective function, allocate the feature vector calculation blocks to multiple computing nodes of the distributed computing framework, and establish a feature vector calculation task in each of the computing nodes. The feature vector calculation task includes the calculation of the similarity of the feature vectors of the corresponding calculation block; Parallelly execute the feature vector calculation tasks on the multiple computing nodes. For the feature vector calculation tasks of each computing node, perform L2 normalization processing on the feature vectors in the corresponding calculation block to obtain normalized feature vectors, and calculate the cosine similarity between the normalized feature vectors by matrix multiplication to obtain the local similarity calculation result; Construct a multi-layer tree-shaped aggregation structure based on the calculation dependency graph. The leaf nodes of the multi-layer tree-shaped aggregation structure correspond to the local similarity calculation results of the multiple computing nodes. Perform an aggregation operation on the local similarity calculation results of multiple child nodes at each layer. The aggregation operation selects the maximum similarity value between the same feature vector pairs as the aggregation result until the aggregation of all levels is completed to obtain the final similarity calculation result.

4. The method according to claim 3, wherein Construct a task scheduling optimization objective function based on the inter-block load balancing metric value and the data transmission volume; according to the optimization result of the task scheduling optimization objective function, allocating the feature vector calculation blocks to multiple computing nodes of the distributed computing framework includes: Construct a calculation dependency adjacency matrix, and the elements of the calculation dependency adjacency matrix represent the dependency relationships between the calculation blocks; based on the calculation dependency adjacency matrix, the block size information of the feature vector calculation blocks and the task allocation status, construct a data transmission volume matrix, and the elements of the data transmission volume matrix are the total data transmission volume between the calculation blocks with dependency relationships; Construct a multi-objective optimization function by combining the inter-block load balancing metric value, the sum of the elements of the data transmission volume matrix and the maximum load of the computing nodes. The multi-objective optimization function includes a load balancing weight coefficient, a communication overhead weight coefficient and a computing load weight coefficient; Generate an initial task allocation plan based on the min-max criterion. Under the conditions of meeting the task integrity constraint, node capacity constraint, and dependency satisfaction constraint, select the plan with the minimum objective function value from the neighborhood solutions of the current allocation plan through iterative optimization as the allocation plan for the next iteration; When the optimization iteration converges or reaches the preset number of iterations, send the optimal allocation plan to each computing node through the distributed scheduler. The optimal allocation plan includes the mapping relationship between computing blocks and computing nodes and the task execution time constraint.

5. The method according to claim 1, wherein Submit the list of files to be replicated to the task scheduler of the distributed file system. The task scheduler dynamically allocates the execution nodes of the replication tasks according to the system resource status, starts a file replication process on the execution nodes. The file replication process performs file replication operations based on the parallel transmission mechanism at the data block level, verifies the data integrity during the replication process, stores the replicated files in the target storage location, and generates a replication task execution report including: Collect the resource status information of the execution nodes in the distributed file system and construct a resource status vector; set a resource weight vector. The weight values of each dimension in the resource weight vector correspond to the importance of each dimension of resources in the resource status vector. Perform weighted calculation on the resource status vector and the resource weight vector to obtain the comprehensive load index of the execution node; Perform feature modeling on the file to be replicated to obtain a file feature vector including file size, file access frequency, and file priority; calculate the resource requirement vector of the replication task based on the file feature vector. The resource requirement vector includes the resource demand quantities of each dimension corresponding to the resource status vector; Sum the product of the reciprocal of the comprehensive load index of the execution node, the reciprocal of the resource status vector and the resource requirement vector, and the file priority in the file feature vector multiplied by the preset balance factor respectively to obtain the objective function value; select the optimal execution node based on the objective function value. The optimal execution node is the execution node with the largest objective function value; Calculate the initial data block size according to the file size in the file feature vector and the preset parallelism, and limit the initial data block size within the range of the preset minimum data block size and maximum data block size to obtain the actual data block size; start a file replication process on the optimal execution node. The file replication process divides the file to be replicated into multiple data blocks according to the actual data block size; Generate a data block-level checksum containing a random salt value and a timestamp for each data block, and generate a file-level checksum based on the data block-level checksum; perform integrity verification on the data transmission process based on the data block-level checksum and the file-level checksum; During the file copying process, the amount of data transferred per unit time is counted to obtain the instantaneous transmission rate, and the cumulative average transmission rate of the entire transmission process is calculated; when the instantaneous transmission rate is lower than the performance threshold determined based on the cumulative average transmission rate, the preset parallelism is adjusted to obtain the updated parallelism, and the data block size is recalculated based on the updated parallelism; at the same time, the objective function value is recalculated based on the latest resource status information and a new optimal execution node is selected; based on the new optimal execution node, performance statistics, resource utilization status, integrity verification results, and execution time of each stage during the transmission process are recorded to generate a task execution report.

6. The method according to claim 5, wherein During the file copying process, the amount of data transferred per unit time is counted to obtain the instantaneous transmission rate, and the cumulative average transmission rate of the entire transmission process is calculated; when the instantaneous transmission rate is lower than the performance threshold determined based on the cumulative average transmission rate, the preset parallelism is adjusted to obtain the updated parallelism, and recalculating the data block size based on the updated parallelism includes: During the file copying process, the number of data blocks transferred and their corresponding data amounts within each unit statistical time window are counted, and the data amount is divided by the duration of the unit statistical time window to obtain the instantaneous transmission rate; the total amount of data transferred from the start time of copying to the current time is counted, and the total amount of data transferred is divided by the elapsed transmission time to obtain the cumulative average transmission rate; The cumulative average transmission rate is multiplied by a preset performance coefficient to obtain the ideal transmission rate, and the ideal transmission rate is corrected based on the current system resource occupancy rate to obtain the performance threshold; the instantaneous transmission rate is compared with the performance threshold; When it is detected that the instantaneous transmission rate is lower than the performance threshold, the difference between the performance threshold and the instantaneous transmission rate is calculated, and the difference is divided by the performance threshold to obtain the performance degradation ratio; the parallelism adjustment coefficient is calculated based on the performance degradation ratio; The preset parallelism is multiplied by the parallelism adjustment coefficient to obtain the updated parallelism; the current resource usage status of the system is obtained, and the maximum parallelism that can be supported is calculated based on the remaining available resources; the updated parallelism is limited within the range of the maximum parallelism to obtain the actual parallelism; The data block size is recalculated based on the actual parallelism, and the remaining size of the file to be copied is divided by the actual parallelism to obtain the new data block size; the new data block size is limited within the range of the preset minimum data block size and maximum data block size to obtain the actual data block size; the data content that has not been transferred is repartitioned according to the actual data block size.

7. The method according to claim 1, characterized in that, The method further includes: Collect the system load factor and historical error rate of the distributed system, calculate the load sensitivity coefficient based on the system load factor, and add the product of the initial retry interval, the load sensitivity coefficient, and the system load factor to obtain the base retry interval; calculate the exponential backoff retry interval based on the base retry interval, the number of retries, the error rate impact factor, and the historical error rate. Statistically analyze the historical data of task status transitions, calculate the transition frequencies between states, and calculate the state transition probabilities based on the transition frequencies; weight the historical state transition probabilities and the state transition probabilities according to the historical weight factor to obtain an updated state transition probability matrix; Calculate the retry priority based on the current retry count, task importance factor, and task error rate of the task. The retry priority decreases as the retry count increases, increases as the task importance factor increases, and decreases as the task error rate increases; calculate the resource reservation ratio according to the ratio of the number of failed tasks to the total number of tasks; Collect the data verification failure rate, and adjust the retry threshold based on the data verification failure rate. The retry threshold increases as the data verification failure rate increases; collect the data inconsistency rate, and adjust the consistency check frequency based on the data inconsistency rate. The consistency check frequency increases as the data inconsistency rate increases; When the data block transmission fails, sort the failed tasks according to the retry priority, initiate retries at the exponential backoff retry intervals, and allocate computing resources within the resource reservation ratio; when the data block verification fails, determine whether to continue retrying based on the retry threshold, and perform data consistency verification according to the consistency check frequency.

8. A file fuzzy copy system based on a big data file cluster for implementing the method according to any one of the preceding claims 1-7, characterized in that, Comprising: A first unit configured to obtain a set of files to be matched, extract features from each file in the set of files to be matched, and construct a set of feature vectors of files to be matched. The set of feature vectors of files to be matched includes the file content feature vectors, file name feature vectors, and file metadata feature vectors of each file. The file content feature vectors are obtained by encoding the file content through a deep learning model, the file name feature vectors are obtained by encoding the file names through a character-level encoding model, and the file metadata feature vectors include the numerical representations of the file size, creation time, and modification time; A second unit configured to divide the set of feature vectors of files to be matched into multiple subtasks based on a distributed computing framework, perform parallel calculations of the feature vector similarity on multiple computing nodes of the distributed computing framework, calculate the similarity scores between each file to be matched and the files in the target file set. The similarity scores are obtained by performing weighted calculations on the file content feature vector similarity, file name feature vector similarity, and file metadata feature vector similarity, and filter the similarity scores according to a preset similarity threshold to generate a list of files to be copied; A third unit configured to submit the list of files to be copied to the task scheduler of the distributed file system. The task scheduler dynamically allocates the execution nodes of the copy tasks according to the system resource status, starts a file copy process on the execution nodes. The file copy process performs file copy operations based on a data block-level parallel transmission mechanism, verifies the data integrity during the copying process, stores the copied files in the target storage location, and generates a copy task execution report.

9. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Methods and decentralized systems that employ distributed machine learning to automatically instantiate and manage distributed applications

    US20230028934A1

  • Reading type examination question generation system and method based on commonsense reasoning

    WO2023225858A1