Multi-file comparison method and device, computer equipment and storage medium

By chunking the files and scoring matrix comparison, the problem of traditional file comparison is solved and efficient file content comparison is achieved.

CN120257983APending Publication Date: 2025-07-04SHENZHEN COMTOP INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510375992.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional document comparison methods are less efficient and it is difficult to efficiently handle the comparison of similar contents of a large number of complex bidding documents and other documents.

Method used

By chunking the multiple files to be compared, a sequence to be compared, and the same block information is compared using the score matrix to determine the target content.

Benefits of technology

Improves file comparison efficiency, reduces complexity, and quickly determines the maximum identical sequence between files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257983A_ABST
    Figure CN120257983A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-file comparison method and device, computer equipment and a storage medium. The method comprises the steps that each to-be-compared file in a plurality of to-be-compared files is subjected to block processing, a to-be-compared sequence corresponding to each to-be-compared file is obtained, and the to-be-compared sequence comprises multiple pieces of block information; the first sequence and the second sequence are compared, a score matrix is obtained, elements in the score matrix are used for representing the number of times that block information in one sequence appears in the other sequence, and the first sequence and the second sequence are any two sequences in the multiple to-be-compared sequences; and determining the same block information in the first sequence and the second sequence based on the score matrix, and determining the target content in the to-be-compared file corresponding to the same block information. By adopting the method, the file comparison efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of document comparison, and particularly to a method, apparatus, computer device, and storage medium for comparing multiple documents. Background Art

[0002] Due to the rapid growth of information volume in modern devices and the continuous improvement of document management requirements, a large number of document comparisons are required. For example, in the field of bidding, there are many involved documents, including bidding documents, tender documents, contract documents, etc. These document contents are complex and have numerous clauses, requiring careful comparison and analysis to detect similar contents in the bids, thereby ensuring the fairness and transparency of the bidding.

[0003] In traditional technologies, usually, one-to-one comparisons are made on the strings or hash values corresponding to multiple documents to obtain the comparison results of each document.

[0004] However, traditional document comparison methods have the problem of low comparison efficiency. Summary of the Invention

[0005] Based on this, it is necessary to provide a method, apparatus, computer device, and storage medium for comparing multiple documents that can improve the comparison efficiency in view of the above technical problems.

[0006] In a first aspect, this application provides a method for comparing multiple documents, and the method includes:

[0007] Performing block processing on each of the multiple documents to be compared to obtain a sequence to be compared corresponding to each of the documents to be compared, where the sequence to be compared includes: a plurality of block information;

[0008] Comparing a first sequence and a second sequence to obtain a score matrix, where the elements in the score matrix are used to represent the number of times the block information in one sequence appears in the other sequence, and the first sequence and the second sequence are any two sequences among the multiple sequences to be compared;

[0009] Based on the score matrix, determining the identical block information in the first sequence and the second sequence, and determining the target content in the documents to be compared corresponding to the identical block information.

[0010] In one embodiment, the comparing the first sequence and the second sequence to obtain a score matrix includes:

[0011] Determining the comparison result of any block information in the first sequence and any block information in the second sequence;

[0012] Based on the comparison result and adjacent matrix elements, determining the matrix element corresponding to the comparison result.

[0013] In one embodiment, determining the matrix element corresponding to the comparison result based on the comparison result and adjacent matrix elements includes:

[0014] If the comparison result is a match, determine the maximum matrix element value from the adjacent matrix elements, increase the single-element value based on the maximum matrix element value of the adjacent matrix elements, and determine it as the matrix element corresponding to the comparison result;

[0015] If the comparison result is a mismatch, determine the maximum matrix element value from the adjacent matrix elements, and determine the maximum matrix element value of the adjacent matrix elements as the matrix element corresponding to the comparison result.

[0016] In one embodiment, determining the target content in the file to be compared corresponding to the identical block information based on the score matrix and determining the identical block information between the first sequence and the second sequence includes:

[0017] Step 1: In the first sequence and the second sequence, obtain two block information corresponding to the maximum matrix element value that has not been backtracked and compared in the score matrix;

[0018] Step 2: Perform backtracking comparison on the two block information corresponding to the maximum matrix element value: If the two block information match, determine the two block information as identical block information, and return to execute Step 1; if the two block information do not match, return to execute Step 1.

[0019] In one embodiment, the block information is a hash value; performing block processing on each file to be compared among multiple files to be compared to obtain a sequence to be compared corresponding to each file to be compared includes:

[0020] Perform block processing on the file to be compared according to a preset length to obtain multiple block contents of the file to be compared;

[0021] Determine the hash value of each block content, and combine the hash values of each file block to obtain the sequence to be compared.

[0022] In one embodiment, the method further includes:

[0023] Perform word segmentation on the file to be compared, and remove irrelevant characters in the file to be compared after word segmentation to obtain the processed file to be compared;

[0024] Performing block processing on each file to be compared among multiple files to be compared to obtain a sequence to be compared corresponding to each file to be compared includes:

[0025] Perform block processing on the processed file to be compared to obtain the sequences to be compared corresponding to the file to be compared.

[0026] In a second aspect, the present application further provides a multi-file comparison device, including:

[0027] A block module, configured to perform block processing on each file to be compared among multiple files to be compared, to obtain the sequences to be compared corresponding to each file to be compared, where the sequences to be compared include: multiple block information;

[0028] A comparison module, configured to compare a first sequence and a second sequence to obtain a score matrix, where the elements in the score matrix are used to represent the number of times the block information in one sequence appears in the other sequence, and the first sequence and the second sequence are any two sequences among the multiple sequences to be compared;

[0029] A determination module, configured to determine the identical block information in the first sequence and the second sequence based on the score matrix, and determine the target content in the file to be compared corresponding to the identical block information.

[0030] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0031] Perform block processing on each file to be compared among multiple files to be compared, to obtain the sequences to be compared corresponding to each file to be compared, where the sequences to be compared include: multiple block information;

[0032] Compare a first sequence and a second sequence to obtain a score matrix, where the elements in the score matrix are used to represent the number of times the block information in one sequence appears in the other sequence, and the first sequence and the second sequence are any two sequences among the multiple sequences to be compared;

[0033] Determine the identical block information in the first sequence and the second sequence based on the score matrix, and determine the target content in the file to be compared corresponding to the identical block information.

[0034] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0035] Perform block processing on each file to be compared among multiple files to be compared, to obtain the sequences to be compared corresponding to each file to be compared, where the sequences to be compared include: multiple block information;

[0036] Compare the first sequence and the second sequence to obtain a scoring matrix, where the elements in the scoring matrix are used to represent the number of times the block information in one sequence appears in the other sequence, and the first sequence and the second sequence are any two sequences among the multiple sequences to be compared;

[0037] Based on the scoring matrix, determine the identical block information in the first sequence and the second sequence, and determine the target content in the file to be compared corresponding to the identical block information.

[0038] In a fifth aspect, the present application also provides a computer program product, including a computer program, which when executed by a processor, implements the following steps:

[0039] Perform a block processing on each file to be compared among multiple files to be compared, to obtain a sequence to be compared corresponding to each file to be compared, where the sequence to be compared includes: multiple block information;

[0040] Compare the first sequence and the second sequence to obtain a scoring matrix, where the elements in the scoring matrix are used to represent the number of times the block information in one sequence appears in the other sequence, and the first sequence and the second sequence are any two sequences among the multiple sequences to be compared;

[0041] Based on the scoring matrix, determine the identical block information in the first sequence and the second sequence, and determine the target content in the file to be compared corresponding to the identical block information.

[0042] The above multi-file comparison method, device, computer device and storage medium perform a block processing on each file to be compared among multiple files to be compared, to obtain a sequence to be compared corresponding to each file to be compared, where the sequence to be compared includes: multiple block information; compare the first sequence and the second sequence to obtain a scoring matrix, where the elements in the scoring matrix are used to represent the number of times the block information in one sequence appears in the other sequence, and the first sequence and the second sequence are any two sequences among the multiple sequences to be compared; based on the scoring matrix, determine the identical block information in the first sequence and the second sequence, and determine the target content in the file to be compared corresponding to the identical block information. By performing a block processing on the file to be compared, the original long sequence problem is decomposed into relatively simple short sequence problems, reducing the complexity of the comparison, thereby improving the efficiency of the comparison, and by backtracking through the scoring matrix obtained by comparing each block, quickly determining the maximum identical sequence between the sequences to be compared, further improving the efficiency of the comparison. Description of the Drawings

[0043] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present application or related technologies. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0044] Figure 1 It is an application environment diagram of the multi-file comparison method in an embodiment;

[0045] Figure 2 It is a schematic flowchart of the multi-file comparison method in an embodiment;

[0046] Figure 3 It is a schematic flowchart of the multi-file comparison method in another embodiment;

[0047] Figure 4 It is a schematic flowchart of the multi-file comparison method in another embodiment;

[0048] Figure 5 It is a schematic flowchart of the multi-file comparison method in another embodiment;

[0049] Figure 6 It is a schematic flowchart of the multi-file comparison method in another embodiment;

[0050] Figure 7 It is a schematic flowchart of the multi-file comparison method in another embodiment;

[0051] Figure 8 It is a structural block diagram of the multi-file comparison device in an embodiment. Detailed implementation manners

[0052] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0053] The multi-file comparison method provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the computer device can be a terminal, and the computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it is used to implement a method for comparing multiple files. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0054] Those skilled in the art can understand that Figure 1 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0055] In one embodiment, as Figure 2 shown, a method for comparing multiple files is provided. Taking the terminal in Figure 1 as an example, it includes:

[0056] S201, perform block processing on each file to be compared among multiple files to be compared, and obtain a sequence to be compared corresponding to each file to be compared. The sequence to be compared includes: multiple block information.

[0057] In the embodiment of the present application, for each file to be compared, the file to be compared can be block-processed by calling a preset script, and the text information in the file to be compared can be segmented into multiple block contents, so as to determine multiple block information according to the multiple block contents, and combine the multiple block information to obtain the sequence to be compared corresponding to each file to be compared.

[0058] Optionally, the chunk information may be the chunk content, or the chunk information may be a string or a value obtained according to the chunk content, etc.

[0059] S202. Compare the first sequence and the second sequence to obtain a scoring matrix, where the elements in the scoring matrix are used to represent the number of times the chunk information in one sequence appears in the other sequence. The first sequence and the second sequence are any two sequences among multiple sequences to be compared.

[0060] In the embodiments of the present application, any two sequences are selected from multiple sequences to be compared as the first sequence and the second sequence for comparison until all sequence combinations are traversed and compared, which means the comparison of multiple sequences to be compared is completed. In the embodiments of the present application, for each comparison, an initial scoring matrix may be constructed first according to the lengths of the first sequence and the second sequence. Further, the first sequence and the second sequence are compared to determine the values of the elements in the scoring matrix, thereby obtaining the scoring matrix.

[0061] Exemplarily, if the first sequence includes m chunk information and the second sequence includes n chunk information, then the length of the first sequence is m, the length of the second sequence is n, and an initial scoring matrix of size m×n is constructed. Further, the number of times each chunk information in the first sequence appears in the first j chunk information in the second sequence may be determined first, where j = 1, 2,..., m. Further, the number of times each chunk information in the second sequence appears in the first i chunk information in the first sequence is determined, where i = 1, 2,..., n. Thus, according to the number of times the chunk information in one sequence appears in the other sequence and the initial scoring matrix, the scoring matrix is obtained.

[0062] S203. Determine the identical chunk information in the first sequence and the second sequence based on the scoring matrix, and determine the target content in the file to be compared corresponding to the identical chunk information.

[0063] In the embodiments of the present application, the identical chunk information in the first sequence and the second sequence is determined according to the magnitudes of the elements in the scoring matrix and the positional relationship of the elements. Further, the identical chunk content is extracted from the files to be compared corresponding to the above two sequences according to the identical chunk information, and the identical chunk content is combined to obtain the target content in the file to be compared.

[0064] In the above application embodiments, each of the multiple files to be compared is block - processed to obtain a sequence to be compared corresponding to each file to be compared. The sequence to be compared includes: multiple block information; the first sequence and the second sequence are compared to obtain a score matrix, and the elements in the score matrix are used to represent the number of times the block information in one sequence appears in another sequence. The first sequence and the second sequence are any two sequences among the multiple sequences to be compared; based on the score matrix, the identical block information in the first sequence and the second sequence is determined, and the target content in the file to be compared corresponding to the identical block information is determined. By block - processing the files to be compared, the problem of the original long sequence is decomposed into relatively simple short - sequence problems, reducing the complexity of comparison, thereby improving the efficiency of comparison. And by backtracking through the score matrix obtained by comparing each block, the maximum identical sequence between the sequences to be compared is quickly determined, further improving the efficiency of comparison.

[0065] In one embodiment, an implementation manner of the above S202 is provided, as Figure 3 shown, the above "comparing the first sequence and the second sequence to obtain a score matrix" includes:

[0066] S301, determining the comparison result of any block information in the first sequence and any block information in the second sequence.

[0067] Among them, the comparison result includes match and mismatch.

[0068] In the embodiments of the present application, any block information in the first sequence can be represented as A[i], and any block information in the second sequence can be represented as B[j], where i = 1, 2,..., n, j = 1, 2,..., m, n is the number of block information in the first sequence, and m is the number of block information in the second sequence. Determine whether any block information in the first sequence is the same as any block information in the second sequence to obtain multiple comparison results, and each comparison result corresponds one - to - one with each matrix element in the score matrix. Among them, if any block information in the first sequence is the same as any block information in the second sequence, the comparison result is a match; if any block information in the first sequence is different from any block information in the second sequence, the comparison result is a mismatch.

[0069] S302, determining the matrix element corresponding to the comparison result based on the comparison result and adjacent matrix elements.

[0070] In the embodiments of the present application, the matrix elements in the first row and the first column of the matrix can both be 0, and the matrix element in the (i + 1)-th row and (j + 1)-th column of the score matrix can be represented as dp[i][j]. According to the order from top to bottom and from left to right, based on each matrix element and the already - determined adjacent matrix elements of this matrix element, the matrix element corresponding to the comparison result is determined.

[0071] Optionally, if the comparison result is a match, the sum of adjacent matrix elements can be incremented by 1 and determined as the matrix element corresponding to the comparison result; if the comparison result is a mismatch, the sum of adjacent matrix elements can be determined as the matrix element corresponding to the comparison result.

[0072] Optionally, in the embodiments of the present application, as Figure 4 shown, the above S302 "determine the matrix element corresponding to the comparison result based on the comparison result and adjacent matrix elements" may include:

[0073] S401, if the comparison result is a match, determine the maximum matrix element value from the adjacent matrix elements, increment the maximum matrix element value of the adjacent matrix elements by the single element value, and determine it as the matrix element corresponding to the comparison result.

[0074] S402, if the comparison result is a mismatch, determine the maximum matrix element value from the adjacent matrix elements, and determine the maximum matrix element value of the adjacent matrix elements as the matrix element corresponding to the comparison result.

[0075] In the embodiments of the present application, an initial score matrix dp is created. If A[i] == B[j], then dp[i + 1][j +1] = dp[i][j] + 1, indicating that when the current elements of the two sequences match, the current longest common subsequence length is the longest common subsequence length of the subsequence without this element plus one; if A[i] != B[j], then dp[i+1][j+1] = max(dp[i+1][j], dp[i][j+1]), indicating that when the current elements of the two sequences do not match, the current longest common subsequence length is the larger value of the two possible sub-problem solutions, that is, either without the current element of sequence A or without the current element of sequence B.

[0076] Exemplarily, when the first sequence is ABCD and the second sequence is AEBD, first create an initial score matrix dp. The size of this matrix is (len(A)+1)×(len(B)+1), where len(A)=4 and len(B)=4, so the size of the initial score matrix is 5x5. Comparing the first character 'A' of the first sequence with all characters of the second sequence, it can be obtained that the first character 'A' in the first sequence matches the first character 'A' in the second sequence. Therefore, dp[1][1] is updated to dp[0][0]+1, which is 1. Continuing the comparison, 'E', 'B', and 'D' in B cannot match 'A'. Therefore, according to the state transition equation, select the maximum value between dp[i][j]'s left dp[i+1][j] and above dp[i][j+1] to fill this matrix element, and fill 1. Further, compare the first character 'B' of the first sequence with all characters of the second sequence. The 'B' in the first sequence does not match the characters 'A' and 'E' in the second sequence. Then select the maximum value between the left dp[i+1][j] and above dp[i][j+1] to fill, that is, fill 1. And the 'B' in the first sequence matches the third character 'B' in the second sequence. Therefore, dp[2][3] is updated to dp[1][2] + 1, which is 2. And the 'B' in the first sequence does not match the character 'D' in the second sequence. Then select the maximum value between the left (dp[i+1][j]) and above (dp[i][j+1]) to fill, that is, fill 2. And so on until the entire dp matrix is filled.

[0077] Exemplarily, comparing the first sequence and the second sequence can be achieved by calling the following code:

[0078] import threading

[0079] from Bio import pairwise2

[0080] from Bio.pairwise2 import format_alignment

[0081] # Read the file and chunk it

[0082] def read_and_chunk_file(file_path, chunk_size):

[0083] # Read the file content and chunk it

[0084] with open(file_path, 'r') as file:

[0085] content = file.read()

[0086] chunks = [content[i:i + chunk_size] for i in range(0, len(content), chunk_size)]

[0087] return chunks

[0088] # Smith-Waterman alignment function

[0089] def smith_waterman_alignment(seq1, seq2):

[0090] alignments = pairwise2.align.localxx(seq1, seq2)

[0091] best_alignment = max(alignments, key=lambda x: x[2])

[0092] return format_alignment(*best_alignment)

[0093] # Parallel alignment function

[0094] def parallel_alignment(file1_chunks, file2_chunks, chunk_index):

[0095] alignment = smith_waterman_alignment(file1_chunks[chunk_index],file2_chunks[chunk_index])

[0096] return chunk_index, alignment

[0097] # Main function

[0098] def main(file1_path, file2_path, chunk_size, num_threads):

[0099] file1_chunks = read_and_chunk_file(file1_path, chunk_size)

[0100] file2_chunks = read_and_chunk_file(file2_path, chunk_size)

[0101] threads = []

[0102] results = []

[0103] for i in range(min(len(file1_chunks), len(file2_chunks))):

[0104] thread = threading.Thread(target=parallel_alignment, args=(file1_chunks, file2_chunks, i))

[0105] threads.append(thread)

[0106] thread.start()

[0107] for thread in threads:

[0108] thread.join()

[0109] results.append(thread.result) # Note: Here, it is necessary to implement the acquisition of thread results. Python threads do not support directly returning results by default

[0110] # Aggregate and process the results

[0111] for chunk_index, alignment in results:

[0112] print(f"Chunk {chunk_index}: {alignment}")

[0113] # Example call

[0114] main('file1.txt', 'file2.txt', 1000, 4)

[0115] In the above embodiments of the application, by determining the comparison result between any block information in the first sequence and any block information in the second sequence, the problem of the original long sequence is decomposed into relatively simple short sequence problems, reducing the complexity of the comparison, thereby improving the comparison efficiency.

[0116] In one embodiment, an implementation of the above S203 is provided. As Figure 5 shown, the above "determining the same block information in the first sequence and the second sequence based on the score matrix, and determining the target content in the file to be compared corresponding to the same block information" includes:

[0117] S501, Step 1: In the first sequence and the second sequence, obtain two block information corresponding to the maximum matrix element value that has not been backtracked and compared in the score matrix.

[0118] In one embodiment, determining the maximum matrix element value that has not been backtracked and compared in the score matrix, that is, the starting point of backtracking and comparison can be the last matrix element dp[i][j] in the score matrix, and determining two block information corresponding to the maximum matrix element value, that is, A[i] and B[j]. Exemplarily, the score matrix can be as shown in Table 1:

[0119] Table 1

[0120]

[0121] S502, Step 2: Perform backtracking and comparison on the two block information corresponding to the maximum matrix element value: If the two block information match, determine the two block information as the same block information, and return to execute Step 1; if the two block information do not match, return to execute Step 1.

[0122] For example, the value of dp[4][4] is 3, indicating that the number of identical block information of "ABCD" and "AEBD" is 3. A[4] and B[4] corresponding to dp[4][4] are both "D", and the two block information match. It is determined that "D" is the same block information. Further, the largest matrix element value in the score matrix that has not been backtracked and is adjacent to dp[4][4] is determined to be dp[3][4], dp[4][3], and dp[3][3]. One of the identical matrix element values ​​is selected, for example, dp[3][3] at the upper left diagonal position of dp[4][4] can be preferentially selected. Comparing A[3] and B[3], there is no match. It is determined that the largest matrix element value in the score matrix that has not been backtracked and is adjacent to dp[3][3] is dp[2][3]. Compare A[2] and B[3], if they match, then "B" is determined to be the same block information. Further, determine that the maximum matrix element value in the score matrix that has not been backtracked and is adjacent to dp[2][3] is dp[1][3], dp[2][2], and dp[1][2]. Select one of the same matrix element values. For example, you can preferentially select dp[1][2], which is the diagonal position to the upper left of dp[2][3]. Compare A[1] and B[2], if they do not match, then determine that the maximum matrix element value in the score matrix that has not been backtracked and is adjacent to dp[1][2] is dp[1][1]. Compare A[1] and B[1], if they match, then "A" is determined to be the same block information.

[0123] Exemplarily, backtracking processing according to the score matrix can be implemented by calling the following code:

[0124] #include<stdio.h>

[0125] #include<stdlib.h>

[0126] #include<string.h>

[0127] / / Returns the maximum of two integers

[0128] int max(int ​​a, int b) {

[0129] return (a > b) ? a : b;

[0130] }

[0131] / / Dynamic programming solves the optimal path problem

[0132] int searchPath(char *X, char *Y, int m, int n) {

[0133] int L[m + 1][n + 1];

[0134] int i, j;

[0135] / / Build L[m + 1][n + 1] to store the length of the path

[0136] for (i = 0; i <= m; i++) {

[0137] for (j = 0; j <= n; j++) {

[0138] if (i == 0 || j == 0)

[0139] L[i][j] = 0;

[0140] else if (X[i - 1] == Y[j - 1])

[0141] L[i][j] = L[i - 1][j - 1] + 1;

[0142] else

[0143] L[i][j] = max(L[i - 1][j], L[i][j - 1]);

[0144] }

[0145] }

[0146] / / L[m][n] contains the length of the optimal path between X[0..m - 1] and Y[0..n - 1]

[0147] return L[m][n];

[0148] }

[0149] / / Print the optimal path, this is an auxiliary function

[0150] void printPath(char *X, char *Y, int m, int n) {

[0151] int index = path(X, Y, m, n);

[0152] char path[index + 1];

[0153] path[index] = '\0'; / / Set the terminator of the string

[0154] int i = m, j = n;

[0155] while (i > 0 && j > 0) {

[0156] if (X[i-1] == Y[j-1]) {

[0157] path[index-1] = X[i-1]; / / If the current character is in the path

[0158] i--; j--; index--; / / Decrease the values

[0159] }

[0160] else if (L[i-1][j] > L[i][j-1])

[0161] i--;

[0162] else

[0163] j--;

[0164] }

[0165] / / Print the path result S

[0166] printf("PATH of %s and %s is %s\n", X, Y, path);

[0167] }

[0168] / / Test code

[0169] int main() {

[0170] char X[] = "AGGTAB";

[0171] char Y[] = "GXTXAYB";

[0172] int m = strlen(X);

[0173] int n = strlen(Y);

[0174] printf("Length of PATH is %d\n", lcs(X, Y, m, n));

[0175] / / If you need to print the PATH, uncomment the following line

[0176] / / printPath(X, Y, m, n);

[0177] return 0;

[0178] }

[0179] In the above application embodiments, according to the size and position of each matrix element, the backtracking comparison path is determined, which improves the efficiency of backtracking based on the score matrix.

[0180] In one embodiment, an implementation manner of the above S201 is provided, and the chunk information is a hash value. As Figure 6 shown, the above "performing chunking processing on each of the multiple files to be compared to obtain a sequence to be compared corresponding to each file to be compared" includes:

[0181] S601, performing chunking processing on the file to be compared according to a preset length to obtain multiple chunk contents of the file to be compared.

[0182] In the embodiments of the present application, the preset length can be a fixed length or a dynamic length. Optionally, the file to be compared is chunked according to a fixed length. For example, the fixed length can be 10KB, 100KB, etc.; or, the dynamic length is determined according to the change of the file content, and then the file to be compared is chunked according to the dynamic length. For example, the file content includes text paragraph boundaries, tables, flowchart text, etc.

[0183] S602, determining the hash value of each chunk content, and combining the hash values of each file chunk to obtain a sequence to be compared.

[0184] In the embodiments of the present application, since the hash matching algorithm has strong adaptability, the hash value corresponding to each chunk content is determined, the hash value corresponding to each chunk content is used as chunk information, and the chunk information is combined to obtain a sequence to be compared.

[0185] In the above application embodiments, by adopting a chunking strategy, the feature values of the file sequence can be effectively extracted, which is beneficial to subsequent feature comparison and duplicate checking.

[0186] In one embodiment, as Figure 7 shown, the above method for comparing multiple files further includes:

[0187] S204, performing word segmentation processing on the file to be compared, and removing irrelevant characters in the file to be compared after word segmentation processing to obtain the processed file to be compared.

[0188] In the embodiments of the present application, two or more files to be compared are determined, and all files to be compared are read into memory or stored in an efficient access data structure, such as using a memory-mapped file. Further, word segmentation is performed on the files to be compared, and interfering, irrelevant characters and words are removed.

[0189] Exemplarily, to perform file word segmentation and remove interfering, irrelevant characters and vocabulary for the file to be compared, it can be achieved by calling the following code:

[0190] import jieba

[0191] jieba.load_userdict('extraDict.txt')

[0192] def stopwordlist():

[0193] stopwords =[line.strip() for line in open('filepath',encoding='utf-8').readlines()]

[0194] return stopwords

[0195] def seg_word(line):

[0196] seg = jieba.cut(line.strip())

[0197] #jieba.cut()

[0198] temp = ''

[0199] wordstop = stopwordlist()

[0200] for word in seg:

[0201] if word not in wordstop:

[0202] if word!="\t":

[0203] temp+=word

[0204] temp+="\n"

[0205] return temp

[0206] def output(inputfilename,outputfilename):

[0207] inputfile = open(inputfilename, encoding='utf-8', mode = "r") #inputfilename file needs to be read, so write "r", and use inputfile to receive

[0208] outputfile = open(outputfilename, encoding='utf-8', mode = "w") #outputfilename file needs to be written, so use "w", and use outputfile to receive

[0209] for line in inputfile.readlines(): #Use a for loop to read inputfile line by line, that is, take out each line and read it

[0210] line_seg = seg_word(line) #Use seg_word() to segment line

[0211] outputfile.write(line_seg)

[0212] inputfile.close

[0213] outputfile.close #Close the two files

[0214] if __name__ == '__main__':

[0215] print("__name__",__name__)

[0216] inputfilename = '03.txt' #This is the existing file to be operated on

[0217] outputfilename = '02.txt' #This is the file obtained after the operation, and the file name can be arbitrary

[0218] output(inputfilename,outputfilename) #output() function, that is, input 'inputfilename' and output 'outputfilename'

[0219] Among them, filepath is the path where the stop word list is located; line.strip() is used to remove the blanks and symbols before and after the text; the for loop is used to read each line of data in the provided stop word list; the open() function is used to open the stop word list, and the readlines() function is used to read; jieba.load_userdict('extraDict.txt') is used to read the custom text extraDict.txt; line is a line of text in inputfile; strip() is used to remove the spaces, symbols in front of a line of text, as well as the spaces and symbols behind the text; stopwordlist() is the function defined above, and wordstop means the stop use of words.

[0220] On this basis, the above S201 "perform chunking processing on each of the multiple files to be compared to obtain a sequence to be compared corresponding to each file to be compared" includes:

[0221] S205, perform chunking processing on the processed file to be compared to obtain a sequence to be compared corresponding to the file to be compared.

[0222] In the embodiment of the present application, for each processed file to be compared, the processed file to be compared can be chunked by calling a preset script, and the text information in the processed file to be compared is split into multiple chunk contents, so as to determine multiple chunk information according to the multiple chunk contents, and the multiple chunk information is combined to obtain a sequence to be compared corresponding to each processed file to be compared.

[0223] Optionally, the chunk information can be the chunk content, or the chunk information can be a string or a value obtained according to the chunk content, etc.

[0224] In the above application embodiment, preprocessing the file to be compared to obtain the processed file to be compared improves the accuracy and efficiency of chunking processing on the file to be compared.

[0225] In one embodiment, a complete multi-file comparison method is provided, including:

[0226] S1, perform word segmentation processing on the file to be compared, and remove the irrelevant characters in the file to be compared after word segmentation processing to obtain the processed file to be compared.

[0227] S2, perform chunking processing on the processed file to be compared according to a preset length to obtain multiple chunk contents of the file to be compared.

[0228] S3, determine the hash value of each chunk content, and combine the hash values of each file chunk to obtain a sequence to be compared.

[0229] S4. Determine the comparison result of any block information in the first sequence and any block information in the second sequence; the elements in the scoring matrix are used to represent the number of times the block information in one sequence appears in another sequence, and the first sequence and the second sequence are any two of the multiple sequences to be compared.

[0230] S5. If the comparison result is a match, determine the maximum matrix element value from the adjacent matrix elements, increase the single-element value based on the maximum matrix element value of the adjacent matrix elements, and determine it as the matrix element corresponding to the comparison result; if the comparison result is a mismatch, determine the maximum matrix element value from the adjacent matrix elements, and determine the maximum matrix element value of the adjacent matrix elements as the matrix element corresponding to the comparison result.

[0231] S6. Step 1: In the first sequence and the second sequence, obtain the two block information corresponding to the maximum matrix element value that has not been backtracked and compared in the scoring matrix.

[0232] S7. Step 2: Perform backtracking comparison on the two block information corresponding to the maximum matrix element value: if the two block information match, determine the two block information as the same block information, and return to execute Step 1; if the two block information do not match, return to execute Step 1.

[0233] In the above multi-file comparison method, each file to be compared among the multiple files to be compared is block-processed to obtain the sequence to be compared corresponding to each file to be compared, and the sequence to be compared includes: multiple block information; the first sequence and the second sequence are compared to obtain a scoring matrix, and the elements in the scoring matrix are used to represent the number of times the block information in one sequence appears in another sequence, and the first sequence and the second sequence are any two of the multiple sequences to be compared; based on the scoring matrix, the same block information in the first sequence and the second sequence is determined, and the target content in the file to be compared corresponding to the same block information is determined. By block-processing the file to be compared, the original long-sequence problem is decomposed into relatively simple short-sequence problems, reducing the complexity of the comparison, thereby improving the comparison efficiency, and by backtracking the scoring matrix obtained by comparing each block, the maximum identical sequence between the sequences to be compared is quickly determined, further improving the comparison efficiency.

[0234] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown in the direction of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0235] Based on the same inventive concept, an embodiment of the present application further provides a multi-file comparison device for implementing the multi-file comparison method described above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the multi-file comparison device provided below can refer to the limitations on the multi-file comparison method in the above text, and will not be repeated here.

[0236] In one embodiment, as Figure 8 shown, a multi-file comparison device is provided, including: a chunking module 10, a comparison module 11, and a backtracking module 12, where:

[0237] The chunking module 10 is used to perform chunking processing on each of the multiple files to be compared, and obtain a sequence to be compared corresponding to each file to be compared. The sequence to be compared includes: multiple chunking information.

[0238] The comparison module 11 is used to compare the first sequence and the second sequence to obtain a score matrix. The elements in the score matrix are used to represent the number of times the chunking information in one sequence appears in the other sequence. The first sequence and the second sequence are any two sequences among the multiple sequences to be compared.

[0239] The determination module 12 is used to determine the identical chunking information in the first sequence and the second sequence based on the score matrix, and determine the target content in the file to be compared corresponding to the identical chunking information.

[0240] In one embodiment, the above comparison module includes: a comparison unit and a first determination unit, where:

[0241] The comparison unit is used to determine the comparison result between any chunking information in the first sequence and any chunking information in the second sequence.

[0242] The first determination unit is used to determine the matrix element corresponding to the comparison result based on the comparison result and the adjacent matrix element.

[0243] In one embodiment, the above-mentioned determining unit is specifically configured to, if the comparison result is a match, determine the maximum matrix element value from adjacent matrix elements, increase the single-element value based on the maximum matrix element value of the adjacent matrix elements, and determine it as the matrix element corresponding to the comparison result; if the comparison result is a mismatch, determine the maximum matrix element value from adjacent matrix elements, and determine the maximum matrix element value of the adjacent matrix elements as the matrix element corresponding to the comparison result.

[0244] In one embodiment, the above-mentioned determining module 12 includes an obtaining unit and a second determining unit, where:

[0245] The obtaining unit is configured to execute step 1: Step 1: In the first sequence and the second sequence, obtain two block information corresponding to the maximum matrix element value that has not been backtracked in the scoring matrix.

[0246] The second determining unit is configured to execute step 2: Perform backtracking comparison on the two block information corresponding to the maximum matrix element value: if the two block information match, determine the two block information as the same block information, and return to execute step 1; if the two block information do not match, return to execute step 1.

[0247] In one embodiment, the above-mentioned blocking module 10 includes: a blocking unit and a third determining unit, where:

[0248] The blocking unit is configured to perform block processing on the file to be compared according to a preset length to obtain multiple block contents of the file to be compared.

[0249] The third determining unit is configured to determine the hash value of each block content, and combine the hash values of each file block to obtain the sequence to be compared.

[0250] In one embodiment, the above-mentioned multi-file comparison device further includes: a word segmentation module, where:

[0251] The word segmentation module is configured to perform word segmentation processing on the file to be compared, and remove irrelevant characters in the file to be compared after word segmentation processing to obtain the processed file to be compared.

[0252] The blocking module is specifically configured to perform block processing on the processed file to be compared to obtain the sequence to be compared corresponding to the file to be compared.

[0253] Each module in the above-mentioned multi-file comparison device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0254] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0255] Perform block processing on each file to be compared among multiple files to be compared, and obtain a sequence to be compared corresponding to each file to be compared. The sequence to be compared includes: multiple block information;

[0256] Compare the first sequence and the second sequence to obtain a score matrix. The elements in the score matrix are used to represent the number of times the block information in one sequence appears in the other sequence. The first sequence and the second sequence are any two sequences among the multiple sequences to be compared;

[0257] Based on the score matrix, determine the identical block information in the first sequence and the second sequence, and determine the target content in the file to be compared corresponding to the identical block information.

[0258] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0259] Determine the comparison result between any block information in the first sequence and any block information in the second sequence;

[0260] Based on the comparison result and the adjacent matrix element, determine the matrix element corresponding to the comparison result.

[0261] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0262] If the comparison result is a match, determine the maximum matrix element value from the adjacent matrix elements, increase the single element value based on the maximum matrix element value of the adjacent matrix elements, and determine it as the matrix element corresponding to the comparison result;

[0263] If the comparison result is a mismatch, determine the maximum matrix element value from the adjacent matrix elements, and determine the maximum matrix element value of the adjacent matrix elements as the matrix element corresponding to the comparison result.

[0264] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0265] Step 1: In the first sequence and the second sequence, obtain two block information corresponding to the maximum matrix element value that has not been backtracked and compared in the score matrix;

[0266] Step 2: Perform backtracking comparison on the two block information corresponding to the maximum matrix element value: If the two block information match, determine the two block information as identical block information, and return to execute Step 1; if the two block information do not match, return to execute Step 1.

[0267] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0268] Chunk the file to be compared according to a preset length to obtain multiple chunk contents of the file to be compared;

[0269] Determine the hash value of each chunk content, and combine the hash values of each file chunk to obtain a sequence to be compared.

[0270] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0271] Perform word segmentation on the file to be compared, and remove irrelevant characters in the file to be compared after word segmentation to obtain the processed file to be compared;

[0272] Chunk each file to be compared among multiple files to be compared to obtain a sequence to be compared corresponding to each file to be compared, including:

[0273] Chunk the processed file to be compared to obtain a sequence to be compared corresponding to the file to be compared.

[0274] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0275] Chunk each file to be compared among multiple files to be compared to obtain a sequence to be compared corresponding to each file to be compared. The sequence to be compared includes: multiple chunk information;

[0276] Compare the first sequence and the second sequence to obtain a score matrix. The elements in the score matrix are used to represent the number of times the chunk information in one sequence appears in the other sequence. The first sequence and the second sequence are any two sequences among multiple sequences to be compared;

[0277] Based on the score matrix, determine the identical chunk information in the first sequence and the second sequence, and determine the target content in the file to be compared corresponding to the identical chunk information.

[0278] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented:

[0279] Determine the comparison result between any chunk information in the first sequence and any chunk information in the second sequence;

[0280] Based on the comparison result and the adjacent matrix element, determine the matrix element corresponding to the comparison result.

[0281] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented:

[0282] If the comparison result is a match, determine the maximum matrix element value from adjacent matrix elements, increase the single-element value based on the maximum matrix element value of the adjacent matrix elements, and determine it as the matrix element corresponding to the comparison result.

[0283] If the comparison result is a mismatch, determine the maximum matrix element value from adjacent matrix elements, and determine the maximum matrix element value of the adjacent matrix elements as the matrix element corresponding to the comparison result.

[0284] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0285] Step 1: In the first sequence and the second sequence, obtain two block information corresponding to the maximum matrix element value in the scoring matrix that has not been backtracked for comparison.

[0286] Step 2: Perform backtracking comparison on the two block information corresponding to the maximum matrix element value: If the two block information match, determine the two block information as the same block information, and return to execute Step 1; if the two block information do not match, return to execute Step 1.

[0287] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0288] Perform block processing on the file to be compared according to a preset length to obtain multiple block contents of the file to be compared;

[0289] Determine the hash value of each block content, and combine the hash values of each file block to obtain the sequence to be compared.

[0290] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0291] Perform word segmentation processing on the file to be compared, and remove irrelevant characters in the file to be compared after word segmentation processing to obtain the processed file to be compared;

[0292] Perform block processing on each file to be compared among multiple files to be compared to obtain a sequence to be compared corresponding to each file to be compared, including:

[0293] Perform block processing on the processed file to be compared to obtain a sequence to be compared corresponding to the file to be compared.

[0294] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0295] Perform block processing on each file to be compared among multiple files to be compared to obtain a sequence to be compared corresponding to each file to be compared, and the sequence to be compared includes: multiple block information;

[0296] Compare the first sequence and the second sequence to obtain a scoring matrix. The elements in the scoring matrix are used to represent the number of times the block information in one sequence appears in the other sequence. The first sequence and the second sequence are any two sequences among multiple sequences to be compared.

[0297] Based on the scoring matrix, determine the identical block information in the first sequence and the second sequence, and determine the target content in the file to be compared corresponding to the identical block information.

[0298] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0299] Determine the comparison result between any block information in the first sequence and any block information in the second sequence;

[0300] Based on the comparison result and the adjacent matrix element, determine the matrix element corresponding to the comparison result.

[0301] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0302] If the comparison result is a match, determine the maximum matrix element value from the adjacent matrix elements, increase the single element value based on the maximum matrix element value of the adjacent matrix elements, and determine it as the matrix element corresponding to the comparison result;

[0303] If the comparison result is a mismatch, determine the maximum matrix element value from the adjacent matrix elements, and determine the maximum matrix element value of the adjacent matrix elements as the matrix element corresponding to the comparison result.

[0304] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0305] Step 1: In the first sequence and the second sequence, obtain two block information corresponding to the maximum matrix element value that has not been backtracked and compared in the scoring matrix;

[0306] Step 2: Perform backtracking comparison on the two block information corresponding to the maximum matrix element value: If the two block information match, determine the two block information as identical block information, and return to execute Step 1; if the two block information do not match, return to execute Step 1.

[0307] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0308] Perform block processing on the file to be compared according to a preset length to obtain multiple block contents of the file to be compared;

[0309] Determine the hash value of each block content, and combine the hash values of each file block to obtain the sequence to be compared.

[0310] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0311] Perform word segmentation on the file to be compared, and remove irrelevant characters in the file to be compared after word segmentation to obtain the processed file to be compared;

[0312] Perform chunking on each of the multiple files to be compared to obtain a comparison sequence corresponding to each file to be compared, including:

[0313] Perform chunking on the processed file to be compared to obtain a comparison sequence corresponding to the file to be compared.

[0314] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0315] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in the present application.

[0316] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for comparing multiple files, characterized in that, The method includes: Performing chunking processing on each of multiple files to be compared, obtaining a sequence to be compared corresponding to each file to be compared, where the sequence to be compared includes: multiple chunk information; Comparing a first sequence and a second sequence to obtain a score matrix, where the elements in the score matrix are used to represent the number of times the chunk information in one sequence appears in the other sequence, and the first sequence and the second sequence are any two sequences among the multiple sequences to be compared; Determining the identical chunk information in the first sequence and the second sequence based on the score matrix, and determining the target content in the file to be compared corresponding to the identical chunk information.

2. The method according to claim 1, wherein The comparing the first sequence and the second sequence to obtain a score matrix includes: Determining the comparison result of any chunk information in the first sequence and any chunk information in the second sequence; Determining the matrix element corresponding to the comparison result based on the comparison result and adjacent matrix elements.

3. The method according to claim 2, wherein The determining the matrix element corresponding to the comparison result based on the comparison result and adjacent matrix elements includes: If the comparison result is a match, determining the maximum matrix element value from the adjacent matrix elements, increasing the single-element value based on the maximum matrix element value of the adjacent matrix elements, and determining it as the matrix element corresponding to the comparison result; If the comparison result is a mismatch, determining the maximum matrix element value from the adjacent matrix elements, and determining the maximum matrix element value of the adjacent matrix elements as the matrix element corresponding to the comparison result.

4. The method according to claim 1, wherein The determining the identical chunk information in the first sequence and the second sequence based on the score matrix, and determining the target content in the file to be compared corresponding to the identical chunk information includes: Step 1: In the first sequence and the second sequence, obtaining two chunk information corresponding to the maximum matrix element value that has not been backtracked and compared in the score matrix; Step 2: Performing backtracking comparison on the two chunk information corresponding to the maximum matrix element value: If the two chunk information match, determining the two chunk information as identical chunk information, and returning to execute Step 1; if the two chunk information do not match, returning to execute Step 1.

5. The method according to claim 1, characterized in that, The chunk information is a hash value; the performing chunking processing on each of multiple files to be compared, obtaining a sequence to be compared corresponding to each file to be compared includes: Chunking the file to be compared according to a preset length to obtain multiple chunk contents of the file to be compared; Determining the hash value of each chunk content, and combining the hash values of each file chunk to obtain the sequence to be compared.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: Performing word segmentation processing on the file to be compared, and removing irrelevant characters in the file to be compared after word segmentation processing to obtain a processed file to be compared; The performing chunking processing on each of multiple files to be compared, obtaining a sequence to be compared corresponding to each file to be compared includes: Performing chunking processing on the processed file to be compared to obtain a sequence to be compared corresponding to the file to be compared.

7. A comparison device for multiple files, characterized in that, The device includes: A chunking module, configured to perform chunking processing on each of multiple files to be compared, so as to obtain a sequence to be compared corresponding to each of the files to be compared, where the sequence to be compared includes: a plurality of chunk information; A comparison module, configured to compare a first sequence and a second sequence to obtain a score matrix, where an element in the score matrix is used to represent the number of times that the chunk information in one sequence appears in the other sequence, and the first sequence and the second sequence are any two sequences among the multiple sequences to be compared; A determination module, configured to determine the identical chunk information in the first sequence and the second sequence based on the score matrix, and determine the target content in the file to be compared corresponding to the identical chunk information.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.