Text duplicate removal method and device, equipment and storage medium
By performing word segmentation and hashing operations on text files, and selecting representative hash values for deduplication, the problem of resource waste and inefficiency caused by duplicate content in text files is solved, achieving efficient text deduplication and pre-trained model training.
Patent Information
- Application Number
- CN202410592403.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-13
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies suffer from resource waste and low deduplication efficiency when training pre-trained models due to the presence of a large amount of duplicate content in text files.
By segmenting the text file to be deduplicated, a set of word groups is generated. M hash generation methods are used to perform hash operations on each set of word groups. M representative target hash values are selected for deduplication, avoiding the need to compare characters one by one.
It improves the deduplication efficiency of text files, reduces resource waste, and enhances the training efficiency of pre-trained models.
Smart Images

Figure CN120951979A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a text deduplication method, apparatus, device, and storage medium. Background Technology
[0002] With the development of artificial intelligence technology, the internet contains a massive amount of text files, many of which contain repetitive content. Directly using text files with repetitive content to perform tasks would result in wasted resources. For example, if the text file used to train a pre-trained model contains repetitive content, directly using the text file to train the model would waste the device's processing resources.
[0003] Therefore, text files need to be deduplicated before use. Currently, this is done by comparing each character in the text file to be deduplicated and then deduplicating the text file based on the comparison results. In practice, it has been found that this method of text deduplication is time-consuming and results in low efficiency. Summary of the Invention
[0004] This application provides a text deduplication method, apparatus, device, and storage medium that can improve the deduplication efficiency of text files.
[0005] This application provides a text deduplication method, including:
[0006] Perform word segmentation on the N text files to be deduplicated, and obtain the word set corresponding to each of the N text files; N is an integer greater than 1;
[0007] Based on M hash generation methods, perform hash operations on each word in each of the above word set to obtain M sets of hash values corresponding to each of the above word set; a set of hash values includes the hash values obtained by performing hash operations on each word in a word set using one hash generation method; M is an integer greater than 1;
[0008] From the M sets of hash values corresponding to each of the above phrase sets, select M target hash values that are representative of each of the above phrase sets; one target hash value is used to represent a set of hash values;
[0009] Based on the M target hash values corresponding to the N word sets, the words in the above N word sets are deduplicated to obtain the deduplicated word set. Based on the above deduplicated word set, a deduplicated text file is generated.
[0010] One embodiment of this application provides a text deduplication device, including:
[0011] The first processing module is used to perform word segmentation on the N text files to be deduplicated, and obtain the word set corresponding to each of the N text files; N is an integer greater than 1.
[0012] The operation module is used to perform hash operations on each word in each of the above word set according to M hash generation methods to obtain M sets of hash values corresponding to each of the above word set; a set of hash values includes the hash values obtained by performing hash operations on each word in a word set using one hash generation method; M is an integer greater than 1;
[0013] The selection module is used to select M representative target hash values for each of the above-mentioned phrase sets from the M sets of hash values corresponding to each phrase set; one target hash value is used to represent a set of hash values.
[0014] The second processing module is used to perform deduplication on the phrases in the N phrase sets according to the M target hash values corresponding to the N phrase sets respectively, to obtain the deduplicated phrase sets, and to generate a deduplicated text file based on the deduplicated phrase sets.
[0015] In one possible implementation, the selection module is further configured to perform the following operations:
[0016] The smallest hash value among the M hash values corresponding to each of the above-mentioned phrase sets is determined as the M representative target hash values for each of the above-mentioned phrase sets; or,
[0017] The maximum hash value from the M hash values corresponding to each of the above phrase sets is determined as the M representative target hash values for each of the above phrase sets; or,
[0018] The average hash value corresponding to each of the M hash values for each of the above phrase sets is determined as the M representative target hash values for each of the above phrase sets.
[0019] In one possible implementation, the second processing module described above is further configured to perform the following operations:
[0020] According to the bucketing rule, the M target hash values corresponding to the i-th word set are bucketed to obtain A hash buckets corresponding to the i-th word set. The bucketing rule is determined based on the M hash generation methods. Each hash bucket includes at least one target hash value, where i is a positive integer less than or equal to M and A is an integer greater than 1.
[0021] Based on the A hash buckets corresponding to the N word sets mentioned above, candidate word set pairs with hash matching relationships are selected from the N word set sets; a candidate word set pair includes two word sets with hash matching relationships.
[0022] Based on the above candidate word set pairs, the word sets in the above N word set are deduplicated to obtain the deduplicated word set.
[0023] In one possible implementation, the number of candidate word pair sets is at least two; the second processing module is also used to perform the following operations:
[0024] Based on the corresponding word groups in the two word groups of each of the above candidate word group sets, determine the first similarity between the two word groups in each of the above candidate word group sets;
[0025] From at least two candidate word pair sets, select the candidate word pair set with the first similarity greater than the target similarity threshold, and use it as the target word pair set;
[0026] Based on the above target word set pairs, the word pairs in the above N word set are deduplicated to obtain the deduplicated word set.
[0027] In one possible implementation, the number of the target phrase set pairs is at least one; the second processing module is further configured to perform the following operations:
[0028] Construct a connected graph based on at least one pair of target word sets; the number of such connected graphs is at least one, and each such connected graph includes at least two nodes; one node corresponds to one word set; nodes corresponding to two word sets belonging to the same pair of target word sets are connected by edges;
[0029] From each node in each of the above connected graphs, select the target nodes that are representative for each of the above connected graphs;
[0030] The set of words corresponding to the target node of each of the above connected graphs and the set of remaining words are determined as the set of words after deduplication; the set of remaining words is the set of words in the above N set of words excluding the set of words in each of the above at least one target set of words.
[0031] In one possible implementation, each node includes data attribute information corresponding to a set of word groups; the second processing module is also used to perform the following operations:
[0032] Based on the data attribute information of each phrase set in the above target phrase set pair, the priority of each phrase set in the above target phrase set pair is determined;
[0033] Based on the priority of each phrase set in the above target phrase set pair, select representative target nodes for each of the above connected graphs from each node in each of the above connected graphs.
[0034] In one possible implementation, the above-described text deduplication device is also used to perform the following operations:
[0035] From the above at least two phrase pairs, select at least one phrase pair for testing as the test phrase pair, and obtain the labeled phrase similarity relationship for each test phrase pair;
[0036] Based on the corresponding word groups in the two word groups of each of the above test word group sets, determine the second similarity between the two word groups in the above test word group set pair;
[0037] Based on the candidate similarity threshold and the second similarity mentioned above, the predicted word similarity relationship between the two word sets in the above test word set pair is determined;
[0038] Based on the above-mentioned labeled word similarity relationships and the above-mentioned predicted word similarity relationships, the accuracy of the above-mentioned candidate similarity thresholds is determined;
[0039] Based on the accuracy of the aforementioned candidate similarity thresholds and the aforementioned candidate similarity thresholds, the aforementioned target similarity thresholds are determined.
[0040] In one possible implementation, the first processing module described above is further configured to perform the following operations:
[0041] Obtain the language categories corresponding to the N text files to be deduplicated;
[0042] Based on the language categories of the N text files, determine the word granularity parameters corresponding to the N text files.
[0043] Based on the word granularity parameters corresponding to the above N text files, word segmentation is performed on the above N text files to obtain the word group sets corresponding to the above N text files.
[0044] In one possible implementation, the above-described text deduplication device is also used to perform the following operations:
[0045] Based on the hash function, perform hash operation on the initial text files in the initial file set to obtain the first hash value corresponding to each of the initial text files in the initial file set.
[0046] Based on the first hash value mentioned above, duplicate initial text files within the initial file set are deleted to obtain candidate text files;
[0047] Based on the hash function described above, the text segments in the candidate text files are hashed to obtain the second hash values corresponding to the text segments in the candidate text files respectively.
[0048] Based on the second hash value mentioned above, duplicate text segments in the candidate text files are deleted, resulting in N text files to be deduplicated.
[0049] One embodiment of this application provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the method described above.
[0050] One embodiment of this application provides a computer storage medium storing a computer program, which, when executed by a processor, performs the steps of the above-described method.
[0051] One aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0052] In this application, each text file is divided into a smaller set of word groups. Based on the word group set and M hash generation methods, M representative target hash values are generated for each text file. Only the M target hash values of each text file need to be compared to perform deduplication on the word group set corresponding to each text file. It is not necessary to compare the characters in each text file one by one, thus improving the efficiency of text file deduplication. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the embodiments of this application or the background art, the accompanying drawings used in the embodiments of this application or the background art will be described below.
[0054] Figure 1 This application provides a text deduplication system based on a text deduplication method.
[0055] Figure 2 This is a flowchart illustrating a text deduplication method provided in an embodiment of this application;
[0056] Figure 3 This is a flowchart illustrating another text deduplication method provided in the embodiments of this application;
[0057] Figure 4 This is a flowchart illustrating another text deduplication method provided in the embodiments of this application;
[0058] Figure 5 This is a flowchart illustrating another text deduplication method provided in the embodiments of this application;
[0059] Figure 6 This is a schematic diagram of target hash value bucketing provided in an embodiment of this application;
[0060] Figure 7 This is a flowchart illustrating another text deduplication method provided in the embodiments of this application;
[0061] Figure 8 This is a schematic diagram of the structure of a text deduplication device provided in an embodiment of this application;
[0062] Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0063] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0064] This application primarily relates to artificial intelligence (AI) technology. AI is the theory, methods, techniques, and application systems that utilize digital computers or computers-controlled machines to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. AI technology is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics.
[0065] Pre-trained models (PTMs), also known as foundational models or large models, refer to deep neural networks (DNNs) with a large number of parameters. These DNNs are trained on massive amounts of unlabeled data, leveraging the function approximation capabilities of large-parameter DNNs to extract common features from the data. Through fine-tuning and efficient parameter fine-tuning techniques, PTMs are suitable for downstream tasks. Therefore, pre-trained models can achieve ideal results in small-shot or zero-shot scenarios. PTMs can be categorized according to the data modality they process, such as language models (ELMO, BERT, GPT), visual models (swin-transformer, ViT, V-MOE), speech models (VALL-E), and multimodal models (ViBERT, CLIP, Flamingo, Gato). Multimodal models refer to models that establish feature representations for two or more data modalities. Pre-trained models are important tools for outputting AI-generated content (AIGC) and can also serve as a general interface connecting multiple specific task models.
[0066] The aforementioned massive amount of unlabeled data may refer to text files obtained by computer devices from web pages. This application provides a text deduplication method that can reduce the repetition rate of text files, thereby improving the training efficiency of pre-trained models.
[0067] To facilitate a clearer understanding of this application, we will first introduce a text deduplication system that implements the text deduplication method of this application, such as... Figure 1 As shown, the text deduplication system includes a server 10 and terminals. In this application, the number of terminals can be one or more. Figure 1 The following example illustrates the concept of two terminals, namely terminal 11 and terminal 12.
[0068] The server 10 is connected to terminals 11 and 12 via a network, enabling data interaction between the server 10 and terminals 11 and 12. The server 10 can obtain an initial set of files from a webpage, perform deduplication on the initial text files in the initial set, and obtain deduplicated text files.
[0069] The initial file set mentioned above can be obtained by terminal 11 from a webpage, or it can be obtained by terminal 12 from a webpage. Server 10 can receive the initial file set sent by terminal 11 or terminal 12, and perform deduplication on the initial text files in the initial file set to obtain deduplicated text files. Specifically, the deduplication process of the initial file set can also be implemented by each terminal.
[0070] It should be noted that the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The server can be a single physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Various user terminals and servers can be directly or indirectly connected via wired or wireless communication; this application does not impose any restrictions on this.
[0071] In one embodiment, via Figure 1 The text deduplication system in this application can implement the text deduplication method described in this application, such as... Figure 2 As shown. Figure 2 This is a flowchart illustrating a text deduplication method provided in an embodiment of this application. Figure 2 The text deduplication process is illustrated using server 10 as an example.
[0072] like Figure 2 As shown, the text deduplication method may include the following steps S21 to S25:
[0073] S21. Obtain the initial file set; Server 10 can obtain the initial text files corresponding to multiple network pages from multiple network pages in the network, and determine the initial text files corresponding to multiple network pages as the initial file set.
[0074] For example, taking the server 10 for deduplicating the training data of the pre-trained model as an example, if the pre-trained model is used in the legal field, the server 10 can obtain multiple text contents related to the legal field from multiple web pages related to the legal field on the network, and determine the multiple text contents as the initial file set.
[0075] It should be noted that each webpage can include text content, images, etc., and each webpage also carries network data information such as Internet Protocol (IP) address, Uniform Resource Locator (URL) data, and domain name. The initial file set includes multiple initial text files, each containing the text content of a webpage. Each initial text file carries data attribute information, which can refer to the network data information of the corresponding webpage, including the IP address, URL data, and domain name of the corresponding webpage.
[0076] Since text content may be repeated between two different web pages, and also within the text content of a single web page, duplicate text paragraphs may exist. Therefore, duplicate initial text files exist in the initial file set. Duplicate initial text files in the initial file set can refer to any of the following situations: 1. At least two initial text files in the initial file set have identical text content; 2. Any initial text file in the initial file set contains duplicate text paragraphs; 3. At least two initial text files in the initial file set have similar text content, meaning their text content has the same semantic meaning.
[0077] For example, web page A1 contains text content B1, and web page A2 contains text content B2, with text content B1 and text content B2 being identical. When server 10 obtains the initial text file C1 corresponding to web page A1 and the initial text file C2 corresponding to web page A2, since initial text file C1 contains text content B1 and initial text file C2 contains text content B2, there is duplicate text content between initial text files C1 and C2. Therefore, the initial text files in the initial file set P determined based on initial text files C1 and C2 are duplicated.
[0078] For example, consider at least two initial text files, including initial text file C3, initial text file C4, and initial text file C5. The text content of initial text file C3 is "He likes to learn mathematics", the text content of initial text file C4 is "He loves learning mathematics", and the text content of initial text file C5 is "He likes to learn Chinese". Initial text files C3 and C4 have the same semantics, initial text files C3 and C5 have different semantics, and initial text files C4 and C5 have different semantics.
[0079] S22. Delete text files with duplicate text content; Server 10 can filter out text files with duplicate text content from the initial file set, perform deduplication on the filtered initial text files, and delete the text files with duplicate text content in the initial file set to obtain candidate text files; the number of candidate text files can be multiple.
[0080] The process of selecting text files with duplicate text content from the initial file set can include any of the following: 1. Server 10 can calculate the first hash value of each initial text file in the initial file set, and select initial text files with the same first hash value from the initial file set based on the first hash value of each initial text file, thus identifying the selected initial text files as text files with duplicate text content; 2. Server 10 can select keywords from each initial text file in the initial file set, and select initial text files with the same keywords from the initial file set based on the keywords of each initial text file, thus identifying the selected initial text files as text files with duplicate text content; 3. Server 10 can randomly select the same number of text segments from each initial text file in the initial file set, and select initial text files with the same selected text segments from the initial file set based on the selected text segments in each initial text file, thus identifying the selected initial text files as text files with duplicate text content.
[0081] S23. Delete duplicate text segments; Server 10 can filter out duplicate text segments from candidate text files, perform deduplication on the filtered text segments, and thus delete the duplicate text segments from the candidate text files, obtaining N text files to be deduplicated; N is an integer greater than 1.
[0082] The process of filtering out duplicate text segments from candidate text files can include any of the following: 1. Server 10 can calculate the second hash value corresponding to each text segment in each candidate text file, and based on the second hash value of each text segment in each candidate text file, filter out text segments with the same second hash value from each candidate text file, and determine the filtered text segments as duplicate text segments; 2. Server 10 can filter out the keywords corresponding to each text segment in each candidate text file, and based on the keywords of each text segment in each candidate text file, filter out text segments with the same keywords from each candidate text file, and determine the filtered text segments as duplicate text segments.
[0083] S24. Delete similar text files; Server 10 can filter out text files with the same semantics from the N text files to be deduplicated, and perform deduplication on the filtered text files to obtain similar text files from the N text files to be deduplicated.
[0084] The process of selecting semantically identical text files from N text files to be deduplicated can include any of the following: 1. Server 10 can perform semantic analysis on each of the N text files, extract key text from each text file, where each key text represents the text content of a text file, and select text files with the same key text from the N text files based on the key text of each text file, thus identifying the selected text files as semantically identical text files; 2. Server 10 can perform word segmentation on each of the N text files to obtain a set of word groups for each text file; and determine the semantically identical text files from the N text files based on the word groups in the set of word groups of each text file.
[0085] S25. Obtain the deduplicated text file; Server 10 performs deduplication processing on the selected text file to obtain the deduplicated filtered text file; Server 10 can determine the deduplicated filtered text file and the unfiltered text files from N text files as the deduplicated text file.
[0086] Please refer to Figure 3 , Figure 3 This application provides a flowchart illustrating another text deduplication method according to an embodiment. This method can be... Figure 1 It can be executed by any terminal in the system, or by... Figure 1 The server in the middle can be used to execute it, or it can be executed by... Figure 1 The terminal and server in the process work together to execute the text deduplication method. The device used to execute this method in this application can be collectively referred to as a computer device. Specifically, the computer device can employ... Figure 3 The method described above executes step S24. Figure 3 Taking the deletion of similar text files using a computer device as an example, the method may include the following steps:
[0087] S31. Generate a set of word groups; computer equipment can perform word segmentation on N text files to obtain a set of word groups corresponding to each of the N text files.
[0088] S32. Select candidate word set pairs; the computer device can use M hash generation methods to perform hash operations on each word in each word set, generating M sets of hash values corresponding to N word sets respectively; the computer device can also select M representative target hash values from the M sets of hash values corresponding to each word set, so that candidate word set pairs in N word sets can be determined based on the M representative target hash values for each word set; a candidate word set pair includes two word sets with hash matching relationship; the number of candidate word set pairs is at least two.
[0089] S33. Determine the target similarity threshold; the computer device can select at least one test word pair from at least two candidate word pair sets for testing; and can determine the target similarity threshold for the candidate word pair set based on the test word pair set.
[0090] S34. Delete similar text files; The computer device can calculate the first similarity between the word sets in each of the candidate word set pairs in at least two candidate word set pairs, select word set pairs from the candidate word set pairs whose first similarity is greater than the target similarity threshold, and perform deduplication processing on the text files corresponding to the word sets in the selected word set pairs, that is, similar text files among the N text files to be deduplicated.
[0091] In this application, each text file is divided into a smaller set of word groups. Based on the word group set and M hash generation methods, M representative target hash values are generated for each text file. Only the M target hash values of each text file need to be compared to perform deduplication on the word group set corresponding to each text file. It is not necessary to compare the characters in each text file one by one, thus improving the efficiency of text file deduplication.
[0092] Further, please see Figure 4 This is a flowchart illustrating another text deduplication method provided in an embodiment of this application. Figure 4 As shown, this method can be derived from... Figure 1 It can be executed by any terminal in the system, or by... Figure 1 The server in the middle can be used to execute it, or it can be executed by... Figure 1 The method involves a terminal and a server working together to execute the text deduplication method. The device used to execute this method in this application can be collectively referred to as a computer device. The method may include the following steps:
[0093] S401. Perform word segmentation on the N text files to be deduplicated, and obtain the word set corresponding to each of the N text files.
[0094] It should be noted that the N text files can be obtained by the computer device after deleting duplicate text files and paragraphs from the initial file set, or they can be obtained by the computer device from web pages. Therefore, each text file corresponds to the text content of a web page, and each text file also carries the data attribute information of the corresponding web page. N is an integer greater than 1.
[0095] It should be noted that the above-mentioned word segmentation processing can refer to splitting a text file into multiple word groups of the same preset length, or it can refer to splitting a text file into multiple word groups based on the semantics of each word or character in the text file. The preset word group length can be a default value set in the computer device, or it can be configured in the computer device by the deduplication operator.
[0096] For example, given N text files, including text file D1 whose text content is "He likes learning math," with a preset phrase length of "3," a computer device can segment "He likes learning math" into phrases E1 "He likes," E2 "Likes to learn," E3 "Likes to learn math," and E4 "Learns math." The computer device can also segment "He likes learning math" into phrases E5 "He likes" and E6 "Learns math." Specifically, in this application, when segmenting phrases, each phrase may include part of the content from the previous phrase, or each phrase may not include the content from the previous phrase.
[0097] For example, among N text files, there is a text file D2 whose text content is "He likes Chinese class". Based on the semantics of the words, the computer device can segment "He likes Chinese class" into the phrase E7 "he", the phrase E8 "likes", and the phrase E9 "Chinese class".
[0098] In this application, after the computer device obtains N text files with deduplication, it can read the text content from each text file, and according to the text content, divide the text content of each text file into multiple word groups; and merge the multiple word groups corresponding to each text file to obtain the word group set corresponding to each text file.
[0099] S402. Based on the M hash generation methods, perform hash operations on each word group in each word group set to obtain the M hash values corresponding to each word group set.
[0100] It should be noted that hash generation methods can include Secure Hash Algorithm 1 (SHA-1), Secure Hash Algorithm 2 (SHA-2), Secure Hash Algorithm 3 (SHA-3), Message-Digest Algorithm 5 (MD5), and so on. A set of hash values consists of hash values obtained by hashing each word in a set of words using a hash generation method; M is an integer greater than 1.
[0101] In this application, the computer device may store a set of hash generation methods, which may include multiple hash generation methods. When the computer device obtains word sets corresponding to N text files, it can randomly select M hash generation methods from the set. Then, using the first hash generation method among the M methods, the hash value of each word in each word set is calculated to obtain the first hash value corresponding to each word set; using the second hash generation method among the M methods, the hash value of each word in each word set is calculated to obtain the second hash value corresponding to each word set; and so on, until the Mth hash generation method among the M methods is used to calculate the hash value of each word in each word set to obtain the Mth hash value corresponding to each word set. Thus, the computer device obtains the M hash values corresponding to each word set.
[0102] It should be noted that the hash generation methods in the above set of hash generation methods can be pre-configured in the computer equipment by the deduplication operator. M can be an integer greater than 1; the value of M can be pre-configured in the computer equipment by the deduplication operator, or the value of M can be determined by the computer equipment based on the number of phrases in the N phrase sets.
[0103] Specifically, a larger value for M indicates that the computer performs more calculations for each set of words. In this case, if the number of words in the N sets of words is large, the computer's processing time will be longer, leading to decreased efficiency. Therefore, a smaller value for M can be chosen when the number of words in the N sets is large, and a larger value for M can be chosen when the number of words in the N sets is small. The computer can be configured with a low value range and a high value range for M. When the number of words in the N sets is greater than or equal to a threshold, a random value is selected from the low value range as M; when the number of words in the N sets is less than the threshold, a random value is selected from the high value range as M.
[0104] For example, the first value range is [150, 200], the second value range is [200, 150], and the word group number threshold is 10000. When the number of word groups in the N word group set is greater than or equal to 10000, a random value is obtained from the first value range as M; when the number of word groups in the N word group set is less than 10000, a random value is obtained from the second value range as M.
[0105] S403. From the M sets of hash values corresponding to each set of words, select M target hash values that are representative of each set of words.
[0106] It should be noted that if two different text files contain duplicate text content, or if they have the same semantics, it indicates that they share the same word groups. When the duplication rate between two text files is high, they contain many identical word groups. In this case, the M hash values corresponding to each text file will have many identical hash values. Comparing each of the M hash values for each text file individually would be computationally time-consuming. Therefore, we can obtain M representative target hash values for each word group from the M hash values corresponding to each text file. The higher the duplication rate between two text files, the higher the duplication rate among the M target hash values; conversely, the lower the duplication rate, the lower the duplication rate among the M target hash values.
[0107] In this application, the computer device can select a first target hash value that is representative of each word set from the first group of hash values corresponding to each word set; and select a second target hash value that is representative of each word set from the second group of hash values corresponding to each word set. This process continues until the computer device can select a Mth target hash value that is representative of each word set from the Mth group of hash values corresponding to each word set.
[0108] It should be noted that a target hash value is used to represent a set of hash values; the selection of the Mth target hash value that is representative of each word set can include any of the following: 1. Obtain the Mth largest hash value from the Mth set of hash values corresponding to each word set; 2. Obtain the Mth smallest hash value from the Mth set of hash values corresponding to each word set; 3. Obtain the Mth median hash value from the Mth set of hash values corresponding to each word set; 4. Obtain the Mth average hash value from the Mth set of hash values corresponding to each word set.
[0109] S404. Based on the M target hash values corresponding to the N word sets, perform deduplication on the word sets in the N word sets to obtain the deduplicated word set. Based on the deduplicated word set, generate a deduplicated text file.
[0110] In this application, a computer device can compare the M target hash values corresponding to any two word sets in the N word set according to the operation order of M hash generation methods, obtain the comparison result between the two word sets in the N word set, perform deduplication processing on the word sets in the N word set according to the comparison result, obtain the deduplicated word set, and generate a deduplicated text file based on the deduplicated word set.
[0111] In this application, each text file is divided into a smaller set of word groups. Based on the word group set and M hash generation methods, M representative target hash values are generated for each text file. Only the M target hash values of each text file need to be compared to perform deduplication on the word group set corresponding to each text file. It is not necessary to compare the characters in each text file one by one, thus improving the efficiency of text file deduplication.
[0112] Further, please see Figure 5 This is a flowchart illustrating another text deduplication method provided in an embodiment of this application. Figure 5 As shown, this method can be derived from... Figure 1 It can be executed by any terminal in the system, or by... Figure 1 The server in the middle can be used to execute it, or it can be executed by... Figure 1The method involves a terminal and a server working together to execute the text processing method. The device used to execute this text processing method in this application can be collectively referred to as a computer device. The method may include the following steps:
[0113] S501. Perform word segmentation on the N text files to be deduplicated, and obtain the word set corresponding to each of the N text files.
[0114] Optionally, based on the hash function, the initial text files in the initial file set are hashed to obtain the first hash value corresponding to each initial text file in the initial file set; based on the first hash value, duplicate initial text files in the initial file set are deleted to obtain candidate text files; based on the hash function, the text segments in the candidate text files are hashed to obtain the second hash value corresponding to each text segment in the candidate text files; based on the second hash value, duplicate text segments in the candidate text files are deleted to obtain N text files to be deduplicated.
[0115] It should be noted that the hash functions mentioned above include Message-Digest Algorithm 5 (MD5) values, Secure Hash Algorithm (SHA), Hash-based Message Authentication Code (HMAC), etc. The initial file set mentioned above can be obtained by the computer device from web pages. The computer device can store the text content obtained from a web page in an initial text file. Each initial text file can include text content (i.e., the text content from the corresponding web page) and data attribute information. The data attribute information of each initial text file can include the Internet Protocol (IP) address of the corresponding web page, the Uniform Resource Locator (URL) data of the corresponding web page, the domain name of the corresponding web page, etc.
[0116] In this application, a computer device can calculate a first hash value of the text content of each initial text file to obtain the first hash value of each initial text file; based on the first hash value of each initial text file, duplicate initial text files are filtered out from the initial file set. The computer device can set a priority list to determine the priority of each initial text file based on the URL corresponding to each initial text file; based on the priority of each initial text file, one initial text file is filtered out from the duplicate initial text files, and the duplicate initial text files other than the filtered initial text file are deleted, resulting in candidate text files. The number of candidate text files can be multiple. At this point, duplicate initial text files in the initial file set have been deleted.
[0117] Furthermore, the computer device can perform paragraph splitting on each candidate text file. Specifically, based on the text content of each candidate text file, the computer device divides each candidate text file into at least one text paragraph according to the paragraph divisions within the text content, resulting in a set of text paragraphs corresponding to each candidate text file; the aforementioned set of text paragraphs includes at least one text paragraph. The computer device can calculate the second hash value of each text paragraph in each set of text paragraphs, and based on the second hash value of each text paragraph in each set of text paragraphs, filter out duplicate text paragraphs from each set of text paragraphs. The computer device can then filter out any text paragraph from the duplicate text paragraphs in each set of text paragraphs, and delete the text paragraphs other than the selected paragraph, resulting in N text files to be deduplicated.
[0118] It should be noted that the aforementioned duplicate initial text files can refer to at least two text files in the initial file set with the same first hash value; the aforementioned duplicate text segments can refer to at least two text segments in the same candidate text file with the same second hash value. A priority list can be set in the computer device. The process of selecting an initial text file from the duplicate initial text files based on the priority of each initial text file can include any of the following: 1. When the priorities of all text files in the duplicate initial text files are unequal, select the initial text file corresponding to the highest priority; 2. When the priorities of all text files in the duplicate initial text files are equal, select the initial text file with the most characters; 3. When the priorities of all text files in the duplicate initial text files are equal, iterate through the priorities of each text file in the duplicate initial text files. If there is only one highest priority among the duplicate initial text files, select the initial text file corresponding to the highest priority; if there are multiple highest priorities among the duplicate initial text files, select the initial text file with the most characters.
[0119] Optionally, obtain the language categories corresponding to the N text files to be deduplicated; determine the word granularity parameters corresponding to the N text files based on the language categories; and perform word segmentation on the N text files based on the word granularity parameters to obtain the word set corresponding to the N text files.
[0120] It should be noted that the language category can include Chinese, English, etc. When the language category of the text file is Chinese, the word granularity parameter corresponding to the text file (i.e., the word granularity parameter for Chinese) is used to instruct the computer device to identify a Chinese character as a word with a length corresponding to one word; when the language category of the text file is Chinese, the word granularity parameter corresponding to the text file (i.e., the word granularity parameter for English) is used to instruct the computer device to identify an English word as a word with a length corresponding to one word, and to perform word segmentation processing on the text file.
[0121] In this application, the computer device can adopt a text classification algorithm to identify N text files and obtain the language categories corresponding to the N text files respectively. When the language category of a certain text file is Chinese, obtain the word granularity parameters for Chinese, determine a Chinese character as a word corresponding to a word length, and form a phrase set according to each word in the text file based on the preset phrase length. When the language category of a certain text file is English, obtain the word granularity parameters for English, determine an English word as a word corresponding to a word length, and form a phrase set according to each word in the text file based on the preset phrase length.
[0122] For example, among the N text files, there are text file B and text file C. The text content of text file E is "他喜欢学习数学", and the text content of text file B is "He enjoys learning mathematics", and the preset phrase length is 3. Since the language type of text file B is Chinese, when the computer device performs word segmentation on text file E, the computer device can determine "他", "喜", "欢", "学", "习", "数", and "学" respectively as words corresponding to a word length; furthermore, according to "他", "喜", "欢", a phrase 1e "他喜欢" with a preset phrase length of 3 can be formed. By performing the above operations in sequence, phrases 2e "喜欢学", 3e "欢学习", 4e "学习数", and 5e "习数学" can be obtained. Since the language type of text file C is English, when the computer device performs word segmentation on text file C, the computer device can determine "He", "enjoys", "learning", and "mathematics" respectively as words corresponding to a word length. Furthermore, according to "He", "enjoys", "learning", a phrase 1d "He enjoyslearning" with a preset phrase length of 3 can be formed. By repeating the above operations, phrase 2d "enjoys learning mathematics" can be obtained.
[0123] S502. Perform a hashing operation on each phrase in each phrase set according to M hashing generation methods to obtain M groups of hash values corresponding to each phrase set.
[0124] S503. Select M target hash values that are representative for each phrase set from the M groups of hash values corresponding to each phrase set.
[0125] Optionally, the minimum hash value among the M hash values corresponding to each of the above-mentioned word set is determined as one of the M representative target hash values for each of the above-mentioned word set; or, the maximum hash value among the M hash values corresponding to each of the above-mentioned word set is determined as one of the M representative target hash values for each of the above-mentioned word set; or, the average hash value among the M hash values corresponding to each of the above-mentioned word set is determined as one of the M representative target hash values for each of the above-mentioned word set.
[0126] In this application, a computer device can filter out the maximum hash value from the first group of hash values corresponding to each word set, and determine the maximum hash value from the first group of hash values corresponding to each word set as the first representative target hash value for each of the aforementioned word sets. This process continues until the maximum hash value from the Mth group of hash values corresponding to each word set is selected, and the maximum hash value from the Mth group of hash values corresponding to each word set is determined as the Mth representative target hash value for each of the aforementioned word sets.
[0127] Alternatively, the computer device can filter out the smallest hash value from the first group of hash values corresponding to each word set, and determine the smallest hash value in the first group of hash values corresponding to each word set as the first representative target hash value for each of the aforementioned word sets. This continues until the smallest hash value in the Mth group of hash values corresponding to each word set is selected, and the smallest hash value in the Mth group of hash values corresponding to each word set is determined as the Mth representative target hash value for each of the aforementioned word sets.
[0128] Alternatively, the computer device can average the hash values of the first group corresponding to each word set to obtain the average hash value of the first group of hash values for each word set. This average hash value is then used as the first representative target hash value for each of the aforementioned word sets. This process continues until the average hash values of the Mth group corresponding to each word set are averaged to obtain the average hash value of the Mth group of hash values for each word set. This average hash value is then used as the Mth representative target hash value for each of the aforementioned word sets.
[0129] S504. According to the bucketing rules, the M target hash values corresponding to the i-th word set are bucketed to obtain A hash buckets corresponding to the i-th word set.
[0130] It should be noted that the bucketing rule is determined based on the M hash generation methods. Each hash bucket includes at least one target hash value, where i is a positive integer less than or equal to M, and A is an integer greater than 1. This bucketing rule is used to indicate which hash bucket the target hash value generated based on each hash generation method should be placed in for the corresponding word group.
[0131] It should be noted that the range of values for the hash generation methods mentioned above can include a random range for the number of hash buckets and a random range for the number of target hash values within a bucket. Specifically, when the number of words in the N word set is less than the word count threshold, it indicates that the number of words in the N word set is relatively small. The computer device can obtain a larger value from the random range for the number of hash buckets, designated as A. The computer device can also obtain a larger value from the random range for the number of target hash values within a bucket, designated as B. That is, each hash bucket includes B target hash values. When the number of words in the N word set is greater than or equal to the word count threshold, it indicates that the number of words in the N word set is relatively large. The computer device can obtain a smaller value from the random range for the number of hash buckets, designated as A. The computer device can also obtain a smaller value from the random range for the number of target hash values within a bucket, designated as B. After obtaining A and B, the computer device multiplies A and B, resulting in a product M.
[0132] In this application, the computer device can, according to the bucketing rule, determine the first target hash value of the first term set obtained based on the first hash generation method as the first target hash value in the first hash bucket corresponding to the first term set. After the first hash bucket corresponding to the first term set has been fully stored, the target hash values of the first term sets not yet placed in hash buckets are stored in other hash buckets; thus obtaining A hash buckets corresponding to the first term set. Furthermore, the computer device can obtain A hash buckets corresponding to each term set through the above operations.
[0133] like Figure 6 As shown, Figure 6 This is a schematic diagram of a target hash value bucketing method provided in an embodiment of this application. Figure 6 In this example, the phrase set includes phrase set F1, phrase set F2 and phrase set F3, and the hash generation method includes hash generation method G1, hash generation method G2, hash generation method G3 and hash generation method G4. Figure 6In the given set of words, the four target hash values corresponding to phrase set F1 are h11 (based on hash generation method G1), h12 (based on hash generation method G2), h13 (based on hash generation method G3), and h14 (based on hash generation method G4). Similarly, the four target hash values corresponding to phrase set F2 are h21 (based on hash generation method G1), h22 (based on hash generation method G2), h23 (based on hash generation method G3), and h24 (based on hash generation method G4). Finally, the four target hash values corresponding to phrase set F3 are h31 (based on hash generation method G1), h32 (based on hash generation method G2), h33 (based on hash generation method G3), and h34 (based on hash generation method G4).
[0134] Furthermore, taking the bucketing process for the four target hash values corresponding to the phrase set F1 as an example, the four target hash values corresponding to the phrase set F1 can be divided into two hash buckets, namely hash bucket k11 and hash bucket k12. Hash bucket k11 is the first hash bucket corresponding to the phrase set F1, and hash bucket k12 is the second hash bucket corresponding to the phrase set F1. Each hash bucket can store two target hash values. Therefore, target hash value h11 can be placed in the first position of hash bucket k11, and target hash value h12 can be placed in the second position of hash bucket k11. Since hash bucket k11 is full, target hash value h13 can be placed in the first position of hash bucket k12, and target hash value h14 can be placed in the second position of hash bucket k12. Similarly, hash buckets k21 and k22 corresponding to the phrase set F2, and hash buckets k31 and k32 corresponding to the phrase set F3 can be obtained.
[0135] S505. Based on the A hash buckets corresponding to the N word sets, select candidate word set pairs with hash matching relationships from the N word set sets.
[0136] It should be noted that a candidate phrase set includes two phrase sets with hash matching relationships; two phrase sets with hash matching relationships can mean that at least one of the A hash buckets corresponding to the two phrase sets has the same target hash value.
[0137] In this application, a computer device can compare the target hash values in the A hash buckets corresponding to any two word sets in N word sets one by one. When it is determined that at least one of the A hash buckets corresponding to the two word sets has the same target hash value, the computer device can determine that the two word sets are candidate word set pairs with hash matching relationship.
[0138] by Figure 6 For example, Figure 6 In the algorithm, phrase set F1 corresponds to hash buckets k11 and k12, phrase set F2 corresponds to hash buckets k21 and k22, and phrase set F3 corresponds to hash buckets k31 and k32. If h11 in hash bucket k11 is equal to h21 in hash bucket k21, and h12 in hash bucket k11 is equal to h22 in hash bucket k21, then the target hash values in hash buckets k11 and k21 are the same. If h11 in hash bucket k11 is equal to h21 in hash bucket k21, but h12 in hash bucket k11 is not equal to h22 in hash bucket k21, then the target hash values in hash buckets k11 and k21 are different.
[0139] Furthermore, if the target hash values in hash buckets k11 and k21 are the same, and the target hash values in hash buckets k12 and k22 are different, it indicates that at least one of the two hash buckets corresponding to phrase set F1 and phrase set F2 has the same target hash value, and phrase set F1 and phrase set F2 are a candidate phrase set pair.
[0140] S506. Based on the candidate word set pairs, perform deduplication on the word sets in the N word set pairs to obtain the deduplicated word set.
[0141] Optionally, from the above at least two word group pairs, at least one word group pair for testing is selected as the test word group pair, and the labeled word group similarity relationship of each test word group pair is obtained; based on the corresponding word groups in the two word groups of each test word group pair, a second similarity between the two word groups in the above test word group pair is determined; based on the candidate similarity threshold and the above second similarity, a predicted word group similarity relationship between the two word groups in the above test word group pair is determined; based on the labeled word group similarity relationship and the above predicted word group similarity relationship, the accuracy of the above candidate similarity threshold is determined; based on the accuracy of the above candidate similarity threshold and the above candidate similarity threshold, the target similarity threshold is determined.
[0142] It should be noted that the second similarity can be obtained based on a similarity algorithm, which may include the Jaccard similarity coefficient, Euclidean distance, cosine similarity, etc. The annotation of phrase similarity relationships can include both similar and dissimilar relationships; the predicted phrase similarity relationships can also include both similar and dissimilar relationships.
[0143] Specifically, when the second similarity of a pair of test word sets is greater than or equal to the candidate similarity threshold, the predicted word similarity relationship between the two word sets of the test word set pair is determined to be a similar relationship; when the second similarity of a pair of test word sets is less than the candidate similarity threshold, the predicted word similarity relationship between the two word sets of the test word set pair is determined to be a dissimilar relationship.
[0144] In this application, after obtaining at least two candidate word pair sets, the computer device can select at least one word pair set as a test word pair set; and use a similarity algorithm to determine the second similarity between the two word pairs in each test word pair set. At this point, the computer device can select a candidate similarity threshold, and based on the candidate similarity threshold, obtain the predicted word similarity relationship for each test word pair set. Simultaneously, the computer device can send test word pair sets to a terminal. The user corresponding to this terminal can be a deduplication operator who receives the labeled word similarity relationship for each test word pair set sent by the terminal. Then, the labeled word similarity relationship and the predicted word similarity relationship for each test word pair set can be compared. When the labeled word similarity relationship and the predicted word similarity relationship of a test word pair set are the same, the prediction result for that test word pair set is determined to be correct; when the labeled word similarity relationship and the predicted word similarity relationship of a test word pair set are different, the prediction result for that test word pair set is determined to be incorrect. Thus, the prediction result for each test word pair set can be obtained, and the accuracy of the candidate similarity threshold can be determined based on the prediction result for each test word pair set.
[0145] Furthermore, when the accuracy of the candidate similarity threshold is greater than or equal to the accuracy threshold, the candidate similarity threshold is determined as the target similarity threshold. When the accuracy of the candidate similarity threshold is less than the accuracy threshold, the candidate similarity threshold can be appropriately reduced to obtain an adjusted candidate similarity threshold. Based on the adjusted candidate similarity threshold, the predicted word similarity relationship of each test word pair is re-determined, thereby determining the accuracy of the adjusted candidate similarity threshold. Based on the accuracy of the adjusted candidate similarity threshold, the target similarity threshold is re-determined. Through threshold filtering, the target similarity threshold for the candidate word set is obtained, ensuring the accuracy of the similarity threshold and improving the deduplication accuracy of the text file.
[0146] Optionally, based on the corresponding word groups in the two word groups of each of the above candidate word group sets, a first similarity between the two word groups in each candidate word group set pair is determined; from at least two candidate word group set pairs, candidate word group set pairs with a first similarity greater than a target similarity threshold are selected as target word group set pairs; based on the above target word group set pairs, the word groups in the above N word group sets are deduplicated to obtain a deduplicated word group set.
[0147] It should be noted that the first similarity score can be obtained based on a similarity algorithm.
[0148] In this application, the computer device, based on a similarity algorithm, traverses the word groups corresponding to the two word groups in each candidate word group set pair to determine the first similarity between the two word groups in each candidate word group set pair; the candidate word group set pair with the first similarity greater than the target similarity threshold is determined as the target word group set pair; based on the above target word group set pair, the target word group set is determined from the target word group set pair; and the target word group set and the word group set other than the word group set in the target word group set pair from the N word group sets are determined as the deduplicated word group set.
[0149] Optionally, a connected graph is constructed based on at least one target word set pair; the number of such connected graphs is at least one, and each such connected graph includes at least two nodes; one node corresponds to one word set; nodes corresponding to two word sets belonging to the same target word set pair are connected by edges; representative target nodes for each such connected graph are selected from each node in each such connected graph; the word set corresponding to the target node of each such connected graph and the remaining word set are determined as the deduplicated word set; the remaining word set is the word set other than the word sets in each of the above N word sets.
[0150] In this application, the computer device can traverse each target word set pair, identifying the two word sets in each target word set pair as two nodes. Simultaneously, during the traversal, if a determined word set is encountered, no further nodes are determined for that determined word set. Next, the nodes corresponding to the two word sets belonging to the same target word set pair can be connected by an edge, thereby obtaining at least one connected graph. From each connected graph, a representative target node can be selected, and the word set corresponding to the target node, as well as the word sets from the N word sets excluding the word sets in the target word set pair, can be determined as the deduplicated word set.
[0151] Optionally, based on the data attribute information of each phrase set in the target phrase set pair, the priority of each phrase set in the target phrase set pair is determined; based on the priority of each phrase set in the target phrase set pair, representative target nodes for each of the connected graphs are selected from each node in each of the connected graphs.
[0152] It should be noted that a text file can include text content and data attribute information; a text file can also include the time when the text file was retrieved. Therefore, the data attribute information of a phrase set can refer to the data attribute information of the corresponding text file.
[0153] Understandably, a computer device can obtain the data attribute information of the text files corresponding to each word set in a target word set pair. Based on the data attribute information of each word set in the target word set pair, the computer device can determine the URL corresponding to each word set in the target word set pair. According to the priority list configured in the computer device, the priority of each word set in the target word set pair can be determined. Based on the priorities of each word set in the target word set pair, when the priorities of the word sets corresponding to each node in the above connected graph are not equal, the node corresponding to the word set with the highest priority in each connected graph is selected as the target node. When the priorities of the word sets corresponding to each node in each connected graph are not equal, the node corresponding to the word set with the highest priority in each connected graph is selected as the target node. When the priorities of the word sets corresponding to each node in each connected graph are equal, the node corresponding to the word set with the most text characters in each connected graph is selected as the target node. When the priorities of the phrase sets corresponding to each node in a connected graph are equal, if there is only one highest priority among all nodes in the connected graph, then the node corresponding to the phrase set with the highest priority in each connected graph is selected as the target node; if there are multiple highest priority among all nodes in the connected graph, then the node corresponding to the phrase set with the most characters among the multiple highest priority phrase sets in each connected graph is selected as the target node. By setting a priority list, in the process of merging multiple nodes, the node with the highest priority is retained according to the specified priority, thereby improving the deduplication accuracy of text files.
[0154] For example, connected Figure 1 The set includes nodes 1, 2, 3, 4, and 5. The priority of the word set corresponding to node 1 is 5, the priority of the word set corresponding to node 2 is 4, the priority of the word set corresponding to node 3 is 1, the priority of the word set corresponding to node 4 is 5, and the priority of the word set corresponding to node 5 is 2. Due to connectivity... Figure 1 The highest priority in the set is 5, and the nodes corresponding to the highest priority phrase set include nodes 1 and 4. If the computer determines that the text number of the phrase set corresponding to node 1 is 54366 and the text number of the phrase set corresponding to node 4 is 67754, since the text number of the phrase set corresponding to node 4 is greater than the text number of the phrase set corresponding to node 1, the computer can determine that node 4 is a connected node. Figure 1 The target node.
[0155] S507. Generate a deduplicated text file based on the deduplicated set of phrases.
[0156] In this application, each text file is divided into a smaller set of word groups. Based on the word group set and M hash generation methods, M representative target hash values are generated for each text file. Only the M target hash values of each text file need to be compared to perform deduplication on the word group set corresponding to each text file. It is not necessary to compare the characters in each text file one by one, thus improving the efficiency of text file deduplication.
[0157] Please see Figure 7 This is a flowchart illustrating another text deduplication method provided in this application embodiment. This method can be... Figure 1 It can be executed by any terminal in the system, or by... Figure 1 The server in the middle can be used to execute it, or it can be executed by... Figure 1 The terminal and server in the application work together to execute the text processing method. The device used to execute the text processing method in this application can be collectively referred to as a computer device.
[0158] like Figure 7 As shown, after obtaining the initial file set 701, if the data volume of the initial file set 701 is large, performing deduplication on the entire initial file set 701 will consume a significant amount of memory resources on the computer device 700, reducing its operating speed and thus lowering the deduplication efficiency for the entire initial file set 701. Therefore, when the computer device 700 obtains an initial file set 701 with a large data volume, it can group the initial file set 701 into multiple initial file subsets. Figure 7 Taking four initial file subsets as an example, namely initial file subset 710, initial file subset 720, initial file subset 730, and initial file subset 740. Each of initial file subsets 710, 720, 730, and 740 can include multiple initial text files. The aforementioned grouping process can refer to extracting a random number of initial text files from the initial file set 700; each initial text file in the initial file set 701 exists only in any one of the initial file subsets 710, 720, 730, and 740, meaning one initial text file corresponds to one initial file subset, and one initial file subset can include multiple initial text files.
[0159] Furthermore, the computer device 700 can send the initial file subset 710 to the computer device 702, the initial file subset 720 to the computer device 703, the initial file subset 730 to the computer device 704, and the initial file subset 740 to the computer device 705. The computer devices 702, 703, 704, and 705 described above can all employ the steps in the above text deduplication method to deduplicate the received initial file subsets, obtaining deduplicated file subsets 712, 722, 732, and 742 respectively. The computer device 700 can acquire the deduplicated file subsets 712, 722, 732, and 742, and merge them to obtain a merged file set 750.
[0160] Furthermore, the computer device 700 can employ the steps described in the above text deduplication method to deduplicate the text files in the merged file set 750, thereby obtaining deduplicated text files. In this way, the memory consumption of a single computer device during the deduplication process can be reduced, and multiple computer devices simultaneously performing deduplication on the initial file set can improve the efficiency of text file deduplication.
[0161] Please see Figure 8 This is a schematic diagram of the structure of a text deduplication device provided in an embodiment of this application. Figure 8 As shown, the text deduplication device may include:
[0162] The first processing module 811 is used to perform word segmentation on the N text files to be deduplicated, and obtain the word set corresponding to each of the N text files; N is an integer greater than 1.
[0163] The operation module 812 is used to perform hash operations on each word in each of the above word set according to M hash generation methods to obtain M sets of hash values corresponding to each of the above word set; a set of hash values includes hash values obtained by performing hash operations on each word in a word set using one hash generation method; M is an integer greater than 1;
[0164] Selection module 813 is used to select M target hash values that are representative of each of the above-mentioned phrase sets from the M sets of hash values corresponding to each of the above-mentioned phrase sets; one target hash value is used to represent a set of hash values;
[0165] The second processing module 814 is used to perform deduplication processing on the word groups in the above N word group sets according to the M target hash values corresponding to the N word group sets respectively, to obtain the deduplicated word group sets, and to generate a deduplicated text file based on the deduplicated word group sets.
[0166] In one possible implementation, the selection module 813 is further configured to perform the following operations:
[0167] The smallest hash value among the M hash values corresponding to each of the above-mentioned phrase sets is determined as the M representative target hash values for each of the above-mentioned phrase sets; or,
[0168] The maximum hash value from the M hash values corresponding to each of the above phrase sets is determined as the M representative target hash values for each of the above phrase sets; or,
[0169] The average hash value corresponding to each of the M hash values for each of the above phrase sets is determined as the M representative target hash values for each of the above phrase sets.
[0170] In one possible implementation, the second processing module 814 is further configured to perform the following operations:
[0171] According to the bucketing rule, the M target hash values corresponding to the i-th word set are bucketed to obtain A hash buckets corresponding to the i-th word set. The bucketing rule is determined based on the M hash generation methods. Each hash bucket includes at least one target hash value, where i is a positive integer less than or equal to M and A is an integer greater than 1.
[0172] Based on the A hash buckets corresponding to the N word sets mentioned above, candidate word set pairs with hash matching relationships are selected from the N word set sets; a candidate word set pair includes two word sets with hash matching relationships.
[0173] Based on the above candidate word set pairs, the word sets in the above N word set are deduplicated to obtain the deduplicated word set.
[0174] In one possible implementation, the number of candidate word pair sets is at least two; the second processing module 814 is further configured to perform the following operations:
[0175] Based on the corresponding word groups in the two word groups of each of the above candidate word group sets, determine the first similarity between the two word groups in each of the above candidate word group sets;
[0176] From at least two candidate word pair sets, select the candidate word pair set with the first similarity greater than the target similarity threshold, and use it as the target word pair set;
[0177] Based on the above target word set pairs, the word pairs in the above N word set are deduplicated to obtain the deduplicated word set.
[0178] In one possible implementation, the number of the target phrase set pairs is at least one; the second processing module 814 is further configured to perform the following operations:
[0179] Construct a connected graph based on at least one pair of target word sets; the number of such connected graphs is at least one, and each such connected graph includes at least two nodes; one node corresponds to one word set; nodes corresponding to two word sets belonging to the same pair of target word sets are connected by edges;
[0180] From each node in each of the above connected graphs, select the target nodes that are representative for each of the above connected graphs;
[0181] The set of words corresponding to the target node of each of the above connected graphs and the set of remaining words are determined as the set of words after deduplication; the set of remaining words is the set of words in the above N set of words excluding the set of words in each of the above at least one target set of words.
[0182] In one possible implementation, each node includes data attribute information corresponding to a set of word groups; the second processing module 814 is further configured to perform the following operations:
[0183] Based on the data attribute information of each phrase set in the above target phrase set pair, the priority of each phrase set in the above target phrase set pair is determined;
[0184] Based on the priority of each phrase set in the above target phrase set pair, select representative target nodes for each of the above connected graphs from each node in each of the above connected graphs.
[0185] In one possible implementation, the above-described text deduplication device is also used to perform the following operations:
[0186] From the above at least two phrase pairs, select at least one phrase pair for testing as the test phrase pair, and obtain the labeled phrase similarity relationship for each test phrase pair;
[0187] Based on the corresponding word groups in the two word groups of each of the above test word group sets, determine the second similarity between the two word groups in the above test word group set pair;
[0188] Based on the candidate similarity threshold and the second similarity mentioned above, the predicted word similarity relationship between the two word sets in the above test word set pair is determined;
[0189] Based on the above-mentioned labeled word similarity relationships and the above-mentioned predicted word similarity relationships, the accuracy of the above-mentioned candidate similarity thresholds is determined;
[0190] Based on the accuracy of the aforementioned candidate similarity thresholds and the aforementioned candidate similarity thresholds, the aforementioned target similarity thresholds are determined.
[0191] In one possible implementation, the first processing module 811 described above is further configured to perform the following operations:
[0192] Obtain the language categories corresponding to the N text files to be deduplicated;
[0193] Based on the language categories of the N text files, determine the word granularity parameters corresponding to the N text files.
[0194] Based on the word granularity parameters corresponding to the above N text files, word segmentation is performed on the above N text files to obtain the word group sets corresponding to the above N text files.
[0195] In one possible implementation, the above-described text deduplication device is also used to perform the following operations:
[0196] Based on the hash function, perform hash operation on the initial text files in the initial file set to obtain the first hash value corresponding to each of the initial text files in the initial file set.
[0197] Based on the first hash value mentioned above, duplicate initial text files within the initial file set are deleted to obtain candidate text files;
[0198] Based on the hash function described above, the text segments in the candidate text files are hashed to obtain the second hash values corresponding to the text segments in the candidate text files respectively.
[0199] Based on the second hash value mentioned above, duplicate text segments in the candidate text files are deleted, resulting in N text files to be deduplicated.
[0200] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0201] In this application, each text file is divided into a smaller set of word groups. Based on the word group set and M hash generation methods, M representative target hash values are generated for each text file. Only the M target hash values of each text file need to be compared to perform deduplication on the word group set corresponding to each text file. It is not necessary to compare the characters in each text file one by one, thus improving the efficiency of text file deduplication.
[0202] Please see Figure 9 , Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Figure 9 As shown, the aforementioned computer device 900 can refer to a server or terminal, including: a processor 901, a network interface 904, and a memory 905. Furthermore, the aforementioned computer device 900 may also include: a user interface 903, and at least one communication bus 902. The communication bus 902 is used to implement communication between these components. In some embodiments, the user interface 903 may include a display screen and a keyboard; optionally, the user interface 903 may also include a standard wired interface or a wireless interface. The network interface 904 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 905 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. The memory 905 may also optionally be at least one storage device located remotely from the aforementioned processor 901. Figure 9 As shown, the memory 905, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and computer programs.
[0203] exist Figure 9 In the computer device 900 shown, the network interface 904 provides network communication functionality; the user interface 903 is mainly used to provide an input interface; and the processor 901 can be used to call computer programs stored in the memory 905 to execute:
[0204] Perform word segmentation on the N text files to be deduplicated, and obtain the word set corresponding to each of the N text files; N is an integer greater than 1;
[0205] Based on M hash generation methods, perform hash operations on each word in each of the above word set to obtain M sets of hash values corresponding to each of the above word set; a set of hash values includes the hash values obtained by performing hash operations on each word in a word set using one hash generation method; M is an integer greater than 1;
[0206] From the M sets of hash values corresponding to each of the above phrase sets, select M target hash values that are representative of each of the above phrase sets; one target hash value is used to represent a set of hash values;
[0207] Based on the M target hash values corresponding to the N word sets, the words in the above N word sets are deduplicated to obtain the deduplicated word set. Based on the above deduplicated word set, a deduplicated text file is generated.
[0208] In one possible implementation, the processor 901 described above can also be used to invoke a computer program stored in the memory 905 to execute:
[0209] The smallest hash value among the M hash values corresponding to each of the above-mentioned phrase sets is determined as the M representative target hash values for each of the above-mentioned phrase sets; or,
[0210] The maximum hash value from the M hash values corresponding to each of the above phrase sets is determined as the M representative target hash values for each of the above phrase sets; or,
[0211] The average hash value corresponding to each of the M hash values for each of the above phrase sets is determined as the M representative target hash values for each of the above phrase sets.
[0212] In one possible implementation, the second processing module described above is further configured to perform the following operations:
[0213] According to the bucketing rule, the M target hash values corresponding to the i-th word set are bucketed to obtain A hash buckets corresponding to the i-th word set. The bucketing rule is determined based on the M hash generation methods. Each hash bucket includes at least one target hash value, where i is a positive integer less than or equal to M and A is an integer greater than 1.
[0214] Based on the A hash buckets corresponding to the N word sets mentioned above, candidate word set pairs with hash matching relationships are selected from the N word set sets; a candidate word set pair includes two word sets with hash matching relationships.
[0215] Based on the above candidate word set pairs, the word sets in the above N word set are deduplicated to obtain the deduplicated word set.
[0216] In one possible implementation, the number of candidate word pair sets is at least two; the processor 901 can also be used to call a computer program stored in memory 905 to execute:
[0217] Based on the corresponding word groups in the two word groups of each of the above candidate word group sets, determine the first similarity between the two word groups in each of the above candidate word group sets;
[0218] From at least two candidate word pair sets, select the candidate word pair set with the first similarity greater than the target similarity threshold, and use it as the target word pair set;
[0219] Based on the above target word set pairs, the word pairs in the above N word set are deduplicated to obtain the deduplicated word set.
[0220] In one possible implementation, the number of the target phrase set pairs is at least one; the processor 901 can also be used to call a computer program stored in the memory 905 to execute:
[0221] Construct a connected graph based on at least one pair of target word sets; the number of such connected graphs is at least one, and each such connected graph includes at least two nodes; one node corresponds to one word set; nodes corresponding to two word sets belonging to the same pair of target word sets are connected by edges;
[0222] From each node in each of the above connected graphs, select the target nodes that are representative for each of the above connected graphs;
[0223] The set of words corresponding to the target node of each of the above connected graphs and the set of remaining words are determined as the set of words after deduplication; the set of remaining words is the set of words in the above N set of words excluding the set of words in each of the above at least one target set of words.
[0224] In one possible implementation, each node includes data attribute information corresponding to a set of word groups; the processor 901 can also be used to call a computer program stored in the memory 905 to execute:
[0225] Based on the data attribute information of each phrase set in the above target phrase set pair, the priority of each phrase set in the above target phrase set pair is determined;
[0226] Based on the priority of each phrase set in the above target phrase set pair, select representative target nodes for each of the above connected graphs from each node in each of the above connected graphs.
[0227] In one possible implementation, the processor 901 described above can also be used to invoke a computer program stored in the memory 905 to execute:
[0228] From the above at least two phrase pairs, select at least one phrase pair for testing as the test phrase pair, and obtain the labeled phrase similarity relationship for each test phrase pair;
[0229] Based on the corresponding word groups in the two word groups of each of the above test word group sets, determine the second similarity between the two word groups in the above test word group set pair;
[0230] Based on the candidate similarity threshold and the second similarity mentioned above, the predicted word similarity relationship between the two word sets in the above test word set pair is determined;
[0231] Based on the above-mentioned labeled word similarity relationships and the above-mentioned predicted word similarity relationships, the accuracy of the above-mentioned candidate similarity thresholds is determined;
[0232] Based on the accuracy of the aforementioned candidate similarity thresholds and the aforementioned candidate similarity thresholds, the aforementioned target similarity thresholds are determined.
[0233] In one possible implementation, the processor 901 described above can also be used to invoke a computer program stored in the memory 905 to execute:
[0234] Obtain the language categories corresponding to the N text files to be deduplicated;
[0235] Based on the language categories of the N text files, determine the word granularity parameters corresponding to the N text files.
[0236] Based on the word granularity parameters corresponding to the above N text files, word segmentation is performed on the above N text files to obtain the word group sets corresponding to the above N text files.
[0237] In one possible implementation, the processor 901 described above can also be used to invoke a computer program stored in the memory 905 to execute:
[0238] Based on the hash function, perform hash operation on the initial text files in the initial file set to obtain the first hash value corresponding to each of the initial text files in the initial file set.
[0239] Based on the first hash value mentioned above, duplicate initial text files within the initial file set are deleted to obtain candidate text files;
[0240] Based on the hash function described above, the text segments in the candidate text files are hashed to obtain the second hash values corresponding to the text segments in the candidate text files respectively.
[0241] Based on the second hash value mentioned above, duplicate text segments in the candidate text files are deleted, resulting in N text files to be deduplicated.
[0242] In this application, each text file is divided into a smaller set of word groups. Based on the word group set and M hash generation methods, M representative target hash values are generated for each text file. Only the M target hash values of each text file need to be compared to perform deduplication on the word group set corresponding to each text file. It is not necessary to compare the characters in each text file one by one, thus improving the efficiency of text file deduplication.
[0243] Furthermore, it should be noted that this application also provides a computer-readable storage medium storing a computer program executed by the aforementioned data processing apparatus. This computer program includes program instructions, which, when executed by the processor, enable the execution of the data processing method described in the corresponding embodiments above. Therefore, further details will not be repeated here. Additionally, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the embodiments of the computer-readable storage medium involved in this application, please refer to the description of the method embodiments of this application.
[0244] As an example, the above program instructions can be deployed and executed on a computer device, or deployed and executed on at least two computer devices in one location, or executed on at least two computer devices distributed in at least two locations and interconnected by a communication network. At least two computer devices distributed in at least two locations and interconnected by a communication network can form a blockchain network.
[0245] The aforementioned computer-readable storage medium may be a data processing apparatus provided in any of the foregoing embodiments or a central storage unit of the aforementioned computer device, such as a hard disk or central storage of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, smart memory card (SMC), secure digital (SD) card, flash card, etc., provided on the computer device. Furthermore, the computer-readable storage medium may include both the central storage unit and external storage devices of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0246] The terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish content in different media, rather than to describe a specific order. Furthermore, the term "comprising," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.
[0247] In practice, the collection and processing of data in this application should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the data subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the data subject.
[0248] This application also provides a computer program product, including a computer program. When executed by a processor, the computer program implements the data processing method and decoding method described in the preceding embodiments, and therefore will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the embodiments of the computer program product involved in this application, please refer to the description of the method embodiments of this application.
[0249] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0250] The methods and related apparatus provided in this application are described with reference to the method flowcharts and / or structural diagrams provided in this application. Specifically, each block of the method flowchart and / or structural diagram, as well as combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable network-connected device to create a machine, such that the instructions, which execute via the processor of the computer or other programmable network-connected device, generate instructions for implementing the process. Figure 1 A schematic diagram of one or more processes and / or structures. Figure 1The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable network-connected device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 A schematic diagram of one or more processes and / or structures. Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable network-connected device to cause a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 A process or multiple processes and / or structures illustrate the steps of the functions specified in one or more boxes.
[0251] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A text deduplication method, characterized in that, include: The N text files to be deduplicated are segmented to obtain the word set corresponding to each of the N text files; N is an integer greater than 1; Based on M hash generation methods, a hash operation is performed on each word in each word set to obtain M hash values corresponding to each word set; a hash value includes the hash value obtained by performing a hash operation on each word in a word set using one hash generation method; M is an integer greater than 1; From the M sets of hash values corresponding to each set of words, select M target hash values that are representative of each set of words; A target hash value is used to represent a set of hash values; Based on the M target hash values corresponding to the N word sets, the words in the N word sets are deduplicated to obtain a deduplicated word set. Based on the deduplicated word set, a deduplicated text file is generated.
2. The method according to claim 1, characterized in that, The step of selecting M representative target hash values for each phrase set from the M sets of hash values corresponding to each phrase set includes: The minimum hash value among the M hash values corresponding to each of the phrase sets is determined as the M representative target hash values for each phrase set; or, The maximum hash value from the M hash values corresponding to each of the phrase sets is determined as the M representative target hash values for each phrase set; or, The average hash value corresponding to each of the M groups of hash values for each of the phrase sets is determined as the M representative target hash values for each of the phrase sets.
3. The method according to claim 1, characterized in that, The step involves deduplicating the word groups in the N word group sets based on the M target hash values corresponding to the N word group sets, resulting in a deduplicated word group set, including: According to the bucketing rule, the M target hash values corresponding to the i-th word set are bucketed to obtain A hash buckets corresponding to the i-th word set; the bucketing rule is determined based on the M hash generation methods, and each hash bucket includes at least one target hash value, where i is a positive integer less than or equal to M, and A is an integer greater than 1; Based on the A hash buckets corresponding to the N word sets, candidate word set pairs with hash matching relationships are selected from the N word set sets; a candidate word set pair includes two word sets with hash matching relationships. Based on the candidate word set pairs, the word groups in the N word set are deduplicated to obtain the deduplicated word set.
4. The method according to claim 3, characterized in that, The number of candidate word pair sets is at least two; the step of deduplicating word pairs in the N word sets to obtain a deduplicated word set includes: Based on the corresponding word groups in the two word groups of each candidate word group set pair, a first similarity between the two word groups in each candidate word group set pair is determined. From at least two candidate word pair sets, select the candidate word pair set with the first similarity greater than the target similarity threshold, and use it as the target word pair set; Based on the target word set pair, the word groups in the N word set are deduplicated to obtain the deduplicated word set.
5. The method according to claim 4, characterized in that, The number of target phrase sets is at least one; the step of deduplicating phrases in the N phrase sets according to the target phrase sets to obtain a deduplicated phrase set includes: A connected graph is constructed based on at least one pair of target word sets; the number of connected graphs is at least one, and each connected graph includes at least two nodes; one node corresponds to one word set; nodes corresponding to two word sets belonging to the same pair of target word sets are connected by edges; From each node in each of the connected graphs, select the target nodes that are representative for each of the connected graphs; The set of words corresponding to the target node of each of the connected graphs and the set of remaining words are determined as the set of words after deduplication; the set of remaining words is the set of words in the N set of words excluding the set of words in each of the at least one target set of words.
6. The method according to claim 5, characterized in that, Each node includes data attribute information for the corresponding set of word groups; The step of selecting representative target nodes for each connected graph from the nodes in each connected graph includes: Based on the data attribute information of each phrase set in the target phrase set pair, the priority of each phrase set in the target phrase set pair is determined; Based on the priority of each phrase set in the target phrase set pair, representative target nodes for each connected graph are selected from each node in each connected graph.
7. The method according to claim 4, characterized in that, The method further includes: From the at least two phrase pairs, select at least one phrase pair for testing as the test phrase pair, and obtain the labeled phrase similarity relationship for each test phrase pair; Based on the corresponding word groups in the two word groups of each test word group set pair, determine the second similarity between the two word groups in the corresponding test word group set pair; Based on the candidate similarity threshold and the second similarity, the predicted word similarity relationship between the two word sets in the test word set pair is determined; The accuracy of the candidate similarity threshold is determined based on the labeled word group similarity relationship and the predicted word group similarity relationship. The target similarity threshold is determined based on the accuracy of the candidate similarity threshold and the candidate similarity threshold.
8. The method according to claim 1, characterized in that, The process involves segmenting the N text files to be deduplicated to obtain a set of word groups corresponding to each of the N text files, including: Obtain the language categories corresponding to the N text files to be deduplicated; Based on the language categories corresponding to the N text files, determine the word granularity parameters corresponding to the N text files respectively; Based on the word granularity parameters corresponding to the N text files, the N text files are segmented to obtain the word set corresponding to the N text files.
9. The method according to claim 1, characterized in that, The method further includes: Based on the hash function, hash operations are performed on the initial text files in the initial file set to obtain the first hash value corresponding to each of the initial text files in the initial file set. Based on the first hash value, duplicate initial text files within the initial file set are deleted to obtain candidate text files; Based on the hash function, perform hash operation on the text segments in the candidate text file to obtain the second hash value corresponding to each text segment in the candidate text file; Based on the second hash value, duplicate text segments in the candidate text files are deleted to obtain N text files to be deduplicated.
10. A text deduplication device, characterized in that, include: The first processing module is used to perform word segmentation on the N text files to be deduplicated, and obtain the word set corresponding to each of the N text files; N is an integer greater than 1; The calculation module is used to perform hash operations on each word in each word set according to M hash generation methods to obtain M sets of hash values corresponding to each word set; a set of hash values includes hash values obtained by performing hash operations on each word in a word set using one hash generation method; M is an integer greater than 1; The selection module is used to select M target hash values that are representative of each phrase set from the M sets of hash values corresponding to each phrase set; A target hash value is used to represent a set of hash values; The second processing module is used to perform deduplication processing on the phrases in the N phrase sets according to the M target hash values corresponding to the N phrase sets respectively, to obtain the deduplicated phrase sets, and to generate a deduplicated text file based on the deduplicated phrase sets.
11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.
12. A computer storage medium, characterized in that, The computer storage medium stores a computer program, which, when executed by a processor, performs the steps of the method as described in any one of claims 1 to 9.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.