Text corpus deduplication processing method and system and storage medium
By randomly sampling, classification and hierarchical clustering of large-scale and multi-field text corpus, partitioning and deduplication, the problem of inefficient deduplication process in the existing technology is solved, and resource conservation and deduplication efficiency is improved.
Patent Information
- Application Number
- CN202510536475.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
When the prior art deals with large-scale and multi-field text corpus, there is a large amount of meaningless computing and storage resources, resulting in inefficient deduplication process.
By randomly sampling the global text corpus to be deduplicated, a subset of text corpus to be deduplicated is obtained, and it is divided into multiple classification sets according to the preset text classification model. Then, the text corpus in each classification collection is clustered hierarchically to form a hierarchical clustering structure, and the text corpus to be deduplicated to be deduplicated is divided into multiple corpus buckets, and the internal and global deduplication is performed.
By refining the deduplication range and partitioning processing, the calculation amount and storage resources are significantly reduced, the deduplication efficiency is improved, and the text corpus deduplication process is optimized.
Smart Images

Figure CN120067337A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text corpus deduplication, and in particular, to a method for deduplicating text corpus, a deduplication processing system, and a storage medium. Background Art
[0002] With the rapid development of the Internet and information technology, the amount of information has increased exponentially, and a large amount of text data has been created, stored, and exchanged. In this process, duplicate or highly similar content will inevitably be generated. These duplicate data not only occupy valuable storage space but may also lead to waste of computing resources. Especially in large-scale data analysis and machine learning model training, duplicate corpus increases the meaningless training time and consumes unnecessary computing power resources. For applications such as search engines and recommendation systems, the existence of duplicate content will affect the user experience, reduce the relevance and diversity of search results, and thus affect the overall performance of the system.
[0003] Text deduplication technology aims to identify and remove redundant text corpus, thereby improving the quality and efficiency of the data set, avoiding overfitting problems caused by duplicate samples, and enhancing the generalization ability of machine learning models. At the same time, text deduplication also helps to optimize the output of information retrieval systems, providing users with more diverse and relevant results. Moreover, during the construction of the corpus, text deduplication can maintain data consistency and structure, enabling subsequent research and development work to be based on a solid and reliable foundation.
[0004] In related technologies, when performing text deduplication, text corpus in various fields are mixed together for processing. Although this method can cover the entire data set, it has significant limitations when dealing with large-scale and multi-field corpus. Texts in different fields may have little semantic similarity, and mixing them together for deduplication operations will result in a large amount of meaningless calculations, and a large amount of storage resources are required to ensure the storage of intermediate files. Summary of the Invention
[0005] The present application provides a method for deduplicating text corpus, a deduplication processing system, and a storage medium to optimize the text corpus deduplication process and save the computing resources and storage resources required for the deduplication process.
[0006] The present application provides a method for deduplicating text corpus, including: randomly sampling the global text corpus to be deduplicated to obtain a subset of the text corpus to be deduplicated; dividing the subset of the text corpus to be deduplicated into multiple classification sets according to a preset text classification model; performing hierarchical clustering on the text corpus in each classification set to obtain a hierarchical clustering structure; dividing the global text corpus to be deduplicated into multiple corpus buckets according to the hierarchical clustering structure; after performing in-bucket deduplication on all corpus buckets, performing global deduplication to obtain the deduplicated text.
[0007] Optionally, according to a preset text classification model, the text corpus subset to be de-duplicated is divided into multiple classification sets, including: determining the attribute vectors of each text corpus in the text corpus subset to be de-duplicated according to the preset text classification model; the attribute vectors include the classification result confidence levels on each classification dimension; and dividing the text corpus subset to be de-duplicated into multiple classification sets according to the attribute vectors according to a preset rule.
[0008] Optionally, determining the attribute vectors of each text corpus in the text corpus subset to be de-duplicated according to a preset text classification model includes: classifying the text corpus in the text corpus subset to be de-duplicated according to a preset plurality of text classifiers to obtain the classification result confidence levels of each text corpus on each classification dimension, and forming the attribute vectors of each text corpus; wherein, the text classifiers correspond one-to-one with the classification dimensions; and the text classifiers include at least one of a language classifier, a domain classifier, a subject classifier, an emotion classifier, a writing style classifier, and a text carrier classifier.
[0009] Optionally, the multiple classification sets include a significant classification set and a fuzzy classification set, and at least one classification dimension is set for the significant classification set; dividing the text corpus subset to be de-duplicated into multiple classification sets according to the attribute vectors according to a preset rule includes: if the attribute vector of the text corpus represents that the classification result confidence level of the text corpus on at least one classification dimension is greater than or equal to a confidence threshold, determining that the text corpus belongs to the significant classification set; if the attribute vector of the text corpus represents that the classification result confidence levels of the text corpus on all classification dimensions are less than the confidence threshold, determining that the text corpus belongs to the fuzzy classification set.
[0010] Optionally, the multiple classification sets include multiple specific significant classification sets, at least one comprehensive significant classification set, and at least one non-specific significant classification set; wherein, each specific significant classification set corresponds to at least one classification dimension; if the attribute vector of the text corpus represents that the classification result confidence level of the text corpus on at least one classification dimension is greater than or equal to a confidence threshold, determining that the text corpus belongs to the significant classification set, including: if the attribute vector of the text corpus represents that the classification result confidence levels of the text corpus on the classification dimensions corresponding to a specific significant classification set are all greater than or equal to the confidence threshold, determining that the text corpus belongs to the corresponding specific significant classification set; if the attribute vector of the text corpus represents that the classification result confidence levels of the text corpus on the classification dimensions corresponding to multiple specific significant classification sets are all greater than or equal to the confidence threshold, determining that the text corpus belongs to the comprehensive significant classification set; if the attribute vector of the text corpus represents that the classification result confidence level of the text corpus on the classification dimensions corresponding to any specific significant classification set cannot be greater than or equal to the confidence threshold at the same time, determining that the text corpus belongs to the non-specific significant classification set.
[0011] Optionally, hierarchical clustering is performed on the text corpora in each classification set to obtain a hierarchical clustering structure, including: determining the semantic vectors of the text corpora in each classification set according to the semantic vector model; hierarchically clustering each classification set according to the semantic vectors of the text corpora according to the clustering model to form multiple clustering clusters, and obtaining a hierarchical clustering structure; wherein, the hierarchical clustering structure includes the dependency relationships between multiple clustering clusters.
[0012] Optionally, the global text corpus to be deduplicated is divided into multiple corpus buckets according to the hierarchical clustering structure, including: marking the cluster numbers of the global text corpus to be deduplicated according to the hierarchical clustering structure; dividing the global text corpus to be deduplicated into different corpus buckets according to the marked cluster numbers of the text corpora.
[0013] Optionally, in-bucket deduplication is performed on all corpus buckets, including: calculating the perplexity of each text corpus in the clustering cluster; using the deduplication algorithm to determine the similarity between the text corpora in each corpus bucket; for the text corpus groups with similarity exceeding the similarity threshold, deduplication is performed to retain the text corpus with the lowest perplexity in the text corpus group and delete the other text corpora in the text corpus group.
[0014] Optionally, the deduplication method further includes: calculating the centroid of the attribute vectors of all text corpora in each clustering cluster to obtain the attribute centroid vector of each clustering cluster; performing global deduplication, including: if there is no intersection between the attribute centroid vector of the clustering cluster in the corpus bucket and the attribute centroid vectors of other clustering clusters globally, it is determined that the global deduplication of the corpus bucket is completed; if there is an intersection between the attribute centroid vector of the clustering cluster in the corpus bucket and the attribute centroid vectors of other clustering clusters, inter-bucket deduplication is performed on the corpus buckets with intersections.
[0015] Optionally, the corpus buckets with intersections include the first corpus bucket and the second corpus bucket; performing inter-bucket deduplication on the corpus buckets with intersections, including: calculating the attribute vectors of each text corpus in the first corpus bucket; calculating the perplexity of each text corpus in the second corpus bucket and the distribution parameters of the attribute vectors of all text corpora; if the attribute vectors of the text corpora in the first corpus bucket conform to the set distribution parameter range of the attribute vectors in the second corpus bucket, deduplication is performed according to the similarity between the semantic vectors of the text corpora in the first corpus bucket and the semantic vectors of the text corpora in the second corpus bucket.
[0016] Optionally, performing deduplication according to the similarity between the semantic vectors of the text corpora and the semantic vectors of the text corpora in the second corpus bucket includes: determining the text corpora in the second corpus bucket whose similarity to the semantic vectors of the text corpora in the first corpus bucket is greater than the set similarity threshold; retaining the text corpus with the lowest perplexity among the text corpora in the second corpus bucket whose similarity to the semantic vectors of the text corpora in the first corpus bucket is greater than the set similarity threshold, and deleting the text corpora in the first corpus bucket.
[0017] Optionally, the deduplication method further includes: when the text corpus increases, dividing the newly added incremental text corpus into corpus buckets according to the hierarchical clustering structure; if the distribution of the attribute vector of the incremental text corpus relative to the attribute vector of the corpus bucket belongs to an outlier, retaining the incremental text corpus; if the distribution of the attribute vector of the incremental text corpus relative to the attribute vector of the corpus bucket does not belong to an outlier, performing deduplication according to the similarity between the incremental text corpus and the text corpus in the belonging corpus bucket.
[0018] The present application provides a deduplication processing system for text corpus, including one or more processors for implementing the foregoing deduplication method of text corpus.
[0019] The present application provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the foregoing deduplication method of text corpus is implemented.
[0020] The deduplication method, deduplication processing system and computer-readable storage medium for text corpus provided by the present application first randomly sample to obtain a subset of the text corpus to be deduplicated to reduce the data volume in some processing processes. Then, based on this subset of the text corpus to be deduplicated, the text corpus in the subset of the text corpus to be deduplicated is divided into multiple classification sets according to a preset classification model, and the texts in different fields are processed separately, which can avoid meaningless cross-field comparisons. Further, the deduplication scope is refined through the hierarchical clustering and corpus bucket division processes, reducing the calculation amount. Finally, deduplication within the bucket is performed first, and then global deduplication is performed, thereby further improving the deduplication efficiency. In this way, the optimization of the deduplication process of the text corpus is realized, which is beneficial to saving the computing resources and storage resources required for the deduplication process. Description of the Drawings
[0021] Figure 1 is a schematic flowchart of a deduplication method for text corpus provided by an embodiment of the present application; Figure 2 is a schematic diagram of the sampling result of the global text corpus to be deduplicated provided by an embodiment of the present application; Figure 3 is a schematic flowchart of a deduplication method for text corpus provided by another embodiment of the present application; Figure 4 is a classification diagram of a subset of the text corpus to be deduplicated provided by an embodiment of the present application; Figure 5 is a partial schematic diagram of hierarchical clustering provided by an embodiment of the present application; Figure 6 is an overall schematic diagram of hierarchical clustering provided by an embodiment of the present application; Figure 7It is a schematic flowchart of a method for deduplicating text corpora provided by another embodiment of the present application; Figure 8-1 It is a schematic diagram of the intersection situation between two corpus buckets provided by an embodiment of the present application; Figure 8-2 It is a schematic diagram of the intersection situation between two corpus buckets provided by another embodiment of the present application; Figure 9 It is a schematic flowchart of a method for deduplicating text corpora provided by another embodiment of the present application; Figure 10 It is a schematic diagram for deduplicating incremental text corpora provided by an embodiment of the present application.
[0022] Reference numerals 10: Global text corpus to be deduplicated; 101: Subset of text corpus to be deduplicated; 102: Remaining text corpus to be deduplicated; 201: First significant classification set; 202: Second significant classification set; 203: Third significant classification set; 204: Fourth significant classification set; 210: Fuzzy classification set; 301: First vector set; 302: First clustering cluster; 303: Second clustering cluster; 304: Third clustering cluster; 305: Fourth clustering cluster; 401: Incremental text corpus. Detailed implementation manners
[0023] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings.
[0024] Combined with Figure 1 As shown, an embodiment of the present application provides a method for deduplicating text corpora, including steps S10 to S50.
[0025] Step S10: Randomly sample the global text corpus to be deduplicated to obtain a subset of the text corpus to be deduplicated.
[0026] As Figure 2 shown, the global text corpus to be deduplicated 10 is divided into the subset of the text corpus to be deduplicated 101 and the remaining text corpus to be deduplicated 102.
[0027] The text corpus can be web text, conversations, papers, books, news, computer code, etc.
[0028] Step S20: According to a preset text classification model, divide the subset of the text corpus to be deduplicated into multiple classification sets.
[0029] Step S30: Hierarchically cluster the text corpora in each classification set to obtain a hierarchical clustering structure.
[0030] Step S40: Divide the global text corpus to be deduplicated into multiple corpus buckets according to the hierarchical clustering structure.
[0031] Step S50: After performing in-bucket deduplication on all the corpus buckets, perform global deduplication to obtain the deduplicated text.
[0032] Using the text deduplication method provided in the embodiments of the present application, first randomly sample to obtain a subset of the text corpus to be deduplicated, so as to reduce the amount of data in some processing procedures. Then, based on this subset of the text corpus to be deduplicated, according to a preset classification model, classify the text corpus in this subset of the text corpus to be deduplicated into multiple classification sets, and separate the text in different fields for processing, which can avoid meaningless cross-field comparisons. Further, refine the deduplication scope through hierarchical clustering and corpus bucket division processes, reducing the computational amount. Finally, perform in-bucket deduplication first and then global deduplication, thereby further improving the deduplication efficiency. In this way, the optimization of the text corpus deduplication process is realized, which is beneficial to saving the computational resources and storage resources required for the deduplication process.
[0033] Combined with Figure 3 As shown in the figure, in some embodiments, the aforementioned step S20: Classify the subset of the text corpus to be deduplicated into multiple classification sets according to a preset text classification model, including steps S201 to S202.
[0034] Step S201: According to a preset text classification model, determine the attribute vectors of each text corpus in the subset of the text corpus to be deduplicated; the attribute vectors include the classification result confidence levels in each classification dimension.
[0035] Step S202: Divide the subset of the text corpus to be deduplicated into multiple classification sets according to the attribute vectors according to a preset rule.
[0036] In this way, dividing the subset of the text corpus to be deduplicated according to the attribute vectors and the preset rules not only improves the efficiency of text processing, but also provides a clear classification basis for subsequent deduplication operations, which is beneficial to improving the overall performance of the text corpus deduplication process.
[0037] Specifically, in some embodiments, according to a preset text classification model, determining the attribute vectors of each text corpus in the subset of the text corpus to be deduplicated includes: Classify the text corpus in the subset of the text corpus to be deduplicated according to a preset plurality of text classifiers to obtain the classification result confidence levels of each text corpus in each classification dimension, and form the attribute vectors of each text corpus. Among them, the text classifiers correspond one-to-one with the classification dimensions. And the text classifiers include at least one of a language classifier, a field classifier, a discipline classifier, an emotion classifier, a writing style classifier, and a text carrier classifier. In this way, the classification result confidence level of a text corpus in each classification temperature can be determined according to the attribute vector of the text corpus, which is convenient for subsequent classification processes based on this.
[0038] Specifically, the language classifier is used to classify the languages used in each text corpus, such as English, Chinese, Spanish, Polish, etc. The domain classifier is used to classify the sub-domains within the knowledge base to which the content described in the text corpus belongs. The subject classifier is used to classify the subjects within the knowledge base to which the content described in the text corpus belongs. The sentiment classifier is used to classify the sentiment in each text corpus, such as positive sentiment, negative sentiment, etc. The writing style classifier is used to classify the writing styles of each text corpus, such as news, technology, spoken language, story, advertisement, etc. The text carrier classifier is used to classify the carriers of each text corpus, such as web pages, papers, news, etc.
[0039] The output result of the text classifier is the classification result confidence expressed in the form of probability, such as 0.6, 0.7, etc. The classification dimensions corresponding to each text classifier and the classification result confidence of each text classifier constitute the attribute vector of each text corpus. In this way, the judgment of the text corpus in each classification dimension can be realized according to the numerical value.
[0040] It can be understood that similar text corpora must have similar attribute vectors. If there are significant differences between two text corpora in any classification dimension of the attribute vector, then these two text corpora must not be duplicates. It should be noted that although the process of calculating the attribute vector of the text corpus using the text classifier will consume certain computing resources, this work will save more computing resources for the subsequent deduplication work of the text corpus.
[0041] After classification, the text corpus subset to be deduplicated is divided into multiple classification sets. Specifically, in some embodiments, the multiple classification sets include a significant classification set and a fuzzy classification set. Among them, the significant classification set is set with at least one classification dimension. According to the attribute vector, the text corpus subset to be deduplicated is divided into multiple classification sets according to a preset rule, including: if the classification result confidence of the text corpus in at least one classification dimension represented by the attribute vector of the text corpus is greater than or equal to the confidence threshold, it is determined that the text corpus belongs to the significant classification set. If the classification result confidence of the text corpus in all classification dimensions represented by the attribute vector of the text corpus is less than the confidence threshold, it is determined that the text corpus belongs to the fuzzy classification set. In this way, by dividing the text corpus subset to be deduplicated into a significant classification set and a fuzzy classification set, text corpora with clear features can be efficiently screened out, and at the same time, it is beneficial to simplify the subsequent processing flow. It not only improves the efficiency of classification and deduplication, but also has strong adaptability and is applicable to a variety of text processing scenarios.
[0042] If the classification result confidence of a text corpus in a classification dimension is greater than or equal to the confidence threshold, for example, the confidence threshold is 0.7, then it is considered that the attribute of the text corpus in this classification dimension is significant. If the attributes of a text corpus are significant in at least one classification dimension, then this text corpus is considered significant, and the text corpus belongs to the significant classification set. If a text corpus is not significant in all classification dimensions, then the text corpus belongs to the fuzzy classification set.
[0043] In some embodiments, the foregoing multiple classification sets include multiple specific significant classification sets, at least one comprehensive significant classification set, and at least one non-specific significant classification set. Among them, each specific significant classification set corresponds to at least one classification dimension. If the attribute vector of a text corpus represents that the classification result confidence of the text corpus in at least one classification dimension is greater than or equal to the confidence threshold, then determining that the text corpus belongs to the significant classification set includes: if the attribute vector of the text corpus represents that the classification result confidence of the text corpus in the classification dimensions corresponding to a specific significant classification set is all greater than or equal to the confidence threshold, then determining that the text corpus belongs to the specific significant classification set that it conforms to. If the attribute vector of the text corpus represents that the classification result confidence of the text corpus in the classification dimensions corresponding to multiple specific significant classification sets is all greater than or equal to the confidence threshold, then determining that the text corpus belongs to the comprehensive significant classification set. If the attribute vector of the text corpus represents that the classification result confidence of the text corpus in the classification dimensions corresponding to any one specific significant classification set cannot be greater than or equal to the confidence threshold at the same time, then determining that the text corpus belongs to the non-specific significant classification set. In this way, it can be ensured that the text corpora with significant attributes can be assigned to and only assigned to one significant classification set, avoiding duplicate classification and at the same time avoiding omission of text corpora.
[0044] Specifically, considering that in the implementation process, a text corpus may meet the classification conditions of multiple significant classification sets. Outside the multiple specific significant classification sets, a comprehensive significant classification set is separately set up to store this type of text corpus. In this way, it can avoid the situation where the same text corpus is classified into multiple significant classification sets and the calculation burden is increased. Further, according to actual needs, the comprehensive significant classification set can be set to one or more. In the case of being set to one, all text corpora that meet the classification conditions of multiple significant classification sets are assigned to this comprehensive significant classification set. In the case of being set to multiple, further detailed division can be made according to the actual situation. For example, text corpora that meet the classification conditions of two significant classification sets at the same time are assigned to one comprehensive significant classification set, and text corpora that meet the classification conditions of more than three significant classification sets at the same time are assigned to another comprehensive significant classification set. It can be set according to specific needs and will not be elaborated here.
[0045] That is, if the significant classification dimensions of a text corpus match all the classification dimensions corresponding to a specific significant classification set, then this text corpus is classified into this specific significant classification set. For example, a specific significant classification set related to the Chinese e-commerce field is preset, and texts with a language dimension classified as Chinese and a writing style dimension of e-commerce with a confidence greater than 0.7 are classified into this category. If the significant classification dimensions of a text corpus match all the classification dimensions corresponding to multiple specific significant classification sets, then this text corpus is classified into the comprehensive significant classification set. If a text corpus is significant in terms of attributes, but its significant classification dimensions do not match any of the classification dimensions of a specific significant classification set and does not belong to any specific significant classification set, then this text corpus is classified into another significant classification set, that is, the non-specific significant classification set.
[0046] As Figure 4 shown, through the above classification process, all text corpora are divided into multiple classification sets, such as the first significant classification set 201, the second significant classification set 202, the third significant classification set 203, the fourth significant classification set 204, and the fuzzy classification set 210.
[0047] In some embodiments, hierarchical clustering is performed on the text corpora in each classification set to obtain a hierarchical clustering structure, including: determining the semantic vectors of the text corpora in each classification set according to the semantic vector model. According to the semantic vectors of each text corpus, hierarchical clustering is performed on each classification set according to the clustering model to form multiple clustering clusters, and a hierarchical clustering structure is obtained. Among them, the hierarchical clustering structure includes the dependency relationships between multiple clustering clusters.
[0048] Here, a pre-trained semantic vector model is used to calculate the semantic vectors of the text corpora in each significant classification set and the fuzzy classification set, and the semantic vector models adopted by different classification sets may be different. Specifically, there are often more professional semantic vector models in some fields of in-depth research. At least part of the specific significant classification sets set here corresponds to the professional semantic vector models related to this field. For the remaining significant classification sets without professional semantic vector models, a general semantic vector model is used. The fuzzy classification set also uses a general semantic vector model. In this way, it not only ensures high-precision semantic understanding in a specific field, but also simplifies the processing process through the wide applicability of the general model, improves the overall processing efficiency, and at the same time enhances the system's adaptability to different fields and languages.
[0049] After determining the semantic vectors of each text corpus, for each classification set, using the clustering algorithm for hierarchical clustering with its semantic vector as the target, and finally determining the tree structure of the clustering model. The clustering model tree includes multiple clusters and the dependency relationships between the clusters. Among them, clustering can be performed according to the distance of the cluster centroids in the vector space.
[0050] For the significant classification set, if the clusters obtained during the clustering process simultaneously meet the two conditions that the text corpus capacity within the cluster is less than the preset capacity threshold and the attribute vectors within the cluster converge within the preset threshold range in each dimension, it is determined that the clustering of this significant classification set ends. Otherwise, continue clustering until these two conditions are met or the clustering level exceeds the preset level. It should be noted that during the real-time process, the preset capacity threshold and the preset threshold range are both preset to appropriate values to prevent over-clustering. Among them, the expected value and standard deviation of the attribute vectors in each cluster are calculated to obtain their distribution parameters. The distribution of the attribute vectors within the cluster can be marked in the form of "expected value ± standard deviation".
[0051] For the fuzzy classification set, if the cluster obtained during the clustering process meets the condition that the text corpus capacity within the cluster is less than the preset capacity threshold, it is determined that the clustering of this fuzzy classification set ends.
[0052] In some embodiments, the number of categories K for each clustering is set to an exponential number of the number of vectors. For example, if the number of vectors is N, then the number of categories for clustering is int[ln(N)+1].
[0053] Take Figure 5 as an example to illustrate the process of hierarchical clustering for a significant text set. First, vectorize the first significant classification set to obtain the first vector set 301. Cluster the first vector set 301 to obtain multiple clusters. Among the multiple clusters obtained in this clustering, the text corpus capacity of the first cluster 302 is greater than the preset capacity threshold and needs to be reclustered. Although the second cluster 303 obtained in this clustering meets the convergence condition in terms of text corpus capacity, that is, it is less than the preset capacity threshold, the standard deviation σ1 in the distribution (u1±σ1) of its attribute vectors in some classification dimensions is too large and does not converge within the preset threshold range. Therefore, the second cluster 303 needs to be reclustered and finally divided into the third cluster 304 and the fourth cluster 305. The attribute distributions of the third cluster 304 and the fourth cluster 305 both meet the convergence conditions, and the hierarchical clustering ends.
[0054] In some embodiments, the global text corpus to be deduplicated is divided into multiple corpus buckets according to a hierarchical clustering structure, including: marking the cluster numbers of the global text corpus to be deduplicated according to the hierarchical clustering structure; dividing the global text corpus to be deduplicated into different corpus buckets according to the marked cluster numbers of each text corpus. After completing the hierarchical clustering, the global text corpus to be deduplicated needs to be clustered according to the established hierarchical clustering structure. And number all the clusters, and each cluster is called a corpus bucket. As Figure 6 shown, the global text corpus to be deduplicated is divided into multiple corpus buckets. So far, the entire hierarchical clustering cluster corpus bucket structure has been constructed.
[0055] Cluster the global text corpus to be deduplicated according to the hierarchical clustering structure, mark the cluster numbers of the clusters, calculate the semantic vector centroids and attribute vector distribution parameters for each cluster, and construct a global cluster semantic centroid vector table and a global cluster attribute vector distribution table.
[0056] Specifically, after constructing the hierarchical clustering cluster corpus bucket structure, by taking the average of the semantic vectors in each corpus bucket, the centroid vector of each corpus bucket is obtained and saved as the global cluster semantic centroid vector table. Calculate the expected value and standard deviation of the attribute vectors in each corpus bucket, so as to obtain the distribution parameters of the attribute vectors of the corresponding clusters, and save the distribution parameters of the attribute vectors of all corpus buckets in the global cluster attribute vector table. Table 1 shows the structure of a global cluster attribute vector table. Among them, the distribution parameters of the attribute vectors can be marked in the form of "expected value ± standard deviation". In the implementation process, the specific technical implementation can use the method of storing separately to save the above vector tables. In this way, it is convenient to directly call relevant data according to the table later. In some embodiments, before performing in-bucket deduplication and global deduplication on the corpus buckets, the global cluster attribute vector distribution table needs to be loaded into memory first.
[0057] Table 1:
[0058] In some embodiments, in-bucket deduplication is performed on all corpus buckets, including: calculating the perplexity of each text corpus in the cluster; using a deduplication algorithm to determine the similarity between text corpora in each corpus bucket; for text corpus groups with similarity exceeding the similarity threshold, perform deduplication to retain the text corpus with the lowest perplexity in the text corpus group and delete other text corpora in the text corpus group. In the implementation process, traverse the text corpora to perform in-bucket deduplication on all corpus buckets. In some embodiments, the deduplication algorithm is the locality-sensitive hashing algorithm (LSH). Using the locality-sensitive hashing algorithm, text corpus groups with similarity exceeding the similarity threshold are obtained, and the text corpus with the smallest perplexity among them is retained, and the others are deleted. The similarity threshold is set to 0.9, for example.
[0059] Further, in at least some embodiments, duplicate removal within the buckets is performed through parallel distributed computing to shorten the overall computing duration. In addition, after duplicate removal is completed, the local sensitive hash vector pool corresponding to each corpus bucket is persistently saved. In this way, when loading subsequently, the entire hash vector pool can be directly loaded bypassing the hashing process.
[0060] In some embodiments, the duplicate removal method further includes step S41 of calculating the centroid of the attribute vectors of all text corpora in each clustering cluster to obtain the attribute centroid vector of each clustering cluster. On this basis, global duplicate removal is performed, including: if there is no intersection between the attribute centroid vector of the clustering cluster in the corpus bucket and the attribute centroid vectors of other clustering clusters globally, it is determined that the global duplicate removal of the corpus bucket is completed; if there is an intersection between the attribute centroid vector of the clustering cluster in the corpus bucket and the attribute centroid vectors of other clustering clusters, inter-bucket duplicate removal is performed on the corpus buckets with intersections. Specifically, as Figure 7 shown, the embodiment of the present application provides a duplicate removal method for text corpora, including steps S10 to step S502.
[0061] Step S10: Randomly sample the global text corpus to be de-duplicated to obtain a subset of the text corpus to be de-duplicated.
[0062] Step S20: According to a preset text classification model, divide the subset of the text corpus to be de-duplicated into multiple classification sets.
[0063] Step S301: Determine the semantic vectors of the text corpora in each classification set according to the semantic vector model.
[0064] Step S302: According to the semantic vectors of each text corpus, hierarchically cluster each classification set according to the clustering model to form multiple clustering clusters and obtain a hierarchical clustering structure.
[0065] Among them, the hierarchical clustering structure includes the dependency relationships between multiple clustering clusters.
[0066] Step S401: Mark the global text corpus to be de-duplicated according to the hierarchical clustering structure.
[0067] Step S402: Divide the global text corpus to be de-duplicated into different corpus buckets according to the cluster numbers marked for each text corpus.
[0068] Step S41: Calculate the centroid of the attribute vectors of all text corpora in each clustering cluster to obtain the attribute centroid vector of each clustering cluster.
[0069] Step S501: Perform duplicate removal within all corpus buckets.
[0070] Step S502: If there is no intersection between the attribute centroid vectors of the clustering clusters in the corpus bucket and the attribute centroid vectors of other clustering clusters globally, it is determined that the global deduplication of the corpus bucket is completed; if there is an intersection between the attribute centroid vectors of the clustering clusters in the corpus bucket and the attribute centroid vectors of other clustering clusters, inter-bucket deduplication is performed on the corpus buckets with intersections.
[0071] Traverse the combination pairs of any two corpus buckets and perform inter-bucket deduplication on any two corpus buckets. The first step of deduplication is to determine whether there is an intersection in the distribution of the attribute vectors of the two corpus buckets. If there is no intersection, then the two buckets do not need to perform inter-bucket deduplication.
[0072] Figure 8-1 and 8-2 Illustrates two cases of judging the intersection of attribute distributions between corpus buckets.
[0073] As Figure 8-1 shown, there is no intersection between the two corpus buckets within the preset threshold range in the classification dimensions of lang_en, that is, the language is English, and sentiment_positve, that is, the text sentiment is positive. Therefore, these two corpus buckets do not need inter-bucket deduplication.
[0074] As Figure 8-2 shown, although the attributes of all dimensions are slightly different, they do not exceed the preset threshold range, and it is determined that there is an intersection. Then these two corpus buckets need to participate in inter-bucket deduplication.
[0075] It should be noted that actual tests show that: under the hierarchical clustering cluster corpus bucket structure described above, most inter-bucket deduplications do not need to be performed. In this way, it is beneficial to save computing resources.
[0076] Taking the corpus buckets with intersections including the first corpus bucket and the second corpus bucket as an example, performing inter-bucket deduplication on the corpus buckets with intersections includes: calculating the attribute vectors of each text corpus in the first corpus bucket; calculating the perplexity of each text corpus in the second corpus bucket and the distribution parameters of the attribute vectors of all text corpora; if the attribute vectors of the text corpora in the first corpus bucket meet the set distribution parameter range of the attribute vectors in the second corpus bucket, deduplication is performed according to the similarity between the semantic vectors of the text corpora in the first corpus bucket and the semantic vectors of the text corpora in the second corpus bucket. In this way, through the screening mechanism of distribution parameters, the accuracy and efficiency of deduplication are improved, effectively avoiding the misdeletion of important information. And through the refined comparison of semantic similarity, duplicate or highly similar texts are effectively identified and removed, significantly improving the accuracy and efficiency of text deduplication.
[0077] Specifically, the text corpus in the first corpus bucket is marked as Ta, and the text corpus in the second corpus bucket is marked as Tb. Calculate whether the attribute vector Va of Ta conforms to the set distribution parameter range of the attribute vectors in the second corpus bucket. If it is not within the preset threshold range, then Ta does not need to be compared with the text corpus in the second corpus bucket subsequently. If the attribute vector Va of Ta conforms to the set distribution parameter range of the attribute vectors in the second corpus bucket, that is, within the preset threshold range, then Ta is compared with the text corpus in the second corpus bucket for duplicate removal.
[0078] Specifically, duplicate removal is performed according to the similarity between the semantic vectors of the text corpus and the semantic vectors of the text corpus in the second corpus bucket, including: determining the text corpus in the second corpus bucket whose similarity with the semantic vector of the text corpus in the first corpus bucket is greater than the set similarity threshold; retaining the text corpus with the lowest perplexity among the text corpus whose similarity with the semantic vector of the text corpus in the first corpus bucket is greater than the set similarity threshold into the second corpus bucket, and deleting the text corpus in the first corpus bucket. Through the duplicate removal strategy based on semantic vector similarity, duplicate or highly similar texts are effectively removed. On this basis, through perplexity comparison, the text corpus with the clearest and most accurate semantic expression is retained, which is beneficial to improving the quality and usability of the obtained text corpus after duplicate removal, and provides a higher-quality data basis for subsequent text processing and analysis. More specifically, Ta and the second corpus bucket use the locality-sensitive hashing algorithm for duplicate removal comparison. If there is a text corpus in the second corpus bucket whose similarity with Ta is greater than the set similarity threshold, then the text corpus with a lower perplexity is retained in the second corpus bucket. When all corpus buckets have been traversed and calculated, all the remaining corpus buckets are the corpus after duplicate removal.
[0079] Combined Figure 9 As shown, the duplicate removal method further includes steps S601 to S603.
[0080] Step S601, when the text corpus increases, the newly added incremental text corpus is assigned to the corpus bucket according to the hierarchical clustering structure.
[0081] Step S602, if the distribution of the attribute vector of the incremental text corpus relative to the attribute vector of the corpus bucket belongs to an outlier, then retain the incremental text corpus.
[0082] Step S603, if the distribution of the attribute vector of the incremental text corpus relative to the attribute vector of the corpus bucket does not belong to an outlier, then duplicate removal is performed according to the similarity between the incremental text corpus and the text corpus in the belonging corpus bucket.
[0083] In this way, through dynamic processing of the newly added incremental text corpus, an efficient and accurate duplicate removal operation is achieved.
[0084] Specifically, the newly added incremental text corpus is divided into corpus buckets according to the hierarchical clustering structure, including: determining the attribute vector of the incremental text corpus according to a preset text classification model. According to its attribute vector, determine the classification set to which the incremental text corpus belongs, for example, marked by the classification number. According to the classification set to which the incremental text corpus belongs and the hierarchical clustering structure, determine the clustering cluster to which it belongs, thus realizing the division of the incremental text corpus into the corpus bucket. Among them, the branch nodes of the clustering clusters are obtained according to the dependency relationship between the clustering clusters in the hierarchical clustering structure, and the specific clustering clusters are further determined to find the corpus bucket. For example, as shown in Figure 10 For the incremental text corpus 401, as shown, according to its attribute vector, the classification number of the classification set to which the incremental text corpus 401 belongs is determined to be C1. On this basis, the branch node C of the clustering cluster is obtained according to the dependency relationship between the clustering clusters in the hierarchical clustering structure 21 , and the specific clustering cluster K is further determined g .
[0085] If the distribution of the attribute vector of the incremental text corpus is an outlier relative to the attribute vector of the corpus bucket to which it belongs, then determine that the incremental text corpus is a non-duplicate text and retain the incremental text corpus. If the distribution of the attribute vector of the incremental text corpus is not an outlier relative to the attribute vector of the corpus bucket to which it belongs, then it is considered that the incremental text corpus may be a duplicate or highly similar text corpus that needs to be further de-duplicated. Among them, the distribution of the attribute vector can be obtained by querying the global clustering cluster attribute vector distribution table according to the clustering cluster number.
[0086] Deduplication is performed according to the similarity between the incremental text corpus and the text corpus in the corpus bucket to which it belongs, including: calculating the perplexity of the incremental text corpus. Based on the de-duplication algorithm, determine the text corpus group whose similarity to the incremental text corpus is higher than the similarity threshold, retain the incremental text corpus and the text corpus with the lowest perplexity in the text corpus group, and delete other text corpora. In this way, not only are duplicate or highly similar texts effectively removed, but also through the optimization of perplexity, it is ensured that the text corpora retained in the corpus library are the clearest and most accurate in semantic expression. In some embodiments, according to the clustering cluster number of the incremental text corpus, query whether the corresponding locality-sensitive hashing pool vector table has been loaded into memory. If not, then load the hashing pool vector table corresponding to the clustering cluster number from the persistent storage into memory to construct the corresponding locality-sensitive vector table. If the number of locality-sensitive hashing pool vector tables loaded into memory exceeds the preset threshold, then recycle the hashing pool vector table that has been used the least in the most recent set time. After retaining the text corpus with the lowest perplexity in the corpus bucket, update the corresponding locality-sensitive hashing pool vector table.
[0087] The present application provides a duplicate removal processing system for text corpora, including one or more processors for implementing the foregoing duplicate removal method for text corpora.
[0088] The present application further provides a computer-readable storage medium storing a program, which, when executed by a processor, implements the foregoing duplicate removal method for text corpora. Here, the computer-readable storage medium may be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, and so on. For example, the computer-readable storage medium may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.
[0089] In the description of the present application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include at least one such feature. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
Claims
1. A method for deduplication of text corpus, characterized in that: include: Randomly sample the global text corpus to be deduplicated to obtain a subset of the text corpus to be deduplicated; According to a preset text classification model, the text corpus subset to be deduplicated is divided into multiple classification sets; Perform hierarchical clustering on the text corpus in each classification set to obtain a hierarchical clustering structure; Dividing the global text corpus to be deduplicated into multiple corpus buckets according to the hierarchical clustering structure; After deduplication is performed within all corpus buckets, global deduplication is performed to obtain the deduplicated text.
2. The deduplication method according to claim 1, characterized in that: According to the preset text classification model, the text corpus subset to be deduplicated is divided into multiple classification sets, including: According to a preset text classification model, determining the attribute vector of each text corpus in the subset of text corpus to be deduplicated; the attribute vector includes the confidence of the classification result on each classification dimension; The text corpus subset to be deduplicated is divided into a plurality of classification sets according to the attribute vector and a preset rule.
3. The deduplication method according to claim 2, characterized in that: The step of determining the attribute vector of each text corpus in the to-be-deduplicated text corpus subset according to the preset text classification model includes: According to a plurality of preset text classifiers, the text corpus in the to-be-deduplicated text corpus subset is classified to obtain the classification result confidence of each text corpus in each classification dimension to form an attribute vector of each text corpus; Among them, the text classifier corresponds one-to-one to the classification dimension; the text classifier includes at least one of a language classifier, a field classifier, a subject classifier, a sentiment classifier, a writing style classifier, and a text carrier classifier.
4. The deduplication method according to claim 2, characterized in that: The multiple classification sets include a significant classification set and a fuzzy classification set, and the significant classification set is provided with at least one classification dimension; The step of dividing the to-be-deduplicated text corpus subset into a plurality of classification sets according to the attribute vector and a preset rule comprises: If the attribute vector of the text corpus represents that the confidence of the classification result of the text corpus in at least one classification dimension is greater than or equal to the confidence threshold, then it is determined that the text corpus belongs to the significant classification set; If the attribute vector of the text corpus represents that the confidence of the classification result of the text corpus in all classification dimensions is less than the confidence threshold, it is determined that the text corpus belongs to the fuzzy classification set.
5. The deduplication method according to claim 4, characterized in that: The multiple classification sets include multiple specific significant classification sets, at least one comprehensive significant classification set and at least one non-specific significant classification set; wherein each specific significant classification set corresponds to at least one classification dimension; If the attribute vector of the text corpus represents that the confidence of the classification result of the text corpus in at least one classification dimension is greater than or equal to the confidence threshold, then determining that the text corpus belongs to the significant classification set includes: If the attribute vector of the text corpus represents that the confidence of the classification result of the text corpus on the classification dimension corresponding to a specific significant classification set is greater than or equal to the confidence threshold, it is determined that the text corpus belongs to the specific significant classification set that meets the requirements; If the attribute vector of the text corpus represents that the confidence of the classification result of the text corpus on the classification dimensions corresponding to the multiple specific significant classification sets is greater than or equal to the confidence threshold, it is determined that the text corpus belongs to the comprehensive significant classification set; If the attribute vector of the text corpus represents that the classification result confidence of the text corpus on the classification dimension corresponding to any specific significant classification set cannot be greater than or equal to the confidence threshold at the same time, it is determined that the text corpus belongs to a non-specific significant classification set.
6. The deduplication method according to claim 1, characterized in that: The hierarchical clustering of the text corpora in each classification set to obtain a hierarchical clustering structure includes: Determine the semantic vectors of the text corpora in each classification set according to the semantic vector model; According to the semantic vectors of each text corpus, each classification set is hierarchically clustered according to the clustering model to form multiple clusters and obtain a hierarchical clustering structure; wherein the hierarchical clustering structure includes dependency relationships between the multiple clusters.
7. The deduplication method according to claim 6, characterized in that: The step of dividing the global text corpus to be deduplicated into a plurality of corpus buckets according to the hierarchical clustering structure includes: Marking the global text corpus to be deduplicated with cluster numbers according to the hierarchical clustering structure; According to the cluster numbers marked on the text corpora, the global text corpora to be deduplicated are divided into different corpus buckets.
8. The deduplication method according to claim 6, characterized in that: Perform deduplication on all corpus buckets, including: Calculate the perplexity of each text corpus in the cluster; Use deduplication algorithms to determine the similarity between text corpora in each corpus bucket; For the text corpus group whose similarity exceeds the similarity threshold, duplicate removal is performed to retain the text corpus with the lowest perplexity in the text corpus group, and other text corpora in the text corpus group are deleted.
9. The deduplication method according to claim 6, characterized in that: The deduplication method further comprises: Calculate the centroid of the attribute vectors of all text corpora in each cluster to obtain the attribute centroid vector of each cluster; The global deduplication process includes: If the attribute centroid vector of the cluster in the corpus bucket does not have any intersection with the attribute centroid vectors of other clusters in the world, it is determined that the global deduplication of the corpus bucket is completed; If the attribute centroid vector of a cluster in a corpus bucket intersects with the attribute centroid vectors of other clusters, the corpus buckets with the intersection are deduplicated between buckets.
10. The deduplication method according to claim 9, characterized in that: The corpus buckets with intersections include a first corpus bucket and a second corpus bucket; Deduplication between corpus buckets with intersections includes: Calculate the attribute vector of each text corpus in the first corpus bucket; Calculate the perplexity of each text corpus in the second corpus bucket and the distribution parameters of the attribute vectors of all text corpora; If the attribute vector of the text corpus in the first corpus bucket meets the set distribution parameter range of the attribute vector in the second corpus bucket, deduplication is performed based on the similarity between the semantic vector of the text corpus in the first corpus bucket and the semantic vector of the text corpus in the second corpus bucket.
11. The deduplication method according to claim 10, characterized in that: The performing deduplication according to the similarity between the semantic vector of the text corpus and the semantic vector of the text corpus in the second corpus bucket includes: Determine text corpora in the second corpus bucket whose semantic vector similarity with the text corpora in the first corpus bucket is greater than a set similarity threshold; The text corpus with the lowest perplexity among the text corpora whose similarity with the semantic vector of the text corpus in the first corpus bucket is greater than a set similarity threshold is retained in the second corpus bucket, and the text corpus in the first corpus bucket is deleted.
12. The deduplication method according to claim 1, characterized in that: The deduplication method further comprises: When the text corpus increases, the newly added incremental text corpus is divided into corpus buckets according to the hierarchical clustering structure; If the attribute vector of the incremental text corpus is an outlier relative to the attribute vector distribution of the corpus bucket, retaining the incremental text corpus; If the attribute vector of the incremental text corpus is not an outlier relative to the attribute vector distribution of the corpus bucket, deduplication is performed based on the similarity between the incremental text corpus and the text corpus in the corpus bucket.
13. A text corpus deduplication processing system, characterized in that: The method comprises one or more processors for implementing the method for deduplicating a text corpus as described in any one of claims 1 to 12.
14. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the method for deduplicating a text corpus as described in any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Event generation method based on text information and related device
CN110209808A
Decision tree-based inquiry distribution method, device, equipment and storage medium
CN113707286A
Data deduplication method and device based on text similarity, storage medium and server
CN114281989A
Target data deduplication method and device, storage medium and electronic device
CN114661702A
Corpus data processing method and device and electronic equipment
CN115391539A