Text Duplication Removal Method, Apparatus, Device and Medium

By calculating text features and similarity information in the preset text library, filtering and removing duplicate text, the problem of unsatisfactory deduplication effect in the prior art is solved, and efficient and accurate deduplication effect is achieved in the case of large data volume and limited CPU resources.

CN114298227BActive Publication Date: 2025-06-27CHINA CONSTRUCTION BANK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111649127.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-06-27
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

In the prior art, when screening online data information, the deduplication effect is not ideal, especially when the data volume is large and the CPU resources are limited, it is difficult to delete duplicate text efficiently and accurately.

Method used

By obtaining the text features of each text in the preset text library, determining the feature vector and generating a feature matrix, calculating the similarity information between texts, filtering out similar texts, and determining and removing the text to be deduplicated based on the similarity information relationship between the second preset quantity and the third preset quantity.

Benefits of technology

In the case of large amount of data and limited CPU resources, the calculation efficiency of duplicate data information is improved, duplicate text is accurately deleted, and the deduplication effect is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114298227B_ABST
    Figure CN114298227B_ABST
Patent Text Reader

Abstract

The present application provides a text deduplication method, apparatus, device and medium, which relates to the technical field of natural language processing. The method includes: obtaining the text features of each text in a preset text library; determining the feature vector of each text according to the text features of each text, and generating a feature matrix between texts in the preset text library according to the feature vector of each text; determining the similarity information between texts in the preset text library according to the feature matrix, and screening out the first preset number of texts similar to each text in the preset text library according to the similarity information; determining and removing the text to be deduplicated in the preset text library according to the relationship between the similarity information of the second preset number of texts and the similarity information of the third preset number of texts. By adopting the technical solution, it is possible to improve the calculation efficiency of duplicate data information and accurately delete duplicate texts in the case of a large amount of data and limited CPU resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of natural language processing, and in particular, to a text deduplication method, apparatus, device, and medium. Background Art

[0002] Currently, the data information on the Internet is relatively complex, and users need to read a large amount of duplicate data information when viewing the required data information. Therefore, how to filter out duplicate data information is a problem that the public is more concerned about.

[0003] Among the current methods for filtering out duplicate data information for users, the deduplication algorithms used are all relatively simple similarity comparison methods, which will result in an unsatisfactory deduplication effect, and users will still see a lot of repetitive content.

[0004] Therefore, there is an urgent need for a text deduplication method that can improve the calculation efficiency of duplicate data information and accurately delete duplicate texts in the case of a large amount of data and limited Central Processing Unit (CPU) resources. Summary of the Invention

[0005] The present application provides a text deduplication method, apparatus, device, and medium, which can improve the calculation efficiency of duplicate data information and accurately delete duplicate texts in the case of a large amount of data and limited CPU resources.

[0006] In a first aspect, the present application provides a text deduplication method, the method comprising:

[0007] Obtaining the text features of each text in a preset text library;

[0008] Determining the feature vector of each text according to the text features of each text, and generating a feature matrix between texts in the preset text library according to the feature vectors of each text;

[0009] Determining the similarity information between texts in the preset text library according to the feature matrix, and screening out a first preset number of texts similar to each text in the preset text library according to the similarity information;

[0010] Determining and removing the text to be deduplicated in the preset text library according to the relationship between the similarity information of a second preset number of texts and the similarity information of a third preset number of texts; wherein the sum of the second preset number and the third preset number is equal to the first preset number.

[0011] In one example, determining and removing the text to be deduplicated in the preset text library according to the relationship between the similarity information of the second preset number of texts and the similarity information of the third preset number of texts includes:

[0012] Determining a list of similar texts for each text according to the similarity information of the second preset number of texts;

[0013] If the similarity information of the third preset number of texts is less than the similarity information of the second preset number of texts, then taking the texts in the list of similar texts as the text to be deduplicated in the preset text library, and removing the text to be deduplicated in the preset text library.

[0014] In one example, the method further includes:

[0015] If the similarity information of the third preset number of texts is not less than the similarity information of the second preset number of texts, then recalculating the similarity information of the third preset number of texts;

[0016] If the recalculated similarity information of the third preset number of texts is greater than the similarity information of the second preset number of texts, then updating the texts in the list of similar texts, and taking the updated texts in the list of similar texts as the text to be deduplicated in the preset text library, and removing the text to be deduplicated in the preset text library.

[0017] In one example, determining the similarity information between texts in the preset text library according to the feature matrix includes:

[0018] Determining a sub-feature matrix between texts in the preset text library according to the feature matrix, and determining the similarity information between the texts according to the sub-feature matrix between the texts; wherein, the number of the sub-feature matrices is at least two.

[0019] In one example, determining the sub-feature matrix between texts in the preset text library according to the feature matrix includes:

[0020] Determining the transposed matrix of the feature matrix according to the feature matrix;

[0021] Determining the sub-feature matrix of the feature matrix according to the feature matrix;

[0022] Determining the sub-feature matrix of the transposed matrix of the feature matrix according to the transposed matrix of the feature matrix;

[0023] Taking the sub-feature matrix of the feature matrix and the sub-feature matrix of the transposed matrix of the feature matrix as the sub-feature matrix between texts in the preset text library.

[0024] In one example, determining the similarity information between the texts according to the sub-feature matrix between the texts includes:

[0025] Determining the first cosine similarity information between each sub-feature matrix of each text of the feature matrix and each sub-feature matrix of each text of the transposed matrix of the feature matrix;

[0026] Determining the second cosine similarity information between the feature matrices according to the first cosine similarity information;

[0027] Determining the similarity information between the texts according to the second cosine similarity information.

[0028] In one example, the text features include text content and text labels, and the text labels represent the feature attributes of the texts; if the text feature is text content, determining the feature vector of each text according to the text feature of each text includes:

[0029] Obtaining the word information in the text content of each text;

[0030] Determining the word vector of the word information according to the word information; wherein, the word vector represents the semantic information of the word information;

[0031] Determining the central vector of each text according to the word vector, and using the central vector of each text as the feature vector of each text.

[0032] In one example, the text features include text content and text labels, and the text labels represent the feature attributes of the texts; if the text feature is text label, determining the feature vector of each text according to the text feature of each text includes:

[0033] Obtaining the frequency information and category information in the text labels of each text;

[0034] Determining the label vector of each text according to the frequency information and the category information, and using the label vector of each text as the feature vector of each text.

[0035] In a second aspect, the present application provides a text deduplication device, and the device includes:

[0036] An acquisition unit, configured to acquire the text features of each text in a preset text library;

[0037] A feature matrix generation unit, configured to determine the feature vector of each text according to the text feature of each text, and generate a feature matrix between the texts in the preset text library according to the feature vector of each text;

[0038] A screening unit, configured to determine similarity information between texts in the preset text library according to the feature matrix, and screen out a first preset number of texts similar to each text in the preset text library according to the similarity information;

[0039] A text to be deduplicated determination unit, configured to determine and remove texts to be deduplicated in the preset text library according to the relationship between the similarity information of a second preset number of texts and the similarity information of a third preset number of texts; wherein, the sum of the second preset number and the third preset number is equal to the first preset number.

[0040] In an example, the text to be deduplicated determination unit includes:

[0041] A similar text list determination module, configured to determine a similar text list for each text according to the similarity information of the second preset number of texts;

[0042] A text to be deduplicated removal module, configured to, if the similarity information of the third preset number of texts is less than the similarity information of the second preset number of texts, use the texts in the similar text list as texts to be deduplicated in the preset text library, and remove the texts to be deduplicated in the preset text library.

[0043] In an example, the text to be deduplicated determination unit includes:

[0044] A calculation module, configured to, if the similarity information of the third preset number of texts is not less than the similarity information of the second preset number of texts, recalculate the similarity information of the third preset number of texts;

[0045] A text update module, configured to, if the recalculated similarity information of the third preset number of texts is greater than the similarity information of the second preset number of texts, update the texts in the similar text list, and use the updated texts in the similar text list as texts to be deduplicated in the preset text library, and remove the texts to be deduplicated in the preset text library.

[0046] The screening unit includes:

[0047] A similarity information determination module, configured to determine a sub-feature matrix between texts in the preset text library according to the feature matrix, and determine the similarity information between the texts according to the sub-feature matrix between the texts; wherein, the number of the sub-feature matrices is at least two.

[0048] The similarity information determination module includes:

[0049] A transposed matrix determination sub-module, configured to determine a transposed matrix of the feature matrix according to the feature matrix;

[0050] A first sub-feature matrix determination sub-module, configured to determine a sub-feature matrix of the feature matrix according to the feature matrix;

[0051] A second sub-feature matrix determination sub-module, configured to determine a sub-feature matrix of the transposed matrix of the feature matrix according to the transposed matrix of the feature matrix;

[0052] A third sub-feature matrix determination sub-module, configured to use the sub-feature matrix of the feature matrix and the sub-feature matrix of the transposed matrix of the feature matrix as the sub-feature matrix between texts in the preset text library.

[0053] A similarity information determination module, including:

[0054] A first cosine similarity information determination sub-module, configured to determine first cosine similarity information between each sub-feature matrix of each text of the feature matrix and each sub-feature matrix of each text of the transposed matrix of the feature matrix;

[0055] A second cosine similarity information determination sub-module, configured to determine second cosine similarity information between the feature matrices according to the first cosine similarity information;

[0056] A similarity information determination sub-module, configured to determine similarity information between the texts according to the second cosine similarity information.

[0057] The text features include text content and text labels, and the text labels represent the feature attributes of the text; if the text feature is text content, the feature matrix generation unit includes:

[0058] A word information acquisition module, configured to acquire word information in the text content of each text;

[0059] A word vector determination module, configured to determine a word vector of the word information according to the word information; wherein, the word vector represents the semantic information of the word information;

[0060] A central vector determination module, configured to determine a central vector of each text according to the word vector, and use the central vector of each text as the feature vector of each text.

[0061] The text features include text content and text labels, and the text labels represent the feature attributes of the text; if the text feature is text label, the feature matrix generation unit includes:

[0062] An acquisition module, configured to acquire the frequency information and category information in the text tags of each text;

[0063] A label vector determination module, configured to determine the label vector of each text according to the frequency information and the category information, and use the label vector of each text as the feature vector of each text.

[0064] In a third aspect, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0065] The memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method described in the first aspect.

[0066] In a fourth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in the first aspect.

[0067] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method described in the first aspect.

[0068] A text deduplication method, device, device and medium provided by the present application, by acquiring the text features of each text in a preset text library; determining the feature vector of each text according to the text features of each text, and generating a feature matrix between texts in the preset text library according to the feature vector of each text; determining the similarity information between texts in the preset text library according to the feature matrix, and screening out the first preset number of texts similar to each text in the preset text library according to the similarity information; determining and removing the texts to be deduplicated in the preset text library according to the relationship between the similarity information of the second preset number of texts and the similarity information of the third preset number of texts; wherein, the sum of the second preset number and the third preset number is equal to the first preset number. By adopting the technical solution, it is possible to improve the calculation efficiency of duplicate data information and accurately delete duplicate texts in the case of a large amount of data and limited CPU resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0070] Figure 1 is a schematic flowchart of a text deduplication method provided by Embodiment 1 of the present application;

[0071] Figure 2 It is a flowchart showing a text deduplication method provided in Embodiment 2 of the present application;

[0072] Figure 3 It is a schematic diagram of a text deduplication device provided in Embodiment 3 of the present application;

[0073] Figure 4 It is a schematic diagram of a text deduplication device provided in Embodiment 4 of the present application;

[0074] Figure 5 It is a block diagram of a terminal device shown according to an exemplary embodiment.

[0075] Through the above-mentioned drawings, the specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and text descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Description of Specific Embodiments

[0076] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0077] A text deduplication method, device, equipment and medium provided by the present application are intended to solve the above technical problems in the prior art.

[0078] Next, the technical solution of the present application and how the technical solution of the present application solves the above technical problems will be described in detail with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. Next, the embodiments of the present application will be described with reference to the drawings.

[0079] Figure 1 It is a flowchart showing a text deduplication method provided in Embodiment 1 of the present application. The following steps are included in Embodiment 1:

[0080] S101. Obtain the text features of each text in the preset text library.

[0081] Exemplarily, the preset text library is a database storing multiple texts. The text features include text content and text tags. The text content is composed of multiple words in the text. After obtaining the text content, word segmentation and word removal are performed on the text content. Among them, word removal includes removing stop words, adverbs, auxiliary words, punctuation marks, prepositions, and some conjunctions. After word segmentation and word removal, valid words are obtained, and the text features are composed of the valid words. The text tags are composed of multiple hierarchical tags. Tags at different levels are in a subordinate relationship, and the tags are independent of each other in pairs. For example, the hierarchical tags can be divided into three layers; among them, the first-layer tags can be finance, stocks, Internet finance, trust, entertainment, movies, TV dramas, and European and American stars; the second-layer tags are high-yield stocks, stable stocks, and movies remade from novels; the third layer is stock code 000000, stock code 000001, and Jane Eyre.

[0082] S102. Determine the feature vector of each text according to the text features of each text, and generate the feature matrix between the texts in the preset text library according to the feature vector of each text.

[0083] In this embodiment, the process of determining the feature vector of each text when the text feature is text content is different from the process of determining the feature vector of each text when the text feature is text tag.

[0084] Specifically, when the text feature is text content, the algorithm that can be adopted to determine the feature vector of each text is Word2vec; when the text feature is text tag, the algorithm that can be adopted to determine the feature vector of each text is tag2vec.

[0085] After obtaining the feature vector of each text, the feature vectors of each text can be combined to obtain the feature matrix between the texts in the preset text library. For example, the number of texts in the preset text library is N, and the feature vector of each text is m i , where the range of i is from 1 to N. Assume that the vector dimension of the feature vector m i is d, then the feature matrix of N texts is M ∈ R N×d .

[0086] S103. Determine the similarity information between the texts in the preset text library according to the feature matrix, and screen out the first preset number of texts similar to each text in the preset text library according to the similarity information.

[0087] In this embodiment, according to the feature matrix between the texts in the preset text library, the sub-feature matrix of the text is determined according to the preset dimension. For example, the feature matrix of N texts is M ∈ R N×d , and the sub-feature matrix can be Mi, where the range of i is 1 - k1, then After obtaining the sub-feature matrix, the similarity information between texts in the preset text library can be calculated based on the sub-feature matrix. Further, the number of sub-feature matrices is the same as the number of preset dimensions.

[0088] Exemplarily, when removing the texts to be de-duplicated in the preset text library, according to the similarity information in the preset text library, the first preset number of texts are screened out; specifically, they are sorted according to the magnitude of the similarity information, and the texts with the similarity information ranked among the first preset number are determined. Among them, if the number of texts in the preset text library is N, then the first preset number of texts screened out is n*k.

[0089] S104. Determine and remove the texts to be de-duplicated in the preset text library according to the relationship between the similarity information of the second preset number of texts and the similarity information of the third preset number of texts; wherein, the sum of the second preset number and the third preset number is equal to the first preset number.

[0090] In this embodiment, it is set that the second preset number of texts is k, then the algorithm used to calculate the similarity information of the second preset number of texts is the WMD algorithm. The algorithm used to calculate the similarity information of the third preset number of texts, that is, the similarity information of n*k - k texts, is the RWMD algorithm.

[0091] Further, the WMD algorithm is a new method for measuring text similarity. For two texts D1 and D2, the words in each text are mapped to the embedding space using the word2vec algorithm, and each word in D1 can find a corresponding word in D2, that is, the distance between each pair of words in the embedding space is found, and the minimum value of the sum of the distances of all word pairs is the WMD.

[0092] The mathematical formula is:

[0093]

[0094]

[0095] Among them, is the weight of word i in text d, c i is the word frequency of word i in text d, c(i, j) is the travel cost of words i and j, c(i, j) = ‖x i -x j ‖, x i , x j are the word vectors of words i and j after embedding respectively. The time complexity of WMD calculation is O(P 3 logP), where P is the number of non-repeated words in the text.

[0096] Specifically, the calculation process of pruning the WMD algorithm, i.e., the RWMD algorithm, is as follows:

[0097] Since the time complexity of the WMD calculation efficiency is O(P 3 logP), where P is the total number of words. In this embodiment, the RWMD algorithm is used to screen out k most similar texts for each text from the preset text library.

[0098] RWMD is based on the WMD objective function. One of the two constraint conditions is removed respectively, and then the minimum value is solved. The maximum value of the two minimum values is used as an approximation of WMD. Therefore, WMD needs to be calculated twice.

[0099] For example, if the second constraint condition is removed, the problem becomes:

[0100]

[0101]

[0102] Obviously, the optimal solution to this problem becomes:

[0103] For a word in text D1, find the most similar word in another text D2 and transfer all to this word, that is:

[0104]

[0105] Use l1(d, d′) and l2(d, d′) to represent the minimum values calculated by removing different constraint conditions respectively. The final minimum value of RWMD is l r (d, d′) = max(l1(d, d′), l2(d, d′)), where the RWMD calculation time complexity is O(P 2 ). RWMD is closer to WMD than the cosine distance of the text center vector.

[0106] In this embodiment, the specific process of determining and removing the texts to be de-duplicated in the preset text library is as follows:

[0107] According to the similarity information, screen out n*k most similar texts from the N texts in the preset text library, where n is a hyperparameter; for each text, calculate the WMD of the text content and the WMD of the text label of the first k texts respectively. According to the formula Obtain the WMD similarity information of the text content and the WMD similarity information of the text label, and perform weighted average according to the WMD similarity information of the text content and the text label to obtain the similarity information of the second preset number of texts.

[0108] For each text, calculate the RWMD of the text content and the RWMD of the text label with the remaining n*K - K texts respectively. According to the formula Obtain the wmd similarity of the text content and the wmd similarity of the text tags, and based on the similarity information of the text content and the text tags, perform weighted averaging to obtain the similarity information of the third preset number of texts.

[0109] After obtaining the similarity information of the second preset number of texts and the similarity information of the third preset number of texts, compare the magnitudes of the similarity information of the second preset number of texts and the similarity information of the third preset number of texts, and determine and remove the texts to be deduplicated in the preset text library.

[0110] A text deduplication method provided by this application includes: obtaining the text features of each text in the preset text library; according to the text features of each text, determining the feature vector of each text, and generating a feature matrix between the texts in the preset text library based on the feature vector of each text; according to the feature matrix, determining the similarity information between the texts in the preset text library, and based on the similarity information, screening out the first preset number of texts similar to each text in the preset text library; according to the relationship between the similarity information of the second preset number of texts and the similarity information of the third preset number of texts, determining and removing the texts to be deduplicated in the preset text library; where the sum of the second preset number and the third preset number is equal to the first preset number. By adopting this technical solution, it is possible to improve the calculation efficiency of duplicate data information and accurately delete duplicate texts in the case of a large amount of data and limited CPU resources.

[0111] Figure 2 It is a schematic flowchart of a text deduplication method provided by Embodiment 2 of this application. The following steps are included in Embodiment 2:

[0112] S201. Obtain the text features of each text in the preset text library.

[0113] Exemplarily, this step can refer to the above step S101 and will not be elaborated further.

[0114] S202. According to the text features of each text, determine the feature vector of each text, and generate a feature matrix between the texts in the preset text library based on the feature vector of each text.

[0115] Exemplarily, this step can refer to the above step S102 and will not be elaborated further.

[0116] S203. According to the feature matrix, determine the similarity information between the texts in the preset text library, and based on the similarity information, screen out the first preset number of texts similar to each text in the preset text library.

[0117] In this embodiment, determining the similarity information between the texts in the preset text library according to the feature matrix includes:

[0118] Determine the sub-feature matrix between texts in the preset text library according to the feature matrix, and determine the similarity information between texts according to the sub-feature matrix between texts; wherein, the number of sub-feature matrices is at least two.

[0119] In this embodiment, according to the feature matrix between texts in the preset text library, the sub-feature matrix of the text is determined according to the preset dimension. For example, the feature matrix of N texts is M∈R N×d , and the sub-feature matrix can be Mi, where the range of i is 1-k1, then After obtaining the sub-feature matrix, the similarity information between texts in the preset text library can be calculated according to the sub-feature matrix. Further, the number of sub-feature matrices is the same as the number of preset dimensions.

[0120] In this embodiment, determining the sub-feature matrix between texts in the preset text library according to the feature matrix includes:

[0121] According to the feature matrix, determine the transpose matrix of the feature matrix; according to the feature matrix, determine the sub-feature matrix of the feature matrix; according to the transpose matrix of the feature matrix, determine the sub-feature matrix of the transpose matrix of the feature matrix; use the sub-feature matrix of the feature matrix and the sub-feature matrix of the transpose matrix of the feature matrix as the sub-feature matrix between texts in the preset text library.

[0122] In this embodiment, it is assumed that there are N texts in the preset text library, and each text vector is m1, where the vector dimension is d. Concatenate m1,..., m N into a large matrix M∈R N×d . If the CPU resources are large enough, the cosine similarity calculation of N texts pairwise can directly use the multiplication of the large matrix of N texts by the transpose of the large matrix of N texts. The specific formula is:

[0123] simi = M * M t ;

[0124] where simi∈R N×N , where simi(k, q) represents the cosine similarity between text k and text q, and the time complexity is O(1).

[0125] Considering the large amount of text data and limited CPU resources, to improve the calculation efficiency, through block matrix multiplication, the M matrix is evenly cut into k1 blocks in column order, that is, a matrix composed of multiple sub-feature matrices is obtained:

[0126]

[0127] Cut the M t matrix evenly into k2 blocks in column order, that is, obtain the sub-feature matrix of the transpose matrix of the feature matrix

[0128] Exemplarily, determining the similarity information between texts according to the sub-feature matrix between texts includes:

[0129] Determining the first cosine similarity information between each sub-feature matrix of each text in the feature matrix and each sub-feature matrix of each text in the transposed matrix of the feature matrix;

[0130] Determining the second cosine similarity information between the feature matrices according to the first cosine similarity information;

[0131] Determining the similarity information between texts according to the second cosine similarity information.

[0132] In this embodiment, calculating the pairwise cosine similarities in the preset text library is transformed into pairwise multiplication of all blocks of the M matrix and all block matrices of the M t matrix, and the time complexity is O(k1*k2), where k1 << N and k2 << N. The specific calculation method refers to the following pseudocode. The first layer of loop runs k1 times, and the second layer of loop runs k2 times. The result of each calculation is the cosine similarity result of the (i, j) block:

[0133]

[0134] where where simi(i, j)(k, q) represents the cosine similarity of the text, where k1 and k2 are hyperparameters that can be adjusted. Among them, simi(i, j) is the first cosine similarity information. According to the first cosine similarity information and the cosine similarity distance The second cosine similarity information of the feature matrix is obtained from the first cosine similarity information between each sub-feature matrix, where simi(i, j) is the d value in the cosine similarity distance. After obtaining the second cosine similarity information, the similarity information between texts is determined according to the second similarity information.

[0135] S204. Determining the list of similar texts for each text according to the similarity information of the second preset number of texts.

[0136] According to the similarity information, n*k of the most similar texts are selected from the N texts in the preset text library, where n is a hyperparameter; the wmd of the text content and the wmd of the text labels of the first k texts are calculated, where k is the second preset number. According to the formula The wmd similarity information of the text content and the wmd similarity information of the text labels are obtained. The wmd similarity information of the text content and the text labels is weighted and averaged to obtain the similarity information of the second preset number of texts, and the second preset number of texts is used as the list of similar texts.

[0137] S205. If the similarity information of the texts of the third preset quantity is less than the similarity information of the texts of the second preset quantity, then use the texts in the similar text list as the texts to be de-duplicated in the preset text library, and remove the texts to be de-duplicated in the preset text library.

[0138] In this embodiment, if the similarity information of the texts of the third preset quantity is less than the similarity information of the texts of the second preset quantity, it indicates that the similarity of the texts of the third preset quantity is less than that of the texts of the second preset quantity. Therefore, the similar text list composed of the texts of the second preset quantity is the relatively similar texts in the preset text library. Therefore, remove the texts in the similar text list.

[0139] S206. If the similarity information of the texts of the third preset quantity is not less than the similarity information of the texts of the second preset quantity, then recalculate the similarity information of the texts of the third preset quantity.

[0140] In this embodiment, if the similarity information of the texts of the third preset quantity is not less than the similarity information of the texts of the second preset quantity, then recalculate the similarity information of the texts of the third preset quantity. When calculating the similarity information of the texts of the third preset quantity, use the WMD algorithm.

[0141] S207. If the recalculated similarity information of the texts of the third preset quantity is greater than the similarity information of the texts of the second preset quantity, then update the texts in the similar text list, and use the texts in the updated similar text list as the texts to be de-duplicated in the preset text library, and remove the texts to be de-duplicated in the preset text library.

[0142] In this embodiment, if the recalculated similarity information of the texts of the third preset quantity is greater than the similarity information of the texts of the second preset quantity, it indicates that there are more similar texts. Therefore, after replacing the texts in the similar text list, remove the texts in the similar text list.

[0143] In an optional embodiment, the text features include text content and text labels, and the text labels represent the characteristic attributes of the texts; if the text feature is text content, according to the text features of each text, determine the feature vector of each text, including:

[0144] Obtain the word information in the text content of each text; according to the word information, determine the word vector of the word information; wherein, the word vector represents the semantic information of the word information; according to the word vector, determine the central vector of each text, and use the central vector of each text as the feature vector of each text.

[0145] In this embodiment, assume that after text d is segmented and stop words are removed, the remaining valid words are w1, w2…, w n, the word frequency corresponding to each word is c1, c2…, c n , the text d can be represented as [d1, d2…, d n , where is the weight of word i in a text, where c i represents the number of times word i appears in text d, and the denominator represents the total number of words in this text (after removing words). After word segmentation, word removal, and the normalized bag-of-words model, the weight of each word in text d can be obtained. Using the word2vec technology to build a language model, map words into a mathematical space to form word embeddings, and the formed word embeddings have rich semantic information. Combining the normalized bag-of-words model with word embeddings to vectorize the text, the specific combination method is shown in the following mathematical formula to form doc embeddings:

[0146] Assume that the word vector of word i is x i , d i is the normalized word frequency of word i, According to the cosine similarity distance, calculate the pairwise cosine similarities of the preset text library, where the cosine similarity distance is:

[0147]

[0148] In an optional embodiment, the text features include text content and text labels, and the text labels characterize the feature attributes of the text; if the text feature is a text label, according to the text features of each text, determine the feature vector of each text, including:

[0149] Obtain the frequency information and category information in the text labels of each text; according to the frequency information and category information, determine the label vector of each text, and use the label vector of each text as the feature vector of each text.

[0150] In this embodiment, count and normalize the frequencies of all labels of the text, and the specific method is similar to the text content word frequency normalization technology. Assume that all labels of text d are t1, t2…, t n , the frequency corresponding to each label is c1, c2…, c n , the text d can be represented as [d1, d2…, d n , where is the weight of label i in text d, where c iDenote the frequency of tag i in text d (the difference in frequencies is mainly reflected in the entity layer, while the tag frequencies in the theme layer and concept layer are 1). Among them, the entity layer, theme layer, and concept layer are three levels of tags, and the hierarchical relationship among the three levels is theme layer > concept layer > entity layer. The denominator represents the total number of tags in this text. Obtain the embeddings of all tags of the text from the tag embedding, and obtain the tag vector of this text by weighted average with the normalized tag frequency and tag weight. The specific calculation method is as follows: Assume the vector of tag i is x i , d i To obtain the normalized frequency of tag i, w i The weight of tag i, the text tag vector can be expressed as:

[0151]

[0152] According to the cosine similarity distance, calculate the pairwise cosine similarities of the tags in the preset text library. Among them, the cosine similarity distance is:

[0153]

[0154] A text deduplication method provided by the present application, by obtaining the text features of each text in the preset text library, determining the feature vector of each text according to the text features of each text, generating a feature matrix between the texts in the preset text library according to the feature vector of each text, determining the similarity information between the texts in the preset text library according to the feature matrix, screening out the first preset number of texts similar to each text in the preset text library according to the similarity information, determining the similar text list of each text according to the similarity information of the second preset number of texts, if the similarity information of the third preset number of texts is less than the similarity information of the second preset number of texts, then taking the texts in the similar text list as the texts to be deduplicated in the preset text library and removing the texts to be deduplicated in the preset text library, if the similarity information of the third preset number of texts is not less than the similarity information of the second preset number of texts, then recalculating the similarity information of the third preset number of texts, if the recalculated similarity information of the third preset number of texts is greater than the similarity information of the second preset number of texts, then updating the texts in the similar text list and taking the updated texts in the similar text list as the texts to be deduplicated in the preset text library and removing the texts to be deduplicated in the preset text library. In this technical solution, by comparing the relationship between the similarity information of the second preset number of texts and the similarity information of the third preset number of texts, the texts to be deduplicated are determined, and this process makes full use of the text content similarity information and text tag similarity information to measure the text similarity, improving the accuracy of removing duplicate texts.

[0155] Figure 3It is a schematic diagram of a text deduplication device provided according to Embodiment 3 of the present application. The device 30 in Embodiment 3 includes:

[0156] An acquisition unit 301, configured to acquire the text features of each text in a preset text library.

[0157] A feature matrix generation unit 302, configured to determine the feature vector of each text according to the text features of each text, and generate a feature matrix between texts in the preset text library according to the feature vectors of each text.

[0158] A screening unit 303, configured to determine the similarity information between texts in the preset text library according to the feature matrix, and screen out the first preset number of texts similar to each text in the preset text library according to the similarity information.

[0159] A text to be deduplicated determination unit 304, configured to determine and remove the text to be deduplicated in the preset text library according to the relationship between the similarity information of the second preset number of texts and the similarity information of the third preset number of texts; wherein, the sum of the second preset number and the third preset number is equal to the first preset number.

[0160] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described device can refer to the corresponding process in the foregoing method embodiment, and will not be elaborated herein.

[0161] Figure 4 It is a schematic diagram of a text deduplication device provided according to Embodiment 4 of the present application. The device 40 in Embodiment 4 includes:

[0162] An acquisition unit 401, configured to acquire the text features of each text in a preset text library.

[0163] A feature matrix generation unit 402, configured to determine the feature vector of each text according to the text features of each text, and generate a feature matrix between texts in the preset text library according to the feature vectors of each text.

[0164] A screening unit 403, configured to determine the similarity information between texts in the preset text library according to the feature matrix, and screen out the first preset number of texts similar to each text in the preset text library according to the similarity information.

[0165] A text to be deduplicated determination unit 404, configured to determine and remove the text to be deduplicated in the preset text library according to the relationship between the similarity information of the second preset number of texts and the similarity information of the third preset number of texts; wherein, the sum of the second preset number and the third preset number is equal to the first preset number.

[0166] In one example, the text to be deduplicated determination unit 404 includes:

[0167] A similar text list determination module 4041, configured to determine a similar text list for each text according to the similarity information of a second preset number of texts.

[0168] A text to be deduplicated removal module 4042, configured to, if the similarity information of a third preset number of texts is less than the similarity information of the second preset number of texts, use the texts in the similar text list as the texts to be deduplicated in the preset text library, and remove the texts to be deduplicated in the preset text library.

[0169] In an example, the text to be deduplicated determination unit 404 includes:

[0170] A calculation module 4043, configured to, if the similarity information of a third preset number of texts is not less than the similarity information of the second preset number of texts, recalculate the similarity information of the third preset number of texts.

[0171] A text update module 4044, configured to, if the recalculated similarity information of the third preset number of texts is greater than the similarity information of the second preset number of texts, update the texts in the similar text list, use the texts in the updated similar text list as the texts to be deduplicated in the preset text library, and remove the texts to be deduplicated in the preset text library.

[0172] A screening unit 403 includes:

[0173] A similarity information determination module 4031, configured to determine a sub-feature matrix between texts in the preset text library according to the feature matrix, and determine the similarity information between texts according to the sub-feature matrix between texts; wherein, the number of sub-feature matrices is at least two.

[0174] A similarity information determination module 4032 includes:

[0175] A transposed matrix determination sub-module 40321, configured to determine the transposed matrix of the feature matrix according to the feature matrix.

[0176] A first sub-feature matrix determination sub-module 40322, configured to determine the sub-feature matrix of the feature matrix according to the feature matrix.

[0177] A second sub-feature matrix determination sub-module 40323, configured to determine the sub-feature matrix of the transposed matrix of the feature matrix according to the transposed matrix of the feature matrix.

[0178] A third sub-feature matrix determination sub-module 40324, configured to use the sub-feature matrix of the feature matrix and the sub-feature matrix of the transposed matrix of the feature matrix as the sub-feature matrix between texts in the preset text library.

[0179] The similarity information determination module 4032 includes:

[0180] The first cosine similarity information determination sub-module 40325 is used to determine the first cosine similarity information between each sub-feature matrix of each text in the feature matrix and each sub-feature matrix of each text in the transposed matrix of the feature matrix.

[0181] The second cosine similarity information determination sub-module 40326 is used to determine the second cosine similarity information between feature matrices according to the first cosine similarity information.

[0182] The similarity information determination sub-module 40327 is used to determine the similarity information between texts according to the second cosine similarity information.

[0183] The text features include text content and text labels, and the text labels represent the feature attributes of the text; if the text feature is text content, the feature matrix generation unit 402 includes:

[0184] The word information acquisition module 4021 is used to acquire the word information in the text content of each text.

[0185] The word vector determination module 4022 is used to determine the word vector of the word information according to the word information; wherein, the word vector represents the semantic information of the word information.

[0186] The central vector determination module 4023 is used to determine the central vector of each text according to the word vector, and use the central vector of each text as the feature vector of each text.

[0187] The text features include text content and text labels, and the text labels represent the feature attributes of the text; if the text feature is text label, the feature matrix generation unit 402 includes:

[0188] The acquisition module 4024 is used to acquire the frequency information and category information in the text labels of each text.

[0189] The label vector determination module 4025 is used to determine the label vector of each text according to the frequency information and category information, and use the label vector of each text as the feature vector of each text.

[0190] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described device can refer to the corresponding process in the foregoing method embodiment, and will not be elaborated herein.

[0191] Figure 5FIG. 0 is a block diagram of a terminal device shown according to an exemplary embodiment, and the device may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0192] Device 500 may include one or more of the following components: processing component 802, memory 804, power component 806, multimedia component 808, audio component 510, input / output (I / O) interface 512, sensor component 514, and communication component 516.

[0193] Processing component 502 generally controls the overall operation of device 500, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. Processing component 502 may include one or more processors 520 to execute instructions to complete all or part of the steps of the above methods. In addition, processing component 502 may include one or more modules to facilitate the interaction between processing component 502 and other components. For example, processing component 502 may include a multimedia module to facilitate the interaction between multimedia component 508 and processing component 502.

[0194] Memory 504 is configured to store various types of data to support the operation of device 500. Examples of these data include instructions for any application or method operating on device 500, contact data, phone book data, messages, pictures, videos, etc. Memory 504 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0195] Power component 506 provides power to various components of device 500. Power component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for device 500.

[0196] The multimedia component 508 includes a screen that provides an output interface between the device 500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 508 includes a front camera and / or a rear camera. When the device 500 is in an operation mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0197] The audio component 510 is configured to output and / or input audio signals. For example, the audio component 510 includes a microphone (MIC) that is configured to receive external audio signals when the device 500 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 504 or transmitted via the communication component 516. In some embodiments, the audio component 510 further includes a speaker for outputting audio signals.

[0198] The I / O interface 512 provides an interface between the processing component 502 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0199] The sensor component 514 includes one or more sensors for providing an assessment of the status of various aspects of the device 500. For example, the sensor component 514 can detect the on / off state of the device 500, the relative positioning of components, such as the display and the keypad of the device 500. The sensor component 514 can also detect a change in the position of the device 500 or a component of the device 500, the presence or absence of user contact with the device 500, the orientation or acceleration / deceleration of the device 500, and the temperature change of the device 500. The sensor component 514 can include a proximity sensor that is configured to detect the presence of nearby objects without any physical contact. The sensor component 514 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 514 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0200] The communication component 516 is configured to facilitate communication, in a wired or wireless manner, between the device 500 and other devices. The device 500 may access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 516 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 516 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0201] In an exemplary embodiment, the device 500 may be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0202] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as a memory 504 including instructions, is also provided. The above instructions may be executed by a processor 520 of the device 500 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a Random Access Memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0203] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a terminal device, enables the terminal device to execute the above text deduplication method of the terminal device.

[0204] This application also discloses a computer program product, including a computer program that, when executed by a processor, implements the method as described in this embodiment.

[0205] The various embodiments of the systems and techniques described above in this application can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0206] The program code for implementing the methods of this application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or electronic device.

[0207] In the context of this application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0208] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0209] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data electronic device), or a computing system including middleware components (e.g., an application electronic device), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0210] A computer system can include a client and an electronic device. The client and the electronic device are generally far from each other and typically interact through a communication network. The relationship between the client and the electronic device is created by computer programs running on respective computers and having a client-server relationship with each other. The electronic device can be a cloud electronic device, also known as a cloud computing electronic device or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The electronic device can also be an electronic device of a distributed system, or an electronic device combined with a blockchain. It should be understood that the various forms of processes shown above can be reordered, added, or deleted steps. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved, and no limitation is made herein.

[0211] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.

[0212] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A text deduplication method, characterized in that, The method includes: Obtaining the text features of each text in a preset text library; wherein, the text features include text content and text labels, and the text labels represent the characteristic attributes of the text; the text labels are composed of multiple hierarchical labels, and the labels at different levels are in a subordinate relationship; Determining the feature vector of each text according to the text features of each text, and generating a feature matrix between the texts in the preset text library according to the feature vector of each text; Determining the similarity information between the texts in the preset text library according to the feature matrix, and screening out the first preset number of texts similar to each text in the preset text library according to the similarity information; Determining and removing the texts to be deduplicated in the preset text library according to the relationship between the similarity information of the second preset number of texts and the similarity information of the third preset number of texts; wherein, the sum of the second preset number and the third preset number is equal to the first preset number.

2. The method according to claim 1, characterized in that, Determining and removing the texts to be deduplicated in the preset text library according to the relationship between the similarity information of the second preset number of texts and the similarity information of the third preset number of texts includes: Determining a list of similar texts for each text according to the similarity information of the second preset number of texts; If the similarity information of the third preset number of texts is less than the similarity information of the second preset number of texts, then using the texts in the list of similar texts as the texts to be deduplicated in the preset text library, and removing the texts to be deduplicated in the preset text library.

3. The method according to claim 2, wherein The method further includes: If the similarity information of the third preset number of texts is not less than the similarity information of the second preset number of texts, then recalculating the similarity information of the third preset number of texts; If the recalculated similarity information of the third preset number of texts is greater than the similarity information of the second preset number of texts, then updating the texts in the list of similar texts, and using the updated texts in the list of similar texts as the texts to be deduplicated in the preset text library, and removing the texts to be deduplicated in the preset text library.

4. The method according to claim 1, wherein Determining the similarity information between the texts in the preset text library according to the feature matrix includes: Determining a sub-feature matrix between the texts in the preset text library according to the feature matrix, and determining the similarity information between the texts according to the sub-feature matrix between the texts; wherein, the number of the sub-feature matrices is at least two.

5. The method according to claim 4, characterized in that, Determining the sub-feature matrix between the texts in the preset text library according to the feature matrix includes: Determining the transpose matrix of the feature matrix according to the feature matrix; Determining the sub-feature matrix of the feature matrix according to the feature matrix; Determining the sub-feature matrix of the transpose matrix of the feature matrix according to the transpose matrix of the feature matrix; Using the sub-feature matrix of the feature matrix and the sub-feature matrix of the transpose matrix of the feature matrix as the sub-feature matrix between the texts in the preset text library.

6. The method according to claim 4, characterized in that, Determining the similarity information between the texts according to the sub-feature matrix between the texts includes: Determine the first cosine similarity information between each sub - feature matrix of each text in the feature matrix and each sub - feature matrix of each text in the transposed matrix of the feature matrix; Determine the second cosine similarity information between the feature matrices according to the first cosine similarity information; Determine the similarity information between the texts according to the second cosine similarity information.

7. The method according to claim 1, characterized in that, If the text feature is text content, according to the text feature of each text, determine the feature vector of each text, including: Obtain the word information in the text content of each text; According to the word information, determine the word vector of the word information; wherein, the word vector represents the semantic information of the word information; According to the word vector, determine the central vector of each text, and use the central vector of each text as the feature vector of each text.

8. The method according to claim 1, wherein If the text feature is a text label, according to the text feature of each text, determine the feature vector of each text, including: Obtain the frequency information and category information in the text label of each text; According to the frequency information and the category information, determine the label vector of each text, and use the label vector of each text as the feature vector of each text.

9. A text deduplication device, characterized in that, The device includes: An acquisition unit, configured to acquire the text features of each text in a preset text library; wherein, the text features include text content and text labels, and the text labels represent the feature attributes of the texts; the text labels are composed of multiple hierarchical labels, and the labels at different levels are in a subordinate relationship; A feature matrix generation unit, configured to determine the feature vector of each text according to the text feature of each text, and generate a feature matrix between texts in the preset text library according to the feature vector of each text; A screening unit, configured to determine the similarity information between texts in the preset text library according to the feature matrix, and screen out the first preset number of texts similar to each text in the preset text library according to the similarity information; A text - to - be - deduplicated determination unit, configured to determine and remove the text - to - be - deduplicated in the preset text library according to the relationship between the similarity information of the second preset number of texts and the similarity information of the third preset number of texts; wherein, the sum of the second preset number and the third preset number is equal to the first preset number.

10. An electronic device, characterized in that, Comprising: A processor, and a memory communicatively connected to the processor; The memory stores computer - executable instructions; The processor executes the computer - executable instructions stored in the memory to implement the method according to any one of claims 1 - 8.

11. A computer-readable storage medium, characterized in that, The computer - readable storage medium stores computer - executable instructions, and when the computer - executable instructions are executed by a processor, they are used to implement the method according to any one of claims 1 - 8.

12. A computer program product, characterized in that, Comprising a computer program, which when executed by a processor implements the method according to any one of claims 1 - 8.

Citation Information

Patent Citations

  • Duplicate checking method, terminal equipment and computer readable storage medium

    CN112948545A

  • Text duplicate checking method and device, equipment and readable storage medium

    CN113641800A

  • Document similarity calculating system and method thereof

    KR1020100064297A