A method, apparatus and device for text deduplication
By splitting the text into entity keywords and descriptive keywords, and calculating sentence vectors using the text classification model, the problem of poor text deduplication in the existing technology is solved, and a more accurate text deduplication effect is achieved.
Patent Information
- Application Number
- CN201910384114.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-05-09
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2039-05-09
AI Technical Summary
In the prior art, when deduplication of text, the similarity is calculated based on the keywords after word segmentation, and the text semantics cannot be accurately characterized, resulting in poor deduplication effect.
By splitting the first feedback text feedback from the target object into entity keywords and description keywords, using the text classification model to determine the word vector, calculate the sentence vector, and similarity with the feedback text in the preset text vector library, text deduplication is achieved.
It improves the accuracy of text deduplication and can calculate text similarity more accurately, thereby achieving accurate and efficient deduplication of text.
Smart Images

Figure CN110162630B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of Internet text analysis, and in particular, to a method, device, and equipment for text deduplication. Background Art
[0002] For a new game or a new version of a game, it will be tested before being officially applied. For example, hundreds of players are recruited to experience the game and feedback on the defects in the game. Usually, multiple players use different expressions or descriptions for the same problem. When counting game defects later, it is necessary to find and extract duplicate feedbacks for different descriptions.
[0003] When the prior art performs text deduplication, it tokenizes the text to be deduplicated; then, directly extracts keywords from the tokenization; next, calculates the similarity between the keywords of two texts; finally, performs text deduplication based on the similarity between the keywords of the texts. In the above-mentioned prior art method for text deduplication, directly using the keywords extracted after tokenization as the basis for calculating the similarity between two texts, due to the single keyword information, it often cannot accurately represent the semantics of the text, and based on the keywords, it is impossible to accurately calculate the similarity between texts, resulting in poor text deduplication effect. Therefore, it is necessary to provide a more effective method for text deduplication to improve the text deduplication effect. Summary of the Invention
[0004] This application provides a method, device, and equipment for text deduplication, which can accurately calculate the similarity between the first feedback text fed back by the target object and the second feedback text in the preset text vector library, thereby improving the accuracy of text deduplication.
[0005] On the one hand, this application provides a method for text deduplication, the method includes:
[0006] Based on the first feedback text fed back by the target object, determine the entity keywords and description keywords in the first feedback text;
[0007] Based on the text classification model, determine the first word vector of the entity keywords and the second word vector of the description keywords;
[0008] Based on the first word vector and the second word vector, determine the sentence vector of the first feedback text;
[0009] Calculate the similarity between the sentence vector of the first feedback text and the sentence vector of the second feedback text in the preset text vector library, where the preset text vector library includes the mapping relationship between the preset second feedback text and the sentence vector;
[0010] Based on the similarity, perform deduplication processing on the first feedback text.
[0011] On the other hand, a text deduplication device is provided, and the device includes:
[0012] A keyword determination module, configured to determine an entity keyword and a description keyword in the first feedback text based on the first feedback text fed back by the target object;
[0013] A word vector determination module, configured to determine a first word vector of the entity keyword and a second word vector of the description keyword based on a text classification model;
[0014] A sentence vector determination module, configured to determine a sentence vector of the first feedback text based on the first word vector and the second word vector;
[0015] A similarity calculation module, configured to calculate a similarity between the sentence vector of the first feedback text and the sentence vector of a second feedback text in a preset text vector library, where the preset text vector library includes a mapping relationship between a preset second feedback text and a sentence vector;
[0016] A deduplication processing module, configured to perform deduplication processing on the first feedback text based on the similarity.
[0017] On the other hand, a text deduplication device is provided, and the device includes: a processor and a memory, where at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the text deduplication method as described above.
[0018] On the other hand, a computer-readable storage medium is provided, where at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the text deduplication method as described above.
[0019] The text deduplication method, device and device provided by this application have the following technical effects:
[0020] Based on the first feedback text fed back by the target object, this application splits the first feedback text into two parts: an entity keyword and a description keyword, that is, the first feedback text is refined and classified, so as to facilitate the text classification model to quickly and accurately determine the first word vector of the entity keyword and the second word vector of the description keyword; then based on the first word vector and the second word vector, the sentence vector of the first feedback text can be accurately obtained; based on the sentence vector, the similarity between the first feedback text and the second feedback text is further accurately calculated, so as to achieve accurate and efficient text deduplication. Description of the Drawings
[0021] To more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0022] Figure 1 It is a schematic diagram of a text deduplication system provided by an embodiment of the present application;
[0023] Figure 2 It is a schematic flowchart of a method for text deduplication provided by an embodiment of the present application;
[0024] Figure 3 It is a schematic flowchart of a method for determining entity keywords and description keywords in the first feedback text based on the first feedback text of the target object provided by an embodiment of the present application;
[0025] Figure 4 It is a schematic flowchart of a method for calculating the weighted average of the first word vector and the second word vector provided by an embodiment of the present application;
[0026] Figure 5 It is a schematic structural diagram for determining entity keywords and description keywords based on the first feedback text provided by an embodiment of the present application;
[0027] Figure 6 It is a schematic diagram of the fastText model architecture provided by an embodiment of the present application;
[0028] Figure 7 It is a schematic diagram of the Huffman tree structure provided by an embodiment of the present application;
[0029] Figure 8 It is a schematic diagram of a display interface for the titles and similarities of five second feedback texts corresponding to the "Saint Seiya" game provided by an embodiment of the present application;
[0030] Figure 9 It is a schematic diagram of a display interface for the titles and similarities of five second feedback texts corresponding to the "PUBG Mobile" game provided by an embodiment of the present application;
[0031] Figure 10 It is another schematic diagram of a display interface for the titles and similarities of five second feedback texts corresponding to the "PUBG Mobile" game provided by an embodiment of the present application;
[0032] Figure 11 It is a schematic structural diagram of a text deduplication device provided by an embodiment of the present application;
[0033] Figure 12 It is a schematic structural diagram of a weighted average calculation sub-module provided by an embodiment of the present application;
[0034] Figure 13 It is a schematic structural diagram of a server provided by an embodiment of the present application. Detailed implementation manners
[0035] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0036] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0037] Please refer to Figure 1 , Figure 1 It is a schematic diagram of a text deduplication system provided by an embodiment of the present application. As Figure 1 shown, the text deduplication system may at least include a server 01 and a client 02.
[0038] Specifically, in the embodiments of this specification, the server 01 may include an independently operating server, or a distributed server, or a server cluster composed of multiple servers. The server 01 may include a network communication unit, a processor, a memory, and so on. Specifically, the server 01 may be used to perform text deduplication processing.
[0039] Specifically, in the embodiments of this specification, the client 02 may include physical devices such as smartphones, desktop computers, tablet computers, laptop computers, digital assistants, and smart wearable devices, or may include software running on the physical devices, such as web pages provided by some service providers to users, or may also be applications provided by these service providers to users. Specifically, the client 02 may be used to query the similarity between feedback texts online.
[0040] The following introduces a method for text deduplication in this application. Figure 2 It is a schematic flowchart of a method for text deduplication provided by an embodiment of this application. This specification provides the method operation steps as described in the embodiment or flowchart, but based on routine or non-creative labor, there may be more or fewer operation steps. The step order listed in the embodiment is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual system or server product executes, it can be executed in the order of the embodiment or as shown in the drawings, or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing). Specifically, as Figure 2 shown, the method may include:
[0041] S201: Based on the first feedback text fed back by the target object, determine the entity keywords and description keywords in the first feedback text.
[0042] In the embodiments of this specification, the target object may include a user, a client. The client may include physical devices such as smartphones, desktop computers, tablet computers, laptop computers, digital assistants, and smart wearable devices, or may include software running on the physical devices, such as web pages provided by some service providers to users, or may also be applications provided by these service providers to users.
[0043] In the embodiments of this specification, the first feedback text may include comments and suggestions of the user for one or more entity objects or virtual objects. For example, the first feedback text may include comments of the user for a physical store, comments of the user for an application program (such as a game application program), or improvement suggestions put forward. The first feedback text may include one or more sentences.
[0044] In the embodiments of this specification, the entity keywords may include nouns, verbs; the description keywords are used to describe the entity keywords, and the description keywords may include adjectives; as Figure 5 shown, based on one first feedback text, one entity keyword and one description keyword can be obtained; for example, the first feedback text is: "I found a bug (fault, vulnerability), and the task reward cannot be received", and its corresponding entity keyword is "task reward", and the description keyword is "cannot be received";
[0045] In the embodiments of this specification, a sentence may include one or more entity keywords, and an entity keyword may be described by one or more descriptive keywords; thus, the first feedback text may include one or more entity keywords and one or more descriptive keywords at the same time.
[0046] Specifically, in the embodiments of this specification, as Figure 3 shown, determining the entity keywords and descriptive keywords in the first feedback text based on the first feedback text for the target object may include:
[0047] S2011: Obtain the original entity keywords and original descriptive keywords in the first feedback text;
[0048] In the embodiments of this specification, the original entity keywords and original descriptive keywords in the first feedback text may be obtained through a preset keyword library;
[0049] Before the step of obtaining the original entity keywords and original descriptive keywords in the first feedback text, it may include:
[0050] Preset a keyword library, the keyword library includes an entity keyword library and a descriptive keyword library, and the keyword library is used to extract the original entity keywords and original descriptive keywords in the first feedback text.
[0051] In the embodiments of this specification, the first feedback text is split into two parts: entity keywords and descriptive keywords, that is, the first feedback text is refined and classified, so as to facilitate the text classification model to quickly and accurately determine the first word vector of the entity keywords and the second word vector of the descriptive keywords.
[0052] S2013: Based on a preset synonym library, perform synonym replacement processing on the original entity keywords and the original descriptive keywords to obtain the standard entity keywords corresponding to the original entity keywords and the standard descriptive keywords corresponding to the original descriptive keywords;
[0053] In the embodiments of this specification, the synonym library is used to normalize different keywords, and the synonym library may include the mapping relationship shown in Table 1 below, and the synonym library may replace associated synonyms with standard synonyms.
[0054] Table 1: Mapping relationship in the synonym library
[0055] Standard synonyms Related synonyms Unable to receive Cannot receive, unable to receive, the reward has not been issued, receiving failed AK AKM, AK47
[0056] In the embodiments of this specification, the keyword library may include a synonym library. After the keyword library extracts the original entity keywords and original description keywords from the first feedback text, synonym replacement is performed on the two types of keywords respectively.
[0057] In the embodiments of this specification, performing synonym replacement before calculating the similarity of the user feedback text by the synonym library can well normalize different expressions of the same concept and synonymous viewpoints, optimize the processing flow, and at the same time improve the accuracy of similarity calculation. The application of the synonym library can expand and recall more keywords with different language expressions, further improving the accuracy of text deduplication.
[0058] S2015: Use the standard entity keyword as the entity keyword and the standard description keyword as the description keyword.
[0059] In the embodiments of this specification, before step S201, the method may further include:
[0060] Determine the first feedback text provided by the target object.
[0061] In the embodiments of this specification, the feedback text of the target object within a preset time period may be used as the first feedback text.
[0062] In the embodiments of this specification, after the step of determining the first feedback text provided by the target object, the method further includes:
[0063] Perform data filtering on the first feedback text.
[0064] In the embodiments of this specification, when the target object provides the first feedback text, there is a corresponding feedback template. For example, when the target object is a game player, its feedback module may include information such as the time when the problem appears in the game, the model of the device where the problem is located, and the game version where the problem is located. In practical applications, data filtering can be performed on this feedback module to delete the invalid information in the first feedback text and only retain the core feedback information of the player (i.e., the information in the feedback template).
[0065] In the embodiments of this specification, by performing data filtering on the first feedback text, the invalid information in the first feedback text can be removed, thereby facilitating the subsequent quick determination of the entity keywords and description keywords in the first feedback text.
[0066] S203: Based on the text classification model, determine the first word vector of the entity keyword and the second word vector of the description keyword.
[0067] In the embodiments of this specification, the text classification model is used to calculate the word vectors of keywords; the text classification model may include word2vec (word to vector, text vectorization), SVM (Support Vector Machine), Logistic Regression, neural network, and fastText model. FastText is a text classifier open-sourced by Facebook AI Research in 2016, and its feature is fast. Compared with other text classification models, such as SVM, Logistic Regression, and neural network models, fastText greatly shortens the training time while maintaining the classification effect. The fastText model inputs a sequence of words (a piece of text or a sentence) and outputs the probabilities that this word sequence belongs to different categories. The words and phrases in the sequence form feature vectors, and the feature vectors are mapped to the intermediate layer through a linear transformation, and the intermediate layer is then mapped to the labels. FastText splits a word into subwords, and uses the average of the subword vectors as the word vector, which can effectively solve the problem of out-of-vocabulary words.
[0068] In the embodiments of this specification, the fastText model includes three parts: model architecture, hierarchical Softmax, and N-gram features. Softmax is a normalized exponential function used to normalize probability values; the conventional Softmax is applied to multi-classification tasks. In this model, hierarchical Softmax essentially transforms the global multi-classification problem into multiple binary classification problems, thereby reducing the computational complexity from O(N) to O(logN);
[0069] N-gram is a concept in the fields of computational linguistics and probability theory, referring to a sequence of N items in a given piece of text or speech. The item can be a syllable, a letter, a word, or a base pair. Usually, N-grams are taken from text or a corpus. When N = 1, it is called unigram, when N = 2, it is called bigram, when N = 3, it is called trigram, and so on.
[0070] such as Figure 6As shown in the figure, the fastText model architecture has three layers, including an input layer, a hidden layer, and an output layer. Among them, X1, X2, X3, ……, Xn correspond to the input layer. The words and phrases in the input layer are composed into feature vectors, and then the feature vectors are mapped to the hidden layer through a linear transformation. The hidden layer solves the maximum likelihood function, and then constructs a Huffman tree according to the weights of each category and the model parameters, and takes the Huffman tree as the output.
[0071] As Figure 7 shown in the figure, a Huffman tree is constructed using the frequencies of the keywords. All leaf nodes are all keywords, and non-leaf nodes are internal parameters. Then the probability P(yj) of y j is calculated by the following formula:
[0072]
[0073] where σ represents the sigmod function, LC represents the left child, f(m) is a specific function (if m = true, then f(m) is 1, otherwise f(m) is -1), θ represents the parameter of the non-leaf node, and X represents the input.
[0074] 1 G of user feedback text can be used as the training text to train the fastText word vector model. Among them, X1, X2, X3, ……, Xn represent the N-gram vectors in a feedback text. Each feature is the average value of the word vectors. The minimum subword length selected is 1, and the maximum subword length is 5. The dimension of the output word vector is 100 dimensions. The minimum subword length and the maximum subword length can also be set according to the actual situation. The dimension represents the features of the words. The more features, the more accurately the words can be distinguished from each other. The dimension here can also be set according to the actual situation. However, if the dimension is too high, the operation efficiency will be reduced.
[0075] Specifically, the technique involved in the fastText word vector model is the introduction of subword-level N-grams features. For the keyword "underwater maze", assuming N takes the value of 2, its bigrams are:
[0076] "<sea", "seabed", "bed maze", "maze", "aze>"
[0077] Among them, "<" and ">" represent the prefix and suffix respectively. We can use these bigrams to represent the keyword "underwater maze", and then use the weighted average of the five bigram subword vectors to represent the word vector of "underwater maze".
[0078] S205: Based on the first word vector and the second word vector, determine the sentence vector of the first feedback text.
[0079] In the embodiments of this specification, determining the sentence vector of the first feedback text based on the first word vector and the second word vector may include:
[0080] S2051: Calculate the weighted average of the first word vector and the second word vector;
[0081] S2053: Determine the weighted average as the sentence vector of the first feedback text.
[0082] In the embodiments of this specification, the method may further include:
[0083] Calculate the first probability weight of the entity keyword, where the first probability weight is used to characterize the probability of the entity keyword appearing in the preset text vector library;
[0084] Calculate the second probability weight of the description keyword, where the second probability weight is used to characterize the probability of the description keyword appearing in the preset text vector library;
[0085] Correspondingly, as Figure 4 shown, calculating the weighted average of the first word vector and the second word vector may include:
[0086] S20511: Based on the first probability weight and the first word vector, determine the weighted word vector of the entity keyword;
[0087] The determining the weighted word vector of the entity keyword based on the first probability weight and the first word vector may include:
[0088] Calculate the product of the first word vector and the first probability weight to obtain a first product;
[0089] Use the first product as the weighted word vector of the entity keyword.
[0090] S20513: Based on the second probability weight and the second word vector, determine the weighted word vector of the description keyword;
[0091] The determining the weighted word vector of the description keyword based on the second probability weight and the second word vector may include:
[0092] Calculate the product of the second word vector and the second probability weight to obtain a second product;
[0093] Use the second product as the weighted word vector of the description keyword.
[0094] In practical applications, the probability weight can be expressed as: Where w is a keyword, and the keyword includes an entity keyword and a description keyword. a is a constant, and a can take the value of 1. p(w) is the probability that the keyword appears in the preset text vector library. If the probability that the keyword appears in the preset text vector library is higher, the weight of the keyword in its corresponding feedback text is lower, and the influence of the keyword on the sentence vector of the feedback text is smaller; on the contrary, the influence of the keyword on the sentence vector of the feedback text is greater. Correspondingly, the weighted word vector can be expressed as: Among them, v w is a word vector.
[0095] S20515: Calculate the average value of the weighted word vectors of the entity keyword and the description keyword to obtain the weighted word vector average value;
[0096] S20517: Use the weighted word vector average value as the weighted average value of the first word vector and the second word vector.
[0097] In practical applications, the calculation formula of the weighted average value of the first word vector and the second word vector is as follows: Among them, v s is the weighted average value, s is the keyword set in the feedback text, and |s| represents the size of the keyword set, that is, the number of keywords in the keyword set.
[0098] In the embodiments of this specification, based on the probability that the keyword appears in the preset text vector library and its corresponding word vector, the weighted word vector of the keyword is calculated, and the average value of the weighted word vectors of all keywords is calculated, so as to obtain a weighted average value of the word vector with a higher accuracy.
[0099] In the embodiments of this specification, the method may further include:
[0100] Determine the first type weight of the entity keyword; the first type weight is used to characterize the importance of the entity keyword;
[0101] Determine the second type weight of the description keyword; the second type weight is used to characterize the importance of the description keyword;
[0102] Correspondingly, the determining the weighted word vector of the entity keyword based on the first probability weight and the first word vector may include:
[0103] Based on the first probability weight, the first type weight and the first word vector, determine the weighted word vector of the entity keyword;
[0104] In the embodiments of this specification, the determining the weighted word vector of the entity keyword based on the first probability weight, the first type weight and the first word vector may include:
[0105] Calculate the product of the first probability weight, the first type weight, and the first word vector to obtain a third product;
[0106] Use the third product as the weighted word vector of the entity keyword.
[0107] Correspondingly, determining the weighted word vector of the description keyword based on the second probability weight and the second word vector may include:
[0108] Determine the weighted word vector of the description keyword based on the second probability weight, the second type weight, and the second word vector.
[0109] In the embodiments of the present specification, determining the weighted word vector of the description keyword based on the second probability weight, the second type weight, and the second word vector may include:
[0110] Calculate the product of the second probability weight, the second type weight, and the second word vector to obtain a fourth product;
[0111] Use the fourth product as the weighted word vector of the entity keyword.
[0112] In practical applications, the type weight can be expressed as: k(t(w)), where w is the keyword, the keyword includes the entity keyword and the description keyword, t represents the type of the keyword, and k represents the weight corresponding to the keyword of type t; correspondingly, the weighted word vector can be expressed as: where, v w is the word vector, is the probability weight, and k(t(w)) is the type weight. Correspondingly, the calculation formula for the weighted average of the first word vector and the second word vector is as follows: where, v s is the weighted average, s is the set of keywords in the feedback text, and |s| represents the size of the keyword set, that is, the number of keywords in the keyword set.
[0113] In the embodiments of the present specification, probability weights and type weights are respectively set for the keywords. On this basis, the weighted word vectors of the keywords are obtained, improving the accuracy of the weighted word vectors; based on the weighted word vectors, the sentence vector accuracy of the first feedback text calculated is also correspondingly improved.
[0114] S207: Calculate the similarity between the sentence vector of the first feedback text and the sentence vector of the second feedback text in the preset text vector library, where the preset text vector library includes the mapping relationship between the preset second feedback text and the sentence vector.
[0115] In the embodiments of this specification, the cosine similarity distance can be used to calculate the similarity between the sentence vector of the first feedback text and the sentence vector of the second feedback text in the preset text vector library. The calculation formula is as follows:
[0116]
[0117] Among them, s1 is the first feedback text, x is the sentence vector corresponding to the first feedback text s1, s2 is the second feedback text, y is the sentence vector corresponding to the second feedback text s2, and θ represents the angle between the sentence vectors x and y; where the full name of sim is similarity, and its meaning is similarity.
[0118] In the embodiments of this specification, based on the obtained sentence vectors with relatively high accuracy, the similarity between different sentence vectors with high accuracy can be obtained, that is, the similarity between the first feedback text and the second feedback text can be obtained.
[0119] S209: Based on the similarity, perform duplicate removal processing on the first feedback text.
[0120] In the embodiments of this specification, the performing duplicate removal processing on the first feedback text based on the similarity may include:
[0121] Determine the first feedback text with a similarity greater than or equal to the preset threshold with the sentence vector of the second feedback text in the preset text vector library as duplicate text; specifically, the preset threshold can be set according to the actual situation. For example, the preset threshold can be set to 80% or 90%.
[0122] Delete the duplicate text.
[0123] In the embodiments of this specification, there may be multiple first feedback texts. Calculate the similarity for each first feedback text respectively, and delete the duplicate texts in the first feedback texts.
[0124] In the embodiments of this specification, the method further includes:
[0125] Determine the first feedback text with a similarity less than the preset threshold with the sentence vector of the second feedback text in the preset text vector library as non-duplicate text;
[0126] Store the non-duplicate text in the preset text vector library.
[0127] In the embodiments of this specification, the preset text vector library may also store the mapping relationship between the feedback text and the sentence vector.
[0128] In the embodiments of this specification, the preset text vector library may further store the mapping relationship between the feedback text and its corresponding title, and at the same time store the mapping relationship between the title of the feedback text and the sentence vector.
[0129] In some embodiments, based on the similarity, the duplicate removal process of the first feedback text can be combined with manual auxiliary judgment. Specifically, the duplicate removal process of the first feedback text based on the similarity may include:
[0130] Obtain the top preset number of sentence vectors in the preset text vector library with the highest to lowest similarity to the sentence vector of the first feedback text;
[0131] Obtain the titles corresponding to the second feedback texts corresponding to the top preset number of sentence vectors;
[0132] Send the mapping relationship between the first feedback text, the titles of the top preset number of second feedback texts, and the similarity to the client;
[0133] Based on the received content, the client user determines whether the first feedback text is a duplicate text and whether there are duplicate texts among the top preset number of second texts;
[0134] When the client user determines that the first feedback text is a duplicate text, delete the duplicate text;
[0135] When the client user determines that the first feedback text is not a duplicate text, store the non-duplicate text in the preset text vector library;
[0136] When the client user determines that there are duplicate texts among the top preset number of second texts, recall them from the preset text vector library.
[0137] The following uses the feedback texts of players of two games, "Saint Seiya" and "PUBG Mobile", to illustrate the text duplicate removal method with manual auxiliary judgment.
[0138] Obtain the top five sentence vectors in the preset text vector library with the highest to lowest similarity to the sentence vector of the first feedback text of the game player;
[0139] In the embodiments of this specification, for the "Saint Seiya" game, obtain the titles and similarity data corresponding to the five second feedback texts with the highest similarity to the first feedback text in the preset text vector library;
[0140] Obtain the titles corresponding to the five second feedback texts;
[0141] Specifically, different second feedback texts may correspond to the same title;
[0142] Send the mapping relationship between the first feedback text, the titles of the five second feedback texts, and the similarity to the client;
[0143] As Figure 8 shown, the display interface of the client shows the mapping relationship between the titles of the five second feedback texts corresponding to the "Saint Seiya" game and the similarity; the display interface shows the titles of the five second feedback texts and the similarity between the second feedback text and the first feedback text; among them, "[Interface] It will freeze when returning after entering the skill upgrade or Eighth Sense interface", "[Galaxy Battle] Galaxy will directly freeze with Deadly Battle", "[Battle] Abnormal jitter appears under high frame rate and high image quality", "[Interface] Clicking on various icons in the interface is invalid" are all the titles corresponding to the feedback texts, and there are two identical titles "[Interface] It will freeze when returning after entering the skill upgrade or Eighth Sense interface" corresponding to two different similarities. It can be seen that the content of the feedback text displayed after clicking on the title is different;
[0144] As Figures 9 - 10 shown, the display interface of the client shows the mapping relationship between the titles of the five second feedback texts corresponding to two different first feedback texts of the "PUBG Mobile" game; the display interface also shows the titles of the five second feedback texts and the similarity between the second feedback text and the first feedback text;
[0145] When the user clicks on the title corresponding to the second feedback text, the second feedback text corresponding to the first five sentence vectors can be obtained;
[0146] Based on the received content, the client user determines whether the first feedback text is a duplicate text and whether there are duplicate texts among the five second texts;
[0147] When the client user determines that the first feedback text is a duplicate text, delete the duplicate text;
[0148] When the client user determines that the first feedback text is a non-duplicate text, store the non-duplicate text in the preset text vector library;
[0149] When the client user determines that there are duplicate texts among the five second texts, trigger "Duplicate" in the display interface, and it can be recalled from the preset text vector library.
[0150] In the embodiments of this specification, "bug" in the display interface refers to a malfunction or vulnerability.
[0151] As can be seen from the technical solutions provided in the embodiments of this specification above, based on the first feedback text fed back by the target object, the embodiments of this specification split the first feedback text into two parts: entity keywords and description keywords, that is, the first feedback text is refined and classified, so as to facilitate the text classification model to quickly and accurately determine the first word vector of the entity keywords and the second word vector of the description keywords; then based on the first word vector and the second word vector, the sentence vector of the first feedback text can be accurately obtained; based on the sentence vector, the similarity between the first feedback text and the second feedback text is further accurately calculated, so as to achieve accurate and efficient duplicate removal of the text.
[0152] The embodiments of this application also provide a device for duplicate removal of text, as Figure 11 shown, the device includes:
[0153] The keyword determination module 1110 can be used to determine the entity keywords and description keywords in the first feedback text based on the first feedback text fed back by the target object;
[0154] The word vector determination module 1120 can be used to determine the first word vector of the entity keywords and the second word vector of the description keywords based on the text classification model;
[0155] The sentence vector determination module 1130 can be used to determine the sentence vector of the first feedback text based on the first word vector and the second word vector;
[0156] The similarity calculation module 1140 can be used to calculate the similarity between the sentence vector of the first feedback text and the sentence vector of the second feedback text in the preset text vector library, and the preset text vector library includes the mapping relationship between the preset second feedback text and the sentence vector;
[0157] The duplicate removal processing module 1150 can be used to perform duplicate removal processing on the first feedback text based on the similarity.
[0158] In some embodiments, the sentence vector determination module 1130 may include:
[0159] The weighted average calculation sub-module is used to calculate the weighted average of the first word vector and the second word vector;
[0160] The sentence vector determination sub-module is used to determine the weighted average as the sentence vector of the first feedback text.
[0161] In some embodiments, the device may further include:
[0162] The first probability weight calculation module is used to calculate the first probability weight of the entity keywords, the
[0163] The first probability weight is used to characterize the probability of the entity keyword appearing in the preset text vector library;
[0164] A second probability weight calculation module, configured to calculate a second probability weight of the description keyword, where the second probability weight is used to characterize the probability of the description keyword appearing in the preset text vector library;
[0165] Correspondingly, as Figure 12 shown, the weighted average calculation sub-module may include:
[0166] A first weighted word vector determination unit 1210, configured to determine a weighted word vector of the entity keyword based on the first probability weight and the first word vector;
[0167] A second weighted word vector determination unit 1220, configured to determine a weighted word vector of the description keyword based on the second probability weight and the second word vector;
[0168] A weighted word vector average determination unit 1230, configured to calculate an average value of the weighted word vectors of the entity keyword and the description keyword to obtain a weighted word vector average value;
[0169] A weighted average determination unit 1240, configured to use the weighted word vector average value as the weighted average of the first word vector and the second word vector.
[0170] In some embodiments, the apparatus may further include:
[0171] A first type weight determination module, configured to determine a first type weight of the entity keyword;
[0172] A second type weight determination module, configured to determine a second type weight of the description keyword;
[0173] Correspondingly, the first weighted word vector determination unit includes:
[0174] A first weighted word vector determination subunit, configured to determine a weighted word vector of the entity keyword based on the first probability weight, the first type weight, and the first word vector;
[0175] The second weighted word vector determination unit includes:
[0176] A second weighted word vector determination subunit, configured to determine a weighted word vector of the description keyword based on the second probability weight, the second type weight, and the second word vector.
[0177] In some embodiments, the keyword determination module may further include:
[0178] A keyword acquisition sub-module, configured to acquire the original entity keywords and original description keywords in the first feedback text;
[0179] A standard keyword acquisition sub-module, configured to perform synonym replacement processing on the original entity keywords and the original description keywords based on a preset synonym library, to obtain standard entity keywords corresponding to the original entity keywords and standard description keywords corresponding to the original description keywords;
[0180] A keyword determination sub-module, configured to use the standard entity keywords as the entity keywords and the standard description keywords as the description keywords.
[0181] In some embodiments, the deduplication processing module may further include:
[0182] A duplicate text determination sub-module, configured to determine a first feedback text whose similarity with the sentence vector of the second feedback text in the preset text vector library is greater than or equal to a preset threshold as a duplicate text;
[0183] A duplicate text deletion sub-module, configured to delete the duplicate text.
[0184] In some embodiments, the apparatus may further include:
[0185] A non-duplicate text determination module, configured to determine a first feedback text whose similarity with the sentence vector of the second feedback text in the preset text vector library is less than the preset threshold as a non-duplicate text;
[0186] A non-duplicate text storage module, configured to store the non-duplicate text in the preset text vector library.
[0187] The apparatus in the apparatus embodiment and the method embodiment are based on the same inventive concept.
[0188] An embodiment of the present application provides a text deduplication device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the text deduplication method provided in the above method embodiment.
[0189] An embodiment of the present application further provides a storage medium, which can be disposed in a terminal to store at least one instruction, at least one program, a code set or an instruction set related to implementing a text deduplication method in a method embodiment. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the text deduplication method provided in the above method embodiment.
[0190] Optionally, in the embodiments of the present specification, the storage medium may be located in at least one of multiple network servers of a computer network. Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media that can store program codes such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.
[0191] The memory described in the embodiments of the present specification can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for functions, etc.; the data storage area can store data created according to the use of the device. In addition, the memory may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory may further include a memory controller to provide the processor with access to the memory.
[0192] The method embodiments for text deduplication provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, a server, or a similar computing device. Taking running on a server as an example, Figure 13 is a hardware structure block diagram of a server for a method of text deduplication provided in the embodiments of the present application. As Figure 13As shown, the server 1300 can vary significantly due to differences in configuration or performance. It can include one or more central processing units (CPUs) 1310 (the processor 1310 can include, but is not limited to, processing devices such as microprocessor MCUs or programmable logic devices FPGAs), a memory 1330 for storing data, and one or more storage media 1320 for storing application programs 1323 or data 1322 (such as one or more mass storage devices). Among them, the memory 1330 and the storage media 1320 can be transient storage or persistent storage. The programs stored in the storage media 1320 can include one or more modules, and each module can include a series of instruction operations on the server. Further, the central processor 1310 can be configured to communicate with the storage media 1320 and execute a series of instruction operations in the storage media 1320 on the server 1300. The server 1300 can also include one or more power supplies 1360, one or more wired or wireless network interfaces 1350, one or more input / output interfaces 1340, and / or one or more operating systems 1321, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSD TM, and so on.
[0193] The input / output interface 1340 can be used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the server 1300. In one instance, the input / output interface 1340 includes a network interface controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the input / output interface 1340 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0194] Those of ordinary skill in the art can understand that Figure 13 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the server 1300 can also include more or fewer components than Figure 13 shown, or have a different configuration from Figure 13 shown.
[0195] As can be seen from the embodiments of the method, apparatus, server, or storage medium for text deduplication provided by the present application above, based on the first feedback text fed back by the target object, the first feedback text is split into two parts: entity keywords and description keywords, that is, the first feedback text is refined and classified, so as to facilitate the text classification model to quickly and accurately determine the first word vector of the entity keywords and the second word vector of the description keywords; then, based on the first word vector and the second word vector, the sentence vector of the first feedback text can be accurately obtained; based on the sentence vector, the similarity between the first feedback text and the second feedback text is further accurately calculated, so as to achieve accurate and efficient text deduplication.
[0196] It should be noted that the above-mentioned sequence of embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above description of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0197] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the apparatus, device, and storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0198] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing related hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.
[0199] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for text deduplication, characterized in that, The method includes: Obtaining the original entity keywords and original description keywords in the first feedback text fed back by the target object through a preset keyword library, where the keyword library includes an entity keyword library and a description keyword library; Performing a synonym replacement process on the original entity keywords and the original description keywords based on a preset synonym library to obtain the standard entity keywords corresponding to the original entity keywords and the standard description keywords corresponding to the original description keywords; Taking the standard entity keywords as entity keywords and the standard description keywords as description keywords; the description keywords are used to describe the entity keywords; Determining a first word vector of the entity keywords and a second word vector of the description keywords based on a text classification model; Determining a first probability weight of the entity keywords and a second probability weight of the description keywords through the following formula: In the formula, a represents a constant, and P(w) represents the probability that the entity keyword or the description keyword appears in a preset text vector library, which does not change with the first feedback text; Determining a first type weight of the entity keywords; the first type weight is used to characterize the importance of the entity keywords; Determining a second type weight of the description keywords; the second type weight is used to characterize the importance of the description keywords; Determining a weighted word vector of the entity keywords based on the first probability weight, the first type weight, and the first word vector; Determining a weighted word vector of the description keywords based on the second probability weight, the second type weight, and the second word vector; Calculating the average value of the weighted word vectors of the entity keywords and the description keywords to obtain an average value of the weighted word vectors; Taking the average value of the weighted word vectors as the weighted average value of the first word vector and the second word vector; and determining the weighted average value as the sentence vector of the first feedback text; Calculating the similarity between the sentence vector of the first feedback text and the sentence vector of a second feedback text in a preset text vector library, where the preset text vector library includes a mapping relationship between the preset second feedback text and the sentence vector; Performing a duplicate removal process on the first feedback text based on the similarity.
2. The method according to claim 1, characterized in that, The performing a duplicate removal process on the first feedback text based on the similarity includes: Determining the first feedback text with a similarity greater than or equal to a preset threshold with the sentence vector of the second feedback text in the preset text vector library as duplicate text; Deleting the duplicate text.
3. The method according to claim 2, characterized in that, The method further includes: Determining the first feedback text with a similarity less than the preset threshold with the sentence vector of the second feedback text in the preset text vector library as non-duplicate text; Storing the non-duplicate text in the preset text vector library.
4. An apparatus for text deduplication, characterized in that, The device includes: A keyword determination module, configured to determine the entity keywords and description keywords in the first feedback text based on the first feedback text fed back by the target object; the description keywords are used to describe the entity keywords; A word vector determination module, configured to determine a first word vector of the entity keyword and a second word vector of the description keyword based on a text classification model; A sentence vector determination module, configured to determine a sentence vector of the first feedback text based on the first word vector and the second word vector; A similarity calculation module, configured to calculate a similarity between the sentence vector of the first feedback text and the sentence vector of a second feedback text in a preset text vector library, where the preset text vector library includes a mapping relationship between a preset second feedback text and a sentence vector; A duplicate removal processing module, configured to perform duplicate removal processing on the first feedback text based on the similarity; The keyword determination module includes: A keyword acquisition sub-module, configured to acquire an original entity keyword and an original description keyword in the first feedback text through a preset keyword library; the keyword library includes an entity keyword library and a description keyword library; A standard keyword acquisition sub-module, configured to perform a synonym replacement process on the original entity keyword and the original description keyword based on a preset synonym library to obtain a standard entity keyword corresponding to the original entity keyword and a standard description keyword corresponding to the original description keyword; A keyword determination sub-module, configured to use the standard entity keyword as the entity keyword and the standard description keyword as the description keyword; A probability weight calculation module, configured to determine a first probability weight of the entity keyword and a second probability weight of the description keyword through the following formula: In the formula, a represents a constant, and P(w) represents the probability that the entity keyword or the description keyword appears in the preset text vector library, which does not change with the first feedback text; A first type weight determination module, configured to determine a first type weight of the entity keyword; A second type weight determination module, configured to determine a second type weight of the description keyword; The sentence vector determination module includes: a weighted average calculation sub-module, configured to calculate a weighted average of the first word vector and the second word vector; a sentence vector determination sub-module, configured to determine the weighted average as the sentence vector of the first feedback text; The weighted average calculation sub-module includes: a first weighted word vector determination unit and a second weighted word vector determination unit; The first weighted word vector determination unit includes: a first weighted word vector determination sub-unit, configured to determine a weighted word vector of the entity keyword based on the first probability weight, the first type weight, and the first word vector; The second weighted word vector determination unit includes: a second weighted word vector determination sub-unit, configured to determine a weighted word vector of the description keyword based on the second probability weight, the second type weight, and the second word vector.
5. The device according to claim 4, characterized in that, The duplicate removal processing module includes: A duplicate text determination sub-module, configured to determine a first feedback text whose similarity with the sentence vector of the second feedback text in the preset text vector library is greater than or equal to a preset threshold as a duplicate text; A duplicate text deletion sub-module, configured to delete the duplicate text.
6. The device according to claim 5, characterized in that, The device further includes: A non-repetitive text determination module, configured to determine a first feedback text whose similarity to the sentence vector of the second feedback text in the preset text vector library is less than the preset threshold as a non-repetitive text; A non-repetitive text storage module, configured to store the non-repetitive text in the preset text vector library.
7. An apparatus for text deduplication, characterized in that, The device includes: a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the text deduplication method according to any one of claims 1-3.
8. A computer-readable storage medium, characterized in that, At least one instruction, at least one program, a code set or an instruction set is stored in the storage medium. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the text deduplication method according to any one of claims 1-3.
Citation Information
Patent Citations
Method and device for matching texts
CN102411583A
A method for computing text similarity by using semantic information
CN109325229A
Method, system and storage medium for improving sentence vector semantics
CN109408802A