Text deduplication method and device, electronic device and readable storage medium
Patent Information
- Application Number
- CN202210356716.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-06
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2042-04-06
AI Technical Summary
[0010]本发明实施例提供了一种文本去重方法及装置、电子设备及可读存储介质,以至少解决由于相关技术中由于线上模型实时推理,准确度以及时效性较差的技术问题
[0023] In this embodiment of the invention, multiple result texts corresponding to the query input are obtained; these result texts are matched against a pre-built thesaurus knowledge base, which is generated based on the prediction results of a pre-trained text deduplication model. The text deduplication model is used to predict semantic repetition based on the textual features, contextual features, and extended features of the result texts; and duplicate texts are filtered out from the multiple result texts based on the matching results of the thesaurus knowledge base. The thesaurus knowledge base, built based on the prediction structure of the pre-trained text deduplication model, achieves the goal of rapid online text deduplication, thereby improving the timeliness and accuracy of recommended search terms and results. This solves the technical problem of poor accuracy and timeliness caused by real-time reasoning in online models in related technologies.
Smart Images

Figure CN114818672B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of database technology, and more specifically, to a text deduplication method and apparatus, electronic device and readable storage medium. Background Technology
[0002] With the development of the internet, users can easily access a vast amount of information, but the cost of filtering out irrelevant information has also increased. Search recommendation technology, based on user-input search keywords or anonymized user information, intelligently pushes relevant information. Life service platforms aggregate massive amounts of life service information and serve users; therefore, search recommendation technology plays a crucial role on these platforms. Because the information on life service platforms is highly homogenized, the text information matched through search recommendation technology needs to be deduplicated to ensure a better user experience. In this scenario, common existing technical solutions fall into the following categories:
[0003] 1. Hash Signature Deduplication: This type of method segments the matching text into word groups, calculates the hash value of each word, and calculates a weighted numeric string based on the hash value of each word, which serves as the hash signature of the matching text. Finally, by calculating the distance between the hash signatures, it determines whether the two texts are duplicates, thereby achieving deduplication of text information.
[0004] 2. Text matching model deduplication: This type of method uses a text matching model to learn from manually labeled duplicate text data, enabling the model to have a certain ability to detect duplicates. Then, it makes inference judgments on the text results recommended by the search to determine whether two texts are duplicates, thereby achieving the deduplication of text information.
[0005] In the process of realizing this invention, the applicant discovered at least the following technical problems in the related art.
[0006] 1. Existing solutions rely solely on textual information for reasoning, and cannot effectively utilize additional context and other information, thus having certain limitations.
[0007] 2. Existing solutions rely on real-time online model inference, resulting in poor performance and difficulty in meeting the diverse performance requirements of related searches, guess what you want to search, and sug searches, thus exhibiting low universality.
[0008] 3. Once the existing solution is launched, the model is only updated during development iterations, resulting in poor timeliness. This deficiency is particularly pronounced in the context of lifestyle service platforms where various merchants and businesses are constantly changing.
[0009] It is evident that no effective solution has yet been proposed in the relevant technologies to address the aforementioned problems. Summary of the Invention
[0010] This invention provides a text deduplication method and apparatus, electronic device and readable storage medium to at least solve the technical problems of poor accuracy and timeliness in related technologies due to real-time reasoning of online models.
[0011] According to one aspect of the present invention, a text deduplication method is provided, comprising: acquiring multiple result texts corresponding to a query input; matching the multiple result texts in a pre-constructed thesaurus, wherein the thesaurus is generated based on the prediction results of a pre-trained text deduplication model, the text deduplication model being used to predict semantic repetition based on the text features, context features, and extended features of the result texts; and filtering out duplicate texts from the multiple result texts based on the matching results of the thesaurus.
[0012] Furthermore, before obtaining the multiple result texts corresponding to the query input, the method further includes: using the text deduplication model, performing semantic repetition prediction based on the text features, context features, and extended features corresponding to the first text data and the second text data respectively, to obtain the prediction results of the first text data and the second text data; if the prediction results indicate that the texts are semantically identical, then the first text data and the second text data are added to the synonym knowledge base.
[0013] Furthermore, the text deduplication model includes a text processing submodule and a compression interaction layer. The text deduplication model performs semantic repetition prediction based on the text features, context features, and extended features corresponding to the first text data and the second text data, respectively. This includes: using the text processing submodule to determine a first vector representation based on the first text features of the first text data and the second text features of the second text data; using the compression interaction layer to determine a second vector representation based on the context features and the extended features; determining a third vector representation based on the text features, context features, and extended features corresponding to the first text data and the second text data, respectively; and determining the prediction result based on the first vector representation, the second vector representation, and the third vector representation.
[0014] Furthermore, the text deduplication model includes a classification layer and a feature enhancement layer. Determining the prediction result based on the first vector representation, the second vector representation, and the third vector representation includes: summing the first vector representation, the second vector representation, and the third vector representation to obtain a fourth vector representation; enhancing the fourth vector representation through the feature enhancement layer to obtain a fifth vector representation; and using the classification layer on the fifth vector representation to determine the prediction results for the first text and the second text.
[0015] Further, if the prediction result is that the text semantics are the same, then the first text data and the second text data are added to the synonym knowledge base, including: determining the synonym knowledge base corresponding to the text semantics based on the text semantics corresponding to the first text data and the second text data, wherein the semantic distance between text pairs in the synonym knowledge base is less than a preset semantic distance threshold; and adding the first text data and the second text data to the synonym knowledge base.
[0016] According to another aspect of the present invention, a text deduplication device is also provided, comprising: an acquisition module for acquiring multiple result texts corresponding to a query input; a matching module for matching the multiple result texts in a pre-built thesaurus, wherein the thesaurus is generated based on the prediction results of a pre-trained text deduplication model, the text deduplication model being used to predict semantic repetition based on the text features, context features, and extended features of the result texts; and a deduplication module for filtering out duplicate texts from the multiple result texts based on the matching results of the thesaurus.
[0017] Furthermore, it also includes: a classification module, used to perform semantic repetition prediction based on the text features, context features, and extended features corresponding to the first text data and the second text data respectively, through the text deduplication model before obtaining the multiple result texts corresponding to the query input, so as to obtain the prediction results of the first text data and the second text data; and a storage module, used to add the first text data and the second text data to the synonym knowledge base if the prediction results are that the text semantics are the same.
[0018] Furthermore, the text deduplication model includes a text processing submodule and a compression interaction layer, wherein the classification module includes: a first determination submodule, used by the text processing submodule to determine a first vector representation based on a first text feature of the first text data and a second text feature of the second text data; a second determination submodule, used by the compression interaction layer to determine a second vector representation based on the context feature and the extended feature; a third determination submodule, used to determine a third vector representation based on the text feature, context feature, and extended feature corresponding to the first text data and the second text data, respectively; and a fourth determination submodule, used to determine the prediction result based on the first vector representation, the second vector representation, and the third vector representation.
[0019] Further, the fourth determining submodule includes: a processing unit, configured to perform vector summation on the first vector representation, the second vector representation, and the third vector representation to obtain a fourth vector representation; a feature enhancement unit, configured to perform feature enhancement on the fourth vector representation through the feature enhancement layer to obtain a fifth vector representation; and a determining unit, configured to determine the prediction results of the first text and the second text through the classification layer on the fifth vector representation.
[0020] Furthermore, the storage module includes: determining the thesaurus corresponding to the text semantics based on the text semantics corresponding to the first text data and the second text data, wherein the semantic distance between text pairs in the thesaurus is less than a preset semantic distance threshold; and adding the first text data and the second text data to the thesaurus.
[0021] According to another aspect of the present invention, an electronic device is also provided, including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the text deduplication method as described above.
[0022] According to another aspect of the present invention, a readable storage medium is also provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the text deduplication method as described above.
[0023] In this embodiment of the invention, multiple result texts corresponding to the query input are obtained; these result texts are matched against a pre-built thesaurus knowledge base, which is generated based on the prediction results of a pre-trained text deduplication model. The text deduplication model is used to predict semantic repetition based on the textual features, contextual features, and extended features of the result texts; and duplicate texts are filtered out from the multiple result texts based on the matching results of the thesaurus knowledge base. The thesaurus knowledge base, built based on the prediction structure of the pre-trained text deduplication model, achieves the goal of rapid online text deduplication, thereby improving the timeliness and accuracy of recommended search terms and results. This solves the technical problem of poor accuracy and timeliness caused by real-time reasoning in online models in related technologies. Attached Figure Description
[0024] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0025] Figure 1This is a schematic diagram of an optional text deduplication method according to an embodiment of the present invention;
[0026] Figure 2 This is a schematic diagram of an optional text deduplication model according to an embodiment of the present invention;
[0027] Figure 3 This is a schematic diagram of another optional text deduplication model according to an embodiment of the present invention;
[0028] Figure 4 This is a schematic diagram of an optional text deduplication device according to an embodiment of the present invention. Detailed Implementation
[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0031] Example 1
[0032] According to embodiments of the present invention, a text deduplication method is provided, such as... Figure 1 As shown, the method includes:
[0033] S102, retrieve multiple result texts corresponding to the query input;
[0034] In this embodiment, the result text can be the recommended text or search results corresponding to the user's query input on the online platform. The query input in this embodiment can be the query terms and / or selected query conditions entered by the user through the search page, or it can be the query terms and / or query conditions generated by the platform based on the user's behavior. Furthermore, the query input can also be the user's selection of relevant recommended terms or notification information on the platform.
[0035] In this embodiment, the system receives user query input for an online platform and then obtains multiple result texts corresponding to the query input. For example, the query input for the user's current query or current browsing session can be obtained through the platform's query entry. Further, the platform performs a search based on the obtained query input and retrieves multiple result texts that match the query input. For example, if the user clicks on "barbecue," each retrieved recommendation can be considered a result text.
[0036] S104, Match multiple result texts in a pre-built thesaurus knowledge base. The thesaurus knowledge base is generated based on the prediction results of a pre-trained text deduplication model. The text deduplication model is used to predict semantic repetition based on the text features, context features, and extended features of the result text.
[0037] In this embodiment, a thesaurus is pre-built, with each thesaurus corresponding to a specific text semantic. Within a thesaurus, the texts are semantically similar, with the semantic similarity exceeding a preset threshold. For example, if text A is a duplicate of text B, and text B is a duplicate of text C, it can be reasonably determined that text A and text C are duplicates. Thus, the pre-built thesaurus enables rapid matching of multiple result texts.
[0038] In this embodiment, the synonym knowledge base is pre-constructed using the prediction results of a text deduplication model. The text deduplication model groups texts from the text data source into pairs, performs classification predictions sequentially, and finally constructs the synonym knowledge base. Furthermore, after the synonym knowledge base is constructed, any new text obtained subsequently is added as a text increment to the semantically similar or identical synonym knowledge base.
[0039] S106, based on the matching results of the synonym knowledge base, filter out duplicate texts from multiple result texts.
[0040] In this embodiment, by matching multiple result texts in the synonym knowledge base, duplicate texts with the same or similar semantics can be filtered out, and the duplicate texts can be removed from the multiple result texts. Then, the filtered result texts are displayed to the user corresponding to the query input.
[0041] It should be noted that in this embodiment, multiple result texts corresponding to the query input are obtained; these result texts are matched in a pre-built thesaurus knowledge base. The thesaurus knowledge base is generated based on the prediction results of a pre-trained text deduplication model, which predicts semantic repetition based on the textual features, contextual features, and extended features of the result texts. Duplicate texts are then filtered out based on the matching results from the thesaurus knowledge base. The thesaurus knowledge base, built based on the prediction structure of the pre-trained text deduplication model, achieves the goal of rapid online text deduplication, thereby improving the timeliness and accuracy of recommended search terms and results. This solves the technical problem of poor accuracy and timeliness caused by real-time reasoning in online models in related technologies.
[0042] Optionally, in this embodiment, before obtaining multiple result texts corresponding to the query input, the process includes, but is not limited to: using a text deduplication model to perform semantic repetition prediction based on the text features, context features, and extended features corresponding to the first text data and the second text data, respectively, to obtain the prediction results of the first text data and the second text data; if the prediction results are that the text semantics are the same, then the first text data and the second text data are added to the synonym knowledge base.
[0043] In the specific implementation of this embodiment, the text deduplication model needs to be trained first.
[0044] In some embodiments, a training sample set is constructed based on the text data of the text recalled from the query input. Each training sample in the training sample set includes at least the following information: text information, context information, and extended information. The text information includes, but is not limited to, the query input and the result text. The context information is constructed based on the correspondence between the text recalled from the query input and the query input on the online platform. The extended information includes, but is not limited to, user information, user interaction information, and user preferences.
[0045] In this embodiment, training samples are constructed based on two sets of training text data. Each training sample includes a first text, a second text, a first context, a second context, first extended information, second extended information, and the semantic similarity between the first and second texts. In some embodiments, each training sample is represented as a seven-tuple <second text, first context, second context, first extended information, second extended information, semantic similarity>.
[0046] In addition, as an optional implementation, since the number of positive samples is often much smaller than that of negative samples in text deduplication tasks, a preliminary threshold control is performed by comparing edit distance and string similarity (jaro winkler) to screen similar samples and filter out most of the negative samples.
[0047] Next, a text deduplication model is trained based on the constructed training sample set. The text information, context information, and extended information corresponding to the first and second training text data are used as model inputs, and the semantic similarity between the first and second training text data is used as the model objective. The text deduplication model is trained until the model converges or the model is iterated to a preset number of times.
[0048] Then, the text features, context features, and extended features corresponding to the first and second text data are input into a pre-trained text deduplication model to obtain the semantic similarity score or the result of whether the first and second text data are semantically similar. If the first and second text data are semantically similar, they are added as synonyms or near-synonyms to a synonym knowledge base with the same semantics.
[0049] If the first text data and the second text data are not semantically similar, then the text corresponding to the first text data and the second text data shall be retained.
[0050] Through the above embodiments, the semantic similarity prediction of the result text corresponding to the query input is performed in advance by the text deduplication model, so as to realize the rapid semantic judgment of the result text of the query input, thereby improving the timeliness of text deduplication.
[0051] Optionally, in this embodiment, the text deduplication model includes a text processing submodule and a compression interaction layer. The text deduplication model performs semantic repetition prediction based on the text features, context features, and extended features corresponding to the first text data and the second text data, respectively. This includes, but is not limited to: determining a first vector representation based on the first text features of the first text data and the second text features of the second text data through the text processing submodule; determining a second vector representation based on the context features and extended features through the compression interaction layer; determining a third vector representation based on the text features, context features, and extended features corresponding to the first text data and the second text data, respectively; and determining a prediction result based on the first vector representation, the second vector representation, and the third vector representation.
[0052] Specifically, in this embodiment, such as Figure 2The diagram illustrates the structure of the text deduplication model 20. The model includes a text processing submodule 210, a compression interaction layer 220, a feature enhancement layer 230, and a classification layer 240. The first text features of the first dataset and the second text features of the second dataset are input into the text processing submodule 210 to obtain a first vector representation. The corresponding context features and extended features of the first and second text data are input into the compression interaction layer 220 to obtain a second vector representation. Then, the vectors corresponding to the text features, context features, and extended features of the first and second text data are concatenated to obtain a third vector representation. Finally, the prediction results for the first and second text data are determined by the feature enhancement layer 230 and the classification layer 240.
[0053] Optionally, in this embodiment, the text deduplication model includes a classification layer and a feature enhancement layer. The prediction result is determined based on the first vector representation, the second vector representation, and the third vector representation, including but not limited to: summing the first vector representation, the second vector representation, and the third vector representation to obtain a fourth vector representation; enhancing the fourth vector representation through the feature enhancement layer to obtain a fifth vector representation; and using the classification layer to determine the prediction results of the first text and the second text based on the fifth vector representation.
[0054] In specific application scenarios, such as Figure 3 The text deduplication model shown has an overall structure of xDeepFm. Contextual features and other extended features are processed by the feature processing layer and then input into the compression interaction layer CIN for display high-order feature combination processing. The deep neural network (DNN) in xDeepFm is replaced with a BERT (Bidirectional Encoder Representation from Transformers) model. The first and second texts are input into the BERT model to obtain the first vector representation.
[0055] In addition, the text deduplication model includes an embedding layer and a compressed interaction network (CIN). First, the context features and extended features are input into the embedding layer in the compressed interaction layer. Then, the output of the embedding layer is input into the CIN to obtain the second vector representation.
[0056] On one hand, a first vector representation is obtained from the text features corresponding to the first text and the second text respectively, and a second vector representation is obtained from the context features and extended features. The first text, the second text, the context features and the extended features are input into the Linear layer to obtain a third vector representation.
[0057] On the other hand, in the Add layer of the text deduplication model, the first vector representation, the second vector representation, and the third vector representation are summed, and then data augmentation is performed through the Mix Up layer to obtain the fourth vector representation.
[0058] Then, the fourth vector representation is classified through the output unit to obtain the semantic similarity or similarity score between the first and second texts. Optionally, the classification layer 240 can be composed of a multilayer perceptron (MLP), which is not limited in this embodiment.
[0059] In the above embodiments, the semantic similarity between the first text data and the second text data is determined by the text features, context features and extended features corresponding to the first text data and the second text data, respectively, through the text deduplication model, thereby improving the accuracy of the prediction results.
[0060] Optionally, in this embodiment, if the prediction result is that the text semantics are the same, the first text data and the second text data are added to the synonym knowledge base, including but not limited to: determining the synonym knowledge base corresponding to the text semantics based on the text semantics corresponding to the first text data and the second text data, wherein the semantic distance between text pairs in the synonym knowledge base is less than a preset semantic distance threshold; and adding the first text data and the second text data to the synonym knowledge base.
[0061] Specifically, in this embodiment, the text deduplication model outputs whether there is a repetition between text pairs of the first and second text data. For example, if text A is repeated with text B, and text B is repeated with text C, it can be reasonably determined that text A and text C are repeated. Since the construction of text pairs is filtered by a distance threshold and does not include all text pair combinations, more repeated texts can be derived based on the transitivity of repetition. However, each deduplication requires traversing this link, which incurs a significant time cost. To improve the deduplication performance and update efficiency of the knowledge base, this embodiment employs a disjoint-set data structure algorithm to maintain a synonym knowledge base in key-value pair format, and performs path compression optimization during offline updates.
[0062] It should be noted that if the text deduplication model undergoes parameter updates or adjustments, the thesaurus will be fully updated; if the text deduplication model is not updated or adjusted, the thesaurus will be updated periodically using incremental updates of synonyms / near-synonyms to ensure that the knowledge base has good coverage and can promptly cover emerging and popular texts.
[0063] This embodiment obtains multiple result texts corresponding to the query input; matches these multiple result texts in a pre-built thesaurus knowledge base, which is generated based on the prediction results of a pre-trained text deduplication model. The text deduplication model is used to predict semantic repetition based on the text features, context features, and extended features of the result texts; and duplicate texts are filtered out from the multiple result texts based on the matching results of the thesaurus knowledge base. The thesaurus knowledge base, built based on the prediction structure of the pre-trained text deduplication model, achieves the goal of fast online text deduplication, thereby improving the timeliness and accuracy of recommended search terms and results. This solves the technical problem of poor accuracy and timeliness caused by real-time reasoning in online models in related technologies.
[0064] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0065] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0066] Example 2
[0067] According to embodiments of the present invention, a text deduplication apparatus for implementing the above-described text deduplication method is also provided, such as... Figure 4 As shown, the device includes:
[0068] 1) Module 40 is used to obtain multiple result texts corresponding to the query input;
[0069] 2) Matching module 42, used to match the multiple result texts in a pre-built thesaurus, wherein the thesaurus is generated based on the prediction results of a pre-trained text deduplication model, and the text deduplication model is used to predict semantic repetition based on the text features, context features and extended features of the result texts.
[0070] 3) Deduplication module 44, used to filter out duplicate texts in the multiple result texts based on the matching results of the synonym knowledge base.
[0071] Optionally, in this embodiment, it further includes:
[0072] 1) The classification module is used to perform semantic repetition prediction based on the text features, context features and extended features corresponding to the first text data and the second text data respectively, through the text deduplication model before obtaining the multiple result texts corresponding to the query input, so as to obtain the prediction results of the first text data and the second text data.
[0073] 2) A storage module, used to add the first text data and the second text data to the synonym knowledge base if the prediction result is that the texts have the same semantics.
[0074] Optionally, in this embodiment, the text deduplication model includes a text processing submodule and a compression interaction layer, wherein the classification module includes:
[0075] 1) A first determining submodule, configured to determine a first vector representation based on a first text feature of the first text data and a second text feature of the second text data through the text processing submodule;
[0076] 2) A second determining submodule, used to determine a second vector representation based on the context features and the extended features through the compressed interaction layer;
[0077] 3) The third determining submodule is used to determine the third vector representation based on the text features, context features, and extended features corresponding to the first text data and the second text data, respectively;
[0078] 4) The fourth determining submodule is used to determine the prediction result based on the first vector representation, the second vector representation and the third vector representation.
[0079] Optionally, in this embodiment, the fourth determining submodule includes:
[0080] 1) A processing unit, configured to perform vector summation on the first vector representation, the second vector representation, and the third vector representation to obtain a fourth vector representation;
[0081] 2) Feature enhancement unit, used to enhance the features of the fourth vector representation through the feature enhancement layer to obtain the fifth vector representation;
[0082] 3) A determining unit, used to determine the prediction results of the first text and the second text by passing the classification layer on the fifth vector representation.
[0083] Optionally, in this embodiment, the storage module 44 includes:
[0084] 1) Based on the text semantics corresponding to the first text data and the second text data, determine the thesaurus corresponding to the text semantics, wherein the semantic distance between text pairs in the thesaurus is less than a preset semantic distance threshold;
[0085] 2) Add the first text data and the second text data to the synonym knowledge base.
[0086] This embodiment obtains multiple result texts corresponding to the query input; matches these multiple result texts in a pre-built thesaurus knowledge base, which is generated based on the prediction results of a pre-trained text deduplication model. The text deduplication model is used to predict semantic repetition based on the text features, context features, and extended features of the result texts; and duplicate texts are filtered out from the multiple result texts based on the matching results of the thesaurus knowledge base. The thesaurus knowledge base, built based on the prediction structure of the pre-trained text deduplication model, achieves the goal of fast online text deduplication, thereby improving the timeliness and accuracy of recommended search terms and results. This solves the technical problem of poor accuracy and timeliness caused by real-time reasoning in online models in related technologies.
[0087] Example 3
[0088] According to an embodiment of the present invention, an electronic device is also provided, including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the text deduplication method as described above.
[0089] Optionally, in this embodiment, the memory is configured to store program code for performing the following steps:
[0090] S1, retrieve multiple result texts corresponding to the query input;
[0091] S2, Match the multiple result texts in a pre-built thesaurus knowledge base, wherein the thesaurus knowledge base is generated based on the prediction results of a pre-trained text deduplication model, and the text deduplication model is used to predict semantic repetition based on the text features, context features and extended features of the result texts.
[0092] S3, based on the matching results of the synonym knowledge base, filter out duplicate texts from the multiple result texts.
[0093] Optionally, specific examples in this embodiment can refer to the examples described in Embodiment 1 above, and will not be repeated here.
[0094] Example 4
[0095] Embodiments of the present invention also provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the text deduplication method described above.
[0096] Optionally, in this embodiment, the readable storage medium is configured to store program code for performing the following steps:
[0097] S1, retrieve multiple result texts corresponding to the query input;
[0098] S2, Match the multiple result texts in a pre-built thesaurus knowledge base, wherein the thesaurus knowledge base is generated based on the prediction results of a pre-trained text deduplication model, and the text deduplication model is used to predict semantic repetition based on the text features, context features and extended features of the result texts.
[0099] S3, based on the matching results of the synonym knowledge base, filter out duplicate texts from the multiple result texts.
[0100] Optionally, the readable storage medium is also configured to store program code for performing the steps included in the method of Embodiment 1 above, which will not be described again in this embodiment.
[0101] Optionally, in this embodiment, the aforementioned readable storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0102] Optionally, specific examples in this embodiment can refer to the examples described in Embodiment 1 above, and will not be repeated here.
[0103] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0104] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0105] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0106] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.
[0107] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0108] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0109] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A text deduplication method, characterized in that, include: Retrieve multiple result texts corresponding to the query input; The multiple result texts are matched in a pre-built thesaurus knowledge base, wherein the thesaurus knowledge base is generated based on the prediction results of a pre-trained text deduplication model. The text deduplication model is used to predict semantic repetition based on the text features, context features, and extended features of the result texts. The context features are constructed based on the correspondence between the texts recalled by the online platform based on the query input and the query input. The extended features include at least one of user information, user interaction information, and user preferences. The text deduplication model includes a text processing submodule and a compression interaction layer, wherein the text processing submodule is a language representation model. The text processing submodule determines a first vector representation based on the first text features of the first text data and the second text features of the second text data. The second vector representation is determined through the compressed interaction layer based on the context features and the extended features; The third vector representation is determined based on the text features, context features, and extended features corresponding to the first text data and the second text data, respectively. The text deduplication model further includes a feature enhancement layer and a classification layer, which sums the first vector representation, the second vector representation, and the third vector representation to obtain a fourth vector representation. The feature enhancement layer is used to enhance the features of the fourth vector representation to obtain the fifth vector representation. The fifth vector representation is classified through the classification layer to determine the prediction results of the first text data and the second text data; The synonym knowledge base uses a disjoint-set data structure algorithm to maintain the key-value pair knowledge base and performs path compression optimization during offline updates; Duplicate texts in the multiple result texts are filtered out based on the matching results from the synonym knowledge base.
2. The method according to claim 1, characterized in that, Before obtaining the multiple result texts corresponding to the query input, the method also includes: The text deduplication model is used to predict semantic repetition based on the text features, context features, and extended features corresponding to the first and second text data, respectively, so as to obtain the prediction results of the first and second text data. If the prediction result indicates that the texts have the same semantic meaning, then the first text data and the second text data are added to the synonym knowledge base.
3. The method according to claim 2, characterized in that, If the prediction result indicates that the texts have the same semantic meaning, then the first text data and the second text data are added to the synonym knowledge base, including: Based on the text semantics corresponding to the first text data and the second text data, the thesaurus corresponding to the text semantics is determined, wherein the semantic distance between text pairs in the thesaurus is less than a preset semantic distance threshold; The first text data and the second text data are added to the synonym knowledge base.
4. The method according to claim 2, characterized in that, Before training the text deduplication model, the following is also included: By using an edit distance and string similarity comparison algorithm to perform preliminary threshold control, similar samples are selected and most negative samples are filtered out in order to construct a training sample set.
5. The method according to claim 1, characterized in that, Also includes: If the text deduplication model updates or adjusts its parameters, the thesaurus will be fully updated. If the text deduplication model is not updated or adjusted, the synonym knowledge base is updated periodically using incremental updates of synonyms.
6. A text deduplication device, characterized in that, include: The retrieval module is used to retrieve multiple result texts corresponding to the query input; The matching module is used to match the multiple result texts in a pre-built thesaurus knowledge base, wherein the thesaurus knowledge base is generated based on the prediction results of a pre-trained text deduplication model, and the text deduplication model is used to perform semantic repetition prediction based on the text features, context features, and extended features of the result texts, wherein the context features are constructed based on the correspondence between the texts recalled by the online platform based on the query input and the query input, and the extended features include at least one of user information, user interaction information, and user preferences; The text deduplication model includes a text processing submodule and a compression interaction layer, wherein the text processing submodule is a language representation model. The text processing submodule determines a first vector representation based on the first text features of the first text data and the second text features of the second text data. The second vector representation is determined through the compressed interaction layer based on the context features and the extended features; The third vector representation is determined based on the text features, context features, and extended features corresponding to the first text data and the second text data, respectively. The text deduplication model further includes a feature enhancement layer and a classification layer, which sums the first vector representation, the second vector representation, and the third vector representation to obtain a fourth vector representation. The feature enhancement layer is used to enhance the features of the fourth vector representation to obtain the fifth vector representation. The fifth vector representation is classified through the classification layer to determine the prediction results of the first text data and the second text data; The synonym knowledge base uses a disjoint-set data structure algorithm to maintain the key-value pair knowledge base and performs path compression optimization during offline updates; The deduplication module is used to filter out duplicate text from the multiple result texts based on the matching results of the thesaurus.
7. The apparatus according to claim 6, characterized in that, Also includes: The classification module is used to perform semantic repetition prediction based on the text features, context features and extended features corresponding to the first text data and the second text data respectively, through the text deduplication model before obtaining the multiple result texts corresponding to the query input, so as to obtain the prediction results of the first text data and the second text data. The storage module is used to add the first text data and the second text data to the synonym knowledge base if the prediction result is that the texts have the same semantics.
8. The apparatus according to claim 7, characterized in that, The storage module includes: Based on the text semantics corresponding to the first text data and the second text data, the thesaurus corresponding to the text semantics is determined, wherein the semantic distance between text pairs in the thesaurus is less than a preset semantic distance threshold; The first text data and the second text data are added to the synonym knowledge base.
9. The apparatus according to claim 7, characterized in that, It also includes a sample construction module, which is used to perform preliminary threshold control by using edit distance and string similarity comparison algorithms before training the text deduplication model, to screen similar samples and filter out most negative samples in order to construct a training sample set.
10. The apparatus according to claim 6, characterized in that, It also includes an update module for: If the text deduplication model updates or adjusts its parameters, the thesaurus will be fully updated. If the text deduplication model is not updated or adjusted, the synonym knowledge base is updated periodically using incremental updates of synonyms.
11. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the text deduplication method as described in any one of claims 1-5.
12. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the text deduplication method as described in any one of claims 1-5.
Citation Information
Patent Citations
Method and device for determining text matching degree, equipment and storage medium
CN112749554A
Associational word deduplication method and device, computer readable storage medium and electronic equipment
CN112765966A
Information acquisition method and device, storage medium and electronic equipment
CN113392094A