A text deduplication method based on a multi-model algorithm and related devices thereof

Through the text deduplication method based on multi-model algorithm, for the duplicate content in massive text, a combination of full deduplication and simhash algorithm is adopted to solve the problems of low deduplication efficiency and high communication cost, and efficient text deduplication and rapid response are achieved.

CN115344685BActive Publication Date: 2025-06-27PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210997003.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-19
Publication Date
2025-06-27
Estimated Expiration
2042-08-19

AI Technical Summary

Technical Problem

In massive texts, too much duplicate content leads to low deduplication efficiency, high communication cost and large space occupancy, making it difficult for the existing technology to quickly extract effective information and respond to user problems.

Method used

The text deduplication method based on a multi-model algorithm is used to obtain the deduplication text of the target seat and determine whether it includes the first type of deduplication text or the second type of deduplication text. For the first type of text to be deduplicated, the second type of text to be deduplicated, the feature text is extracted and hashed calculated using the simhash algorithm to obtain the feature fingerprint value, and the duplicate is retrieved and deduplicated through binning.

Benefits of technology

It improves text deduplication efficiency, reduces communication costs, reduces space, and improves the work efficiency and customer response capabilities of the seats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115344685B_ABST
    Figure CN115344685B_ABST
Patent Text Reader

Abstract

This application belongs to the field of big data technology and relates to a text deduplication method based on a multi-model algorithm, including obtaining the text to be deduplicated of a target seat; determining whether the text to be deduplicated includes the first type of text to be deduplicated or the second type of text to be deduplicated; if the text to be deduplicated includes the first type of text to be deduplicated, performing full-scale deduplication on the first type of text to be deduplicated; if the text to be deduplicated includes the second type of text to be deduplicated, extracting the feature text of the second type of text to be deduplicated through the simhash algorithm and performing hash calculation to obtain a feature fingerprint value; binning the feature fingerprint value, and respectively performing retrieval deduplication on the multiple segments of query feature fingerprint values obtained after binning. This application also provides a text deduplication device, a computer device, and a storage medium based on a multi-model algorithm. This application can reduce communication costs, is beneficial to improving retrieval efficiency, reducing the storage consumption of machines, and quickly responding to customer needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data technology, and in particular, to a text deduplication method, device, computer device, and storage medium based on a multi-model algorithm. Background Art

[0002] In big data, there is a vast amount of text, and a large amount of duplicate text content is included. In the current service industry, usually, the agent will actively communicate with the user to ensure the effective progress of the business. However, in the communication, there is a large amount of business information, and it is difficult for the agent to quickly obtain the historical chat information of each user for timely reply. Moreover, the chat information has various types, including long content, short content, and different language forms, etc. Especially when there is too much duplicate text content in the historical chat information, manual deduplication by intelligent agents leads to the inability to quickly extract effective information, and the reply does not correspond to the problem consulted by the user. And when replying to the user's question, a lot of information needs to be input or copied and pasted for reply. It can be seen that the existing technology has problems of low deduplication efficiency, high communication cost, and large space occupation when dealing with a vast amount of text with excessive duplicate content. Summary of the Invention

[0003] The purpose of the embodiments of this application is to propose a text deduplication method, device, computer device, and storage medium based on a multi-model algorithm, which can improve the deduplication efficiency, reduce the communication cost, and reduce the occupied space when dealing with a vast amount of text with excessive duplicate content.

[0004] To solve the above technical problems, the embodiments of this application provide a text deduplication method based on a multi-model algorithm, and adopt the following technical solutions:

[0005] Obtain the text to be deduplicated of the target agent;

[0006] Judge whether the text to be deduplicated includes the first type of text to be deduplicated or the second type of text to be deduplicated;

[0007] If the text to be deduplicated includes the first type of text to be deduplicated, perform full deduplication on the first type of text to be deduplicated;

[0008] If the text to be deduplicated includes the second type of text to be deduplicated, extract the feature text of the second type of text to be deduplicated through the simhash algorithm and perform hash calculation to obtain a feature fingerprint value;

[0009] Perform binning on the feature fingerprint value, and respectively perform retrieval deduplication on the multiple segments of feature fingerprint values obtained after binning.

[0010] Further, before the step of obtaining the text to be deduplicated of the agent, the following steps are further included:

[0011] Accessing the underlying database through the computing engine, loading the historical communication texts of all agents stored in the underlying database into the computing memory, wherein the historical communication texts of all agents include the to-be-deduplicated texts of the target agent, and each historical communication text includes an identification tag of the corresponding agent;

[0012] The historical communication texts of all the agents in the computing memory are divided based on the identification tags to obtain the historical communication texts of each agent.

[0013] Furthermore, the step of determining whether the to-be-deduplicated text includes the first category of to-be-deduplicated text or the second category of to-be-deduplicated text specifically includes:

[0014] Determine whether the to-be-deduplicated text of the target agent includes text whose length is less than a preset length threshold;

[0015] Determine whether the to-be-deduplicated text of the target seat includes non-Chinese text;

[0016] If the to-be-deduplicated texts of the target agent include texts whose length does not reach the preset length threshold / the to-be-deduplicated texts of the target agent include the non-Chinese texts, it is determined that the to-be-deduplicated texts include the first category of to-be-deduplicated texts;

[0017] If the to-be-deduplicated texts of the target agent include texts whose length reaches the preset length threshold, it is determined that the to-be-deduplicated texts include the second type of to-be-deduplicated texts.

[0018] Furthermore, the first type of text to be deduplicated is fully deduplicated:

[0019] The non-Chinese text / the text whose length is less than a preset length threshold is fully deduplicated through a collection container, and completely repeated text in the first category of text to be deduplicated is filtered.

[0020] Furthermore, the step of extracting the characteristic text of the second type of text to be deduplicated by the simhash algorithm and performing hash calculation to obtain the characteristic fingerprint value specifically includes:

[0021] The second type of text to be deduplicated is segmented by the word segmenter in the simhash algorithm to obtain a plurality of text phrases;

[0022] Based on the API interface of the word segmenter, extract the plurality of feature texts from the plurality of text phrases by using the TF-IDF algorithm;

[0023] Calculate hash values ​​of the plurality of feature texts respectively based on a hash function;

[0024] Weight the hash value of each of the feature texts, and merge and reduce the dimensionality of the weighted results based on the order of multiple feature texts to obtain the feature fingerprint value.

[0025] Further, the step of binning the feature fingerprint value and separately retrieving and removing duplicates from the multiple segments of the to-be-query feature fingerprint values obtained after binning specifically includes:

[0026] Determine the unit bin length based on the length of the feature fingerprint value, and bin the feature fingerprint value according to the unit bin length to obtain multiple segments of the to-be-query feature fingerprint values;

[0027] Retrieve each segment of the to-be-query feature fingerprint value. If no identical to-be-de-duplicated feature fingerprint value is found, retain the second type of to-be-de-duplicated text corresponding to which no identical feature fingerprint value is found;

[0028] If it is retrieved that there is any segment identical to multiple segments of the to-be-query feature fingerprint values, calculate the difference points between the two identical segments of the to-be-query feature fingerprint values based on the feature fingerprint value;

[0029] If the difference points do not meet the preset difference points, remove duplicates from the second type of to-be-de-duplicated text corresponding to the two segments of the to-be-query feature fingerprint values whose difference points do not meet the preset difference points.

[0030] Further, after the step of binning the feature fingerprint value and separately retrieving and removing duplicates from the multiple segments of the to-be-query feature fingerprint values obtained after binning, it further includes:

[0031] Create a retrieval database for the target seat, and store the text after removing duplicates from the target seat into the retrieval database, where the text after removing duplicates from the target seat includes the first type of to-be-de-duplicated text after full-scale duplicate removal and the second type of to-be-de-duplicated text after binning and retrieving and removing duplicates.

[0032] To solve the above technical problems, an embodiment of the present application further provides a text duplicate removal device based on a multi-model algorithm, which adopts the following technical solutions:

[0033] An acquisition module, configured to acquire the text to be de-duplicated of the target seat;

[0034] A judgment module, configured to judge whether the text to be de-duplicated includes the first type of to-be-de-duplicated text or the second type of to-be-de-duplicated text;

[0035] A first duplicate removal module, configured to, if the text to be de-duplicated includes the first type of to-be-de-duplicated text, perform full-scale duplicate removal on the first type of to-be-de-duplicated text;

[0036] A calculation module, configured to, if the text to be deduplicated includes the second type of text to be deduplicated, extract the feature text of the second type of text to be deduplicated through the simhash algorithm and perform hash calculation to obtain a feature fingerprint value;

[0037] A second deduplication module, configured to bin the feature fingerprint values and perform retrieval deduplication on the multiple segments of feature fingerprint values obtained after binning respectively.

[0038] To solve the above technical problems, an embodiment of the present application further provides a computer device, which adopts the following technical solutions:

[0039] It includes a memory and a processor. Computer-readable instructions are stored in the memory. When the processor executes the computer-readable instructions, the steps of the text deduplication method based on a multi-model algorithm described in any of the above embodiments are implemented.

[0040] To solve the above technical problems, an embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solutions:

[0041] Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are executed by a processor, the steps of the text deduplication method based on a multi-model algorithm described in any of the above embodiments are implemented.

[0042] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects: By obtaining the text to be deduplicated of the target agent, identifying whether the text to be deduplicated includes the first type of text to be deduplicated or the second type of text to be deduplicated, performing full-scale deduplication on the included first type of text to be deduplicated, and for the included second type of text to be deduplicated, extracting the feature text of the second type of text to be deduplicated through the simhash algorithm and performing hash calculation to obtain a feature fingerprint value, binning the feature fingerprint values, and performing retrieval deduplication on the multiple segments of feature fingerprint values obtained after binning respectively. Therefore, in a large sample data with excessive text duplicate content, multiple sets of text deduplication methods can reduce the text content retrieved by the agent. When the agent replies to information, it does not need to input the text in full each time, which is beneficial to improving the retrieval efficiency, reducing the storage consumption of the machine, and quickly responding to customers. Description of the Drawings

[0043] To more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0044] Figure 1is an exemplary system architecture diagram to which the present application can be applied;

[0045] Figure 2 is a flowchart of an embodiment provided by a text deduplication method based on a multi-model algorithm according to the present application;

[0046] Figure 3 is Figure 2 a flowchart of a specific embodiment before step 201 in;

[0047] Figure 4 is Figure 2 a flowchart of a specific embodiment of step 202 in;

[0048] Figure 5 is Figure 2 a flowchart of a specific embodiment of step 204 in;

[0049] Figure 6 is Figure 2 a flowchart of a specific embodiment of step 205 in;

[0050] Figure 7 is a flowchart of another embodiment provided by a text deduplication method based on a multi-model algorithm according to the present application;

[0051] Figure 8 is a schematic structural diagram of an embodiment provided by a text deduplication device based on a multi-model algorithm according to the present application;

[0052] Figure 9 is a schematic structural diagram of an embodiment of a computer device according to the present application. Detailed Description of the Invention

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.

[0054] References to "embodiments" in this specification mean that specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and is not necessarily referring to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0055] To enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0056] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0057] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0058] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0059] The server 105 may be a server providing various services, such as a background server supporting the pages displayed on the terminal devices 101, 102, 103.

[0060] The server 105 may be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0061] It should be noted that the text deduplication method based on multi-model algorithms provided by the embodiments of the present application is generally executed by the server / terminal device. Correspondingly, the text deduplication method device based on multi-model algorithms is generally set in the server / terminal device.

[0062] It should be understood that Figure 1 the numbers of the terminal devices 101, 102, 103, the network 104 and the server 105 in Figure 1 are merely illustrative. According to the implementation requirements, there can be any number of terminal devices 101, 102, 103, network 104 and server 105.

[0063] Continuing to refer to Figure 2 , a flowchart of an embodiment of a text deduplication method based on a multi-model algorithm according to the present application is shown. The text deduplication method based on a multi-model algorithm includes the following steps:

[0064] Step S201, obtaining the text to be deduplicated of the target agent.

[0065] In this embodiment, an electronic device (such as the server / terminal device shown in Figure 1 ) on which a text deduplication method based on a multi-model algorithm runs can obtain the text to be deduplicated of the above-mentioned target agent and perform various data transmissions through a wired connection method or a wireless connection method. It should be noted that the above-mentioned wireless connection method may include, but is not limited to, 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other currently known or future-developed wireless connection methods.

[0066] In this embodiment, the above-mentioned target agent may refer to a certain agent determined from multiple customer services and currently needing to communicate with a customer. The above-mentioned text to be deduplicated may refer to multiple historical communication texts corresponding to the target agent that need to be deduplicated. The text to be deduplicated may include various types of texts. For example, according to a preset length limit, it can be divided into long texts and short texts, and according to the language, it can be divided into Chinese texts and non-Chinese texts, etc. The historical communication texts of all agents can be stored in the form of a database. Therefore, the database storing the historical communication texts of each agent can be directly accessed to retrieve the text to be deduplicated of the target agent.

[0067] Step S202, determining whether the text to be deduplicated includes a first type of text to be deduplicated or a second type of text to be deduplicated.

[0068] In this embodiment, the text to be deduplicated is divided into the first type of text to be deduplicated and the second type of text to be deduplicated. Among them, the first type of text to be deduplicated may include non-Chinese texts and short texts. The non-Chinese texts may refer to texts whose content includes English, and the short texts may refer to Chinese texts whose length does not reach the preset length threshold. The second type of text to be deduplicated may refer to Chinese texts whose length reaches the preset length threshold. By identifying multiple texts in the text to be deduplicated and judging the language type and text length of each text, it can be identified whether each text belongs to the first type of text to be deduplicated or the second type of text to be deduplicated. After identifying all the texts to be deduplicated of the target seat, all the texts to be deduplicated can be classified into two categories respectively, that is, the above-mentioned first type of text to be deduplicated and the second type of text to be deduplicated.

[0069] Step S203, if the text to be deduplicated includes the first type of text to be deduplicated, perform full deduplication on the first type of text to be deduplicated.

[0070] In this embodiment, the above full deduplication may be to perform full deduplication on non-Chinese texts / texts with a length less than the preset length threshold through a set container, and filter out the texts that are completely repeated in the first type of text to be deduplicated. Among them, full deduplication may be to query whether there are two or more completely identical texts in the first type of text to be deduplicated. For the completely identical texts, the set container can filter out the completely identical texts and only allow one to enter the set container, and other different first type of texts to be deduplicated can enter the set container.

[0071] Step S204, if the text to be deduplicated includes the second type of text to be deduplicated, extract the feature text of the second type of text to be deduplicated through the simhash algorithm and perform hash calculation to obtain the feature fingerprint value.

[0072] In this embodiment, when it is recognized that there is the second type of text to be deduplicated in the text to be deduplicated, the feature extraction of the second type of text to be deduplicated can be performed to obtain the above-mentioned feature text. Among them, the feature text may be the core words extracted after word segmentation of the second type of text to be deduplicated. The core words may refer to the words with high semantic value for the text, and the semantics of the second type of text to be deduplicated can be expressed more accurately through the core words.

[0073] In this application, for a large amount of text to be deduplicated, the SimHash algorithm is adopted. The main work of the SimHash algorithm can be understood as dimensionality reduction of the text to generate a fingerprint. By comparing the Hamming distance of the fingerprints of different texts, the similarity between two texts can be judged. Among them, the Hamming distance refers to the number of different bits in the corresponding bits of two legal codes in information coding, which is called the code distance. The number of different bits in the corresponding bits of two codewords is called the Hamming distance between the two codewords. In this embodiment, the fingerprint refers to the above-mentioned feature fingerprint value (SimHash value). Through the SimHash algorithm, the extracted feature text can be subjected to hash calculation to obtain the feature fingerprint value, and each feature text corresponds to a string of feature fingerprint values. The feature fingerprint value is a string of codes, for example: 0111010000111100.

[0074] Step S205: Bin the feature fingerprint values, and retrieve and deduplicate the multiple segments of feature fingerprint values to be queried obtained after binning respectively.

[0075] In this embodiment, the above binning may refer to segmenting the feature fingerprint values. After dividing them into several segments, each segment is deduplicated separately. After all the second types of text to be deduplicated of the target seat are binned, retrieval can be performed based on all the feature fingerprint values obtained after binning to determine whether there are the same parts in each segment of the feature fingerprint values to be queried after segmentation of the same feature fingerprint value. Only when no same parts are retrieved in each segment of the feature fingerprint values to be queried in the same feature fingerprint value can the corresponding text be retained; if the same parts are retrieved in any segment of the feature fingerprint values to be queried in the feature fingerprint value, it can be judged that the text corresponding to the feature fingerprint value may have the same semantic content and needs to be deduplicated. Specifically, deduplication can be performed based on the difference points of the feature fingerprint values.

[0076] Specifically, deduplication is performed on the first type of text to be deduplicated and the second type of text to be deduplicated respectively. The deduplicated text can be stored in a database. For the deduplicated text, when the seat replies to the user's information, it does not need to input the full text each time, but only needs to input part of the text to quickly retrieve the text communicated in the past through the ES interface. In this way, the data resources can be fully utilized to improve the work efficiency of the seat.

[0077] In an embodiment of the present invention, by obtaining the text to be deduplicated of a target seat, it is identified whether the text to be deduplicated includes a first type of text to be deduplicated or a second type of text to be deduplicated. For the included first type of text to be deduplicated, full deduplication is performed. For the included second type of text to be deduplicated, the feature text of the second type of text to be deduplicated is extracted through the simhash algorithm and hashing calculation is performed to obtain a feature fingerprint value. The feature fingerprint value is binned, and the multiple segments of query feature fingerprint values obtained after binning are respectively retrieved and deduplicated. Therefore, in a large sample data with excessive text duplicate content, multiple sets of text deduplication methods can reduce the text content retrieved by the seat. When the seat replies to information, it does not need to input the text in full each time, reducing the communication cost, facilitating the improvement of retrieval efficiency, reducing the storage consumption of the machine, and quickly responding to customer needs.

[0078] In some alternative implementation manners, such as Figure 3 shown, Figure 3 is a flowchart of a specific implementation manner before step 201 in Figure 2 . Before step 201, the above electronic device can also be used to perform the following steps:

[0079] Step S301, access the underlying database through a computing engine, and load all the historical communication texts of all seats stored in the underlying database into the computing memory. Among them, the historical communication texts of all seats include the text to be deduplicated of the target seat, and each historical communication text contains an identification label of the corresponding seat.

[0080] Specifically, before deduplication, all the historical communication texts of all seats can be stored in the underlying database. In this embodiment, the underlying database can refer to a Hive database, a mainstream component for storing big data. The spark computing engine can be used to access the underlying Hive database, and the historical communication texts of each seat are respectively loaded into the computing memory. The computing memory can be understood as a carrier for storing texts during the computing process. Among them, the historical communication texts of all seats include the text to be deduplicated of the target seat, and each historical communication text contains an identification label of the corresponding seat. Different seat identities can be distinguished through the identification label. For example: different seats correspond to different numbers, and each historical communication text carries the number of the seat, and the number is used as a label for distinction. As a possible embodiment, in order to avoid ineffective occupation of space, the data in the underlying database can be regularly cleaned.

[0081] Step S302, divide all the historical communication texts of all seats in the computing memory based on the identification label to obtain the historical communication texts of each seat.

[0082] Specifically, after being loaded into the computing memory, the historical communication texts of all seats can be divided based on the identification tags of the corresponding seats included in each historical communication text, and the historical communication texts of each seat can be independently distinguished. Among all the seats, the target seat in the above embodiment is included.

[0083] In this embodiment, by loading the historical communication texts of all seats into the computing memory and independently distinguishing the historical communication texts of all seats based on the identification tags in the computing memory, it is beneficial to quickly retrieve the texts to be de-duplicated of the target seat.

[0084] In some alternative implementation manners, such as Figure 4 shown, Figure 4 is Figure 2 a flowchart of a specific embodiment manner of step 202 in

[0085] Step S2021, determine whether the texts to be de-duplicated of the target seat include texts with a length less than a preset length threshold.

[0086] Specifically, the length of each text to be de-duplicated can be judged, and the texts with a length less than the preset length threshold can be screened. Among them, considering the input cost and repetition rate of the communication content, the preset length threshold can be set to 6, or other values can also be set.

[0087] Step S2022, determine whether the texts to be de-duplicated of the target seat include non-Chinese texts.

[0088] Specifically, for different language types, it can be identified whether the texts to be de-duplicated of the target seat contain non-Chinese texts, and the non-Chinese texts can be screened. The non-Chinese texts can include texts in other forms such as English texts.

[0089] Step S2023, if the texts to be de-duplicated of the target seat include texts with a length not reaching the preset length threshold / the texts to be de-duplicated of the target seat include non-Chinese texts, then determine that the texts to be de-duplicated include the first type of texts to be de-duplicated.

[0090] Step S2024, if the texts to be de-duplicated of the target seat include texts with a length reaching the preset length threshold, then determine that the texts to be de-duplicated include the second type of texts to be de-duplicated.

[0091] Specifically, when the target agent's to-be-deduplicated texts include texts whose length does not reach a preset length threshold / the target agent's to-be-deduplicated texts include non-Chinese texts, it is determined that the to-be-deduplicated texts include the first category of to-be-deduplicated texts; if the to-be-deduplicated texts include texts whose length reaches a preset length threshold, it is determined that the to-be-deduplicated texts include the second category of to-be-deduplicated texts, and the first category of to-be-deduplicated texts and the second category of to-be-deduplicated texts are independently distinguished and deduplicated using different deduplication methods. In the process of identifying the text length and identifying the non-Chinese text, the texts can be identified sequentially based on the text communication time, or the texts can be identified randomly.

[0092] In this embodiment, the text to be deduplicated is divided into a first category of text to be deduplicated and a second category of text to be deduplicated by distinguishing the text length and language type of the text to be deduplicated. Using multiple sets of text deduplication methods in large samples with excessive text repetition and different types can reduce the text content retrieved by the agents. The agents do not need to enter the full text every time when replying to the information, which reduces the communication cost, is conducive to improving the retrieval efficiency, reducing the storage consumption of the machine, and quickly responding to customer needs.

[0093] In some optional implementations, such as Figure 5 As shown, Figure 5 for Figure 2 The step 204 executed by the electronic device specifically includes the following steps:

[0094] Step S2041, segmenting the second type of text to be deduplicated by the word segmenter in the simhash algorithm to obtain multiple text phrases.

[0095] Specifically, the simhash algorithm includes five steps: word segmentation, hashing, weighting, merging, and dimensionality reduction. First, a word segmenter can be used to deduplicate the second type of text to be deduplicated, and each second type of text to be deduplicated can be divided into several text phrases. The above word segmenter can use the jieba word segmenter, which has its own API interface.

[0096] Step S2042: Based on the API interface of the word segmenter, multiple feature texts are extracted from the multiple text phrases through the TF-IDF algorithm.

[0097] Specifically, in the word segmentation process, based on the API interface in the jieba word segmenter, the TF-IDF algorithm (term frequency-inverse document frequency) can be used to extract feature text from text phrases. Feature text can be the core word in multiple text phrases, which has higher semantic value. After obtaining multiple feature texts, a weight can be assigned to each feature text. Among them, the weight can be set to different classes according to the different feature texts, and each class corresponds to a weight. Among them, the TF-IDF algorithm is a commonly used weighting technology for information retrieval and data mining. TF is term frequency (TermFrequency) and IDF is inverse document frequency index (Inverse Document Frequency).

[0098] Step S2043, respectively calculating hash values ​​of the plurality of feature texts based on the hash function.

[0099] Specifically, the hash value of each feature text can be calculated through the hash function. The hash value is an n-bit signature composed of the binary number 01. For example, the hash value of height is 0101, and the hash value of size is 1011. In this way, each feature text can be hashed into a hash value in binary encoding form.

[0100] Step S2044, weighting the hash value of each feature text, and merging and reducing the dimension of the weighted results based on the order of multiple feature texts to obtain a feature fingerprint value.

[0101] Specifically, based on the hash value, all feature texts are weighted (W), that is, W = hash (hash value) × weight (weight), and when 1 is encountered, the hash value and weight are positively multiplied, and when 0 is encountered, the hash value and weight are negatively multiplied. For example: the hash value of weight is 0110, and the weight is 4, then the weighted result is -4, 4, 4, -4. Then the weighted results of each feature text of the same second category of deduplication text are accumulated to become a sequence string. Then the dimension is reduced, if it is greater than 0, it is set to 1, otherwise it is set to 0, so as to obtain the simhash value of the second category of deduplication text. Finally, the similarity of different texts can be judged according to the Hamming distance of the simhash values. For example: the weighted result of the second category of deduplication text a is 9, -9, 1, -1. The 01 string obtained by dimension reduction is: "1, 0, 1, 0", which is the simhash value of the second category of deduplication text a.

[0102] In this embodiment, each second type of text to be deduplicated is tokenized, hashed, weighted, merged, and dimensionality-reduced through the SimHash algorithm, and finally the SimHash value of each second type of text to be deduplicated is obtained. The SimHash value for each second type of text to be deduplicated can be retrieved through binning operations. In a large sample of data with excessive duplicate text content, multiple text deduplication methods can reduce the text content retrieved by the agent. When the agent replies with information, it does not need to input the full text each time, reducing the communication cost, facilitating the improvement of retrieval efficiency, reducing the storage consumption of the machine, and quickly responding to customer needs.

[0103] In some alternative implementation manners, such as Figure 6 shown, Figure 6 for Figure 2 is a flowchart of a specific embodiment of step 205 in

[0104] Step S2051: Determine the unit bin length based on the length of the feature fingerprint value, and bin the feature fingerprint value according to the unit bin length to obtain multiple segments of feature fingerprints to be queried.

[0105] In this embodiment, each feature fingerprint value will form a 64-bit fingerprint code after being processed by the SimHash algorithm. To improve the retrieval efficiency, the unit bin length can be determined according to the length of the feature fingerprint value. For example: the feature fingerprint value is a 64-bit fingerprint code, divided into 4 segments, then the unit bin length is 16, so 4 segments of fingerprints to be queried will be obtained. Among them, the reason for choosing to divide into 4 segments is that if the difference points between two different feature fingerprint values are less than or equal to 3, then they are considered similar, so only one text corresponding to the feature fingerprint value is retained. After binning, multiple segments of feature fingerprints to be queried can be obtained.

[0106] Step S2052: Retrieve each segment of the feature fingerprints to be queried. If no identical feature fingerprints to be deduplicated are found, retain the second type of text to be deduplicated corresponding to the feature fingerprints for which no identical ones are found.

[0107] Specifically, after binning all the second type of text to be deduplicated for the target agent, for each segment of the feature values to be queried of a certain text in the second type of text to be deduplicated, if no identical feature fingerprints to be deduplicated are retrieved for each segment of the feature values to be queried, then retain the second type of text to be deduplicated corresponding to the feature fingerprints for which no identical ones are found, that is, it is recognized that there is no duplicate text for this text.

[0108] Step S2053: If any segment is found to be the same as the multiple segments of feature fingerprints to be queried, calculate the difference points between the two identical segments of feature fingerprints to be queried based on the feature fingerprint value.

[0109] Similarly, for each segment of the query feature values of a piece of text in the second type of text to be deduplicated, if any one of the multiple segments of query feature fingerprint values of the same text is retrieved as the same, the difference points between the two segments of query feature fingerprint values that are the same are calculated based on the feature fingerprint values. Among them, the difference point is the Hamming distance between the two segments of query feature fingerprint values.

[0110] Step S2054, if the difference point does not meet the preset difference point, the second type of text to be deduplicated corresponding to the two segments of query feature fingerprint values whose difference points do not meet the preset difference point is deduplicated.

[0111] In this embodiment, the above preset difference point can be set to 3, that is, when the difference point is less than or equal to 3, it can be considered that the two query feature fingerprint values for calculating the difference point are duplicates and need to be deduplicated. Therefore, the second type of text to be deduplicated corresponding to the two segments of query feature fingerprint values whose difference points do not meet the preset difference point can be deduplicated. Similarly, based on the above embodiments, deduplication can be completed for all the texts to be deduplicated for each agent seat.

[0112] In this embodiment, by binning the feature fingerprint values, multiple segments of query feature fingerprint values obtained after binning are retrieved and deduplicated respectively, and for the case where the same query feature fingerprint values are retrieved, deduplication judgment is performed through the difference points. In a large sample data with excessive duplicate text content, multiple sets of text deduplication methods can reduce the text content retrieved by the agent seat. Binning can perform deduplication more accurately. The agent seat does not need to input the full text each time when replying to information, which is beneficial to improving the retrieval efficiency, reducing the storage consumption of the machine, and quickly responding to customers.

[0113] In some alternative implementation manners, after step 205, the above electronic device can also be used to perform the following steps:

[0114] Create a retrieval database for the target agent seat, and store the deduplicated text of the target agent seat in the retrieval database, where the deduplicated text of the target agent seat includes the first type of text to be deduplicated after full deduplication and the second type of text to be deduplicated after binning and retrieval deduplication.

[0115] Specifically, a one-to-one retrieval database can be created. After completing the deduplication of the first type of text to be deduplicated and the second type of text to be deduplicated of the target agent seat, the deduplicated text data can be stored in the retrieval database of the target agent seat. The same applies to other agent seats. After completing the text deduplication for each agent seat, the deduplicated data of each can be stored in the retrieval database of the corresponding agent seat. Creating a one-to-one retrieval database can be used for the customer service to quickly retrieve the historical communication text through the ES interface, which is beneficial to making full use of data resources and improving the work efficiency of the agent seat.

[0116] In some possible embodiments, referring to Figure 7 as shown, Figure 7 is the overall flowchart of another text deduplication method based on a multi-model algorithm provided in this embodiment.

[0117] First, the underlying Hive database can be accessed through the spark computing engine, and the historical communication texts of all agents stored in the underlying database are loaded into the computing memory. The agents are divided according to the identification tags of the agents, and the first type of text to be deduplicated (non-Chinese texts and texts with a length less than the preset length threshold) and the second type of text to be deduplicated (texts with a length reaching the preset length threshold) are obtained respectively. Full-scale deduplication is performed on the first type of text to be deduplicated; for the second type of text to be deduplicated, text tokenization, feature text extraction (core word extraction), hash value calculation, weighting, merging and dimensionality reduction processing (text fingerprint), binning of feature fingerprint values (fingerprint binning), and retrieval deduplication are performed. After completing the deduplication operations on the first type of text to be deduplicated and the second type of text to be deduplicated for all agents, a one-to-one retrieval database can be created to store the deduplicated text data of each agent respectively for agent retrieval (text landing). The agent only needs to input part of the text, and then can quickly retrieve the historical communication text through the ES interface for automatic completion and quickly reply to the customer. Therefore, in a large sample data with excessive text duplicates, multiple sets of text deduplication methods can reduce the text content retrieved by the agent. When the agent replies to information, it does not need to input the text in full each time, reducing the communication cost, being beneficial to improving the retrieval efficiency, reducing the storage consumption of the machine, and quickly responding to customer needs.

[0118] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned method embodiments. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), etc., or a random access memory (RAM), etc.

[0119] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially in the direction of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this document, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0120] Further reference is made to Figure 8 , as an implementation of the method shown above Figure 2 , an embodiment of a text deduplication device based on a multi-model algorithm is provided in this application. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0121] As shown in Figure 8 , Figure 8 is a schematic structural diagram of a text deduplication device based on a multi-model algorithm provided in this embodiment. The device 800 includes:

[0122] An acquisition module 801, configured to acquire the text to be deduplicated of a target seat.

[0123] A judgment module 802, configured to judge whether the text to be deduplicated includes the first type of text to be deduplicated or the second type of text to be deduplicated.

[0124] A first deduplication module 803, configured to perform full-scale deduplication on the first type of text to be deduplicated if the text to be deduplicated includes the first type of text to be deduplicated.

[0125] A calculation module 804, configured to extract the feature text of the second type of text to be deduplicated by the simhash algorithm and perform hash calculation to obtain a feature fingerprint value if the text to be deduplicated includes the second type of text to be deduplicated.

[0126] A second deduplication module 805, configured to perform binning on the feature fingerprint value and perform retrieval deduplication on each of the multiple segments of the to-be-query feature fingerprint values obtained after binning.

[0127] In this embodiment, the target agent obtained by the obtaining module 801 may refer to a certain agent determined from multiple customer service representatives and currently needed to communicate with a customer. The above text to be de-duplicated may refer to multiple historical communication texts corresponding to the target agent that need to be de-duplicated. The text to be de-duplicated may include various types of texts. For example, according to length, it can be divided into long texts and short texts, and according to language, it can be divided into Chinese texts and non-Chinese texts, etc. The historical communication texts of all agents can be stored in the form of a database. Therefore, the database storing the historical communication texts of each agent can be directly accessed to retrieve the text to be de-duplicated of the target agent.

[0128] More specifically, when the above determination module 802 makes a determination, the text to be de-duplicated can be divided into a first type of text to be de-duplicated and a second type of text to be de-duplicated. Among them, the first type of text to be de-duplicated may include non-Chinese texts and short texts. Non-Chinese texts may refer to texts whose content includes English, and short texts may refer to Chinese texts whose length does not reach the preset length threshold. The second type of text to be de-duplicated may refer to Chinese texts whose length reaches the preset length threshold. By identifying multiple texts in the text to be de-duplicated through the determination module 802 and judging the language type and text length of each text, it can be identified whether each text belongs to the first type of text to be de-duplicated or the second type of text to be de-duplicated. After the determination module 802 has identified all the text to be de-duplicated of the target agent, all the text to be de-duplicated can be classified into two categories, that is, the above-mentioned first type of text to be de-duplicated and the second type of text to be de-duplicated.

[0129] More specifically, the above full-scale de-duplication by the first de-duplication module 803 can be to perform full-scale de-duplication on non-Chinese texts / texts with a length less than the preset length threshold through a set container, and filter out the texts that are completely repeated in the first type of text to be de-duplicated. Among them, full-scale de-duplication can be to query whether there are two or more texts that are exactly the same in the first type of text to be de-duplicated. For texts that are exactly the same, the set container can filter out the exactly the same texts, and only allow one to enter the set container, and other different first type of text to be de-duplicated can enter the set container.

[0130] More specifically, when it is recognized that there is a second type of text to be de-duplicated in the text to be de-duplicated, the above calculation module 804 can extract features from the second type of text to be de-duplicated to obtain the above feature text. Among them, the feature text can be the core words extracted after word segmentation of the second type of text to be de-duplicated. The core words can refer to words with high semantic value for the text, and the semantic of the second type of text to be de-duplicated can be expressed more accurately through the core words.

[0131] In this application, for a large amount of text to be deduplicated, the calculation module 804 thus adopts the SimHash algorithm. The main work of the SimHash algorithm can be understood as dimensionality reduction of the text to generate a fingerprint. By comparing the Hamming distance of the fingerprints of different texts, the similarity between two texts can be judged. Among them, the Hamming distance refers to the number of bits with different encodings in the corresponding bits of two legal codes in information encoding, and the number of bits with different bit values of two codewords is called the Hamming distance between the two codewords. In this embodiment, the fingerprint refers to the above-mentioned feature fingerprint value (SimHash value). Through the SimHash algorithm, the extracted feature text can be subjected to hash calculation to obtain the feature fingerprint value, and each feature text corresponds to a string of feature fingerprint values. The feature fingerprint value is a string of codes, for example: 0111010000111100.

[0132] Furthermore, through the second deduplication module 805, a binning operation is first performed. Binning can refer to segmenting the feature fingerprint value. After dividing it into several segments, each segment is deduplicated separately. When all the second-type texts to be deduplicated of the target seat have been binned, retrieval can be performed based on all the feature fingerprint values obtained after binning to determine whether there are identical parts in each segment of the to-be-query feature fingerprint value of the same feature fingerprint value. Only when no identical parts are retrieved in each segment of the to-be-query feature fingerprint value in the same feature fingerprint value can the corresponding text be retained; if identical parts are retrieved in any segment of the to-be-query feature fingerprint value in the feature fingerprint value, it can be determined that the text corresponding to the feature fingerprint value may have the same semantic content, and the second deduplication module 805 needs to perform deduplication, specifically based on the difference points of the feature fingerprint value.

[0133] In the embodiment of this application, the acquisition module 801 acquires the text to be deduplicated of the target seat, and based on the judgment module 802, it identifies whether the text to be deduplicated includes the first-type text to be deduplicated or the second-type text to be deduplicated. For the included first-type text to be deduplicated, full deduplication is performed by the first deduplication module 803; for the included second-type text to be deduplicated, the SimHash algorithm provided by the calculation module 804 is used to extract the feature text of the second-type text to be deduplicated and perform hash calculation to obtain the feature fingerprint value. The second deduplication module 805 bins the feature fingerprint value, and separately retrieves and deduplicates the multiple segments of to-be-query feature fingerprint values obtained after binning. Therefore, in a large sample data with excessive text duplicate content, multiple sets of text deduplication methods can reduce the text content retrieved by the seat. The seat does not need to input the text in full each time when replying to information, reducing the communication cost, which is beneficial to improving the retrieval efficiency, reducing the storage consumption of the machine, and quickly responding to customer needs.

[0134] In some alternative implementation manners of this embodiment, the above device further includes:

[0135] A loading module, which is used to access an underlying database through a computing engine and load historical communication texts of all seats stored in the underlying database into computing memory. Among them, the historical communication texts of all seats include the text to be de-duplicated of a target seat, and each historical communication text contains an identification label of the corresponding seat.

[0136] A partitioning module, which is used to partition the historical communication texts of all seats in the computing memory based on the identification label to obtain the historical communication texts of each seat.

[0137] In this embodiment, the historical communication texts of all seats are loaded into the computing memory through the loading module, and the historical communication texts of all seats in the computing memory are independently distinguished based on the identification label by the partitioning module, which is beneficial to quickly retrieve the text to be de-duplicated of the target seat.

[0138] In some optional implementation manners of this embodiment, the judgment module 802 includes a judgment sub-module and a determination sub-module. Among them:

[0139] The judgment sub-module is used to judge whether the text to be de-duplicated of the target seat includes text with a length less than a preset length threshold.

[0140] The judgment sub-module is also used to judge whether the text to be de-duplicated of the target seat includes non-Chinese text.

[0141] The determination sub-module is used to determine that the text to be de-duplicated includes the first type of text to be de-duplicated if the text to be de-duplicated of the target seat includes text with a length not reaching the preset length threshold / the text to be de-duplicated of the target seat includes non-Chinese text.

[0142] The determination sub-module is also used to determine that the text to be de-duplicated includes the second type of text to be de-duplicated if the text to be de-duplicated of the target seat includes text with a length reaching the preset length threshold.

[0143] In this embodiment, the text length and language type of the text to be de-duplicated are distinguished by the judgment sub-module and the determination sub-module, and the text to be de-duplicated is divided into the first type of text to be de-duplicated and the second type of text to be de-duplicated. Using multiple sets of text de-duplication methods in large samples with excessive text repetition and different types can reduce the text content retrieved by the seat. The seat does not need to input the text in full each time when replying to information, which reduces the communication cost, is beneficial to improving the retrieval efficiency, reducing the storage consumption of the machine, and quickly responding to customer needs.

[0144] In some optional implementation manners of this embodiment, the first de-duplication module 803 is also used to perform full de-duplication on non-Chinese text / text with a length less than the preset length threshold through a set container, and filter out the texts that are exactly the same in the first type of text to be de-duplicated.

[0145] In some alternative implementation manners of this embodiment, the calculation module 804 includes a word segmentation sub-module, a feature extraction sub-module, a hash calculation sub-module, and a weighting and dimensionality reduction processing sub-module.

[0146] The word segmentation sub-module is used to segment the second type of text to be deduplicated through the word segmenter in the simhash algorithm to obtain multiple text phrases.

[0147] The feature extraction sub-module is used to extract multiple feature texts from the multiple text phrases based on the API interface provided by the word segmenter through the TF-IDF algorithm.

[0148] The hash calculation sub-module is used to calculate the hash values of the multiple feature texts respectively based on the hash function.

[0149] The weighting and dimensionality reduction processing sub-module is used to weight the hash value of each feature text, and merge and perform dimensionality reduction processing on the weighted results based on the order of the multiple feature texts to obtain the feature fingerprint value.

[0150] In this embodiment, each second type of text to be deduplicated is segmented, hashed, weighted, merged, and dimensionally reduced through the simhash algorithm, and finally the simhash value of each second type of text to be deduplicated is obtained. The simhash value of each second type of text to be deduplicated can be retrieved through binning operations. In a large sample data with too much duplicate text content, multiple text deduplication methods can reduce the text content retrieved by the agent. When the agent replies to information, it does not need to input the text in full each time, reducing the communication cost, which is beneficial to improving the retrieval efficiency, reducing the storage consumption of the machine, and quickly responding to customer needs.

[0151] In some alternative implementation manners of this embodiment, the second deduplication module 805 includes a binning sub-module, a retrieval sub-module, a difference point calculation sub-module, and a deduplication sub-module. Among them:

[0152] The binning sub-module is used to determine the unit binning length based on the length of the feature fingerprint value, and bin the feature fingerprint value according to the unit binning length to obtain multiple segments of feature fingerprints to be queried.

[0153] The retrieval sub-module is used to retrieve each segment of the feature fingerprint to be queried. If no identical feature fingerprint to be deduplicated is found, the second type of text to be deduplicated corresponding to the feature fingerprint where no identical one is found is retained.

[0154] The difference point calculation sub-module is used to calculate the difference points between two identical segments of the feature fingerprints to be queried based on the feature fingerprint value if any identical segment is found among the multiple segments of the feature fingerprints to be queried.

[0155] A deduplication sub-module, configured to perform deduplication on the second type of text to be deduplicated corresponding to two segments of query feature fingerprint values where the difference points do not meet the preset difference points if the difference points do not meet the preset difference points.

[0156] In this embodiment, by binning the feature fingerprint values, multiple segments of query feature fingerprint values obtained after binning are respectively retrieved and deduplicated. And for the case where the same query feature fingerprint values are retrieved, deduplication judgment is performed through the difference points. In large sample data with excessive text duplicate content, multiple sets of text deduplication methods can reduce the text content retrieved by the agent. Binning can perform deduplication more accurately. The agent does not need to input the full text each time when replying to information, which is beneficial to improving the retrieval efficiency, reducing the storage consumption of the machine, and quickly responding to customers.

[0157] In some optional implementation manners of this embodiment, it further includes a creation module, configured to create a retrieval database for the target agent, and store the text after deduplication of the target agent into the retrieval database, where the text after deduplication of the target agent includes the first type of text to be deduplicated after full deduplication and the second type of text to be deduplicated after retrieval and deduplication by binning.

[0158] To solve the above technical problems, the embodiments of the present application further provide a computer device. For details, please refer to Figure 9 , Figure 9 which is the basic structural block diagram of the computer device in this embodiment.

[0159] The computer device 90 includes a memory 901, a processor 902, and a network interface 903 that are communicatively connected to each other through a system bus. It should be noted that only the computer device 90 with components 901-903 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0160] The computer device can be a desktop computer, a notebook, a palm computer, a cloud server and other computing devices. The computer device can perform human-computer interaction with the user through a keyboard, a mouse, a remote control, a touchpad or a voice control device and other means.

[0161] The memory 901 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 901 may be an internal storage unit of the computer device 90, such as the hard disk or memory of the computer device 90. In other embodiments, the memory 901 may also be an external storage device of the computer device 90, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 90. Of course, the memory 901 may also include both the internal storage unit and the external storage device of the computer device 90. In this embodiment, the memory 901 is generally used to store the operating system and various application software installed on the computer device 90, such as computer-readable instructions of the text deduplication method based on the multi-model algorithm. In addition, the memory 901 can also be used to temporarily store various data that have been output or will be output.

[0162] In some embodiments, the processor 902 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 902 is generally used to control the overall operation of the computer device 90. In this embodiment, the processor 902 is used to run the computer-readable instructions stored in the memory 901 or process data, such as running the computer-readable instructions of the text deduplication method based on the multi-model algorithm.

[0163] The network interface 903 may include a wireless network interface or a wired network interface, and the network interface 903 is generally used to establish a communication connection between the computer device 90 and other electronic devices.

[0164] In an embodiment of the present invention, by obtaining the text to be deduplicated of a target seat, it is identified whether the text to be deduplicated includes the first type of text to be deduplicated or the second type of text to be deduplicated. For the included first type of text to be deduplicated, full-scale deduplication is performed. For the included second type of text to be deduplicated, the feature text of the second type of text to be deduplicated is extracted through the simhash algorithm and hashing calculation is performed to obtain a feature fingerprint value. The feature fingerprint value is binned, and the multiple segments of queryable feature fingerprint values obtained after binning are respectively retrieved for deduplication. Therefore, in a large sample data with excessive text duplicate content, multiple sets of text deduplication methods can reduce the text content retrieved by the seat. When the seat replies to information, it does not need to input the text in full each time, reducing the communication cost, being beneficial to improving the retrieval efficiency, reducing the storage consumption of the machine, and quickly responding to customer needs.

[0165] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor, so that at least one processor executes the steps of the text deduplication method based on a multi-model algorithm as described above.

[0166] In an embodiment of the present invention, by obtaining the text to be deduplicated of a target seat, it is identified whether the text to be deduplicated includes the first type of text to be deduplicated or the second type of text to be deduplicated. For the included first type of text to be deduplicated, full-scale deduplication is performed. For the included second type of text to be deduplicated, the feature text of the second type of text to be deduplicated is extracted through the simhash algorithm and hashing calculation is performed to obtain a feature fingerprint value. The feature fingerprint value is binned, and the multiple segments of queryable feature fingerprint values obtained after binning are respectively retrieved for deduplication. Therefore, in a large sample data with excessive text duplicate content, multiple sets of text deduplication methods can reduce the text content retrieved by the seat. When the seat replies to information, it does not need to input the text in full each time, reducing the communication cost, being beneficial to improving the retrieval efficiency, reducing the storage consumption of the machine, and quickly responding to customer needs.

[0167] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), including several instructions to enable a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the text deduplication method based on a multi-model algorithm of each embodiment of the present application.

[0168] Obviously, the embodiments described above are only a part of the embodiments of this application, rather than all of them. The preferred embodiments of this application are shown in the accompanying drawings, but they do not limit the patent scope of this application. This application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure that makes use of the content of this application's specification and drawings, directly or indirectly applied in other related technical fields, is similarly within the scope of patent protection of this application.

Claims

1. A text deduplication method based on a multi-model algorithm, characterized in that Including the following steps: Obtain the text to be deduplicated of the target seat; Determine whether the text to be deduplicated includes the first type of text to be deduplicated or the second type of text to be deduplicated; If the text to be deduplicated includes the first type of text to be deduplicated, perform full deduplication on the first type of text to be deduplicated; If the text to be deduplicated includes the second type of text to be deduplicated, extract the feature text of the second type of text to be deduplicated through the simhash algorithm and perform hash calculation to obtain a feature fingerprint value; Perform binning on the feature fingerprint value, and respectively perform retrieval deduplication on multiple segments of the to-be-query feature fingerprint values obtained after binning; The step of determining whether the text to be deduplicated includes the first type of text to be deduplicated or the second type of text to be deduplicated specifically includes: Determine whether the text to be deduplicated of the target seat includes text with a length less than a preset length threshold; Determine whether the text to be deduplicated of the target seat includes non-Chinese text; If the text to be deduplicated of the target seat includes text with a length not reaching the preset length threshold or the text to be deduplicated of the target seat includes the non-Chinese text, determine that the text to be deduplicated includes the first type of text to be deduplicated; If the text to be deduplicated of the target seat includes text with a length reaching the preset length threshold, determine that the text to be deduplicated includes the second type of text to be deduplicated; The step of performing binning on the feature fingerprint value and respectively performing retrieval deduplication on multiple segments of the to-be-query feature fingerprint values obtained after binning specifically includes: Determine the unit binning length based on the length of the feature fingerprint value, perform binning on the feature fingerprint value according to the unit binning length to obtain multiple segments of the to-be-query feature fingerprint values; Perform retrieval on each segment of the to-be-query feature fingerprint value. If no identical to-be-deduplicated feature fingerprint value is retrieved, retain the second type of text to be deduplicated corresponding to the feature fingerprint values for which no identical ones are retrieved; If it is retrieved that there is any segment identical to multiple segments of the to-be-query feature fingerprint values, calculate the difference points between the two segments of the to-be-query feature fingerprint values that are identical based on the feature fingerprint value; If the difference points do not meet the preset difference points, perform deduplication on the second type of text to be deduplicated corresponding to the two segments of the to-be-query feature fingerprint values for which the difference points do not meet the preset difference points; Among them, after completing the deduplication operation on the first type of text to be deduplicated and the second type of text to be deduplicated of all seats, create a one-to-one retrieval database to respectively store the deduplicated text data of each seat for seat retrieval.

2. The text deduplication method based on a multi-model algorithm according to claim 1, wherein Before the step of obtaining the text to be deduplicated of the target seat, the following steps are further included: Access the underlying database through a computing engine, and load the historical communication texts of all seats stored in the underlying database into the computing memory. Among them, the historical communication texts of all seats include the text to be deduplicated of the target seat, and each historical communication text contains an identification label of the corresponding seat; Based on the identification label, divide the historical communication texts of all seats in the computing memory to obtain the historical communication texts of each seat.

3. The text deduplication method based on a multi-model algorithm according to claim 1, wherein The step of performing full deduplication on the first type of text to be deduplicated specifically includes: The non-Chinese text or the text whose length is less than a preset length threshold is fully deduplicated through a collection container, and completely repeated text in the first category of text to be deduplicated is filtered.

4. The text deduplication method based on a multi-model algorithm according to claim 1, characterized in that, The step of extracting the characteristic text of the second type of text to be deduplicated by the simhash algorithm and performing hash calculation to obtain the characteristic fingerprint value specifically includes: The second type of text to be deduplicated is segmented by the word segmenter in the simhash algorithm to obtain a plurality of text phrases; Based on the API interface of the word segmenter, extract the plurality of feature texts from the plurality of text phrases by using the TF-IDF algorithm; Calculate hash values ​​of the plurality of feature texts respectively based on a hash function; The hash value of each feature text is weighted, and the weighted results are merged and reduced in dimension based on the order of multiple feature texts to obtain the feature fingerprint value.

5. The text deduplication method based on a multi-model algorithm according to claim 1, wherein After the step of binning the characteristic fingerprint values ​​and respectively searching and removing duplicates of the multiple segments of characteristic fingerprint values ​​to be queried obtained after binning, the method further includes: A retrieval database of the target agent is created, and the deduplicated texts of the target agent are stored in the retrieval database, wherein the deduplicated texts of the target agent include the first category of texts to be deduplicated after full deduplication and the second category of texts to be deduplicated after deduplication by retrieval in boxes.

6. A text deduplication device based on a multi-model algorithm, characterized in that The text deduplication device based on the multi-model algorithm implements the steps of the text deduplication method based on the multi-model algorithm as described in any one of claims 1 to 5, and the text deduplication device based on the multi-model algorithm comprises: The acquisition module is used to obtain the text to be deduplicated of the target agent; A judgment module, used for judging whether the to-be-deduplicated texts include the first category of to-be-deduplicated texts or the second category of to-be-deduplicated texts; A first deduplication module, configured to perform full deduplication on the first category of text to be deduplicated if the text to be deduplicated includes the first category of text to be deduplicated; A calculation module, configured to extract feature text of the second type of text to be deduplicated by using a simhash algorithm and perform hash calculation to obtain a feature fingerprint value if the text to be deduplicated includes the second type of text to be deduplicated; The second deduplication module is used to bin the characteristic fingerprint values, and retrieve and deduplication the multiple segments of characteristic fingerprint values ​​to be queried obtained after binning.

7. A computer device, characterized in that, It comprises a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the text deduplication method based on a multi-model algorithm as described in any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the text deduplication method based on a multi-model algorithm as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Text processing method and device, computing device and medium

    CN110765756A

  • Real-time data deduplication method and system based on sentence-level indexes

    CN112527948A