Text duplicate removal method and related device
By encoding and classifying the coded sets of text, the problem of low text deduplication efficiency in the prior art is solved, and the effect of efficient deduplication is achieved.
Patent Information
- Application Number
- CN202510618518.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The deduplication deduplication method in the prior art is inefficient and cannot effectively process duplicate text in large amounts of text data.
By encoding each text in the current source computing node, the text coded value is obtained and divided into several types of encoding sets. The similarity of text coded value in the same type of encoding set is higher than the similarity of text coded value between different types of encoding sets. For various encoding sets, the text corresponding to the encoding set in the current source computing node is deduplicated, and the text to be deduplicated is determined based on its similarity to the text encoding value of other texts in the encoding set.
Through this method, the calculation overhead is reduced, the deduplication efficiency is improved, and duplicate text in a large amount of text data can be effectively processed.
Smart Images

Figure CN120144548A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of natural language processing, and in particular, to a text deduplication method, a text deduplication device, a target communication system, an electronic device, and a computer-readable storage medium. Background Art
[0002] Text deduplication technology is one of the core technologies in the fields of natural language processing (NLP) and data management. Its main goal is to identify and eliminate duplicate texts from a text collection. This technology is widely used in scenarios such as search engine optimization, data warehouse cleaning, copyright infringement detection, and preprocessing of machine learning training datasets. With the explosive growth of Internet data, efficient text deduplication algorithms are crucial for reducing storage costs, improving model training efficiency, and ensuring data quality.
[0003] The text deduplication method in the related art stores texts in a database and realizes deduplication by using the query, comparison, and deletion methods of traditional databases. This method has low deduplication efficiency. Summary of the Invention
[0004] The present application provides a text deduplication method, a text deduplication device, a target communication system, an electronic device, and a computer-readable storage medium, which can solve the problem of low deduplication efficiency of the text deduplication method in the related art.
[0005] The present application provides a text deduplication method, including: respectively performing a first encoding on each text in the current source computing node to obtain the text encoding values of each text; dividing the text encoding values of each text into several types of encoding sets, where the similarity between the text encoding values within the same type of encoding set is higher than the similarity between different text encoding values in different types of encoding sets; for each type of encoding set, performing deduplication on the texts corresponding to the encoding set in the current source computing node, where the texts to be deduplicated are determined based on the similarity between the texts to be deduplicated and the text encoding values of other texts in the encoding set where they are located.
[0006] The present application provides a text deduplication method, including: receiving at least one type of encoding set sent by at least one source computing node, where the encoding set is obtained by the source computing node dividing the text encoding values of each text in the source computing node, and the text encoding values of each text are obtained by the source computing node respectively encoding each text; combining the same type of encoding sets sent by each source computing node; for each type of combined encoding set, obtaining the similarity between the text encoding values in the combined encoding set; based on the similarity, determining the texts to be deduplicated in the combined encoding set, and notifying the source computing node where the texts to be deduplicated are located to delete the texts to be deduplicated.
[0007] The present application provides a text deduplication device, including: an encoding module, a partitioning module, and a deduplication module. The encoding module is configured to perform primary encoding on each text in the current source computing node respectively to obtain the text encoding values of each text; the partitioning module is configured to partition the text encoding values of each text into several categories of encoding sets, wherein the similarity between the text encoding values within the same category of encoding set is higher than the similarity between different text encoding values in different categories of encoding sets; the deduplication module is configured to, for each category of encoding sets, deduplicate the texts corresponding to the encoding sets in the current source computing node, wherein the texts to be deduplicated are determined based on the similarity between the texts to be deduplicated and the text encoding values of other texts in the encoding sets where they are located.
[0008] The present application provides a text deduplication device, including: a receiving module, a combining module, an obtaining module, and a determining module. The receiving module is configured to receive at least one category of encoding sets sent by at least one source computing node, and the encoding sets are obtained by the source computing node partitioning the text encoding values of each text in the source computing node, and the text encoding values of each text are obtained by the source computing node encoding each text respectively; the combining module is configured to combine the same category of encoding sets sent by each source computing node; the obtaining module is configured to, for each category of the combined encoding sets, obtain the similarity between the text encoding values in the combined encoding sets; the determining module is configured to, based on the similarity, determine the texts to be deduplicated in the combined encoding sets, and notify the source computing node where the texts to be deduplicated are located to delete the texts to be deduplicated.
[0009] The present application provides a target communication system, including several nodes, and the several nodes include at least one source computing node and at least one similarity computing node, and the source computing node and the similarity computing node are used for the above method.
[0010] The present application provides an electronic device, including a memory and a processor, and the processor is configured to execute program instructions stored in the memory to implement the above method.
[0011] The present application provides a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, the above method is implemented.
[0012] In the above solution, the text encoding values of each text in the current source computing node are obtained by performing one-time encoding on each text; the text encoding values of each text are divided into several classes of encoding sets, and the similarity between the text encoding values within the same class of encoding set is higher than the similarity between different text encoding values in different classes of encoding sets; for the texts corresponding to the encoding sets in each class of encoding sets, duplicate removal is performed, and the texts to be de-duplicated are determined based on the similarity between the texts to be de-duplicated and the text encoding values of other texts in the encoding sets where they are located. Therefore, in this application, instead of directly calculating the similarity of each text and performing duplicate removal based on the similarity of each text, each text is first encoded into a text encoding value and then divided into several classes of encoding sets, and then, taking the encoding sets as units, duplicate removal is performed based on the similarity between the text encoding values in the encoding sets. Therefore, the computing overhead can be reduced and the duplicate removal efficiency can be improved.
[0013] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit this application. Brief Description of the Drawings
[0014] The accompanying drawings here are incorporated into the specification and constitute a part of this specification. These drawings show embodiments consistent with this application and are used together with the specification to illustrate the technical solutions of this application.
[0015] Figure 1 is a schematic structural diagram of the target communication system provided by this application; Figure 2 is a schematic structural diagram of a specific example of the target communication system provided by this application; Figure 3 is a schematic structural diagram of another specific example of the target communication system provided by this application; Figure 4 is a schematic flowchart of the first embodiment of the text duplicate removal method provided by this application; Figure 5 is a schematic flowchart of the second embodiment of the text duplicate removal method provided by this application; Figure 6 is a schematic flowchart of the third embodiment of the text duplicate removal method provided by this application; Figure 7 is a schematic flowchart of the fourth embodiment of the text duplicate removal method provided by this application; Figure 8 is a schematic diagram of the encoding set allocation for each node of the target communication system provided by this application; Figure 9 is another schematic diagram of the encoding set allocation for each node of the target communication system provided by this application; Figure 10 is yet another schematic diagram of the encoding set allocation for each node of the target communication system provided by this application; Figure 11It is a schematic flowchart of the fifth embodiment of the text deduplication method provided by this application; Figure 12 It is a schematic diagram of the large model pre-training system based on PYTORCH of this application; Figure 13 It is a schematic diagram of a specific example of the text deduplication method provided by this application; Figure 14 It is a schematic diagram of the text encoding value of the text of this application; Figure 15 It is a schematic diagram of the encoding set obtained by dividing 4 nodes of this application; Figure 16 It is a comparison schematic diagram of the Jaccard similarity between the text encoding values of the text and the true similarity between the texts; Figure 17 It is a schematic structural diagram of an embodiment of the text deduplication device of this application; Figure 18 It is a schematic structural diagram of another embodiment of the text deduplication device of this application; Figure 19 It is a schematic structural diagram of an embodiment of the electronic device of this application; Figure 20 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of this application. Detailed implementation manners
[0016] Next, in combination with the accompanying drawings of the specification, the solutions of the embodiments of this application will be described in detail.
[0017] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures, interfaces, and technologies are set forth in order to provide a thorough understanding of this application.
[0018] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more than two. In addition, the term "at least one" in this article represents any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.
[0019] Figure 1 It is a schematic structural diagram of the target communication system provided by this application. As Figure 1As shown in the figure, the target communication system includes several nodes. The nodes have the ability to deduplicate text and can be physical machines, virtual machines, or container instances, etc. The physical machines can be electronic devices such as computers, mobile phones, servers, etc. with the ability to perform text deduplication calculations. The nodes can communicate with each other.
[0020] The several nodes include at least one source computing node and at least one similarity computing node. The same node can be both a source computing node and a similarity computing node.
[0021] Figure 2 It is a schematic structural diagram of a specific example of the target communication system provided by this application. As Figure 2 shown, each of the several nodes is both a source computing node and a similarity computing node. For example, the several nodes include one node, and this node is both a source computing node and a similarity computing node. Another example is that the several nodes include nodes A to D, and nodes A to D are both source computing nodes and similarity computing nodes.
[0022] Figure 3 It is a schematic structural diagram of another specific example of the target communication system provided by this application. As Figure 3 shown, some of the several nodes are source computing nodes, and some of the other nodes are similarity computing nodes. For example, the several nodes include nodes A to D, nodes A to B are source computing nodes, and nodes C to D are similarity computing nodes.
[0023] The source computing nodes store text. In some embodiments, there is one source computing node among the several nodes. In this case, all the text is stored in this source computing node. In some embodiments, there can also be multiple source computing nodes among the several nodes. In this case, all the text is stored in the source computing nodes in a uniform or non-uniform distribution.
[0024] The target communication system can be a large model pre-training system based on PYTORCH, a specially built communication system, or other existing communication systems.
[0025] Based on the above target communication system, the text deduplication method provided by this application can be implemented. The following introduces the embodiments of the text deduplication method provided by this application. It should be noted that the text deduplication process of each source computing node for the text it stores is similar. In the following embodiments of this application, only one source computing node (the current source computing node) is used as an example to introduce the text deduplication method provided by this application.
[0026] Figure 4 It is a schematic flowchart of the first embodiment of the text deduplication method provided by this application. As Figure 4 shown, in this embodiment, the text deduplication method may include the following steps: S110: Encode each text in the current source computing node once to obtain the text encoding values of each text.
[0027] The execution subject of this embodiment is the current source computing node.
[0028] In some embodiments, from the perspective of content, the text can be comments, opinions, conversations, logs, news, etc. Each text in the current source computing node refers to each text stored in the current source computing node. From the perspective of application scenarios, the text can be pre-training text for a large model, text to be queried in a database, etc.
[0029] In some embodiments, the method of encoding once can be hash encoding, such as Minhash, simHash. Or, the method of encoding once can be a statistics-based encoding method, such as the bag-of-words model, TF-IDF. Or, the method of encoding once can be word embedding-based encoding, such as static word vectors, context-dependent word vectors. Or, the method of encoding once can be deep learning-based semantic encoding, etc.
[0030] In some embodiments, each text can be directly encoded once to obtain the text encoding values of each text.
[0031] In some embodiments, each text can be divided into several text segments and then encoded once. In this case, the text encoding value of the text includes the encoded sub-values of each text segment.
[0032] In some embodiments, the text encoding value of the text further includes the location information of the text. The location information of the text includes the first address information of the current source computing node where the text is located, and / or the second address information of the text in the current source computing node. The first address information is used to identify the location of the current source computing node where the text is located, and the second address information is used to identify the storage location of the text in the current source computing node.
[0033] S120: Divide the text encoding values of each text into several classes of encoding sets.
[0034] Among them, the similarity between the text encoding values within the same class of encoding set is higher than the similarity between different text encoding values in different classes of encoding sets.
[0035] The similarity between text encoding values can represent the similarity between the corresponding texts. Moreover, the similarity between text encoding values is positively correlated with the similarity between the corresponding texts. Specifically, the higher the similarity between text encoding values, the higher the similarity between the corresponding texts. Conversely, the lower the similarity between text encoding values, the lower the similarity between the corresponding texts. For example, the higher the similarity between the text encoding value a of text 1 and the text encoding value b of text 2, the higher the similarity between text 1 and text 2.
[0036] In some embodiments, the same text encoding values can be grouped into the same class of encoding sets, and different text encoding values can be grouped into different classes of encoding sets.
[0037] In some embodiments, based on the similarity between the text encoding values of each text, the text encoding values of each text can be clustered to obtain several classes of encoding sets. Thus, the text encoding values with a similarity greater than or equal to the similarity threshold are grouped into the same class of encoding sets, and the text encoding values with a similarity less than the similarity threshold are grouped into different classes of encoding sets.
[0038] It can be understood that, compared with the method of directly grouping different text encoding values into different classes of encoding sets, clustering based on similarity can reduce the number of classes of encoding sets.
[0039] In some embodiments, S120 includes S121 - S122. S121: For each text, perform secondary encoding on the text encoding value of the text to obtain the class encoding value of the text. S122: Divide the text encoding values corresponding to the texts with the difference degree between class encoding values less than the preset difference degree into the same encoding set. The method of secondary encoding can be hash encoding or other encoding methods. The method of secondary encoding can be the same as or different from the method of primary encoding. The difference degree less than the preset difference degree can be divided into two cases. One case is that the difference degree is equal to 0, that is, the class encoding values are the same. The other case is that the difference degree is greater than 0 and less than the preset difference degree, that is, the class encoding values are similar.
[0040] In some embodiments, S122 includes: Divide the text encoding values corresponding to the texts with the same class encoding value into the same class of encoding sets.
[0041] It can be understood that, compared with the method of directly grouping the text encoding values corresponding to the texts with the same class encoding value into different classes of encoding sets, the method of grouping the text encoding values corresponding to the texts with the same and similar class encoding values into different classes of encoding sets can reduce the number of classes of encoding sets.
[0042] S130: For each class of encoding sets, remove duplicates from the texts corresponding to the encoding sets in the current source computing node.
[0043] Among them, the text to be deduplicated is determined based on the similarity between the text to be deduplicated and the text encoding values of other texts in the encoding set where it is located.
[0044] In some embodiments, the text to be deduplicated corresponding to each type of encoding set is determined by the current source computing node.
[0045] In some embodiments, the text to be deduplicated corresponding to at least one type of encoding set is determined by the current source computing node. The text to be deduplicated corresponding to other types of encoding sets is determined by other source computing nodes / similarity computing nodes. When the text to be deduplicated corresponding to the same type of encoding set is determined by the same source computing node / similarity computing node, the fewer the number of categories of the encoding set, the fewer the number of source computing nodes / similarity computing nodes required to determine the text to be deduplicated.
[0046] In some embodiments, the text to be deduplicated corresponding to each type of encoding set is determined by other source computing nodes / similarity computing nodes.
[0047] In the above solution, each text in the current source computing node is encoded once to obtain the text encoding value of each text; the text encoding values of each text are divided into several types of encoding sets, and the similarity between the text encoding values within the same type of encoding set is higher than the similarity between different text encoding values in different types of encoding sets; for the text corresponding to each type of encoding set, deduplication is performed on the text to be deduplicated, and the text to be deduplicated is determined based on the similarity between the text to be deduplicated and the text encoding values of other texts in the encoding set where it is located. Therefore, in this application, the similarity of each text is not directly calculated and deduplication is performed based on the similarity of each text. Instead, each text is encoded into a text encoding value once and then divided into several types of encoding sets. Then, taking the encoding set as a unit, deduplication is performed based on the similarity between the text encoding values in the encoding set. Therefore, the calculation overhead can be reduced and the deduplication efficiency can be improved.
[0048] Figure 5 It is a schematic flowchart of the second embodiment of the text deduplication method provided by this application. This embodiment is a further expansion of S110. In this embodiment, the text is divided into several text segments and then encoded once. As Figure 5 shown, in this embodiment, S110 may include the following steps: S210: For each text, divide the text into several text segments.
[0049] In some embodiments, the text can be divided into several text segments in units of paragraphs, sentences, words, characters, etc.
[0050] In some embodiments, the text can be divided into several text segments by using methods such as NGram and sliding window.
[0051] S220: Encode each text segment of the text once to obtain the encoded sub-values of each text segment of the text.
[0052] S230: Use the encoded sub-values of each text segment of the text to obtain the text encoding value of the text.
[0053] In some embodiments, the encoded sub-values of each text segment of the text can be combined to obtain the text encoding value of the text.
[0054] In some embodiments, the encoded sub-values of each text segment of the text and the location information of the text can be combined to obtain the text encoding value of the text. Wherein, the location information of the text includes the first address information of the current source computing node where the text is located and the second address information of the text in the current source computing node. Alternatively, the location information of the text only includes the first address information of the current source computing node where the text is located.
[0055] It can be understood that when the text encoding value of the text includes the first address information and the second address information of the text, it is supported to send the encoding set to other source computing nodes / similarity computing nodes to determine the text to be deduplicated through other source computing nodes / similarity computing nodes.
[0056] Different from the foregoing embodiments, through the implementation of this embodiment, the present application divides each text into several text segments and then encodes them once, which can improve the accuracy of the text encoding value for text expression. In addition, the location information of the text can be carried in the text encoding value, and it is supported to send the encoding set to other source computing nodes / similarity computing nodes to determine the text to be deduplicated through other source computing nodes / similarity computing nodes.
[0057] Further, in some embodiments, the communication system where the current source computing node is located is a target communication system, and several source computing nodes in the target communication system all contain several texts, and the current source computing node is one of the several source computing nodes. Based on this, S130 can be extended as follows: In S130, at least one type of encoding set among several types of encoding sets can be used as the first type of encoding set, and / or, at least one type of encoding set among several types of encoding sets can be used as the second type of encoding set. The text to be deduplicated corresponding to the first type of encoding set is determined by the current source computing node, and the text to be deduplicated corresponding to the second type of encoding set is determined by other source computing nodes / similarity computing nodes.
[0058] In some embodiments, all types of encoding sets among several types of encoding sets are used as the first type of encoding set.
[0059] In some embodiments, among several types of encoding sets, one type of encoding set is used as the first type of encoding set, and the other types of encoding sets are used as the second type of encoding set.
[0060] In some embodiments, among several types of encoding sets, a first preset number of types of encoding sets are used as the first type of encoding set, and the other types of encoding sets are used as the second type of encoding set. The first preset number is greater than 1 and less than the total number of types of the encoding sets. For example, if there are 4 types of encoding sets in total, the first preset number can be 2 or 3.
[0061] In some embodiments, among several types of encoding sets, a second preset number of types of encoding sets are used as the second type of encoding set, and the other types of encoding sets are used as the first type of encoding set. The second preset number is greater than 1 and less than the total number of types of the encoding sets. For example, if there are 4 types of encoding sets in total, the second preset number can be 2 or 3.
[0062] In some embodiments, all types of encoding sets among several types of encoding sets are used as the second type of encoding set.
[0063] It can be understood that in the case of using at least one type of encoding set as the second type of encoding set, it is necessary to send the second type of encoding set to other source computing nodes / similar computing nodes, so that the other source computing nodes / similar computing nodes can determine the text to be deduplicated corresponding to the second type of encoding set, and receive the relevant data (second address information or text encoding value) of the text to be deduplicated determined by the other source computing nodes / similar computing nodes. Since the data volume of the second type of encoding set / text encoding value / second address information is smaller than that of the text, compared with directly transmitting the text between the current source computing node and other source computing nodes / similar computing nodes, the present application only transmits the second type of encoding set / text encoding value / second address information between nodes, and the text will always be stored only in the source computing node and will not be transmitted between source computing nodes, which can greatly reduce the data volume transmitted between nodes, thereby improving the communication efficiency and further improving the deduplication efficiency. In addition, the number of times the node reads and writes the text stored by it is reduced, and the storage space occupied by the transmitted data on the node is reduced.
[0064] Figure 6 It is a schematic flowchart of the third embodiment of the text deduplication method provided by the present application. This embodiment is a further expansion of S130. As Figure 6 shown, in this embodiment, S130 may include the following steps: S310: Use at least one type of encoding set as the first type of encoding set.
[0065] In some embodiments, all types of encoding sets can be used as the first type of encoding set.
[0066] In some embodiments, the first preset number of classes of coding sets may be used as the first class of coding sets.
[0067] In some embodiments, one of the classes of coding sets may be used as the first class of coding sets.
[0068] S320: For each first class of coding sets, combine the first class of coding sets in each source computing node, and calculate the similarity between the text coding values in the combined first class of coding sets.
[0069] The first class of coding sets in each source computing node includes the first class of coding sets retained by the current source computing node and the first class of coding sets received by the current source computing node from other source computing nodes.
[0070] In some embodiments, the first class of coding sets in each source computing node is all the classes of coding sets divided by the corresponding source computing node. In this case, it can be regarded that all the classes of coding sets divided by each source computing node are aggregated to the current source computing node, and the current source computing node determines the text to be de-duplicated.
[0071] In some embodiments, the first class of coding sets in each source computing node is the first preset number of classes of coding sets divided by the corresponding source computing node. In this case, it can be regarded that the first preset number of classes of coding sets divided by each source computing node are aggregated to the current source computing node, and the current source computing node determines the text to be de-duplicated.
[0072] In some embodiments, the first class of coding sets in each source computing node is one of the classes of coding sets divided by the corresponding source computing node. In this case, it can be regarded that one of the classes of coding sets divided by each source computing node is aggregated to the current source computing node, and the current source computing node determines the text to be de-duplicated. In this case, it can be regarded that the same class of coding sets divided by each source computing node is aggregated to the current source computing node, and the current source computing node determines the text to be de-duplicated.
[0073] In some embodiments, when the first class of coding sets of each source computing node only includes one class of coding sets, the coding sets of this class of each source computing node can be combined to obtain the combination result of this first class of coding sets.
[0074] In some embodiments, when the first type of coding sets of each source computing node include multiple types of coding sets, the various types of first type of coding sets of each source computing node can be combined respectively to obtain the combined results of the various types of first type of coding sets. For example, if the first type of coding sets of each source computing node include the coding set of category 1 and the coding set of category 2, the coding sets of category 1 of each source computing node can be combined to obtain the combined result of category 1, and the coding sets of category 2 of each source computing node can be combined to obtain the combined result of category 2.
[0075] S330: Use the texts corresponding to the text coding values whose similarities in the combined first type of coding sets meet the similarity requirements as the duplicate text group.
[0076] In some embodiments, that the similarity meets the similarity requirements can be that the similarity between pairwise text coding values is greater than the similarity threshold. For example, there are 100 texts in the combined first type of coding sets, and the similarity between pairwise of "50 of these texts" is greater than or equal to the similarity threshold, while the similarity between pairwise of "the other 50 texts" is less than the similarity threshold. Then, "50 of these texts" are used as the duplicate text group.
[0077] In some embodiments, that the similarity meets the similarity requirements can also be that the difference between the similarity of pairwise of some text coding values and the similarity of pairwise of other text coding values is greater than the difference threshold. For example, there are 100 texts in the combined first type of coding sets, the similarity between pairwise of "50 of these texts" is greater than the first similarity, the similarity between pairwise of "the other 50 texts" is less than the second similarity, the first similarity is greater than the second similarity, and the difference between the second similarity and the first similarity is greater than the difference threshold. Then, "50 of these texts" are used as the duplicate text group.
[0078] S340: Select at least one text from the duplicate text group as the first text to be de-duplicated, delete the first text to be de-duplicated located at the current source computing node, and notify the source computing nodes where the non-local texts are located to delete the non-local texts.
[0079] Among them, the non-local text is the first text to be de-duplicated that is not located at the current source computing node.
[0080] Each text in the duplicate text group is a duplicate text. For duplicate texts, only one needs to be retained in each source computing node. Therefore, at least one text needs to be selected from the duplicate text group as the first text to be de-duplicated for deletion in the corresponding source computing node.
[0081] Since the duplicate text group is determined based on the first type of coding set retained by the current source computing node and the received first type of coding set, the duplicate text group may contain the first text to be de-duplicated corresponding to the retained first type of coding set, or may contain the first text to be de-duplicated corresponding to the received first type of coding set. The first text to be de-duplicated corresponding to the retained first type of coding set is stored in the current source computing node and belongs to local text. The first text to be de-duplicated corresponding to the received first type of coding set is not stored in the current source computing node but in other source computing nodes and belongs to non-local text. Therefore, local text can be directly deleted, and for non-local text, the source computing node where it is located needs to be notified to delete it.
[0082] Different from the foregoing embodiments, in this embodiment, the first type of coding sets divided by each source computing node are aggregated to the current source computing node to determine the first text to be de-duplicated corresponding to the first type of coding set through the current source computing node, thereby saving computing overhead.
[0083] Figure 7 It is a schematic flowchart of the fourth embodiment of the text de-duplication method provided by this application. This embodiment is a further expansion of S130. As Figure 7 shown, in this embodiment, S130 may include the following steps: S410: Use at least one type of coding set as the second type of coding set.
[0084] In some embodiments, all types of coding sets may be used as the second type of coding set.
[0085] In some embodiments, a second preset number of types of coding sets may be used as the second type of coding set.
[0086] In some embodiments, one of the types of coding sets may be used as the second type of coding set.
[0087] S420: For each second type of coding set, send the second type of coding set of the current source computing node to the similarity calculation node in the target communication system, so that the similarity calculation node combines the second type of coding sets of each source computing node, and based on the similarity between the text coding values in the combined second type of coding set, determines the second text to be de-duplicated corresponding to the combined second type of coding set, and notifies the source computing node where the second text to be de-duplicated is located to delete the second text to be de-duplicated.
[0088] In some embodiments, the similarity calculation node is a source computing node. Alternatively, the similarity calculation node is not a source computing node.
[0089] In some embodiments, if there are multiple second - type coding sets, each second - type coding set of the current source computing node can be sent to different similarity - computing nodes respectively. Alternatively, some of the second - type coding sets of the current source computing node can be sent to the same similarity - computing node. Alternatively, all of the second - type coding sets of the current source computing node can be sent to the same similarity - computing node.
[0090] In some embodiments, when the similarity - computing node is not the source computing node, the second - type coding sets of each source computing node include the second - type coding sets sent by each source computing node received by the similarity - computing node. The similarity - computing node can combine the second - type coding sets of each source computing node to obtain a combined second - type coding set; based on the similarity between the text coding values in the combined second - type coding set, determine the second text to be deduplicated (non - local) corresponding to the combined second - type coding set, and notify the source computing node where it is located to delete it.
[0091] In some embodiments, when the similarity - computing node is the source computing node, the second - type coding sets of each source computing node include the second - type coding sets reserved by the similarity - computing node and the second - type coding sets sent by other source computing nodes (including the current source computing node). The similarity - computing node can combine the second - type coding sets of each source computing node to obtain a combined second - type coding set; based on the similarity between the text coding values in the combined second - type coding set, determine the second text to be deduplicated (local and non - local) corresponding to the combined second - type coding set, directly delete the local second text to be deduplicated, and notify the source computing node where the non - local second text to be deduplicated is located to delete it.
[0092] In some embodiments, when the second - type coding sets of each source computing node include only one type of coding set, the similarity - computing node can combine the coding sets of this type of each source computing node to obtain a combined result of this type of second - type coding set.
[0093] In some embodiments, when the second - type coding sets of each source computing node include multiple types of coding sets, the similarity - computing node can combine each type of second - type coding set of each source computing node respectively to obtain a combined result of each type of second - type coding set. For example, if the second - type coding sets of each source computing node include a coding set of category 3 and a coding set of category 4, the coding sets of category 3 of each source computing node can be combined to obtain a combined result of category 3, and the coding sets of category 4 of each source computing node can be combined to obtain a combined result of category 4.
[0094] Different from the foregoing embodiments, in this embodiment, the second type of encoding sets obtained by partitioning each source computing node are aggregated to the same similarity calculation node, so as to determine the second text to be deduplicated through the similarity calculation node, thereby saving computing overhead.
[0095] To facilitate the understanding of S310 - S340 and S410 - S420, the determination of the text to be deduplicated (the first text to be deduplicated, the second text to be deduplicated) is described below in the form of three specific examples.
[0096] Example 1: All the class encoding sets obtained by partitioning the current source computing node are used as the first type of encoding sets. The first type of encoding sets of each source computing node are aggregated to the current source computing node, and the current source computing node determines the first text to be deduplicated.
[0097] Figure 8 is a schematic diagram of the encoding set allocation of each node in the target communication system of this application. As Figure 8 shown, the target communication system includes nodes A - D. Nodes A - D are both source computing nodes and similarity calculation nodes, and node A is the current source computing node. Node A partitions the encoding sets A1 - A4 of categories 1 - 4, node B partitions the encoding sets B1 - B4 of categories 1 - 4, node C partitions the encoding sets C1 - C4 of categories 1 - 4, and node D partitions the encoding sets D1 - D4 of categories 1 - 4. Nodes A - D all use A1 - A4, B1 - B4, C1 - C4, D1 - D4 as the first type of encoding sets, and nodes B - D send the 4 - type encoding sets B1 - B4, C1 - C4, D1 - D4 to node A respectively. Node A combines the encoding sets A1, B1, C1, D1 of category 1 to obtain the combination result of category 1, combines the encoding sets A2, B2, C2, D2 of category 2 to obtain the combination result of category 2, combines the encoding sets A3, B3, C3, D3 of category 3 to obtain the combination result of category 3, and combines the encoding sets A4, B4, C4, D4 of category 4 to obtain the combination result of category 4; the first text to be deduplicated is determined respectively based on the combination results of categories 1 - 4.
[0098] Example 2: One of the class encoding sets obtained by partitioning the current source computing node is used as the first type of encoding set, and the other class encoding sets are used as the second type of encoding sets. The first type of encoding sets of each source computing node are aggregated to the current source computing node, and the current source computing node determines the first text to be deduplicated. The second type of encoding sets of each source computing node are aggregated to the similarity calculation node, and the similarity calculation node determines the second text to be deduplicated.
[0099] Figure 9 is another schematic diagram of the encoding set allocation of each node in the target communication system of this application. Different from Figure 8 that is, inFigure 9 Among them, nodes A - D respectively use the coding sets A1, B1, C1, D1 of category 1 as the first - type coding sets, and use A2 - A4, B2 - B4, C2 - C4 as the second - type coding sets. The first - type coding sets A1, B1, C1, D1 are aggregated to node A, and node A determines the corresponding first text to be de - duplicated. The coding sets A2, B2, C2, D2 of category 2 in the second - type coding sets are aggregated to node B, and node B determines the corresponding second text to be de - duplicated. The coding sets A3, B3, C3, D3 of category 3 in the second - type coding sets are aggregated to node C, and node C determines the corresponding second text to be de - duplicated. The coding sets A4, B4, C4, D4 of category 4 in the second - type coding sets are aggregated to node D, and node D determines the corresponding second text to be de - duplicated.
[0100] Example 3: All the class coding sets obtained by partitioning the current source computing nodes are used as the second - type coding sets, and the first - type coding sets of each source computing node are aggregated to the similarity computing node, and the similarity computing node determines the second text to be de - duplicated.
[0101] It can be understood that in Example 2 and Example 3 and similar situations, the fewer the number of categories of the coding sets, the fewer the number of source computing nodes / similarity computing nodes required to determine the text to be de - duplicated.
[0102] Figure 10 is another schematic diagram of the coding set allocation of each node in the target communication system of the present application. Different from Figure 8 and Figure 9 In Figure 10 the target communication system further includes nodes E - F, and nodes E - F are similarity computing nodes.
[0103] Nodes A - D respectively use the encoding sets A1 - A4, B1 - B4, C1 - C4, D1 - D4 of categories 1 - 4 as the second - type encoding sets. The encoding sets A1 - A2, B1 - B2, C1 - C2, D1 - D2 of categories 1 - 2 in the second - type encoding sets are aggregated to node E. Node E combines the encoding sets A1, B1, C1, D1 of category 1 to obtain the combined result of category 1, combines the encoding sets A2, B2, C2, D2 of category 2 to obtain the combined result of category 2, and determines the corresponding second text to be de - duplicated respectively based on the combined result of category 1 and the combined result of category 2. The encoding sets A3 - A4, B3 - B4, C3 - C4, D3 - D4 of categories 3 - 4 in the second - type encoding sets are aggregated to node F. Node F combines the encoding sets A3, B3, C3, D3 of category 3 to obtain the combined result of category 3, combines the encoding sets A4, B4, C4, D4 of category 4 to obtain the combined result of category 4, and determines the corresponding second text to be de - duplicated respectively based on the combined result of category 3 and the combined result of category 4.
[0104] Further, in some embodiments, in S340, the source computing node where the non - local text is located is notified to delete the non - local text, or in S420, the source computing node where the second text to be de - duplicated is located is notified to delete the second text to be de - duplicated, including S510 - S530 (not shown in the figure). S510: Use the non - local text or the second text to be de - duplicated as the target text to be de - duplicated. S520: From the text encoding value of the target text to be de - duplicated, obtain the first address information of the source computing node where the target text to be de - duplicated is located and the second address information of the target text to be de - duplicated in the source computing node. S530: According to the first address information, send the second address information to the source computing node where the target text to be de - duplicated is located to notify the source computing node to delete the text corresponding to the second address information.
[0105] In some embodiments, different from S530, the text encoding value can be sent to the source computing node where the target text to be de - duplicated is located to notify the source computing node to delete the text corresponding to the second address information.
[0106] In some embodiments, the text de - duplication method provided in this application is executed by several nodes of a distributed target communication system. Among them, if the similarity calculation nodes are all source computing nodes, the several nodes are several source computing nodes. If the similarity calculation nodes can be non - source computing nodes, the several nodes can include several source computing nodes and similarity calculation nodes.
[0107] Based on this, the text de - duplication method further includes: when it is necessary to send data to other nodes, convert the data to be sent into a format supported by the target communication system. The distributed target communication system can be a large - model pre - training system based on PYTORCH, HADOOP, SPARK, etc.
[0108] For example, the text deduplication method provided in this application is executed by several nodes under the large model pre-training system based on PYTORCH, and the several nodes include the current source computing node. Based on this, the text deduplication method further includes: in the case of needing to send data to other nodes, converting the data to be sent into a format supported by the large model pre-training system based on PYTORCH.
[0109] The data to be sent can be an encoding set, the second address information of the target deduplicated text in the source computing node, the text encoding result, etc. Specifically, after dividing to obtain several categories of encoding sets, in the case of needing to send the encoding set, convert the encoding set into a format supported by the large model pre-training system based on PYTORCH and then send it. After determining the target deduplicated text, convert the second address information or the text encoding result of the target deduplicated text in the source computing node into a format supported by the large model pre-training system based on PYTORCH and then send it.
[0110] It can be understood that the large model pre-training system based on PYTORCH is a communication system for realizing large model pre-training. When the text deduplication method is executed by several nodes under the large model pre-training system based on PYTORCH, on the one hand, there is no need to additionally build a target communication system, reducing the cost of building the target communication system. On the other hand, the large model pre-training system based on PYTORCH has built-in efficient distributed communication algorithms, including different communication protocols, which can use the CPU for data processing or the GPU for data processing (data processing across multiple nodes), the data communication is stable, and the communication strategy can separately call the built-in P2P interface design. The underlying layer of PYTORCH can use C++ for algorithm efficiency optimization, the serialization scheme can be custom-designed, any complex data structure can be compatible, and the number of nodes can be freely expanded. Therefore, the communication efficiency is high, the stability is high, and the compatibility is high.
[0111] Figure 11 It is a schematic flowchart of the fifth embodiment of the text deduplication method provided in this application. As Figure 11 shown, in this embodiment, the text deduplication method may include the following steps: S610: Receive at least one type of encoding set sent by at least one source computing node.
[0112] The encoding set is obtained by the source computing node dividing the text encoding values of each text in the source computing node, and the text encoding value of each text is obtained by the source computing node encoding each text respectively.
[0113] The execution subject of this embodiment is the similarity calculation node.
[0114] S620: Combine the same type of encoding sets sent by each source computing node.
[0115] For example, receive the encoding set of category 1 and the encoding set of category 2 sent by each source computing node, combine the encoding sets of category 1 sent by each source computing node to obtain a combined result of category 1, and combine the encoding sets of category 2 sent by each source computing node to obtain a combined result of category 2.
[0116] S630: For each combined encoding set, obtain the similarity between each text encoding value in the combined encoding set.
[0117] For example, obtain the similarity between each text encoding value in the combined result of category 1, and obtain the similarity between each text encoding value in the combined result of category 2.
[0118] S640: Based on the similarity, determine the text to be deduplicated in the combined encoding set, and notify the source computing node where the text to be deduplicated is located to delete the text to be deduplicated.
[0119] In some embodiments, notifying the source computing node where the text to be deduplicated is located to delete the text to be deduplicated includes: obtaining the first address information of the source computing node where the text to be deduplicated is located and the second address information of the text to be deduplicated in the source computing node from the text encoding value of the text to be deduplicated; according to the first address information, sending the second address information to the source computing node where the text to be deduplicated is located to notify the source computing node to delete the text corresponding to the second address information.
[0120] For other detailed descriptions related to this embodiment, please refer to the previous embodiments and will not be elaborated here.
[0121] Different from the foregoing embodiments, in this embodiment, the same type of encoding sets of each source computing node can be aggregated, and the same type of encoding sets of each source computing node can be combined to obtain each combined encoding set, and the text to be deduplicated is determined based on the similarity between each text encoding value in each combined encoding set. Therefore, the computing overhead can be reduced.
[0122] For ease of understanding, the text deduplication method provided in this application is described below in the form of a specific example: With reference to Figure 12 , Figure 12 is a schematic diagram of the large model pre-training system based on PYTORCH in this application. As Figure 12 shown, the system includes nodes 1-N, and nodes 1-N are all source computing nodes. The text data is evenly divided into data blocks 1-N and distributed and stored in nodes 1-N.
[0123] With reference to Figure 13 , Figure 13It is a schematic diagram of a specific example of the text deduplication method provided by this application. As Figure 13 shown, the text deduplication method includes: 1. Each node respectively performs a hash encoding on the text stored in it to obtain the text encoding values of each text.
[0124] 1) For a text stored in a node, the node divides it into 128 text segments according to the NGram method.
[0125] 2) Respectively perform a hash encoding on the 128 text segments of the text to obtain the encoded sub-values 0 - 128 of the 128 text segments of the text.
[0126] 3) Combine the encoded sub-values 0 - 128 of the 128 text segments of the text, the first address information NodelID of the text, and the second address information InnerID of the text to obtain the text encoding value of the text. NodelID identifies the node where the text is located. InnerID identifies the position of the text in the node. Figure 14 It is a schematic diagram of the text encoding value of the text of this application.
[0127] 2. Each node respectively divides the text encoding values of the text stored in it into N types of encoding sets. The number of encoding set categories = the total number of nodes N.
[0128] For each node, the node respectively performs a secondary hash encoding on the text encoding values of the text stored in it to obtain the category encoding values of each text. The node divides the text encoding values of each text based on the category encoding values, where the text encoding values corresponding to texts with a difference degree between category encoding values less than the preset difference degree are divided into the same encoding set.
[0129] Figure 15 It is a schematic diagram of the encoding sets obtained by dividing 4 nodes in this application. As Figure 15 shown, there are a total of nodes 1 - 4 (N = 4). Node 1 divides the text encoding values into encoding sets 11, 12, 13, 14 of categories 1 - 4 based on the similarity between the text encoding values of the text stored in it. Similarly, node 2 divides to obtain encoding sets 21, 22, 23, 24 of categories 1 - 4, node 3 divides to obtain encoding sets 31, 32, 33, 34 of categories 1 - 4, and node 4 divides to obtain encoding sets 41, 42, 43, 44 of categories 1 - 4.
[0130] 3. Each node summarizes the same type of encoding sets to the same node and different types of encoding sets to different nodes according to the SCATTER strategy of the ring communication scheme of PYTORCH.
[0131] Before summarizing various coding sets, each node needs to convert the various coding sets obtained by its division into the TENSOR data structure supported by PYTORCH for communication and transmission between different nodes.
[0132] Under the SCATTER strategy of the ring communication scheme in PYTORCH, each node communicates directly only with its predecessor node and successor node.
[0133] Continue to refer to Figure 15 , node 1 retains the coding set 11 of category 1 it obtained and receives the coding sets 21, 31, and 41 of category 1 sent by nodes 2-4. Node 2 retains the coding set 22 of category 2 it obtained and receives the coding sets 12, 32, and 42 of category 2 sent by nodes 1, 3-4. Node 3 retains the coding set 33 of category 3 it obtained and receives the coding sets 13, 23, and 43 of category 3 sent by nodes 1-2 and 4. Node 4 retains the coding set 44 of category 4 it obtained and receives the coding sets 14, 24, and 34 of category 4 sent by nodes 1-3.
[0134] Thus, the coding sets of category 1 divided by nodes 1-4 are summarized to node 1, the coding sets of category 2 divided by nodes 1-4 are summarized to node 2, the coding sets of category 3 divided by nodes 1-4 are summarized to node 3, and the coding sets of category 4 divided by nodes 1-4 are summarized to node 4.
[0135] When the text volume is large enough, it can be ensured that the number of text coding values received by each node is basically the same, so that the load is balanced.
[0136] 4. Each node determines the duplicate text group.
[0137] Node 1 combines the 4 coding sets of category 1 to obtain the combined result of category 1; based on the Jaccard similarity between the text coding values in the combined result of category 1, it determines the duplicate text group corresponding to the combined result of category 1.
[0138] Node 2 combines the 4 coding sets of category 2 to obtain the combined result of category 2; based on the Jaccard similarity between the text coding values in the combined result of category 2, it determines the duplicate text group corresponding to the combined result of category 2.
[0139] Node 3 combines the 4 coding sets of category 3 to obtain the combined result of category 3; based on the Jaccard similarity between the text coding values in the combined result of category 3, it determines the duplicate text group corresponding to the combined result of category 3.
[0140] Node 4 combines the encoding sets of 4 categories of type 4 to obtain the combined result of category 4; based on the Jaccard similarity between the text encoding values in the combined result of category 4, the duplicate text group corresponding to the combined result of category 4 is determined.
[0141] The calculation formula for the Jaccard similarity between the text encoding values of two texts in the combined result is as follows: X = {x1, x2, x3,...x128}; Y = {y1,y2,y3,...y128}; ; Among them, X represents the text encoding value of one of the texts in the combined result, and x1 represents the first encoded sub-value of X. Y represents the text encoding value of another text, and y1 represents the first encoded sub-value of Y. represents the Jaccard similarity between X and Y, represents the number of encoded sub-values in the intersection of X and Y, represents the number of encoded sub-values in the union of X and Y.
[0142] Figure 16 is a comparison schematic diagram of the Jaccard similarity between the text encoding values of texts and the true similarity between texts. As Figure 16 shown, parti and partj represent the indexes of text i and text j respectively, linei and linej represent the line numbers of text i and text j in the source text data respectively, and value represents the specific content of text i and text j. real dis represents the true similarity between texts, and minhash represents the Jaccard similarity between the text encoding values of texts. It can be seen from Figure 16 that minhash can better represent the true similarity between texts.
[0143] In the case where the Jaccard similarity between the text encoding values of two texts in the combined result is greater than the Jaccard similarity threshold, these two texts are determined as duplicate texts.
[0144] 5. Each node directly deletes the local texts in the duplicate text group, and for the non-local texts in the duplicate text group, notifies the node where they are located to delete them.
[0145] The duplicate text group determined by Node 1 contains 8 duplicate texts, and duplicate texts 1-7 to be de-duplicated are selected from the 8 duplicate texts.
[0146] Node 1 determines the NodelID and InnerID of the 7 repeated texts respectively based on the text encoding values of the repeated texts 1-7. The NodelIDs of the repeated texts 1-7 represent Node 1, Node 1, Node 1, Node 2, Node 2, Node 3, and Node 4 respectively. That is, the repeated texts 1-3 are stored in Node 1 (local text), the repeated texts 4-5 are stored in Node 2 (non-local text), the repeated text 6 is stored in Node 3 (non-local text), and the repeated text 7 is stored in Node 4 (non-local text).
[0147] Node 1 directly deletes the repeated texts 1-3.
[0148] After converting the InnerIDs of the repeated texts 4-5, 6, and 7 into the format supported by the large model pre-training system based on PYTORCH, Node 1 sends them to Node 2, Node 3, and Node 4 respectively, so that Node 2 deletes the repeated texts 4-5, Node 3 deletes the repeated text 6, and Node 4 deletes the repeated text 7.
[0149] It can be understood that in the related art, the text deduplication method based on the distributed communication system determines the texts to be deduplicated based on the similarity between texts. Therefore, the texts need to be transmitted between nodes in the communication system. As the amount of text data increases, nodes need to be added to improve the deduplication efficiency. However, at the same time, the communication volume between nodes will increase exponentially, resulting in a double decline in communication efficiency and stability.
[0150] In the above specific example of this application, all texts are evenly divided among all nodes of PYTORCH. For each node, the node performs a hash encoding on each text therein to form a text encoding value; the node performs a secondary hash encoding on each text encoding value to form a category encoding value; the node divides the text encoding values with the same category encoding value into the same class encoding set. The same class encoding sets are aggregated to the same node, and different class encoding sets are aggregated to different nodes to determine the texts to be deduplicated. The text encoding value carries the NodelID and InnerID of the corresponding text. After each node determines the texts to be deduplicated, it sends the InnerID of the texts to be deduplicated to the NodelID node.
[0151] On the one hand, instead of directly calculating the similarity between texts, calculating the similarity between text encoding values can improve the deduplication efficiency and save computational overhead.
[0152] Furthermore, in order to determine the texts to be deduplicated, the text encoding values and InnerIDs are transmitted between nodes. Therefore, the amount of data to be transmitted is small, which can further improve the deduplication efficiency.
[0153] Furthermore, the same class encoding sets are aggregated to the same node to determine the nodes to be deduplicated, which can further improve the deduplication efficiency.
[0154] Further, the text deduplication method is implemented based on PYTORCH, without the need to additionally build a target communication system. In addition, due to the performance of PYTORCH itself, it can improve communication efficiency, stability, and scalability, and further improve the deduplication efficiency.
[0155] Figure 17 It is a schematic structural diagram of an embodiment of the text deduplication device of the present application. As Figure 17 shown, the text deduplication device 70 includes an encoding module 71, a partitioning module 72, and a deduplication module 73.
[0156] The encoding module 71 is configured to perform a first encoding on each text in the current source computing node respectively to obtain the text encoding values of each text.
[0157] The partitioning module 72 is configured to partition the text encoding values of each text into several categories of encoding sets, where the similarity between the text encoding values within the same category of encoding set is higher than the similarity between different text encoding values in different categories of encoding sets.
[0158] The deduplication module 73 is configured to perform deduplication on the texts corresponding to the encoding sets in the current source computing node for each category of encoding sets, where the texts to be deduplicated are determined based on the similarity between the texts to be deduplicated and the text encoding values of other texts in the encoding sets where they are located.
[0159] For other detailed descriptions of the text deduplication device 70 in this embodiment, please refer to the previous embodiments and will not be elaborated here.
[0160] Figure 18 It is a schematic structural diagram of another embodiment of the text deduplication device of the present application. As Figure 18 shown, the text deduplication device 80 includes a receiving module 81, a combining module 82, an obtaining module 83, and a determining module 84.
[0161] The receiving module 81 is configured to receive at least one category of encoding sets sent by at least one source computing node, where the encoding sets are obtained by the source computing node partitioning the text encoding values of each text in the source computing node, and the text encoding values of each text are obtained by the source computing node encoding each text respectively.
[0162] The combining module 82 is configured to combine the same category of encoding sets sent by each source computing node; The obtaining module 83 is configured to obtain the similarity between the text encoding values in the combined encoding sets for each combined category of encoding sets.
[0163] The determining module 84 is configured to determine the texts to be deduplicated in the combined encoding sets based on the similarity, and notify the source computing node where the texts to be deduplicated are located to delete the texts to be deduplicated.
[0164] For other detailed descriptions of the text deduplication device 80 in this embodiment, please refer to the previous embodiments and will not be elaborated here.
[0165] Figure 19 It is a schematic structural diagram of an embodiment of an electronic device of the present application. As Figure 19 shown, the electronic device 90 includes a memory 91 and a processor 92. The processor 92 is used to execute program instructions stored in the memory 91 to implement the steps in any of the above method embodiments. In a specific implementation scenario, the electronic device 90 may include, but is not limited to: a microcomputer, a server. In addition, the electronic device 90 may also include carrier devices such as a laptop computer, a tablet computer, etc., which are not limited here.
[0166] Specifically, the processor 92 is used to control itself and the memory 91 to implement the steps in any of the above method embodiments. The processor 92 may also be referred to as a CPU (Central Processing Unit). The processor 92 may be an integrated circuit chip with signal processing capabilities. The processor 92 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 92 may be implemented jointly by integrated circuit chips.
[0167] Please refer to Figure 20 , Figure 20 It is a schematic structural diagram of an embodiment of a computer-readable storage medium of the present application. The computer-readable storage medium 100 has program instructions 101 stored thereon. When the program instructions 101 are executed by the processor, the steps in any of the above method embodiments are implemented.
[0168] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. Its specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be elaborated here.
[0169] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be elaborated in this article.
[0170] In several embodiments provided by the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. In another image position, the couplings or direct couplings or communication connections shown or discussed among each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0171] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
Claims
1. A text deduplication method, characterized in that: include: Encode each text in the current source computing node once respectively to obtain a text encoding value of each text; Dividing the text code values of each of the texts into several types of code sets, wherein the similarity between the text code values in the same type of code set is higher than the similarity between different text code values in different types of code sets; For each type of the coding set, the text corresponding to the coding set in the current source computing node is deduplicated, wherein the text to be deduplicated is determined based on the similarity between the text coding values of the text to be deduplicated and other texts in the coding set.
2. The method according to claim 1, characterized in that The step of encoding each text in the current source computing node to obtain a text encoding value of each text includes: For each of the texts, dividing the text into a plurality of text segments; Encoding each text segment of the text once respectively to obtain a coding subvalue of each text segment of the text; The text encoding value of the text is obtained by using the encoding sub-values of each text segment of the text.
3. The method according to claim 2, characterized in that The one-time encoding method is hash encoding; And / or, the obtaining of the text encoding value of the text by using the encoding sub-values of each text segment of the text comprises: The encoding sub-values of each text segment of the text and the position information of the text are combined to obtain the text encoding value of the text, wherein the position information of the text includes the first address information of the current source computing node where the text is located and the second address information of the text at the current source computing node.
4. The method according to any one of claims 1 to 3, characterized in that: The text encoding values of the texts are divided into several types of encoding sets, including: For each of the texts, the text code value of the text is re-encoded to obtain the category code value of the text; The text encoding values corresponding to the texts whose difference between the category encoding values is less than a preset difference are divided into the same encoding set.
5. The method according to claim 4, characterized in that The secondary encoding method is hash encoding; And / or, dividing the text code values corresponding to the texts whose difference between the category code values is less than a preset difference into the same code set includes: The text encoding values corresponding to the texts with the same category encoding values are divided into the same category of encoding sets.
6. The method according to claim 1, characterized in that The communication system where the current source computing node is located is a target communication system, a plurality of source computing nodes in the target communication system all contain a plurality of texts, and the current source computing node is one of the plurality of source computing nodes; For each type of the code set, deduplication of text corresponding to the code set in the current source computing node includes: At least one type of the coding set is used as a first type of coding set. For each of the first type of coding sets, the first type of coding sets in each of the source computing nodes are combined, and the similarity between the text coding values in the combined first type of coding set is calculated. The texts corresponding to the text coding values in the combined first type of coding set whose similarities meet the similarity requirements are used as a repeated text group, and at least one text is selected from the repeated text group as a first text to be deduplicated, the first text to be deduplicated located at the current source computing node is deleted, and the source computing node where the non-local text is located is notified to delete the non-local text, wherein the non-local text is the first text to be deduplicated that is not located at the current source computing node; and / or, At least one type of the coding set is used as a second type of coding set. For each of the second type of coding sets, the second type of coding set of the current source computing node is sent to the similarity computing node in the target communication system, so that the similarity computing node combines the second type of coding sets of each of the source computing nodes, and based on the similarity between the text coding values in the combined second type of coding set, determines the second text to be deduplicated corresponding to the combined second type of coding set, and notifies the source computing node where the second text to be deduplicated is located to delete the second text to be deduplicated.
7. The method according to claim 6, characterized in that Among the several types of code sets, one type of code set is used as the first type of code set, and the other types of code sets are used as the second type of code sets; And / or, the similarity calculation node is the source calculation node; And / or, notifying a source computing node where the non-local text is located to delete the non-local text, or notifying a source computing node where the second to-be-deduplicated text is located to delete the second to-be-deduplicated text, includes: Using the non-local text or the second to-be-deduplicated text as a target deduplicated text; Acquire, from the text encoding value of the target deduplicated text, first address information of a source computing node where the target deduplicated text is located and second address information of the target deduplicated text in the source computing node; According to the first address information, the second address information is sent to the source computing node where the target deduplicated text is located, so as to notify the source computing node to delete the text corresponding to the second address information.
8. The method according to claim 1, characterized in that: The text is a pre-trained text of a large model; and / or The text deduplication method is executed by several nodes under the large model pre-training system based on PYTORCH, and the several nodes include the current source computing node; the method also includes: When data needs to be sent to other nodes, the data to be sent is converted into a format supported by the PYTORCH-based large model pre-training system.
9. A text deduplication method, characterized in that: include: Receiving at least one type of encoding set sent by at least one source computing node, wherein the encoding set is obtained by dividing the text encoding values of each text in the source computing node by the source computing node, and the text encoding values of each text are obtained by encoding each text by the source computing node respectively; Combining the same type of code sets sent by each of the source computing nodes; For each type of the combined code set, obtaining the similarity between the text code values in the combined code set; Based on the similarity, the text to be deduplicated in the combined encoding set is determined, and the source computing node where the text to be deduplicated is located is notified to delete the text to be deduplicated.
10. The method according to claim 9, characterized in that The notifying the source computing node where the to-be-deduplicated text is located to delete the to-be-deduplicated text includes: From the text encoding value of the to-be-deduplicated text, obtain first address information of a source computing node where the to-be-deduplicated text is located and second address information of the to-be-deduplicated text in the source computing node; According to the first address information, the second address information is sent to the source computing node where the to-be-deduplicated text is located, so as to notify the source computing node to delete the text corresponding to the second address information.
11. A text deduplication device, characterized in that: include: The encoding module is used to encode each text in the current source computing node once to obtain the text encoding value of each text; A division module, used for dividing the text code values of each text into several types of code sets, wherein the similarity between the text code values in the same type of code set is higher than the similarity between different text code values in different types of code sets; A deduplication module is used to deduplicate the text corresponding to the coding set in the current source computing node for each type of coding set, wherein the text to be deduplicated is determined based on the similarity between the text coding value of the text to be deduplicated and other texts in the coding set.
12. A text deduplication device, characterized in that: include: A receiving module, used for receiving at least one type of coding set sent by at least one source computing node, wherein the coding set is obtained by dividing the text coding values of each text in the source computing node by the source computing node, and the text coding values of each text are obtained by encoding each text by the source computing node respectively; A combining module, used for combining the same type of code sets sent by each of the source computing nodes; An acquisition module, for acquiring, for each type of the combined code set, the similarity between the text code values in the combined code set; A determination module is used to determine the to-be-deduplicated text in the combined encoding set based on the similarity, and to notify the source computing node where the to-be-deduplicated text is located to delete the to-be-deduplicated text.
13. A target communication system, characterized in that: The method comprises a plurality of nodes, wherein the plurality of nodes comprises at least one source computing node and at least one similarity computing node, wherein the source computing node is used to execute the method according to any one of claims 1 to 8, and the similarity computing node is used to execute the method according to any one of claims 9 to 10.
14. An electronic device, characterized in that: The invention comprises a memory and a processor, wherein the processor is used to execute program instructions stored in the memory to implement the method according to any one of claims 1 to 10.
15. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Text duplicate removal method and device and storage medium
CN113688629A
Text deduplication method and device based on text modal self-supervision
CN115357690A
Text classification method, model training method, related device and electronic equipment
CN115658903A
Text duplicate removal method and device, electronic equipment and storage medium
CN116341566A
Similar text aggregation method and device, equipment and storage medium thereof
CN117093717A