Text Similarity Judgment Method, Device and Storage Medium

By extracting segment features and key words from texts to determine similarity, the method addresses adaptability and accuracy issues in text similarity judgment, enhancing precision in varied text structures.

CN114757299BActive Publication Date: 2025-07-15CHINA CONSTRUCTION BANK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210469090.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-07-15
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

The existing text similarity judgment methods are poorly adaptable, especially when text paragraph position changes and sentence pattern changes, the accuracy is low.

Method used

By extracting the paragraph characteristics of each paragraph of the text, the first similarity between paragraphs is determined, and when the first similarity is greater than the threshold, the keywords of the paragraph are further extracted, the second similarity of the paragraph is determined based on the keyword, and finally the text similarity is determined based on the first similarity and the second similarity.

Benefits of technology

It improves the accuracy of text similarity judgment, especially when text paragraph position conversion and sentence pattern transformation, it can more accurately judge text similarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114757299B_ABST
    Figure CN114757299B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus and storage medium for text similarity judgment, which relates to the technical field of data processing. The method extracts the paragraph features of each paragraph of two texts for text similarity judgment, and then, based on the paragraph features, determines the first similarity between each paragraph of the two texts. If the first similarity between paragraphs is greater than a threshold, keywords of each paragraph are further extracted, and based on the keywords, the second similarity of each paragraph is determined. Thus, according to the first similarity and the second similarity, the text similarity of the two texts is determined, solving the problem of poor adaptability of the existing text similarity judgment, such as improving the accuracy of text similarity judgment when the positions of text paragraphs are swapped or the sentence patterns are changed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a method, apparatus, and storage medium for judging text similarity. Background Art

[0002] With the rapid development of the Internet and the advent of the big data era, the amount of various text data has increased exponentially, and there are also various forms of reference, which puts forward higher requirements for the recognition of similar content and the accuracy of similarity judgment.

[0003] In related technologies, text similarity judgment refers to the measurement of the similarity between two texts, which has a wide range of applications in multiple fields. For example, in information retrieval, similarity can be used to identify similar words and improve the recall rate. Existing text similarity judgment usually analyzes the similarity by using the sentences in each paragraph of the text.

[0004] However, the adaptability of existing text similarity judgment is poor. For example, when the positions of text paragraphs change or the sentence patterns change, the accuracy of text similarity judgment is relatively low. Summary of the Invention

[0005] This application provides a method, apparatus, and storage medium for judging text similarity to solve the problems of poor adaptability of existing text similarity judgment and relatively low accuracy of text similarity judgment.

[0006] In a first aspect, an embodiment of this application provides a method for judging text similarity, including:

[0007] Determine a first text and a second text for which text similarity is to be judged, and extract the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text;

[0008] Based on the paragraph feature of paragraph i of the first text and the paragraph feature of paragraph j of the second text, determine a first similarity between paragraph i and paragraph j, where paragraph i is any paragraph in the first text, i = 1, 2,..., m, m is an integer, m is determined according to the total number of paragraphs in the first text, paragraph j is any paragraph in the second text, j = 1, 2,..., n, n is an integer, and n is determined according to the total number of paragraphs in the second text;

[0009] If the first similarity is greater than a first preset threshold, extract the keywords of paragraph i and paragraph j respectively;

[0010] Based on the keywords of paragraph i and paragraph j, determine a second similarity between paragraph i and paragraph j;

[0011] Determine the text similarity between the first text and the second text according to the first similarity and the second similarity.

[0012] In a possible implementation, the determining the first similarity between paragraph i of the first text and paragraph j of the second text based on the paragraph feature of paragraph i of the first text and the paragraph feature of paragraph j of the second text includes:

[0013] Perform word segmentation on the paragraph feature of paragraph i and the paragraph feature of paragraph j respectively to obtain a first cluster and a second cluster;

[0014] Calculate the intersection-over-union ratio of the first cluster and the second cluster, and determine the first similarity between paragraph i and paragraph j based on the intersection-over-union ratio.

[0015] In a possible implementation, before calculating the intersection-over-union ratio of the first cluster and the second cluster, it further includes:

[0016] Determine the non-intersecting words between paragraph i and paragraph j according to the first cluster and the second cluster;

[0017] The calculating the intersection-over-union ratio of the first cluster and the second cluster includes:

[0018] If the number of negative words in the non-intersecting words between paragraph i and paragraph j is even, calculate the intersection-over-union ratio of the first cluster and the second cluster.

[0019] In a possible implementation, before extracting the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text, it further includes:

[0020] Extract the text features of the first text and the text features of the second text respectively;

[0021] Determine the third similarity between the first text and the second text based on the text features of the first text and the text features of the second text;

[0022] The extracting the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text includes:

[0023] If the third similarity is greater than a second preset threshold, extract the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text.

[0024] In a possible implementation, the determining the text similarity between the first text and the second text according to the first similarity and the second similarity includes:

[0025] Determine the text similarity between the first text and the second text according to the first similarity, the second similarity, and the third similarity.

[0026] In a possible implementation, the determining the text similarity between the first text and the second text according to the first similarity, the second similarity, and the third similarity includes:

[0027] According to the first text and the second text, respectively determine the first coefficient corresponding to the first similarity, the second coefficient corresponding to the second similarity, and the third coefficient corresponding to the third similarity;

[0028] Based on the first similarity, the first coefficient, the second similarity, the second coefficient, the third similarity, and the third coefficient, obtain the paragraph similarity between the first text and the second text;

[0029] According to the paragraph similarity between the first text and the second text, determine the text similarity between the first text and the second text.

[0030] In a possible implementation, the determining the text similarity between the first text and the second text according to the paragraph similarity between the first text and the second text includes:

[0031] According to each paragraph of the first text and each paragraph of the second text, determine the paragraph weight corresponding to the paragraph similarity;

[0032] Based on the paragraph similarity and the paragraph weight corresponding to the paragraph similarity, determine the text similarity between the first text and the second text.

[0033] In a possible implementation, before extracting the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text, it further includes:

[0034] Judge whether the type of the first text is a preset text type, and judge whether the type of the second text is the preset text type;

[0035] The extracting the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text includes:

[0036] If the type of the first text is the preset text type and the type of the second text is the preset text type, then extract the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text.

[0037] Second aspect, an embodiment of the present application provides a text similarity judgment device, including:

[0038] A first feature extraction module, configured to determine a first text and a second text for text similarity judgment, and extract paragraph features of each paragraph of the first text and paragraph features of each paragraph of the second text;

[0039] A first similarity determination module, configured to determine a first similarity between paragraph i of the first text and paragraph j of the second text based on the paragraph feature of paragraph i of the first text and the paragraph feature of paragraph j of the second text, where paragraph i is any paragraph in the first text, i = 1, 2,..., m, m is an integer, m is determined according to the total number of paragraphs in the first text, paragraph j is any paragraph in the second text, j = 1, 2,..., n, n is an integer, n is determined according to the total number of paragraphs in the second text;

[0040] A second feature extraction module, configured to extract keywords of paragraph i and paragraph j respectively if the first similarity is greater than a first preset threshold;

[0041] A second similarity determination module, configured to determine a second similarity between paragraph i and paragraph j based on the keywords of paragraph i and paragraph j;

[0042] A text similarity judgment module, configured to determine the text similarity between the first text and the second text according to the first similarity and the second similarity.

[0043] In a possible implementation manner, the first similarity determination module is specifically configured to:

[0044] Perform word segmentation processing on the paragraph feature of paragraph i and the paragraph feature of paragraph j respectively to obtain a first cluster and a second cluster;

[0045] Calculate the intersection-union ratio of the first cluster and the second cluster, and determine the first similarity between paragraph i and paragraph j based on the intersection-union ratio.

[0046] In a possible implementation manner, the first similarity determination module is specifically configured to:

[0047] Determine the non-intersection vocabulary between paragraph i and paragraph j according to the first cluster and the second cluster;

[0048] If the number of negative words in the non-intersection vocabulary between paragraph i and paragraph j is an even number, calculate the intersection-union ratio of the first cluster and the second cluster.

[0049] In a possible implementation manner, the first feature extraction module is specifically configured to:

[0050] Extract the text features of the first text and the text features of the second text respectively;

[0051] Based on the text features of the first text and the text features of the second text, determine the third similarity between the first text and the second text;

[0052] If the third similarity is greater than the second preset threshold, extract the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text.

[0053] In a possible implementation manner, the text similarity judgment module is specifically configured to:

[0054] Determine the text similarity between the first text and the second text according to the first similarity, the second similarity, and the third similarity.

[0055] In a possible implementation manner, the text similarity judgment module is specifically configured to:

[0056] According to the first text and the second text, respectively determine the first coefficient corresponding to the first similarity, the second coefficient corresponding to the second similarity, and the third coefficient corresponding to the third similarity;

[0057] Based on the first similarity, the first coefficient, the second similarity, the second coefficient, the third similarity, and the third coefficient, obtain the paragraph similarity between the first text and the second text;

[0058] Determine the text similarity between the first text and the second text according to the paragraph similarity between the first text and the second text.

[0059] In a possible implementation manner, the text similarity judgment module is specifically configured to:

[0060] According to each paragraph of the first text and each paragraph of the second text, determine the paragraph weight corresponding to the paragraph similarity;

[0061] Based on the paragraph similarity and the paragraph weight corresponding to the paragraph similarity, determine the text similarity between the first text and the second text.

[0062] In a possible implementation manner, the first feature extraction module is specifically configured to:

[0063] Determine whether the type of the first text is a preset text type, and determine whether the type of the second text is the preset text type;

[0064] If the type of the first text is the preset text type and the type of the second text is the preset text type, then extract the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text.

[0065] In a third aspect, an embodiment of the present application provides a text similarity judgment device, including:

[0066] A processor;

[0067] A memory; and

[0068] A computer program;

[0069] Wherein, the computer program is stored in the memory and is configured to be executed by the processor, and the computer program includes instructions for executing the method described in the first aspect.

[0070] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and the computer program causes the server to execute the method described in the first aspect.

[0071] In a fifth aspect, an embodiment of the present application provides a computer program product, including computer instructions, and the computer instructions are executed by the processor to execute the method described in the first aspect.

[0072] The text similarity judgment method, device and storage medium provided by the embodiments of the present application, this method extracts the paragraph features of each paragraph of two texts for text similarity judgment, and then, based on the paragraph features, determines the first similarity between each paragraph of the two texts. If the first similarity between paragraphs is greater than the threshold, further extract the keywords of each paragraph, and based on the keywords, determine the second similarity of each paragraph. Thus, according to the first similarity and the second similarity, determine the text similarity of the two texts, and solve the problem of poor adaptability of the existing text similarity judgment, such as improving the accuracy of text similarity judgment when the positions of text paragraphs are swapped and the sentence patterns are transformed. Description of the Drawings

[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0074] Figure 1 Schematic diagram of the text similarity judgment system architecture provided by the embodiments of the present application;

[0075] Figure 2 Flow schematic diagram of a text similarity judgment method provided by the embodiments of the present application;

[0076] Figure 3 Flow schematic diagram of another text similarity judgment method provided by the embodiments of the present application;

[0077] Figure 4 Flow schematic diagram of yet another text similarity judgment method provided by the embodiments of the present application;

[0078] Figure 5 Schematic diagram of the structure of a text similarity judgment device provided by the embodiments of the present application;

[0079] Figure 6 Schematic diagram of the basic hardware architecture of a text similarity judgment device provided by the embodiments of the present application. Detailed implementation manners

[0080] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0081] In the specification, claims and above-mentioned accompanying drawings of the present application, terms such as "first", "second", "third" and "fourth" (if any) are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0082] In addition, in the technical solutions of the present application, the collection, storage, use, processing, transmission, provision and disclosure of information such as financial data or user data comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0083] In related technologies, text similarity judgment refers to the measurement of the similarity between two texts, which has wide applications in multiple fields. For example, in information retrieval, similarity can be used to identify similar words to improve the recall rate; in the scenario of automatic question answering, similarity can be used to calculate the matching degree between the questions asked by users in natural language and the questions in the corpus, and return the answer corresponding to the question with the highest matching degree as the response, etc. Existing text similarity judgment usually analyzes the similarity by using the sentences in each paragraph of the text. However, the adaptability of existing text similarity judgment is poor. For example, when the positions of text paragraphs change or the sentence patterns are transformed, the accuracy of text similarity judgment is relatively low.

[0084] To solve the above problems, the embodiments of the present application propose a text similarity judgment method. By extracting the paragraph features of each paragraph of two texts, determining the first similarity between each paragraph of the two texts, and based on the first similarity between each paragraph, further extracting the keywords of each paragraph, and based on these keywords, determining the second similarity of each paragraph. Thus, according to the first similarity and the second similarity, the text similarity of the two texts is determined, solving the problem of poor adaptability of existing text similarity judgment and improving the accuracy of text similarity judgment.

[0085] Optionally, a text similarity judgment method provided by the present application can be applicable to Figure 1 the schematic diagram of the text similarity judgment system architecture shown in Figure 1 As shown in

[0086] In the specific implementation process, the receiving device 101 can be an input / output interface or a communication interface, and can be used to receive texts for text similarity judgment, etc.

[0087] The processing device 102 can obtain the texts for text similarity judgment through the above-mentioned receiving device 101. Furthermore, it extracts the paragraph features of each paragraph of the text, determines the first similarity between each paragraph of the text based on the paragraph features, and based on the first similarity between each paragraph, extracts the keywords of each paragraph, and based on these keywords, determines the second similarity of each paragraph. Thus, according to the first similarity and the second similarity, the text similarity between the texts is determined, improving the accuracy of text similarity judgment.

[0088] The display device 103 can be used to display the above-mentioned first similarity, second similarity, text similarity, etc.

[0089] The display device can also be a touch display screen, which is used to receive user instructions while displaying the above-mentioned content to realize interaction with users.

[0090] It should be understood that the above processing device can be implemented by a processor reading instructions in a memory and executing the instructions, or can be implemented by chip circuits.

[0091] The above system is only an exemplary system, and can be set according to application requirements during specific implementation.

[0092] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the architecture of the text similarity judgment system. In other feasible embodiments of the present application, the above architecture may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or different component arrangements, which can be specifically determined according to the actual application scenario and are not limited herein. Figure 1 The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0093] In addition, the system architecture described in the embodiments of the present application is to more clearly illustrate the technical solutions of the embodiments of the present application, and does not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0094] The following takes several embodiments as examples to describe the technical solutions of the present application, and the same or similar concepts or processes may not be repeated in some embodiments.

[0095] Figure 2 FIG. is a schematic flow chart of a text similarity judgment method provided by an embodiment of the present application. The execution subject of this embodiment can be Figure 1 the processing device in, and the specific execution subject can be determined according to the actual application scenario, and the embodiments of the present application do not make special limitations on this. As Figure 2 shown, the text similarity judgment method provided by the embodiments of the present application may include the following steps:

[0096] S201: Determine a first text and a second text for which text similarity is to be judged, and extract the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text.

[0097] Here, after the above processing device determines the first text and the second text for which text similarity is to be judged, it can first judge whether the type of the first text is a preset text type, and judge whether the type of the second text is the above preset text type. If the type of the first text is the above preset text type, and the type of the second text is the above preset text type, then extract the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text.

[0098] Among them, the above-mentioned preset text type can be determined according to the actual situation, such as files of the txt type. In the embodiments of the present application, the above-mentioned processing device first determines the type of the text for which text similarity is to be judged. If it is a preset text type, the subsequent steps are directly executed. If it is not a preset text type, it is converted into a preset text type and then the subsequent steps are executed to facilitate subsequent text processing.

[0099] Optionally, when the above-mentioned processing device extracts the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text, a pre-trained paragraph feature extraction model can be used for extraction. Here, the above-mentioned paragraph feature extraction model is used to extract the paragraph features of each paragraph of the text. The paragraph feature can be understood as a paragraph summary, the main content of the paragraph, etc.

[0100] S202: Based on the paragraph feature of paragraph i of the first text and the paragraph feature of paragraph j of the second text, determine the first similarity between paragraph i and paragraph j, where paragraph i is any paragraph in the first text, i = 1, 2,..., m, m is an integer, m is determined according to the total number of paragraphs in the first text, paragraph j is any paragraph in the second text, j = 1, 2,..., n, n is an integer, and n is determined according to the total number of paragraphs in the second text.

[0101] Exemplarily, the above-mentioned processing device can perform word segmentation processing on the paragraph feature of paragraph i and the paragraph feature of paragraph j respectively to obtain a first cluster and a second cluster. Furthermore, calculate the intersection ratio of the first cluster and the second cluster, and based on this intersection ratio, determine the first similarity between paragraph i and paragraph j.

[0102] For example, the above-mentioned first cluster and second cluster are respectively denoted as E1 and E2. When the above-mentioned processing device calculates the intersection ratio of the first cluster and the second cluster, it can calculate the intersection of E1 and E2, denoted as E ∩ , and calculate the union of E1 and E2, denoted as E ∪ , and then calculate the intersection ratio of the two clusters sim = E ∩ / E ∪ .

[0103] Assume that E1 = {A1, A2, A3,...} and E2 = {B1, B2, B3,...}, the total number of words in E1 is Count A , the total number of words in E2 is Count B , calculate the intersection of E1 and E2, that is, if a certain word exists in both sets, then the word is added to the intersection, denoted as E ∩ , and at the same time, the number of words in the intersection increases by 1. The union refers to the number of words remaining after removing duplicates in the two sets of E1 and E2, denoted as E ∪, the intersection over union (IoU) calculates the ratio of the number of words in the intersection to the number of words in the union.

[0104] For example, E1 = {"friend", "Xiaoming", "Xiaohong", "become"},

[0105] E2 = {"friend", "Xiaoming", "Xiaohua", "is"}, then

[0106] E ∪ = {"Xiaoming", "Xiaohong", "Xiaohua", "become", "is", "friend"}, the number of words in the union is 6, and E ∩ = {"friend", "Xiaoming"}, the number of words in the intersection is 2, and the IoU of the two clusters is 2 / 6 = 1 / 3.

[0107] After calculating the above IoU, the above processing device can use the above IoU as the first similarity between paragraph i and paragraph j.

[0108] Among them, before calculating the IoU of the above first cluster and the second cluster, the above processing device also considers first determining the non-intersecting words of paragraph i and paragraph j according to the above first cluster and the second cluster. If the number of negative words in the non-intersecting words of paragraph i and paragraph j is even, then calculate the IoU of the above first cluster and the second cluster.

[0109] Here, the above processing device considers the situation of negative words appearing in the non-intersecting words between clusters. If the total number of negative words in the non-intersecting words of paragraph i and paragraph j is even, it indicates that the possibility of paragraph i and paragraph j being similar is very high. Then, further calculate the IoU of the above first cluster and the second cluster, and based on this IoU, determine the first similarity between paragraph i and paragraph j. Thus, subsequent text similarity judgment is performed according to this first similarity.

[0110] S203: If the above first similarity is greater than the first preset threshold, then extract the keywords of paragraph i and paragraph j respectively.

[0111] Here, if the above first similarity is greater than the first preset threshold, it indicates that the possibility of paragraph i and paragraph j being similar is very high. To further improve the accuracy of subsequent text similarity judgment, when the above first similarity is greater than the first preset threshold, the above processing device further extracts the keywords of paragraph i and paragraph j. Among them, the above first preset threshold can be determined according to the actual situation, such as 1 / 2.

[0112] Optionally, when the above processing device extracts the keywords of paragraph i and paragraph j, it can use a pre-trained keyword extraction model for extraction. Here, the above keyword extraction model is used to extract the keywords of each paragraph of the text. The keyword can be determined according to the actual situation, such as a word with a relatively high correlation with the paragraph content.

[0113] S204: Determine the second similarity between paragraph i and paragraph j based on the keywords in paragraph i and paragraph j.

[0114] Exemplarily, the above processing device can perform word segmentation on the keywords of paragraph i and the keywords of paragraph j respectively to obtain a third cluster and a fourth cluster. Further, calculate the intersection-to-union ratio of the third cluster and the fourth cluster, and based on this intersection-to-union ratio, determine the second similarity between paragraph i and paragraph j. For example, the above processing device takes the above intersection-to-union ratio as the second similarity between paragraph i and paragraph j.

[0115] Wherein, before calculating the intersection-to-union ratio of the above third cluster and fourth cluster, the above processing device can also determine the keyword non-intersection vocabulary between paragraph i and paragraph j according to the third cluster and the fourth cluster. If the number of negative vocabulary in the non-intersection vocabulary of paragraph i and paragraph j is even, then calculate the intersection-to-union ratio of the third cluster and the fourth cluster.

[0116] Here, the above processing device considers the situation where negative words appear in the non-intersection vocabulary between clusters. If the total number of negative vocabulary in the non-intersection vocabulary between paragraph i and paragraph j is even, it indicates that the possibility of similarity between paragraph i and paragraph j is very high. Here, further calculate the intersection-to-union ratio of the above third cluster and fourth cluster, and based on this intersection-to-union ratio, determine the second similarity between paragraph i and paragraph j. Thus, subsequent text similarity judgment is performed according to this second similarity.

[0117] S205: Determine the text similarity between the first text and the second text according to the above first similarity and second similarity.

[0118] In the embodiments of the present application, the above processing device can respectively determine a first coefficient corresponding to the above first similarity and a second coefficient corresponding to the above second similarity according to the above first text and second text. Further, based on the above first similarity, first coefficient, second similarity, and second coefficient, obtain the paragraph similarity between the above first text and the second text. Thus, according to this paragraph similarity, determine the text similarity between the above first text and the second text. For example, the above processing device multiplies the above first similarity, first coefficient, second similarity, and second coefficient, and obtains the paragraph similarity between the above first text and the second text based on the multiplication result.

[0119] Wherein, the above processing device can determine the paragraph weight corresponding to the above paragraph similarity according to each paragraph of the above first text and each paragraph of the second text, and then, based on the above paragraph similarity and the above paragraph weight, determine the text similarity between the above first text and the second text. For example, the above processing device multiplies the above paragraph similarity by the above paragraph weight, and determines the text similarity between the above first text and the second text based on the multiplication result.

[0120] In the embodiment of the present application, by extracting the paragraph features of each paragraph of two texts for text similarity judgment, and then, based on the paragraph features, determining the first similarity between each paragraph of the two texts. If the first similarity between paragraphs is greater than a threshold, keywords of each paragraph are further extracted, and based on the keywords, the second similarity of each paragraph is determined. Thus, according to the first similarity and the second similarity, the text similarity of the two texts is determined, solving the problem of poor adaptability of the existing text similarity judgment, such as improving the accuracy of text similarity judgment when the positions of text paragraphs are swapped or the sentence patterns are changed.

[0121] In addition, before the above processing device extracts the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text, it also considers first extracting the text features of the first text and the text features of the second text respectively, and based on the text features, determining the third similarity between the first text and the second text. When the third similarity is greater than a second preset threshold, the steps of extracting the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text are further executed, further improving the accuracy of subsequent text similarity judgment. Moreover, when determining the text similarity between the first text and the second text subsequently, in addition to considering the first similarity and the second similarity, the above processing device can also consider the third similarity, that is, according to the first similarity, the second similarity and the third similarity, determining the text similarity between the first text and the second text, which also improves the accuracy of text similarity judgment. Figure 3 The flow diagram of another text similarity judgment method provided by the embodiment of the present application is as Figure 3 shown, and the method may include:

[0122] S301: Determine a first text and a second text for text similarity judgment, and extract the text features of the first text and the text features of the second text respectively.

[0123] Here, the above processing device may use a pre-trained text feature extraction model to extract the text features of the first text and the text features of the second text. Among them, the above text feature extraction model is used to extract the paragraph features of the text. The text features can be understood as text summaries, main contents of the text, etc.

[0124] S302: Based on the text features of the first text and the text features of the second text, determine the third similarity between the first text and the second text.

[0125] Exemplarily, the above processing device may perform word segmentation processing on the text features of the first text and the text features of the second text respectively to obtain a fifth cluster and a sixth cluster. Then, calculate the intersection-union ratio of the fifth cluster and the sixth cluster, and based on the intersection-union ratio, determine the third similarity between the first text and the second text.

[0126] Among them, before calculating the intersection-over-union ratio of the above-mentioned fifth cluster and sixth cluster, the above-mentioned processing device also first determines the non-intersecting words of the full text of the first text and the second text according to the above-mentioned fifth cluster and sixth cluster. If the number of negative words in the non-intersecting words of the full text of the first text and the second text is even, then calculate the intersection-over-union ratio of the above-mentioned fifth cluster and sixth cluster.

[0127] Here, the above-mentioned processing device considers the situation where negative words appear in the non-intersecting words between clusters. If the total number of negative words in the non-intersecting words of the first text and the second text is even, it indicates that the possibility of similarity between the first text and the second text is very high. Here, further calculate the intersection-over-union ratio of the above-mentioned fifth cluster and sixth cluster, and based on this intersection-over-union ratio, determine the third similarity between the first text and the second text. Thus, subsequent text similarity judgment is performed according to this third similarity.

[0128] S303: If the above-mentioned third similarity is greater than the second preset threshold, extract the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text.

[0129] Here, if the above-mentioned third similarity is greater than the second preset threshold, it indicates that the possibility of similarity between the first text and the second text is very high. To further improve the accuracy of subsequent text similarity judgment, when the above-mentioned third similarity is greater than the second preset threshold, the above-mentioned processing device further extracts the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text.

[0130] S304: Based on the paragraph feature of paragraph i of the first text and the paragraph feature of paragraph j of the second text, determine the first similarity between paragraph i and paragraph j, where paragraph i is any paragraph in the first text, i = 1, 2, …, m, m is an integer, m is determined according to the total number of paragraphs in the first text, paragraph j is any paragraph in the second text, j = 1, 2, …, n, n is an integer, n is determined according to the total number of paragraphs in the second text.

[0131] S305: If the above-mentioned first similarity is greater than the first preset threshold, extract the keywords of paragraph i and paragraph j respectively.

[0132] S306: Based on the keywords of paragraph i and paragraph j, determine the second similarity between paragraph i and paragraph j.

[0133] Among them, the implementation manners of steps S304 - S306 are the same as those of the above-mentioned steps S202 - S204, and will not be elaborated here.

[0134] S307: Determine the text similarity between the first text and the second text according to the above-mentioned first similarity, second similarity and third similarity.

[0135] Exemplarily, the above processing device may respectively determine a first coefficient corresponding to a first similarity, a second coefficient corresponding to a second similarity, and a third coefficient corresponding to a third similarity according to the above first text and second text. Furthermore, based on the above first similarity, first coefficient, second similarity, second coefficient, third similarity, and third coefficient, the paragraph similarity between the first text and the second text is obtained, and according to this paragraph similarity, the text similarity between the first text and the second text is determined. For example, the above processing device multiplies the above first similarity, first coefficient, second similarity, second coefficient, third similarity, and third coefficient, and obtains the paragraph similarity between the first text and the second text based on the multiplication result.

[0136] Here, the above first coefficient may be determined according to the number of negative words in the non-intersecting words corresponding to the above paragraph, the above second coefficient may be determined according to the number of negative words in the keyword non-intersecting words corresponding to the above paragraph, and the above third coefficient may be determined according to the number of negative words in the non-intersecting words of the above full text.

[0137] For example, the paragraph similarity between the above first text and the second text may be determined according to the expression:

[0138] S 段落 = S 第三相似度 *(2 - 2 全文非交集否定词%2 ) * S 第一相似度 *(2 - 2 段落非交集否定词%2 )

[0139] * S 第二相似度 *(2 - 2 段落关键词非交集否定词%2 )

[0140] It is determined that, where S 段落 represents the above paragraph similarity, S 第一相似度 represents the above first similarity, 2 - 2 段落非交集否定词%2 represents the above first coefficient, S 第二相似度 represents the above second similarity, 2 - 2 段落关键词非交集否定词%2 represents the above second coefficient, S 第三相似度 represents the above third similarity, 2 - 2 全文非交集否定词%2 represents the above third coefficient.

[0141] Optionally, the above processing device may determine the paragraph weight corresponding to the above paragraph similarity according to each paragraph of the above first text and each paragraph of the second text. Thus, based on the above paragraph similarity and the above paragraph weight, the text similarity between the first text and the second text is determined. For example, the above processing device multiplies the above paragraph similarity by the above paragraph weight, and determines the text similarity between the first text and the second text based on the multiplication result.

[0142] Before extracting the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text in the embodiments of the present application, it is also considered to first extract the text features of the first text and the text features of the second text respectively. Based on the text features, the third similarity between the first text and the second text is determined. When the third similarity is greater than the second preset threshold, the steps of extracting the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text are further executed, which further improves the accuracy of subsequent text similarity judgment. Moreover, when determining the text similarity between the first text and the second text subsequently, in addition to considering the first similarity and the second similarity, the embodiments of the present application can also consider the third similarity, that is, according to the first similarity, the second similarity and the third similarity, the text similarity between the first text and the second text is determined, which also improves the accuracy of text similarity judgment and solves the problem of poor adaptability of existing text similarity judgment.

[0143] Here, as Figure 4 shown, Figure 4 Another text similarity judgment schematic diagram is given, in which the content identical or similar to Figure 2 and Figure 3 the embodiments is referred to the above, and will not be repeated here. As Figure 4As shown, after the above processing device determines the first text and the second text for text similarity judgment, it first determines whether the type of the first text is a preset text type and whether the type of the second text is the preset text type. If the type of the first text is the above preset text type and the type of the second text is the above preset text type, it directly executes the subsequent steps; otherwise, it performs text preprocessing to convert the type of the text that does not meet the requirements into the above preset text type. Then, the above processing device can respectively extract the text features of the first text and the second text, and based on the text features of the first text and the text features of the second text, determine the third similarity between the first text and the second text. If the third similarity is greater than the second preset threshold, it further extracts the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text. Furthermore, the above processing device determines the first similarity between paragraph i and paragraph j based on the paragraph feature of paragraph i in the first text and the paragraph feature of paragraph j in the second text, where paragraph i is any paragraph in the first text, i = 1, 2, …, m, m is an integer, m is determined according to the total number of paragraphs in the first text, paragraph j is any paragraph in the second text, j = 1, 2, …, n, n is an integer, n is determined according to the total number of paragraphs in the second text. If the first similarity is greater than the first preset threshold, the above processing device can further extract the keywords of paragraph i and paragraph j respectively, and based on the keywords of paragraph i and paragraph j, determine the second similarity between paragraph i and paragraph j. According to the above first similarity, second similarity and third similarity, the text similarity between the first text and the second text is determined.

[0144] Among them, taking the determination of the above first similarity as an example, when the above processing device determines the first similarity between paragraph i and paragraph j, it can perform word segmentation processing on the paragraph feature of paragraph i and the paragraph feature of paragraph j respectively to obtain the first cluster and the second cluster. Furthermore, it calculates the intersection-union ratio of the first cluster and the second cluster, and based on this intersection-union ratio, determines the first similarity between paragraph i and paragraph j.

[0145] Here, before the above processing device calculates the intersection-union ratio of the first cluster and the second cluster, it can also determine the non-intersection vocabulary between paragraph i and paragraph j according to the first cluster and the second cluster. If the number of negative vocabulary in the non-intersection vocabulary between paragraph i and paragraph j is even, it calculates the intersection-union ratio of the first cluster and the second cluster.

[0146] In summary, compared with the prior art, in the embodiment of the present application, by extracting the text features of the above first text and the text features of the second text, based on the text features, determining the third similarity between the first text and the second text, based on the third similarity, further extracting the paragraph features of each paragraph of the two texts, determining the first similarity between each paragraph of the two texts, and based on the first similarity between each paragraph, extracting the keywords of each paragraph, and based on the keywords, determining the second similarity of each paragraph, thereby, according to the first similarity, the second similarity and the third similarity, determining the text similarity of the two texts, solving the problem of poor adaptability of the existing text similarity judgment, such as when the positions of text paragraphs are swapped, the sentence patterns are changed, and negative words appear in non-intersecting words, improving the accuracy of text similarity judgment.

[0147] Corresponding to the text similarity judgment method in the above embodiment, Figure 5 It is a schematic structural diagram of a text similarity judgment device provided in an embodiment of the present application. For the convenience of description, only the parts related to the embodiment of the present application are shown. Figure 5 It is a schematic structural diagram of a text similarity judgment device provided in an embodiment of the present application. The text similarity judgment device 50 includes: a first feature extraction module 501, a first similarity determination module 502, a second feature extraction module 503, a second similarity determination module 504, and a text similarity judgment module 505. Here, the text similarity judgment device may be the above processing device itself, or a chip or integrated circuit that implements the functions of the processing device. It should be noted here that the division of the first feature extraction module, the first similarity determination module, the second feature extraction module, the second similarity determination module, and the text similarity judgment module is only a logical function division, and physically the two may be integrated or independent.

[0148] Among them, the first feature extraction module 501 is used to determine the first text and the second text for text similarity judgment, and extract the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text.

[0149] The first similarity determination module 502 is used to determine the first similarity between paragraph i of the first text and paragraph j of the second text based on the paragraph feature of paragraph i of the first text and the paragraph feature of paragraph j of the second text, where paragraph i is any paragraph in the first text, i = 1, 2,..., m, m is an integer, m is determined according to the total number of paragraphs in the first text, paragraph j is any paragraph in the second text, j = 1, 2,..., n, n is an integer, and n is determined according to the total number of paragraphs in the second text.

[0150] The second feature extraction module 503 is configured to extract the keywords of the paragraph i and the paragraph j respectively if the first similarity is greater than a first preset threshold.

[0151] The second similarity determination module 504 is configured to determine a second similarity between the paragraph i and the paragraph j based on the keywords of the paragraph i and the paragraph j.

[0152] The text similarity judgment module 505 is configured to determine the text similarity between the first text and the second text according to the first similarity and the second similarity.

[0153] In a possible implementation manner, the first similarity determination module 502 is specifically configured to:

[0154] Perform word segmentation processing on the paragraph features of the paragraph i and the paragraph features of the paragraph j respectively to obtain a first cluster and a second cluster;

[0155] Calculate the intersection-over-union ratio of the first cluster and the second cluster, and determine the first similarity between the paragraph i and the paragraph j based on the intersection-over-union ratio.

[0156] In a possible implementation manner, the first similarity determination module 502 is specifically configured to:

[0157] Determine the non-intersecting words between the paragraph i and the paragraph j according to the first cluster and the second cluster;

[0158] If the number of negative words in the non-intersecting words between the paragraph i and the paragraph j is even, calculate the intersection-over-union ratio of the first cluster and the second cluster.

[0159] In a possible implementation manner, the first feature extraction module 501 is specifically configured to:

[0160] Extract the text features of the first text and the text features of the second text respectively;

[0161] Determine a third similarity between the first text and the second text based on the text features of the first text and the text features of the second text;

[0162] If the third similarity is greater than a second preset threshold, extract the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text.

[0163] In a possible implementation manner, the text similarity judgment module 505 is specifically configured to:

[0164] Determine the text similarity between the first text and the second text according to the first similarity, the second similarity, and the third similarity.

[0165] In a possible implementation, the text similarity determination module 505 is specifically configured to:

[0166] According to the first text and the second text, respectively determine the first coefficient corresponding to the first similarity, the second coefficient corresponding to the second similarity, and the third coefficient corresponding to the third similarity;

[0167] Based on the first similarity, the first coefficient, the second similarity, the second coefficient, the third similarity, and the third coefficient, obtain the paragraph similarity between the first text and the second text;

[0168] According to the paragraph similarity between the first text and the second text, determine the text similarity between the first text and the second text.

[0169] In a possible implementation, the text similarity determination module 505 is specifically configured to:

[0170] According to each paragraph of the first text and each paragraph of the second text, determine the paragraph weight corresponding to the paragraph similarity;

[0171] Based on the paragraph similarity and the paragraph weight corresponding to the paragraph similarity, determine the text similarity between the first text and the second text.

[0172] In a possible implementation, the first feature extraction module 501 is specifically configured to:

[0173] Determine whether the type of the first text is a preset text type, and determine whether the type of the second text is the preset text type;

[0174] If the type of the first text is the preset text type and the type of the second text is the preset text type, then extract the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text.

[0175] The device provided in the embodiments of the present application can be used to execute the technical solutions of the above method embodiments, and its implementation principles and technical effects are similar, which will not be elaborated here in the embodiments of the present application.

[0176] Optionally, Figure 6 Schematically provide a possible basic hardware architecture diagram of the text similarity determination device of the present application.

[0177] See Figure 6, the text similarity judgment device includes at least one processor 601 and a communication interface 603. Further optionally, it may also include a memory 602 and a bus 604.

[0178] Among them, in the text similarity judgment device, the number of processors 601 can be one or more, Figure 6 only one processor 601 is shown for illustration. Optionally, the processor 601 can be a central processing unit (CPU), a graphics processing unit (GPU), or a digital signal processor (DSP). If the text similarity judgment device has multiple processors 601, the types of the multiple processors 601 can be different or the same. Optionally, the multiple processors 601 of the text similarity judgment device can also be integrated into a multi-core processor.

[0179] The memory 602 stores computer instructions and data; the memory 602 can store the computer instructions and data required to implement the above-mentioned text similarity judgment method provided by this application. For example, the memory 602 stores instructions for implementing the steps of the above-mentioned text similarity judgment method. The memory 602 can be any one or any combination of the following storage media: non-volatile memory (such as Read-Only Memory (ROM), Solid State Disk (SSD), Hard Disk Drive (HDD), optical disc), volatile memory.

[0180] The communication interface 603 can provide information input / output for the at least one processor. It can also include any one or any combination of the following devices: network interfaces (such as Ethernet interfaces), devices with network access functions such as wireless network cards.

[0181] Optionally, the communication interface 603 can also be used for data communication between the text similarity judgment device and other computing devices or terminals.

[0182] Further optionally, Figure 6 a thick line is used to represent the bus 604. The bus 604 can connect the processor 601 to the memory 602 and the communication interface 603. In this way, through the bus 604, the processor 601 can access the memory 602 and can also perform data interaction with other computing devices or terminals using the communication interface 603.

[0183] In this application, the text similarity judgment device executes the computer instructions in the memory 602, enabling the text similarity judgment device to implement the above-mentioned text similarity judgment method provided by this application, or enabling the text similarity judgment device to deploy the above-mentioned text similarity judgment device.

[0184] From a logical function division perspective, for example, as Figure 6 shown, the memory 602 may include a first feature extraction module 501, a first similarity determination module 502, a second feature extraction module 503, a second similarity determination module 504, and a text similarity judgment module 505. The inclusion here only involves that when the instructions stored in the memory are executed, they can respectively implement the functions of the first feature extraction module, the first similarity determination module, the second feature extraction module, the second similarity determination module, and the text similarity judgment module, rather than being limited to a physical structure.

[0185] This application provides a computer-readable storage medium, and the computer program product includes computer instructions, and the computer instructions direct a computing device to execute the above-mentioned text similarity judgment method provided by this application.

[0186] This application provides a computer program product, including computer instructions, and the computer instructions are executed by a processor to perform the above-mentioned text similarity judgment method.

[0187] This application provides a chip, including at least one processor and a communication interface, and the communication interface provides information input and / or output for the at least one processor. Further, the chip may further include at least one memory for storing computer instructions. The at least one processor is used to call and run the computer instructions to execute the above-mentioned text similarity judgment method provided by this application.

[0188] In several embodiments provided by this application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the device or unit may be in an electrical, mechanical or other form.

[0189] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0190] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.

Claims

1. A method for judging text similarity, characterized in that, Including: Determine a first text and a second text for which text similarity is to be judged; Extract the text features of the first text and the text features of the second text respectively; Based on the text features of the first text and the text features of the second text, determine a third similarity between the first text and the second text; If the third similarity is greater than a second preset threshold, then extract the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text; Based on the paragraph features of paragraph i of the first text and the paragraph features of paragraph j of the second text, determine a first similarity between paragraph i and paragraph j, where paragraph i is any paragraph in the first text, i = 1, 2, …, m, m is an integer, m is determined according to the total number of paragraphs in the first text, paragraph j is any paragraph in the second text, j = 1, 2, …, n, n is an integer, n is determined according to the total number of paragraphs in the second text; If the first similarity is greater than a first preset threshold, then extract the keywords of paragraph i and paragraph j respectively; Based on the keywords of paragraph i and paragraph j, determine a second similarity between paragraph i and paragraph j; According to the first text and the second text, determine a first coefficient corresponding to the first similarity, a second coefficient corresponding to the second similarity, and a third coefficient corresponding to the third similarity respectively; Based on the first similarity, the first coefficient, the second similarity, the second coefficient, the third similarity and the third coefficient, obtain the paragraph similarity between the first text and the second text; According to the paragraph similarity between the first text and the second text, determine the text similarity between the first text and the second text.

2. The method according to claim 1, characterized in that The determining the first similarity between paragraph i and paragraph j based on the paragraph features of paragraph i of the first text and the paragraph features of paragraph j of the second text includes: Perform word segmentation processing on the paragraph features of paragraph i and the paragraph features of paragraph j respectively to obtain a first cluster and a second cluster; Calculate the intersection-union ratio of the first cluster and the second cluster, and based on the intersection-union ratio, determine the first similarity between paragraph i and paragraph j.

3. The method according to claim 2, wherein Before calculating the intersection-union ratio of the first cluster and the second cluster, it further includes: According to the first cluster and the second cluster, determine the non-intersecting words between paragraph i and paragraph j; The calculating the intersection-union ratio of the first cluster and the second cluster includes: If the number of negative words in the non-intersecting words between paragraph i and paragraph j is even, then calculate the intersection-union ratio of the first cluster and the second cluster.

4. The method according to claim 1, wherein The determining the text similarity between the first text and the second text according to the paragraph similarity between the first text and the second text includes: According to each paragraph of the first text and each paragraph of the second text, determine the paragraph weights corresponding to the paragraph similarity; Determine the text similarity between the first text and the second text based on the paragraph similarity and the paragraph weight corresponding to the paragraph similarity.

5. The method according to any one of claims 1 to 3, characterized in that, Before extracting the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text, it further includes: Determine whether the type of the first text is a preset text type, and determine whether the type of the second text is the preset text type; The extraction of the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text includes: If the type of the first text is the preset text type and the type of the second text is the preset text type, then extract the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text.

6. A text similarity judgment device, characterized in that, It includes: A first feature extraction module, used to determine the first text and the second text for text similarity judgment, and extract the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text; A first similarity determination module, used to determine the first similarity between paragraph i of the first text and paragraph j of the second text based on the paragraph feature of paragraph i of the first text and the paragraph feature of paragraph j of the second text, where paragraph i is any paragraph in the first text, i = 1, 2, …, m, m is an integer, m is determined according to the total number of paragraphs in the first text, paragraph j is any paragraph in the second text, j = 1, 2, …, n, n is an integer, n is determined according to the total number of paragraphs in the second text; A second feature extraction module, used to extract the keywords of paragraph i and paragraph j respectively if the first similarity is greater than a first preset threshold; A second similarity determination module, used to determine the second similarity between paragraph i and paragraph j based on the keywords of paragraph i and paragraph j; A text similarity judgment module, used to determine the text similarity between the first text and the second text according to the first similarity and the second similarity; The first feature extraction module is specifically used for: Extract the text features of the first text and the text features of the second text respectively; Based on the text features of the first text and the text features of the second text, determine the third similarity between the first text and the second text; If the third similarity is greater than a second preset threshold, then extract the paragraph features of each paragraph of the first text and the paragraph features of each paragraph of the second text; The text similarity judgment module is specifically used for: According to the first text and the second text, determine the first coefficient corresponding to the first similarity, the second coefficient corresponding to the second similarity, and the third coefficient corresponding to the third similarity respectively; Based on the first similarity, the first coefficient, the second similarity, the second coefficient, the third similarity and the third coefficient, obtain the paragraph similarity between the first text and the second text; Determine the text similarity between the first text and the second text according to the paragraph similarity between the first text and the second text.

7. The device according to claim 6, characterized in that The first similarity determination module is specifically configured to: Perform word segmentation processing on the paragraph features of the paragraph i and the paragraph features of the paragraph j respectively to obtain a first cluster and a second cluster; Calculate the intersection-to-union ratio of the first cluster and the second cluster, and determine the first similarity between the paragraph i and the paragraph j based on the intersection-to-union ratio.

8. The device according to claim 7, characterized in that, The first similarity determination module is specifically configured to: Determine the non-intersecting vocabulary between the paragraph i and the paragraph j according to the first cluster and the second cluster; If the number of negative words in the non-intersecting vocabulary between the paragraph i and the paragraph j is even, calculate the intersection-to-union ratio of the first cluster and the second cluster.

9. A text similarity judgment device, characterized in that, Comprising: A processor; A memory; And A computer program; Wherein, the computer program is stored in the memory and is configured to be executed by the processor, and the computer program includes instructions for executing the method according to any one of claims 1-5.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program causes the server to execute the method according to any one of claims 1-5.

11. A computer program product, characterized in that, Including computer instructions, the computer instructions are executed by the processor to execute the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Bidding document similarity calculation method and device

    CN111160445A

  • Text similarity determination method and device and storage medium

    CN111737997A