Text similarity detection method, readable medium and electronic device

By constructing the ‘text-keyword’ graph, using TextRank and RWR algorithms to calculate text similarity, the problems of high resource occupation, slow speed and insufficient accuracy in the existing methods are solved, and efficient and accurate text similarity detection is achieved.

WO2025148190A1PCT designated stage expired Publication Date: 2025-07-17EBAOTECH CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/088863
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-10
Filing Date
2024-04-19
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

The existing text similarity calculation methods occupy a lot of resources, have slow calculation speed, and lack reliability in the results, which fail to effectively capture the internal structure information of the text.

Method used

By constructing a ‘text-keyword’ graph, the keywords are extracted using the TextRank algorithm and the correlation score between keywords and text is calculated based on the RWR algorithm, and the text similarity is directly calculated without vectorization operations.

Benefits of technology

The text similarity calculation process is simple and easy to implement, with less resource occupancy, high calculation speed and high accuracy, improving the accuracy of text similarity detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024088863_17072025_PF_FP_ABST
    Figure CN2024088863_17072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the field of text processing. Disclosed are a text similarity detection method, a readable medium and an electronic device, which ensure that a text similarity calculation process is simple and easily implemented, occupies few resources, and has a relatively high calculation speed and a relatively high level of calculation accuracy. In the method, keywords of each of a plurality of pieces of candidate text may be extracted, and a "text-keyword" graph may be constructed on the basis of the inclusion relationship between the candidate text and the keywords. That is, the "text-keyword" graph is used for representing the inclusion relationship between the candidate text and the keywords. Then, for a piece of input text, a keyword set of the input text may be extracted. Then, a correlation score between each keyword in the keyword set and each piece of candidate text is calculated on the basis of the "text-keyword" graph. Next, a similarity score between the input text and each piece of candidate text is calculated on the basis of the correlation score. Finally, pieces of candidate text with relatively high similarity scores are taken as text similar to the input text.
Need to check novelty before this filing date? Find Prior Art

Description

Text similarity detection method, readable medium and electronic device

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 10, 2024, with application number 202410039257.1 and application name “A text similarity detection method, readable medium and electronic device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of text processing technology, and in particular to a text similarity detection method, a readable medium, and an electronic device. Background Art

[0003] In natural language processing (NLP) tasks, we often need to determine whether two texts are similar and calculate the degree of similarity between them. For example, when preprocessing a corpus, we can identify and delete duplicate texts from a large amount of corpus based on text similarity.

[0004] At present, text similarity calculation is usually based on the method of combining text vectorization with language model, such as the language model can be an n-gram model (a statistical language model). Specifically, the method can pre-process the text, including operations such as word segmentation, removal of stop words, and stem extraction to obtain a vocabulary composed of words in the text. Then, the frequency of occurrence of each word in the text is calculated using the n-gram model, and it is expressed as a vector, which can be regarded as the representation of the text in the n-gram space. It can be understood that, usually in order to better compare the similarity of each text, the vector of a text can be normalized with the vector of other texts. Furthermore, methods such as cosine similarity or Euclidean distance can be used to calculate the similarity of the vector of each text in the n-gram space with the vector of other texts.

[0005] However, when calculating text similarity, vectorization methods require first loading the model into the device's memory and then vectorizing the input text. This consumes a lot of resources and, as the model size increases, the execution speed slows down. Consequently, the text similarity calculation process may consume more resources and be slower. Furthermore, these vectorization methods only consider the semantic information of the text, but lack the ability to capture the text's inherent structural information, resulting in unreliable text similarity calculation results.

[0006] Summary of the Invention

[0007] The embodiments of the present application provide a text similarity detection method, a readable medium, and an electronic device, which ensure that the text similarity calculation process is simple and easy to implement, occupies few resources, has a high calculation speed, and has a high calculation accuracy.

[0008] In the first aspect, an embodiment of the present application provides a text similarity detection method, which includes: obtaining an input text; extracting keywords from the input text to obtain a keyword set; determining a node graph between a plurality of preset candidate texts and a plurality of candidate keywords, wherein the plurality of candidate keywords are keywords extracted from the plurality of candidate texts, the node graph uses the plurality of candidate texts as one type of node, the plurality of keywords as another type of node, the inclusion relationship between the plurality of candidate texts and the plurality of candidate keywords as an edge, and the first score (i.e., TextRank score) of a candidate keyword of a candidate text as the weight of the edge between the candidate text and the candidate keyword; according to the node graph, calculating the relevance score between each keyword in the keyword set and each candidate text; calculating the similarity score between the input text and each candidate text according to the relevance score between each keyword in the keyword set and each candidate text; taking the k candidate texts with the highest similarity score as similar texts of the input text, where k is a positive integer, such as k is 1 or 2. At this point, the above-mentioned node graph is the "text-keyword" graph mentioned below. In this way, the method does not need to load model resources, greatly reduces resources, and ensures calculation speed. Furthermore, this method does not require vectorization of the text, but instead considers the correlation between keywords within each text, thereby improving the accuracy of text similarity detection. In other words, the text similarity calculation process in this application is simple and easy to implement, consumes few resources, has a high calculation speed, and has high calculation accuracy.

[0009] In a possible implementation of the first aspect, the relevance score between each keyword in the keyword set and each candidate text is iteratively calculated using the following formula: l (c,t)=a(1-a) l W c,t +R l-1 (c, t); where R l (c, t) is the correlation score between keyword c and text node t, l is the number of iterations of the random walk with restart (RWR) algorithm expressed in the formula, a is the restart factor and a∈(0, 1), W c,trepresents the transition probability between keyword c and text node t, C represents the keyword set, c is a keyword in the keyword set C, and t is a candidate text in multiple candidate texts. It can be understood that the present application can use the RWR method to calculate the relevance score between the keywords of the input text and each candidate text, which can take into account the local similarity and global similarity between each node in the "text-keyword" graph, so that the calculated relevance score between the keyword and the text is more accurate.

[0010] In a possible implementation of the first aspect, the similarity score between the input text and each candidate text is calculated using the following formula: in, represents the similarity score between text p and text t, and S(c) is the weight of keyword c in text p, calculated using TextRank (a text ranking algorithm). p is the input text, C represents the keyword set, c is a keyword in keyword set C, and t is a text in multiple candidate texts. It can be understood that the edges between text nodes and keyword nodes in the node graph have weights, which are used to subsequently calculate the relevance scores between the keyword and each candidate text. For example, the weights between text nodes and keyword nodes can be calculated using the TextRank algorithm, meaning that a "text-keyword" graph can be generated based on the TextRank algorithm.

[0011] In a possible implementation of the first aspect, a node graph is obtained based on the following method: obtaining multiple candidate texts; using a word segmenter to segment each candidate text in the multiple candidate texts based on the word segmentation dictionary and stop words corresponding to the multiple candidate texts to obtain the words of each candidate text; determining the co-occurrence relationship between the words of each candidate text according to the set window size, and constructing a word graph of the words in each candidate text based on the co-occurrence relationship between the words; extracting the candidate keywords in the word graph of each candidate text and the first score of the corresponding candidate keywords, wherein the first score is used to reflect the importance of a keyword to a text; constructing a node graph based on the inclusion relationship between each candidate text and the corresponding candidate keywords, and the first score of each candidate keyword in each candidate text; wherein the weight between a candidate text and a connected candidate keyword in the node graph is the first score of the candidate keyword to the candidate text. It can be understood that the word segmentation dictionary and stop words corresponding to the multiple candidate texts can be the word segmentation dictionary and stop words of the text field (such as the insurance field) to which these candidate texts belong, including the default word segmentation dictionary and default stop words, the custom word segmentation dictionary and the content in the custom stop words. For example, the node graph generation process in the present application can be executed in an offline stage of the electronic device.

[0012] In a possible implementation of the first aspect, the candidate keywords of a candidate text are the top m words with the largest first scores among all words in the word graph of the candidate text. The present application may sort the words in the word graph of the candidate text from largest to smallest based on the first scores of all words, and select the top m words with the largest scores as the candidate keywords for the candidate text.

[0013] In a possible implementation of the first aspect above, keywords are extracted from the input text to obtain a keyword set, including: based on the word segmentation dictionary and stop words corresponding to the input text, using a word segmenter to segment the input text to obtain the words of the input text; determining the co-occurrence relationship between the words of the input text according to the set window size, and constructing a word graph of the words in the input text based on the co-occurrence relationship between the words; obtaining the first score of all words in the word graph of the input text, wherein the first score is used to reflect the importance of a keyword to a text; adding the first m words with the largest first scores among all words in the word graph of the input text to the keyword set of the input text, where m is a positive integer. It can be understood that the word segmentation dictionary and stop words corresponding to the input text can be the word segmentation dictionary and stop words of the text field to which the input text belongs (such as the insurance field), including the default word segmentation dictionary and default stop words, the custom word segmentation dictionary and the content in the custom stop words. For example, the similarity text calculation process of the input text can also be executed in the online stage. For example, the present application may sort the words from large to small according to the first scores of all words in the word graph of the input text, and select the first m words with larger scores to add to the keyword set of the input text.

[0014] In a possible implementation of the first aspect, the first score is calculated in the following manner:

[0015] Among them, S(v i ) represents node v i The first score, S(v j ) represents node v j The first fraction, In(v i ) indicates v i The entry node set, Out(v j ) indicates v j The outgoing node set, node v i and v j They are two word nodes in the word graph of a text, node v i and v j The weight of the edge between them is denoted as w ji , d is the damping coefficient and ranges from 0 to 1, V k′ is another node in the text, w jk′For node V k′ and node v j It can be understood that the TextRank algorithm calculates the score of each node by iteration until the node score converges. In each iteration, the algorithm first calculates the score S(v i ), and then use the node score as the node weight for the next iteration to calculate the new weight S(v i ) until the node score converges. Thus, the weight of the node after convergence is used as the first score of the node.

[0016] In a possible implementation of the first aspect, the input text and the candidate texts belong to the same field. For example, the input text and the candidate texts both belong to the insurance field.

[0017] In a second aspect, an embodiment of the present application provides a readable medium having instructions stored thereon. When the instructions are executed on an electronic device, the electronic device executes the text similarity detection method in the above-mentioned first aspect and any possible implementation of the first aspect.

[0018] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the processors of the electronic device, for executing the text similarity detection method in the above-mentioned first aspect and any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] FIG1 shows a block diagram of text similarity calculation according to some embodiments of the present application;

[0020] FIG2 is a schematic diagram showing a flow chart of a text similarity detection method according to some embodiments of the present application;

[0021] FIG3 is a schematic diagram showing a process of extracting keywords from an input text according to some embodiments of the present application;

[0022] FIG4 shows a schematic diagram of a process for generating a “text-keyword” graph according to some embodiments of the present application;

[0023] FIG5 shows a schematic structural diagram of a mobile phone according to some embodiments of the present application;

[0024] FIG6 shows a schematic structural diagram of an electronic device according to some embodiments of the present application;

[0025] FIG7 shows a block diagram of a system according to some embodiments of the present application. DETAILED DESCRIPTION

[0026] Illustrative embodiments of the present application include, but are not limited to, a text similarity detection method, a readable medium, and an electronic device.

[0027] In order to facilitate those skilled in the art to understand the solutions in the embodiments of the present application, some concepts and terms involved in the embodiments of the present application are explained below.

[0028] TextRank algorithm: is a text ranking algorithm used to extract keywords and phrases from text. It regards each word in the text as a node, and constructs a word network by calculating the similarity between words, and then calculates the TextRank score of each word (which can be recorded as the first score), and finally extracts important words and phrases. Specifically, in a word network containing n nodes, the TextRank algorithm initializes a weight or score 1 / n for each word. After the iterative formula converges multiple times, the TextRank score of each word will change. The words with the highest scores are selected and can be used as the optimal and suboptimal keywords for the text, and so on. Usually only the top 3-5 are needed. Among them, n is a positive integer. In addition, the keywords obtained after executing the TextRank algorithm will carry the corresponding scores, which can be used to initialize the weights of the edges in the "text-keyword" graph.

[0029] Co-occurrence refers to the frequency or probability of certain words or concepts appearing together in a text. Co-occurrence can be measured by counting the number of times two words co-occur in a text. If two words frequently co-occur in a text, then they have a strong co-occurrence relationship. Conversely, if two words rarely co-occur, then their co-occurrence relationship is weak.

[0030] Window parameter: Also known as the window size, it specifies the scope of the text context considered when calculating co-occurrence relationships. Specifically, the window size specifies the maximum distance between two words in the text. Only when the distance between two words is less than or equal to the window size are the two words considered to co-occur.

[0031] Random walk with restart (RWR) algorithm: used to calculate the similarity between nodes in a network. The basic idea of ​​this method is to start from the target node, randomly walk to the neighboring node with a certain probability, and then randomly walk to the neighboring node of the neighboring node with the same probability until a certain maximum number of steps is reached. In this process, each time a new node is randomly walked to, the access counter of the node is increased by one. Then, the algorithm performs a reverse random walk starting from the target node in the same way until the maximum number of steps is reached. Finally, the access counters of the forward and reverse random walks are added to obtain the similarity score (or correlation score) of each node. The advantage of the RWR algorithm is that it can consider the local and global similarities between nodes, and can handle the similarity calculation problem in large-scale networks.

[0032] As mentioned above, existing text similarity detection methods consume a lot of resources, have slow calculation speeds, and lack reliability in calculation results.

[0033] Based on this, the present application provides a text similarity detection method that can extract keywords from each candidate text among multiple candidate texts and construct a "text-keyword" graph based on the inclusion relationship between the candidate texts and the keywords. Specifically, this "text-keyword" graph is used to characterize the inclusion relationship between the candidate texts and the keywords. Then, for an input text, a set of keywords can be extracted from the input text. Based on the "text-keyword" graph, a correlation score is calculated between each keyword in the keyword set and each candidate text. Furthermore, a similarity score is calculated between the input text and each candidate text based on each correlation score. Finally, the candidate text with the highest similarity score is designated as the text similar to the input text. This method eliminates the need to load model resources, significantly reducing resources and ensuring computational speed. Furthermore, this method eliminates the need to vectorize the text, instead considering the correlation between keywords within each text. This method is therefore beneficial for improving the accuracy of text similarity detection. The text similarity calculation process in the present application is simple and easy to implement, consumes few resources, and achieves high computational speed and accuracy.

[0034] This application can construct a "text-keyword" graph using the inclusion relationship between each candidate text and the extracted keyword as an edge. It is understood that the "text-keyword" graph can also be described by other names, such as a node graph of candidate texts and keywords, a node graph of the relationship between candidate texts and keywords, etc., which is not specifically limited in this embodiment of the application.

[0035] In some embodiments, the "text-keyword" graph is a directed weighted graph, i.e., a node architecture graph with direction and weight. Specifically, the "text-keyword" graph contains two types of nodes, namely text nodes and keyword nodes, and contains one type of edge, i.e., an edge connecting the text node and the keyword node, and the edge between them describes the relationship that "the text contains keywords". In addition, the edge between the text node and the keyword node has a weight, which is used to subsequently calculate the relevance score of the keyword and each candidate text. For example, the weight between the text node and the keyword node can be calculated using the TextRank algorithm, i.e., the "text-keyword" graph can be generated based on the TextRank algorithm.

[0036] In some embodiments, the present application can use the RWR method to calculate the relevance scores between the keywords of the input text and each candidate text. In this way, the present application can take into account the local similarity and global similarity between each node in the "text-keyword" graph, making the calculated relevance scores between keywords and texts more accurate.

[0037] The following is a brief description of the generation and application scenarios of the "text-keyword" graph in this application in conjunction with the accompanying drawings. Referring to Figure 1, a text similarity calculation block diagram is provided in an embodiment of this application. The block diagram includes an offline phase and an online phase. The offline phase is mainly used to generate a "text-keyword" graph corresponding to multiple candidate texts, and the online phase is mainly used to determine texts similar to the input text from the "text-keyword" graph.

[0038] Specifically, in the offline phase shown in FIG1 , taking the candidate texts text 1, text 2, and text 3 as examples, keywords can be extracted from each candidate text. For example, keywords 1, 2, and 3 are extracted from text 1, keywords 1 and 3 are extracted from text 2, and keywords 2 and 3 are extracted from text 3. Furthermore, based on the inclusion relationship between the three text nodes text 1, text 2, and text 3 and the three keyword nodes keywords 1, 2, and 3, a “text-keyword” graph 00 can be constructed. Taking text 2 as an example, the edge between text 2 and keyword 1 in the “text-keyword” graph 00 in FIG1 (represented by a black line segment) indicates that text 2 and keyword 1 have an inclusion relationship, and the edge between text 2 and keyword 3 (represented by a black line segment) indicates that text 2 and keyword 3 have an inclusion relationship, while text 2 and keyword 2 do not have an inclusion relationship.

[0039] In some embodiments, a text can contain multiple keywords, and a keyword can be contained in multiple texts. That is, a text node in the "text-keyword" graph can be connected to one or more keyword nodes, and a keyword node can be connected to one or more text nodes. For example, if keyword 2 has a containment relationship with both text 1 and text 3, then keyword 2 can establish edges with both text 1 and text 3.

[0040] It is understood that the candidate texts for actually constructing the "text-keyword" graph 00 are not limited to the texts 1 to 3 shown in FIG1 , and may include more candidate texts. Accordingly, the "text-keyword" graph 00 also includes the inclusion relationship between other candidate texts and other keywords, which is not specifically limited in this application.

[0041] In the online stage shown in Figure 1, for the input text, one or more keywords of the input text can be extracted. Then calculate the correlation scores (or similarity scores) of all the keywords of the input text and all the candidate texts in the "text-keyword" figure 00. Subsequently, based on the correlation scores of all the keywords of the input text and each candidate text, calculate the similarity scores between the input text and each candidate text, and output the top k candidate texts with higher similarity scores, that is, output the top k similar texts. Among them, the value of k can be set according to actual needs, such as k taking a value between 3 and 5, and this application does not make specific restrictions on this.

[0042] It is understood that the "offline" in FIG1 refers to the electronic device performing text similarity detection disconnecting from the Internet, also known as offline mode or offline. In addition, the "online" in FIG1 refers to the electronic device establishing a connection with the Internet, also known as online.

[0043] In some other embodiments, the process of generating the "text-keyword" graph and the process of calculating similar texts of the input text in FIG1 can also be performed in other stages. For example, the process of generating the "text-keyword" graph in this application can be performed in an online stage, and the process of calculating similar texts of the input text can also be performed in an offline stage. This application does not make specific limitations on this.

[0044] In some embodiments, the electronic device that generates the “text-keyword” graph in this application (i.e., the generating device) and the electronic device that uses the “text-keyword” graph to perform text similarity detection (referred to as the detecting device) can be the same electronic device or different electronic devices, and this application does not make specific restrictions on this. For example, when two processes are executed by different electronic devices, the detecting device can first obtain the “text-keyword” graph from the generating device, and then perform text similarity detection based on the “text-keyword” graph. In the following embodiments, the generating device and the detecting device are mainly used as an example to illustrate the same electronic device.

[0045] The execution subject of the text similarity detection method provided in the embodiment of the present application can be an electronic device, or a module or unit in the electronic device for executing the method.

[0046] Next, the text similarity detection method provided in an embodiment of the present application is described in detail with reference to FIG. 2 . The method may be performed by an electronic device and includes the following steps:

[0047] S201: Obtain input text.

[0048] In some embodiments, the input text may be input in text form or voice form, which is not specifically limited in this application. When the input text is input in voice form, the electronic device may convert the input voice into text to obtain the input text.

[0049] Furthermore, the text similarity detection method provided in the embodiments of the present application can be applied to any field requiring similar text detection. For example, in the insurance field, for an input text in the insurance field, information related to the insurance field can be retrieved from a known text library. It is understood that the candidate text and the input text typically belong to the same field, such as insurance.

[0050] In some embodiments, the text similarity detection method provided by the present application can be applied to various text similarity detection scenarios, such as related question retrieval or document retrieval. For example, in a web search such as Baidu™ search, a question text is entered, and a "Related Search" column will appear on the returned page, and text similar to the question text will be displayed in the "Related Search", thereby helping users find related questions. For another example, in the intelligent customer service "You may want to ask", after the user enters the question text, the intelligent customer service can output questions similar to the question text. In a question-and-answer system, some classic questions and corresponding answers are usually prepared. When the user's question is very similar to a classic question, the system directly returns the prepared answer. For another example, in document retrieval, the user enters a query text, and the query device will return documents containing keywords. Of course, the text similarity detection scenarios that can be applied to the present application include but are not limited to the above examples, and can also be any other achievable scenarios.

[0051] S202: Extract keywords from the input text to obtain a keyword set.

[0052] In some embodiments, the present application may use the TextRank algorithm to extract one or more keywords from the input text, and the one or more keywords may constitute a keyword set.

[0053] Referring to FIG. 3, the process of extracting keywords from the input text in S202 will be described. Specifically, the above S202 can be implemented through the following S2021 to S2025:

[0054] S2021: Load the word segmentation dictionary and stop words corresponding to the input text.

[0055] Among them, the word segmentation dictionary and stop words can be the dictionary and stop words in the text field to which the input text belongs, such as the dictionary and stop words in the insurance field.

[0056] For example, the above word segmentation dictionary can be the default dictionary or custom dictionary for word segmentation. For example, the default dictionary can be the default dictionary in the field to which the input text belongs. For example, in the insurance query field, the default dictionary can be the dictionary in the insurance field. The custom dictionary can be set in combination with the field to which the input text belongs and can be customized by relevant technical personnel, which is beneficial to improving the word segmentation accuracy of the default dictionary.

[0057] The above stop words are the default stop words or custom stop words for filtering useless words in the text, such as words that have no obvious effect on reflecting the text semantics, such as "de", "this", etc. Specifically, for each sentence (or text), the electronic device can perform word segmentation and词性标注 on it, and then剔除 the stop words, only retaining words of specified词性, such as nouns, verbs, adjectives, etc.

[0058] S2022: Based on the loaded word segmentation dictionary and stop words, use a word segmenter to perform word segmentation on the input text to obtain the words of the input text.

[0059] For example, the above word segmenter can be the jieba word segmenter (a Chinese word segmentation tool), or it can be other word segmentation tools. This application does not make specific limitations on this. Specifically, this application can first perform word segmentation on the input text based on the loaded word segmentation dictionary, and then remove the stop words in the word segmentation result based on the loaded stop words to obtain the final word segmentation result.

[0060] S2023: Determine the co-occurrence relationship between the words of the input text according to the set window size, and based on the co-occurrence relationship between the words, construct a word graph of the words in the input text.

[0061] Among them, this application can construct a word graph of the words in the input text by taking the words as nodes and the co-occurrence relationship as edges based on the co-occurrence relationship between the words.

[0062] It should be noted that there is an unclear expression "词性标注" in the original text, which should be a more specific term in the relevant field. And "剔除" should be a more accurate Chinese word, here it is translated as "remove" for the purpose of translation. You can adjust according to the actual situation.It can be understood that the above window size can be the window size used to extract keywords using the TextRank algorithm, which is used to establish the co-occurrence relationship of word segmentations. Generally speaking, the choice of window size needs to be adjusted according to the specific text length, text type, and research purpose. Some common window sizes include 2, 3, 5, etc. In practical applications, the window size usually needs to be experimented and tuned to determine the most appropriate window size. At this point, the word graph of the input text can be processed using the TextRank algorithm later.

[0063] Specifically, this application constructs a word graph G of the input text based on the co-occurrence relationship within a sliding window of a set window size.<V,E> Here, E represents the set of nodes consisting of words in the input text, and E represents the set of edges constructed based on the co-occurrence relationship between these words. Furthermore, an edge between any two nodes is constructed using a co-occurrence relationship only if their corresponding words co-occur in a window of length K, i.e., a maximum of K words co-occur, with K typically set to 2. This means that the window size can be 2.

[0064] S2024: Obtain the TextRank scores of all words in the word graph of the input text based on the TextRank algorithm.

[0065] In this case, a TextRank score is used to reflect the importance of a word (such as a keyword) to the input text. As you can understand, in the TextRank algorithm, each node (i.e., sentence or word) has a weight (i.e., TextRank score) that represents the importance of the node. The weight of a node is determined by its relevance to other nodes and the importance of the node itself.

[0066] In some embodiments, the TextRank score of a word, ie, the TextRank score, can be defined as S(v i ), and calculated by the following formula (1):

[0067] Among them, S(v i ) represents node v i The TextRank score, S(v j ) represents node v j TextRank score, In(v i ) indicates v i The entry node set, that is, the node v i All connected nodes. Out(v j ) indicates v j The set of outgoing nodes, i.e., from node v j All nodes that can be reached by starting from Node v i and v jThey are two word nodes in the word graph of a text, node v i and v j The weight of the edge between them is denoted as w ji . That is w ji Represents node v j and node v i The correlation between them can be calculated by cosine similarity and other methods. d is the damping coefficient or attenuation coefficient, ranging from 0 to 1, and generally takes a value of 0.85. In addition, V k′ Can be another node in the text, and w jk′ For node V k′ and node v j The weight between .

[0068] In some embodiments, the present application can initialize the weight of each node in the word graph based on the above formula (1), such as initializing it to 1 or 1 / n (n is the number of nodes in the word graph), and then iteratively calculate each weight until convergence, and obtain the weight of each node after iterative calculation, that is, obtain the TextRank score of each node.

[0069] Specifically, the TextRank algorithm calculates the score of each node by iteration until the node scores converge. In each iteration, the algorithm first calculates the score S(v i ), and then use the node score as the node weight for the next iteration to calculate the new weight S(v i ) until the node score converges. Thus, the weight of the node after convergence is used as the TextRank score of the node.

[0070] S2025: Add the top m words with the largest TextRank scores among all words in the word graph of the input text to the keyword set of the input text, where m is a positive integer.

[0071] In some embodiments, the present application can sort the words from large to small according to the TextRank scores of all words in the word graph of the input text, and select the top m words with larger scores to add to the keyword set of the input text.

[0072] Here, m is a positive integer, and its specific value can be set according to actual needs and is not specifically limited.

[0073] Next, the text similarity detection method in FIG. 3 will be described in detail from S203 to S205 .

[0074] S203: Based on the “text-keyword” graph, calculate the relevance score between each keyword in the keyword set of the input text and each candidate text.

[0075] The “text-keyword” graph has inclusion relationships between multiple candidate texts and multiple keywords, and also has weights between multiple candidate texts and the included keywords.

[0076] It can be understood that before the electronic device obtains the user's input text, it can pre-generate a "text-keyword" graph consisting of multiple candidate texts and their keywords. The generation process of the "text-keyword" graph will be described in detail below and will not be repeated here.

[0077] In some embodiments, it is assumed that the input text is p, the keyword set of the input text p is C, and the keyword in the keyword set C is c. Moreover, the candidate text in the "text-keyword" graph is t, that is, the text node in the "text-keyword" graph is t. Then, on the "text-keyword" graph, the keywords included in the keyword set C of the input text p can be located / matched. For each keyword c, the RWR algorithm is iteratively executed to obtain the correlation score between the keyword c and a text node t, which is recorded as R l (c, t), the iterative process is shown by the following formula (2): R l (c,t)=a(1-a) l W c,t +R l-1 (c,t) (2)

[0078] Where l is the number of iterations of the RWR algorithm, a is the restart factor and a∈(0,1), W c,t represents the transition probability between keyword c and text node t.

[0079] S204: Calculating similarity scores between the input text and each candidate text based on the correlation scores between each keyword in the input text and each candidate text.

[0080] In some embodiments, the present application can perform a weighted summation of the relevance scores between each keyword and each candidate text based on the RWR algorithm and the TextRank algorithm to calculate the similarity score between the input text p and a candidate text t. For example, the present application can define the similarity score between the input text p and a candidate text t as shown in the following formula (3):

[0081] Among them, S(c) is the weight of keyword c of input text p calculated by TextRank algorithm, that is, TextRank score.

[0082] S205: The first k candidate texts with higher similarity scores are used as similar texts of the input text.

[0083] It can be understood that the value of k is a positive integer, for example, k is 2. The specific value can be determined according to actual needs and is not specifically limited to this.

[0084] For example, if the input text is "What is the division of labor among departments in reviewing business contracts?", the top two candidate texts with higher similarity scores may be Rank 1: "What is the main division of labor among departments in reviewing procurement contracts?" and Rank 2: "Which departments / individuals need to approve business contracts?".

[0085] In this way, the text similarity detection method in this application does not need to load model resources, nor does it need to perform vectorization operations on the text, so it can reduce resource usage and improve similarity detection results. Specifically, the method uses the TextRank algorithm to extract keywords from the text based on the word graph, and uses the RWR algorithm to calculate the correlation between the keywords of the input text and the candidate text based on the "text-keyword graph", and then combines the TextRank algorithm and the RWR algorithm to calculate the similarity between the input text and the candidate text. In this way, the present application can realize text similarity detection based on the calculation of the word graph and the "text-keyword" graph, which is simple and easy to implement and takes up fewer resources.

[0086] In some embodiments, the “text-keyword” graph between candidate texts and keywords in the present application may be generated based on the TextRank algorithm.

[0087] 4 shows a process for generating a "text-keyword" graph, i.e., a node graph, provided for application implementation. The execution subject of the process may be an electronic device, and the process includes the following steps:

[0088] S401: Acquire a candidate text library.

[0089] In some embodiments, the candidate text library can be collected from the internet or from internal company documents. For example, internal company documents can include company training materials, such as employee handbooks, insurance information, InsureMO (a knowledge base in the insurance field), legal information, etc. Furthermore, the candidate text can be from any field, such as insurance.

[0090] The electronic device may generate a “text-keyword” graph corresponding to the candidate text library when first acquiring the candidate text library.

[0091] S402: Based on the word segmentation dictionary and stop words corresponding to the candidate text library, use a word segmenter to segment each candidate text in the candidate text library to obtain words of each candidate text.

[0092] For example, the word segmentation dictionary and stop words corresponding to the candidate text library may be the dictionary and stop words in the text field to which the candidate text belongs, such as the insurance field.

[0093] In some embodiments, the word segmentation dictionary and stop words corresponding to the candidate text library can be the same or have the same field as the word segmentation dictionary and stop words corresponding to the input text above. In addition, the above-mentioned word segmenter can be a jieba word segmenter or other word segmenters, and this application does not make specific limitations on this.

[0094] S403: Determine the co-occurrence relationship between words in each candidate text in the candidate text library according to the set window size, and construct a word graph of the words in each candidate text based on the co-occurrence relationship between the words, using the words as nodes and the co-occurrence relationship as edges.

[0095] Specifically, this application constructs a candidate word (i.e., a word in a candidate text) word graph G′=<V′,E′> Where V′ represents the set of nodes consisting of candidate words, and E′ represents the set of edges constructed based on co-occurrence relationships. Furthermore, an edge between any two nodes is constructed using co-occurrence relationships only if their corresponding words co-occur in a window of length K, i.e., a maximum of K words co-occur, with K typically set to 2. This means the window size can be 2.

[0096] S404: Extract keywords and corresponding TextRank scores of keywords from the word graph of each candidate text based on the TextRank algorithm.

[0097] Among them, the keywords in the candidate text in this application can also be called candidate keywords, and the embodiments of this application do not specifically limit this name.

[0098] In some embodiments, the TextRank scores of all words in the word graph of each candidate text can be obtained based on the TextRank algorithm, and the m words with higher TextRank scores for each candidate text are used as keywords. The TextRank score of a word can be the weight of the word in the word graph.

[0099] In some embodiments, the present application can iteratively calculate the TextRank score of each word in each candidate text based on the above formula (1), and select keywords for each candidate text based on these TextRank scores. Specifically, the present application can initialize the weight of each node in the word graph of the candidate text based on the above formula (1), such as initializing it to 1 or 1 / n, and then iteratively calculate each weight until convergence, obtaining the weight of each node after the iterative calculation, that is, obtaining the TextRank score of each node.

[0100] It is understandable that the present application can process each candidate text in the candidate text library separately to extract the keywords in each candidate text and the TextRank scores (ie, weights) between each keyword and each candidate text.

[0101] S405: Construct a “text-keyword” graph corresponding to the candidate text library based on the inclusion relationship between each candidate text and the corresponding keyword, and the TextRank score of each keyword in each candidate text.

[0102] In some embodiments, the present application may regard candidate texts and keywords as two different nodes, and add each candidate text and each corresponding keyword to the node set of the "text-keyword" graph; based on the "inclusion" relationship, connect the text node and the keyword node, and add the edge between each candidate text and the corresponding keyword to the edge set of the "text-keyword" graph; use the TextRank score of a keyword in a candidate text obtained by the TextRank algorithm as the weight between the text node where the candidate text is located and the node where the keyword it contains is located in the "text-keyword" graph, so as to obtain the weight of each edge in the "text-keyword" graph. That is, the present application can construct a node set of the "text-keyword" graph corresponding to the candidate text library based on text and keywords as two nodes, the inclusion relationship between candidate text and keywords as the edge of text nodes and keyword nodes, and the TextRank score of a keyword in a candidate text as the weight between the candidate text and the keyword.

[0103] In some implementations, when text in the candidate text library is subsequently updated, the electronic device can update the text nodes in the "text-keyword" graph and detect whether the keywords need to be updated. In this case, whether the keywords need to be updated is determined by checking whether the keywords extracted from the new text already exist in the "text-keyword" graph. Specifically, if the newly generated keyword does not exist in the "text-keyword" graph, the keyword can be added to the "text-keyword" graph.

[0104] In addition, based on the keywords of the candidate text generated by the TextRank algorithm, this application can also be manually intervened to make corrections to improve the accuracy of keyword extraction and thus improve the accuracy of the "text-keyword" graph.

[0105] It can be understood that the "text-keyword" in this application is a node graph composed of multiple candidate texts and corresponding multiple keywords, and the multiple keywords are multiple candidate keywords extracted from multiple candidate texts. Specifically, the node graph is based on the inclusion relationship between each candidate text and the corresponding candidate keyword, as well as the TextRank score of each candidate keyword in each candidate text. Among them, each candidate keyword is extracted from the word graph of each candidate text. At this time, the word graph of the words in each candidate text can determine the co-occurrence relationship between the words of each candidate text according to the set window size, and is constructed based on the co-occurrence relationship between the words. In addition, the word graph of each candidate text is based on the word segmentation dictionary and stop words corresponding to the multiple candidate texts, and each candidate text in the multiple candidate texts is segmented using a word segmenter to obtain the words of each candidate text.

[0106] In this way, the text similarity detection method provided by the present application can generate a "text-keyword" graph between multiple candidate texts and their keywords based on the TextRank algorithm. The "text-keyword" graph not only has the inclusion relationship between each candidate text and each keyword, but also has the weight between each candidate text and each keyword. Therefore, the present application can determine the similarity score between the input text and each candidate text based on the keywords of the input text based on the "text-keyword" graph, thereby realizing text similarity detection based on graph calculation.

[0107] The electronic devices applicable to this application can be any electronic device with text processing capabilities, such as tablet computers, wearable electronic devices, in-vehicle electronic devices, augmented reality (AR) devices, virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), intelligent question-answering robots, servers, etc. In this case, the execution subject of the text similarity detection method provided in this application can be any electronic device with text processing capabilities.

[0108] Next, the structure of the electronic device provided in the embodiments of the present application is described.

[0109] In some embodiments, the electronic device provided by the present application is a mobile phone as an example for description.

[0110] As shown in Figure 5, the mobile phone 10 may include a processor 110, a power module 140, a memory 180, a mobile communication module 130, a wireless communication module 120, a sensor module 190, an audio module 150, a camera 170, an interface module 160, a button 101 and a display screen 102, etc.

[0111] It should be understood that the illustrated structure of the embodiment of the present invention does not constitute a specific limitation on the mobile phone 10. In other embodiments of the present application, the mobile phone 10 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0112] The processor 110 may include one or more processing units, for example, a processing module or processing circuit such as a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a microprocessor (MCU), an artificial intelligence (AI) processor, or a programmable logic device (FPGA). Different processing units may be independent devices or integrated into one or more processors. A storage unit may be provided in the processor 110 for storing instructions and data. In some embodiments, the storage unit in the processor 110 is a cache memory 180. For example, the memory 180 may be used to store a candidate text library and a "text-keyword" graph corresponding to the candidate text library. The processor 110 may use the RWR algorithm to calculate the similarity score between the input text and each candidate text based on the "text-keyword" graph.

[0113] The power module 140 may include a power supply, a power management component, and the like. The power supply may be a battery. The power management component manages the charging of the power supply and the supply of power to other modules. In some embodiments, the power management component includes a charging management module and a power management module. The charging management module receives charging input from a charger; the power management module connects the power supply, the charging management module, and the processor 110. The power management module receives input from the power supply and / or the charging management module to power the processor 110, the display 102, the camera 170, and the wireless communication module 120.

[0114] The mobile communication module 130 may include, but is not limited to, an antenna, a power amplifier, a filter, a low noise amplifier (LNA), and the like. The mobile communication module 130 may provide wireless communication solutions for mobile phone 10, including 2G / 3G / 4G / 5G. The mobile communication module 130 may receive electromagnetic waves through the antenna, filter, amplify, and perform other processing on the received electromagnetic waves, and transmit them to the modem processor for demodulation. The mobile communication module 130 may also amplify the signals modulated by the modem processor and convert them into electromagnetic waves for radiation via the antenna. In some embodiments, at least some of the functional modules of the mobile communication module 130 may be located in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 130 may be located in the same device as at least some of the modules of the processor 110. Wireless communication technologies may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), and the like.

[0115] The wireless communication module 120 may include an antenna and transmit and receive electromagnetic waves via the antenna. The wireless communication module 120 may provide wireless communication solutions for the mobile phone 10, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), and the like. The mobile phone 10 may communicate with the network and other devices through wireless communication technologies.

[0116] In some embodiments, the mobile communication module 130 and the wireless communication module 120 of the mobile phone 10 may also be located in the same module.

[0117] The display screen 102 is used to display a human-computer interaction interface, images, videos, etc. The display screen 102 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.

[0118] The sensor module 190 may include a proximity sensor, a pressure sensor, a gyro sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, and the like.

[0119] The audio module 150 is used to convert digital audio information into analog audio signal output, or convert analog audio input into digital audio signal. The audio module 150 can also be used to encode and decode audio signals. In some embodiments, the audio module 150 can be provided in the processor 110, or some functional modules of the audio module 150 can be provided in the processor 110. In some embodiments, the audio module 150 can include a speaker, an earpiece, a microphone, and a headphone jack.

[0120] Camera 170 is used to capture still images or video. The lens generates an optical image of an object and projects it onto a photosensitive element. The photosensitive element converts the optical signal into an electrical signal, which is then passed to an image signal processing unit (ISP) for conversion into a digital image signal. Mobile phone 10 implements its camera function through the ISP, camera 170, video codec, graphics processing unit (GPU), display 102, and application processor.

[0121] The interface module 160 includes an external memory interface, a universal serial bus (USB) interface, and a subscriber identification module (SIM) card interface. The external memory interface can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the mobile phone 10. The external memory card communicates with the processor 110 via the external memory interface to implement data storage. The USB interface is used for communication between the mobile phone 10 and other electronic devices. The subscriber identification module card interface is used to communicate with the SIM card installed in the mobile phone 1010, for example, to read the phone number stored in the SIM card or write the phone number to the SIM card.

[0122] In some embodiments, the mobile phone 10 further includes buttons 101, a motor, and an indicator. The buttons 101 may include a volume button, an on / off button, and the like. The motor is used to vibrate the mobile phone 10, for example, when a user's mobile phone 10 is called, to prompt the user to answer the call. The indicator may include a laser pointer, a radio frequency indicator, an LED indicator, and the like.

[0123] In some embodiments, the electronic device provided by the present application is a laptop, tablet computer, intelligent question-answering robot or other electronic devices as an example for description. Referring to FIG6 , a schematic diagram of the structure of an electronic device provided by an embodiment of the present application is shown.

[0124] Specifically, as shown in FIG6 , the electronic device 10a includes: a processor 1101, a memory 1102, a communication interface 1103, a bus 1104, and a display screen 1105. The processor 1101, the memory 1102, the communication interface 1103, and the display screen 1105 are interconnected via the bus 1104. The bus 1104 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus 1104 may be divided into an address bus, a data bus, a control bus, and the like. For ease of illustration, FIG6 shows only one thick line, but this does not mean that there is only one bus or one type of bus.

[0125] In one example, display screen 1105 can receive user input text and display text similarity detection results for the input text, such as text similar to the input text. Memory 1102 is used to store, for example, a candidate text library and a corresponding "text-keyword" graph. Processor 1101 is used to calculate similarity scores between the input text and each candidate text based on the "text-keyword" graph using the RWR algorithm. Communication interface 1103 is used to obtain the candidate text library from other devices, such as a server.

[0126] In addition, the electronic device provided in this application may include a hardware system. In this case, the execution subject of the text similarity detection method in this application may be the hardware system in the electronic device.

[0127] 7 is a schematic diagram of a hardware system in an electronic device according to the present application. In one embodiment, the hardware system 1400 shown in FIG7 may include one or more processors 1404, a system control logic 1408 connected to at least one of the processors 1404, a system memory 1412 connected to the system control logic 1408, a non-volatile memory (NVM) 1416 connected to the system control logic 1408, and a network interface 1420 connected to the system control logic 1408.

[0128] In some embodiments, processor 1404 may include one or more single-core or multi-core processors. In some embodiments, processor 1404 may include any combination of general-purpose processors and specialized processors (e.g., graphics processors, application processors, baseband processors, etc.). In embodiments where system 1400 employs an enhanced base station (evolved node B, eNB) 101 or a radio access network (RAN) controller 102, processor 1404 may be configured to execute various embodiments, for example, one or more of the multiple embodiments shown in Figures 2-4.

[0129] In some embodiments, system control logic 1408 may include any suitable interface controller to provide any suitable interface to at least one of processors 1404 and / or any suitable device or component in communication with system control logic 1408 .

[0130] In some embodiments, the system control logic 1408 may include one or more memory controllers to provide an interface to the system memory 1412. The system memory 1412 may be used to load and store data and / or instructions. In some embodiments, the memory 1412 of the system 1400 may include any suitable volatile memory, such as a suitable dynamic random access memory (DRAM).

[0131] NVM / memory 1416 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, NVM / memory 1416 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as at least one of a hard disk drive (HDD), a compact disc (CD) drive, and a digital versatile disc (DVD) drive.

[0132] NVM / storage 1416 may include a portion of the storage resources on the device on which system 1400 is installed, or it may be accessible to the device but not necessarily part of the device. For example, NVM / storage 1416 may be accessed over a network via network interface 1420 .

[0133] In particular, system memory 1412 and NVM / storage 1416 may include, respectively, a temporary copy and a permanent copy of instructions 1424. Instructions 1424 may include instructions that, when executed by at least one of processors 1404, cause system 1400 to implement the methods illustrated in Figures 2-4. In some embodiments, instructions 1424, hardware, firmware, and / or software components thereof may additionally or alternatively reside in system control logic 1408, network interface 1420, and / or processor 1404.

[0134] The network interface 1420 may include a transceiver for providing a radio interface for the system 1400 to communicate with any other suitable devices (such as a front-end module, an antenna, etc.) via one or more networks. In some embodiments, the network interface 1420 may be integrated with other components of the system 1400. For example, the network interface 1420 may be integrated with at least one of the processor 1404, the system memory 1412, the NVM / storage 1416, and a firmware device (not shown) having instructions. When at least one of the processors 1404 executes the instructions, the system 1400 implements the methods shown in Figures 2-4.

[0135] The network interface 1420 may further include any suitable hardware and / or firmware to provide a multiple-input multiple-output radio interface. For example, the network interface 1420 may be a network adapter, a wireless network adapter, a telephone modem, and / or a wireless modem.

[0136] In one embodiment, at least one of the processors 1404 may be packaged together with logic for one or more controllers of the system control logic 1408 to form a system in a package (SiP). In one embodiment, at least one of the processors 1404 may be integrated on the same die with logic for one or more controllers of the system control logic 1408 to form a system on chip (SoC).

[0137] System 1400 may further include input / output (I / O) devices 1432. I / O devices 1432 may include a user interface to enable a user to interact with system 1400, and peripheral component interfaces to enable peripheral components to interact with system 1400. In some embodiments, system 1400 may further include sensors for determining at least one of environmental conditions and location information related to system 1400.

[0138] In some embodiments, the user interface may include, but is not limited to, a display (e.g., an LCD display, a touch screen display, etc.), a speaker, a microphone, one or more cameras (e.g., a still image camera and / or a video camera), a flashlight (e.g., an LED flash), and a keyboard.

[0139] In some embodiments, the peripheral component interface may include, but is not limited to, a non-volatile memory port, an audio jack, and a power interface.

[0140] In some embodiments, the sensors may include, but are not limited to, a gyroscope sensor, an accelerometer, a proximity sensor, an ambient light sensor, and a positioning unit. The positioning unit may also be part of or interact with the network interface 1420 to communicate with components of a positioning network (e.g., Global Positioning System (GPS) satellites).

[0141] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0142] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0143] Program code can be implemented with a high-level programming language or an object-oriented programming language to communicate with the processing system. Where necessary, program code can also be implemented in assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any particular programming language. In either case, the language can be a compiled language or an interpreted language.

[0144] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments can also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, instructions can be distributed over a network or through other computer-readable media. Therefore, machine-readable media can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to floppy disks, optical disks, optical discs, read-only memories (CD-ROMs), magneto-optical disks, read-only memories (ROMs), random access memories (RAMs), erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in electrical, optical, acoustic, or other forms of propagation signals. Therefore, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0145] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or order may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of a structural or method feature in a particular figure does not imply that such feature is required in all embodiments, and in some embodiments, such features may not be included or may be combined with other features.

[0146] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems raised by this application. In addition, in order to highlight the innovative part of this application, the above-mentioned device embodiments of this application do not introduce units / modules that are not closely related to solving the technical problems raised by this application. This does not mean that other units / modules do not exist in the above-mentioned device embodiments.

[0147] It should be noted that in the examples and description of this patent, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "including a" does not exclude the presence of other identical elements in the process, method, article or device that includes the element.

[0148] Although the present application has been shown and described with reference to certain preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the application.

Claims

1. A text similarity detection method, characterized in that The method includes: Obtain the input text; Extract keywords from the input text to obtain a keyword set; Determine a node graph between a plurality of preset candidate texts and a plurality of candidate keywords, where the plurality of candidate keywords are keywords extracted from the plurality of candidate texts. The node graph uses the plurality of candidate texts as one type of node, the plurality of keywords as another type of node, the inclusion relationship between the plurality of candidate texts and the plurality of candidate keywords as edges, and the first score of a candidate keyword of a candidate text as the weight of the edge between the candidate text and the candidate keyword; According to the node graph, calculate the correlation scores between each keyword in the keyword set and each candidate text; According to the correlation scores between each keyword in the keyword set and each candidate text, calculate the similarity scores between the input text and each candidate text; Take the k candidate texts with the highest similarity scores as the similar texts of the input text, where k is a positive integer.

2. The method according to claim 1, wherein The relevance scores between each keyword in the keyword set and each of the candidate texts are calculated iteratively through the following formula: R l (c, t) = a(1 - a) l W c,t + R l-1 (c, t); where R l (c, t) is the relevance score between the keyword c and the text node t, l is the number of iterations of the restart random walk RWR algorithm represented by the formula, a is the restart factor and a ∈ (0, 1), W c,t represents the transition probability between the keyword c and the text node t, C represents the keyword set, c is a keyword in the keyword set C, and t is a candidate text among the multiple candidate texts.

3. The method according to claim 2, wherein The similarity score between the input text and each of the candidate texts is calculated by the following formula: Among them, Denote the similarity score between text p and text t. S(c) is the weight of keyword c of text p calculated by the TextRank algorithm. p is the input text, C represents the keyword set, c is a keyword in the keyword set C, and t is a text in the plurality of candidate texts.

4. The method according to claim 1, wherein The node graph is obtained based on the following method: Obtain the plurality of candidate texts; Based on the word segmentation dictionary and stop words corresponding to the plurality of candidate texts, use a word segmenter to segment each candidate text in the plurality of candidate texts to obtain the words of each candidate text; Determine the co-occurrence relationship between the words of each candidate text according to the set window size, and based on the co-occurrence relationship between the words, construct a word graph of the words in each candidate text; Extract keywords from the word graph of each candidate text to obtain the plurality of candidate keywords and the first score of each candidate keyword, where the first score is used to reflect the importance of a keyword to a text; Construct the node graph according to the inclusion relationship between each candidate text and the corresponding candidate keyword, and the first score of each candidate keyword in each candidate text; Among them, the weight between a candidate text and a connected candidate keyword in the node graph is the first score of the candidate keyword for the candidate text.

5. The method according to claim 4, characterized in that, The candidate keywords of a candidate text are the first m words with relatively large first scores among all the words in the word graph of the candidate text, where m is a positive integer.

6. The method according to claim 1, wherein The extracting keywords from the input text to obtain a keyword set includes: Based on the word segmentation dictionary and stop words corresponding to the input text, use a word segmenter to segment the input text to obtain the words of the input text; Determine the co-occurrence relationship between the words of the input text according to the set window size, and based on the co-occurrence relationship between the words, construct a word graph of the words in the input text; Obtain the first scores of all the words in the word graph of the input text, where the first score is used to reflect the importance of a keyword to a text; Add the first m words with relatively large first scores among all the words in the word graph of the input text to the keyword set of the input text, where m is a positive integer.

7. The method according to any one of claims 1 to 6, characterized in that, The first fraction is calculated as follows: Among them, S(v i ) represents the first score of node v i . S(v j ) represents the first score of node v j . In(v i ) represents the set of in-nodes of v i . Out(v j ) represents the set of out-nodes of v j . Nodes v i and v j are respectively two word nodes in the word graph of a text. The weight of the edge between nodes v i and v j is denoted as w ji . d is the damping coefficient and its value range is from 0 to 1. V k′ is another node in the said text. w jk′ is the weight between node V k′ and node v j .

8. The method according to claim 1, wherein The field to which the input text belongs is the same as the fields to which the multiple candidate texts belong.

9. A readable medium, characterized in that, The readable medium stores instructions that, when executed on an electronic device, cause the electronic device to perform the text similarity detection method according to any one of claims 1 to 8.

10. An electronic device, characterized in that, Comprising: A memory for storing instructions executed by one or more processors of an electronic device, and a processor, which is one of the processors of the electronic device, for performing the text similarity detection method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • A TextRank-based keyword extraction method and device

    CN109918660A

  • Text retrieval method and device

    CN110019668A

  • Text vector generation method based on unsupervised graph neural network structure

    CN110705260A

  • Keyword extraction method and device

    CN112926310A

  • Electronic equipment device, method for calculating similarity between texts, and program

    JP2005122515A