Multilingual text similarity comparison method

By constructing undirected graphs and mapped dictionaries of multilingual comparison dictionaries, the problem of low efficiency in multilingual text similarity comparison is solved, and efficient cross-language text comparison is achieved, which is suitable for complex documents with multilingual cross-sections.

CN120068842AActive Publication Date: 2025-05-30TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510525210.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

It is difficult for the prior art to efficiently perform similarity comparison of multilingual texts, especially in complex documents with multilingual crossovers, where existing methods have long response time, slow speed and high cost.

Method used

By constructing an undirected graph and a mapped dictionary based on a multilingual comparison dictionary, segmenting the text into multiple alignment units, and determining the target mapping set of terms based on the undirected graph and mapped dictionary, cross-language text similarity comparison is achieved.

Benefits of technology

It realizes efficient similarity comparison of multilingual texts, reduces development costs, is highly applicable, and can break language barriers in academic creation and avoid malicious plagiarism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068842A_ABST
    Figure CN120068842A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of natural language processing, particularly relates to a multilingual text similarity comparison method, and aims to solve the problems of high difficulty and low efficiency in comparison of multilingual crossed complex texts. The method comprises the steps of obtaining an undirected graph constructed by a contrast dictionary between different languages and a mapping dictionary constructed by nodes and edges; splitting the first text and the second text into text comparison units; determining a target mapping set corresponding to each lexical item after word segmentation of the text comparison unit through an undirected graph and a mapping dictionary; determining a candidate comparison unit of any text comparison unit in the first text in the second text through the target mapping set; and performing similarity comparison according to the coverage proportion of the hit identifiers in any text comparison unit and the candidate comparison unit. According to the method, multi-language crossed complex text comparison can be efficiently realized, language barriers are broken, cross-language malicious plagiarism is avoided, and globalized academic ecology is maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] By comparing the similarity of texts in different languages, tasks such as cross - language plagiarism detection and cross - language retrieval can be achieved. The existing methods for comparing the similarity of texts in different languages mainly utilize the semantic associations between different languages.

[0003] For example, by using a machine translation model based on deep learning to align the corpus data of two texts in different languages to be compared, the translation from the source - language text to the target - language text is realized, so as to convert two texts in different languages into texts in the same language, and then the similarity of texts in the same language is compared.

[0004] For example, based on the aligned corpus of two texts in different languages to be compared, by separately performing word segmentation and pre - processing on the source text and the target text, each word or phrase in the source - language text is mapped to the corresponding word or phrase in the target language, and then the similarity between the two texts is calculated.

[0005] The comparison method based on the aligned corpus can only compare two languages simultaneously and cannot adapt to the comparison of complex documents with multi - language intersections. And for large - scale multi - language similarity comparison, the efficiency requirement for text comparison is extremely high. The existing deep - learning - based methods have a long comparison response time, slow speed, and high comparison cost. Summary of the Invention

[0006] To solve the above problems in the prior art, that is, the problem of high difficulty and low efficiency in comparing complex texts with multi - language intersections, in the first aspect of this application, a method for comparing the similarity of multi - language texts is proposed. The method includes: Step S10: Obtain an undirected graph constructed based on a multi - language control dictionary set. Among them, any language in the multi - language control dictionary set corresponds to at least one control dictionary. The nodes of the undirected graph are words in any language, and the edges of the undirected graph are the control relationships between nodes in different languages; Step S20: Obtain a mapping dictionary constructed from a mapping set of nodes and nodes, where the mapping set includes the identifiers of each edge connected by the nodes; Step S30: Respectively perform text segmentation on the first text and the second text to be compared to generate multiple text comparison units; Step S40: Based on the nodes, perform word segmentation on each text comparison unit to obtain multiple terms, and determine the target mapping set corresponding to any term based on the undirected graph and the mapping dictionary; Step S50: Based on the target mapping set, determine the candidate comparison unit of any text comparison unit in the first text in the second text, where the identifier that hits in both any text comparison unit and the candidate comparison unit is the hit identifier; Step S60: Determine the first hit word of the hit identifier in any text comparison unit and the second hit word of the hit identifier in the candidate comparison unit; Step S70: Perform a similarity comparison between any text comparison unit and the candidate comparison unit according to the first coverage ratio of the first hit word in any text comparison unit and the second coverage ratio of the second hit word in the candidate comparison unit.

[0007] In some embodiments, obtaining an undirected graph constructed based on a multilingual comparison dictionary set includes: Taking words in any language as nodes, and connecting two nodes with a control relationship based on the comparison dictionary to generate a reference edge; Obtain the number of languages to which other nodes connected to any node belong, and the number of reference edges connected to any node; When the number of languages and the number of reference edges are the same, determine that the other nodes are translation pairs; Connect the nodes that are translation pairs to each other to generate extended edges; Generate an undirected graph from the nodes, reference edges, and extended edges.

[0008] In some embodiments, obtaining an undirected graph constructed based on a multilingual comparison dictionary set further includes: Determine the translation weight of the reference edge; Based on the translation weight of the reference edge and a preset attenuation coefficient, determine the translation weight of the extended edge.

[0009] In some embodiments, when the number of languages and the number of reference edges are the same, determining that the other nodes are translation pairs includes: When the number of languages and the number of reference edges are the same, and the number of reference edges spanned between the other nodes is less than or equal to a preset cross-edge threshold, determine that the other nodes are translation pairs.

[0010] In some embodiments, based on a target mapping set, determining a candidate comparison unit in a second text for any text comparison unit in a first text includes: Taking the text comparison unit as a level, establish an inverted index for the second text according to the identifiers in each target mapping set, and determine the text comparison unit corresponding to any identifier in the second text; Number the text comparison units corresponding to each identifier in the second text and form a unit number set; Count the unit numbers corresponding to any reference identifier, where the reference identifier is any identifier in the target mapping set corresponding to the term in the first text; Based on the number of times the unit number appears in the statistical process and the corresponding translation weight, determine the similarity ratio of any unit number; Select candidate comparison units according to the similarity ratio.

[0011] In some embodiments, determining the similarity ratio of any unit number based on the number of occurrences of the unit number in the statistical process and the corresponding translation weight includes: When the unit number appears, accumulate the corresponding translation weights to obtain an accumulated weight; Calculate the ratio of the accumulated weight to the number of elements in the unit number set as the similarity ratio.

[0012] In some embodiments, selecting candidate comparison units according to the similarity ratio includes: When the similarity ratio of any unit number is greater than or equal to a preset similarity threshold, select the text comparison unit corresponding to any unit number as the candidate comparison unit.

[0013] In some embodiments, determining the target mapping set corresponding to any term based on the undirected graph and the mapping dictionary includes: Select the node where any term is located in the undirected graph, and the edges connected by the node based on the languages of the first text and the second text to form a subgraph; Generate a target mapping set based on the nodes in the subgraph and the identifiers of the edges connected by the nodes.

[0014] In some embodiments, obtaining the first coverage ratio and the second coverage ratio includes: Obtain the first union of the first hit words in any text comparison unit, and count the number of characters in the first union as the first hit coverage length of the first hit words in any text comparison unit; Calculate the ratio of the first hit coverage length to the number of characters in any text comparison unit as the first coverage ratio; Obtain the second union of the second hit words in the candidate comparison units, and count the number of characters in the second union as the second hit coverage length of the second hit words in the candidate comparison units; Calculate the ratio of the second hit coverage length to the number of characters in the candidate comparison unit as the second coverage ratio.

[0015] In some embodiments, performing a similarity comparison between any text comparison unit and the candidate comparison unit includes: When the first coverage ratio of the first hit words in any text comparison unit is greater than the first preset ratio and the second coverage ratio of the second hit words in the candidate comparison unit is greater than the second preset ratio, determine that any text comparison unit is similar to the candidate comparison unit.

[0016] Advantages of this application: (1) First, an undirected graph constructed by obtaining a comparison dictionary between different languages and a mapping dictionary constructed by nodes and edges are used as tools for similarity comparison between different languages. Then, the first text and the second text are split into small text comparison units to facilitate similarity comparison. Further, through the undirected graph and the mapping dictionary, the target mapping set corresponding to each term after the text comparison unit is segmented is determined. The identifier of each edge in the target mapping set represents the comparison relationship of the term between different languages, and the term can be located in the first text and the second text respectively through the identifier. Then, the candidate comparison unit of any text comparison unit in the first text in the second text is determined through the target mapping set, realizing the preliminary positioning of similar text segments. Finally, similarity comparison is performed according to the coverage ratio of the hit identifiers in any text comparison unit and the candidate comparison unit, completing the cross-language text similarity comparison. Since the undirected graph includes multiple languages, the present application can directly perform similarity comparison on multi-language texts based on the undirected graph, and can realize complex text comparison across multiple languages without translation conversion, breaking language barriers in academic creation, avoiding malicious cross-language plagiarism, and maintaining the global academic ecosystem.

[0017] (2) Since the undirected graph and the mapping dictionary are generated and stored in advance, they can be directly called during similarity comparison, without involving high-parameter models based on deep learning, belonging to basic numerical operations, improving the efficiency of cross-language text similarity comparison and reducing the development cost.

[0018] (3) In addition, the undirected graph can be expanded and assembled by increasing or decreasing the comparison dictionary, and does not require major changes to the method structure, which can meet the needs of different users, has strong applicability and a wide application range. Brief Description of the Drawings

[0019] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives and advantages of the present application will become more obvious: Figure 1 is a schematic flowchart of a multi-language text similarity comparison method provided by an embodiment of the present application; Figure 2 is a schematic diagram of an undirected graph provided by an embodiment of the present application; Figure 3 is a step flowchart of a multi-language text similarity comparison method provided by an embodiment of the present application; Figure 4 is another schematic diagram of an undirected graph provided by an embodiment of the present application; Figure 5 is another step flowchart of a multi-language text similarity comparison method provided by an embodiment of the present application; Figure 6 It is another flowchart of steps of a multi - language text similarity comparison method provided by an embodiment of the present application; Figure 7 It is another flowchart of steps of a multi - language text similarity comparison method provided by an embodiment of the present application; Figure 8 It is a system block diagram of a multi - language text similarity comparison system provided by an embodiment of the present application; Figure 9 It is a schematic structural diagram of a computer system of a server for implementing the method, system, and device embodiments of the present application. Detailed implementation manners

[0020] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.

[0021] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.

[0022] A multi - language text similarity comparison method according to the first embodiment of the present application includes steps S10 - S70, as Figure 1 shown, the method includes: Step S10, obtaining an undirected graph constructed based on a multi - language control dictionary set, wherein any language in the multi - language control dictionary set corresponds to at least one control dictionary, the nodes of the undirected graph are words in any language, and the edges of the undirected graph are control relationships between nodes in different languages.

[0023] Optionally, a control dictionary refers to a reference tool that arranges the vocabulary and phrases of two or more languages in correspondence for translation and understanding between different languages, such as a Chinese - English dictionary, a Chinese - French dictionary, etc.

[0024] In the embodiments of the present application, the control dictionary set includes multiple control dictionaries, and any language corresponds to at least one control dictionary.

[0025] As an example, the control dictionary set M includes a Chinese - English dictionary A_B, a Chinese - Russian dictionary A_C, an English - Russian dictionary B_C, and a Russian - Italian dictionary C_D, which can be denoted as:

[0026] In the above example, for Chinese, there are a Chinese-English dictionary A_B and a Chinese-Russian dictionary A_C; for English, there are a Chinese-English dictionary A_B and an English-Russian dictionary B_C; for Russian, there are an English-Russian dictionary B_C and a Russian-Italian dictionary C_D; and for Italian, there is a Russian-Italian dictionary C_D.

[0027] Further, construct a translation comparison undirected graph according to the set of comparison dictionaries.

[0028] As an example, the Chinese-English dictionary A_B includes comparison pairs: a1-b1, a1-b2, where a1 is a Chinese word and b1 and b2 are English words respectively.

[0029] The Chinese-Russian dictionary A_C includes a comparison pair: a2-c3, where a2 is a Chinese word and c3 is a Russian word.

[0030] The English-Russian dictionary B_C includes comparison pairs: b1-c1, b2-c2, b2-c3, where c1, c2, and c3 are Russian words respectively.

[0031] The Russian-Italian dictionary C_D includes comparison pairs: c1-d3, c2-d1, c3-d2, c3-d5, where d1, d2, d3, and d5 are Italian words respectively.

[0032] It can be understood that the Chinese word a1 can form comparison pairs with the English words b1 and b2 respectively, that is, the Chinese word a1 can be translated into the English words b1 and b2 respectively.

[0033] Further, use the words in any language as nodes and the comparison relationships between the nodes in different languages as edges to generate an undirected graph of the set of comparison dictionaries.

[0034] As an example, according to the comparison pairs included in each comparison dictionary in the above set of comparison dictionaries M, generate an undirected graph as shown in Figure 2 shown, where each edge is corresponding to a digital identifier for the purpose of differentiating the edges.

[0035] In other embodiments, the comparison dictionary may also include different variants within the same language, such as the comparison between simplified Chinese and traditional Chinese, or the comparison of professional terms in different fields.

[0036] Step S20: Obtain a mapping dictionary constructed from the nodes and the mapping set of the nodes, where the mapping set includes the identifiers of each edge connected to the nodes.

[0037] Optionally, for any node, there is at least one edge connected to it, such as Figure 2As shown, taking node a1 as an example, in an undirected graph, node b1 and node b2 are connected. The identifier of the edge between node a1 and node b1 is 0, and the identifier of the edge between node a1 and node b2 is 2. Then the mapping set of node a1 is [0, 2].

[0038] Further, obtain the mapping sets of each node and construct a mapping dictionary with the corresponding nodes. Taking Figure 2 the undirected graph as an example, the obtained mapping dictionary is as follows: { a1: [0, 2] a2: [3] b1: [0, 1] b2: [2, 4, 5] c1: [1, 8] c2: [4, 6] c3: [3, 5, 7, 9] d1: [6] d2: [7] d3: [8] d5: [9] } It should be noted that the mapping set of a node is the statistics of the identifiers of the edges connected to the node. The mapping dictionary composed of each node and its mapping set is an icon conversion of the undirected graph, which is convenient for searching and positioning.

[0039] Step S30: Perform text segmentation on the first text and the second text to be compared respectively, and generate multiple text comparison units.

[0040] Optionally, text segmentation is the process of dividing continuous text into smaller and more meaningful paragraphs or units according to certain rules or standards. It plays an important role in the fields of natural language processing, document processing, information retrieval, etc.

[0041] In the embodiments of the present application, text segmentation is performed according to actual needs. For example, text segmentation is performed according to rules such as sentences, fragments, paragraphs, and chapters, and different text segmentation tools are selected according to different rules to implement.

[0042] It should be noted that the first text and the second text can either use different languages or the same language, that is, the present application can perform cross-language text similarity comparison or same-language text similarity comparison.

[0043] As an example, perform text segmentation on the first text L in Chinese according to paragraphs, and obtain the text comparison unit set {L 0 , L 1 ,..., L N}, where L 0,L 1 ,...,L N are different text comparison units respectively; segment the second text R in English by paragraphs to obtain a set of text comparison units {R 0 ,R 1 ,...,R N}, where R 0 ,R 1 ,...,R N are different text comparison units respectively.

[0044] Step S40, segment each text comparison unit based on the nodes to obtain multiple terms, and determine the target mapping set corresponding to any term based on the undirected graph and the mapping dictionary.

[0045] Optionally, for any text comparison unit of the first text and the second text, segment the text by matching any text comparison unit with the nodes.

[0046] As an example, the original text of any text comparison unit is: He came to Shanghai Jiao Tong University. The words corresponding to the nodes include: He, came, Shanghai, Jiao Tong, University, Shanghai Jiao Tong University. By matching with the nodes, the segmentation result is: He / came / Shanghai / Shanghai Jiao Tong University / Jiao Tong / University. The terms obtained after segmentation include: He, came, Shanghai, Jiao Tong, University, Shanghai Jiao Tong University.

[0047] It should be noted that the above example is only a simple example. In other embodiments, the text comparison unit for segmentation can be a longer or shorter paragraph.

[0048] Furthermore, determine the target mapping set corresponding to any term by determining the identifiers of the edges connected by the node corresponding to any term in the undirected graph.

[0049] As an example, by searching the undirected graph, it is determined that the identifier of the edge connected by the node corresponding to "He" in the undirected graph is 0, the identifier of the edge connected by the node corresponding to "came" in the undirected graph is 1, the identifier of the edge connected by the node corresponding to "Shanghai" in the undirected graph is 2, the identifier of the edge connected by the node corresponding to "Shanghai Jiao Tong University" in the undirected graph is 3, the identifier of the edge connected by the node corresponding to "Jiao Tong" in the undirected graph is 4, and the identifier of the edge connected by the node corresponding to "University" in the undirected graph is 5.

[0050] Then the target mapping set corresponding to the term "He" is {0}, the target mapping set corresponding to the term "came" is {1}, the target mapping set corresponding to the term "Shanghai" is {2}, the target mapping set corresponding to the term "Shanghai Jiao Tong University" is {3}, the target mapping set corresponding to the term "Jiao Tong" is {4}, and the target mapping set corresponding to the term "University" is {5}.

[0051] Step S50: Based on the target mapping set, determine the candidate matching unit in the second text for any text matching unit in the first text. Among them, the identifier that is hit in both any text matching unit and the candidate matching unit is the hit identifier.

[0052] Optionally, the undirected graph edge connects two nodes with a control relationship. After determining multiple target mapping sets included in each text matching unit, for any text matching unit in the first text, through the identifiers of each edge in the target mapping set, the corresponding candidate matching unit can be found in the second text.

[0053] As an example, for any text matching unit to be selected in the second text, count the proportion of the number of identifiers that are hit in both any text matching unit and the to-be-selected matching unit among the number of identifiers included in the to-be-selected matching unit. When this proportion is greater than a certain threshold, determine that the to-be-selected text matching unit is the candidate matching unit of any text matching unit.

[0054] Furthermore, denote the identifier corresponding to the term that is hit in both any text matching unit and the candidate matching unit as the hit identifier.

[0055] It can be understood that the term corresponding to the hit identifier is the identifier that exists in both text matching units, and it is the basis for comparing the similarity of the two text matching units.

[0056] In the embodiment of the present application, the hit identifier is the identifier of the edge that exists in both any text matching unit and the candidate matching unit, that is, the number of hit identifiers is less than or equal to the number of identifiers included in the candidate matching unit, and also less than or equal to the number of identifiers in any text matching unit.

[0057] Step S60: Determine the first hit term of the hit identifier in any text matching unit, and the second hit term of the hit identifier in the candidate matching unit.

[0058] Optionally, through the one-to-one correspondence between the edge and the node, the corresponding terms can be determined in any text matching unit and the candidate matching unit respectively through the identifier, that is, the first hit term in any text matching unit and the second hit term in the candidate matching unit.

[0059] Step S70: Compare the similarity between any text matching unit and the candidate matching unit according to the first coverage ratio of the first hit term in any text matching unit and the second coverage ratio of the second hit term in the candidate matching unit.

[0060] Optionally, the more the hit words cover in the text, the more corresponding similar word items there are. By calculating the covered length of the hit words in the two text comparison units, the hit ratio of the hit words can be determined, which is used to determine the similarity degree of the two text comparison units.

[0061] As an example, in the case where both the first coverage ratio and the second coverage ratio are relatively large, that is, there are more similar word items in the two text comparison units, it is determined that any one of the text comparison units is similar to the candidate comparison unit.

[0062] For example, a similarity threshold is set. In the case where both the first coverage ratio and the second coverage ratio are greater than or equal to the similarity threshold, it is determined that any one of the text comparison units is similar to the candidate comparison unit.

[0063] To more clearly illustrate a multi-language text similarity comparison method of the present application, the following details each step in the embodiments of the present application.

[0064] As a possible implementation manner, please refer to Figure 3 , in the process of constructing an undirected graph in step S10, the following steps are included: Step S101, use the words in any language as nodes, and connect two nodes with a control relationship based on the control dictionary to generate a reference edge.

[0065] As an example, an undirected graph as shown in Figure 2 is generated for the control dictionary set M.

[0066] The process of generating an undirected graph as shown in Figure 2 can refer to the relevant description in step S10, and will not be elaborated here.

[0067] Furthermore, determine the translation weight of the reference edge.

[0068] It should be noted that the two nodes connected by the reference edge are two different language words that can be directly translated, and the control relationship is the most accurate. Determine the translation weight of the reference edge as a translation reference.

[0069] As an example, set the translation weight of the reference edge to 1.

[0070] Step S102, obtain the number of languages to which the other nodes connected to any node belong, and the number of reference edges connected to any node.

[0071] Take Figure 2Taking the node b1 in [[]] as an example, the other nodes connected to the node b1 are a1 and c1. Among them, the language to which a1 belongs is Chinese, the language to which c1 belongs is Russian, and the number of languages to which the other nodes connected to the node b1 belong is 2; there are two reference edges connected to the node b1, and the corresponding identifiers are 0 and 1 respectively, that is, the number of reference edges is also 2.

[0072] Taking Figure 2 the node c3 in [[]] as another example, the other nodes connected to the node c3 are a2, b2, d2 and d5. Among them, the language to which a2 belongs is Chinese, the language to which b2 belongs is English, and the languages to which d2 and d5 belong are Italian. The number of languages to which the other nodes connected to the node c3 belong is 3; there are 4 reference edges connected to the node c3, and the corresponding identifiers are 3, 5, 7 and 9 respectively, that is, the number of reference edges is 4.

[0073] Step S103, when the number of languages and the number of reference edges are the same, determine that the other nodes are translation pairs with each other.

[0074] It can be understood that when the number of languages and the number of reference edges are the same, several nodes connected by the node can be translated into the same meaning, and the mutual translation of the other several nodes can be realized through this node.

[0075] Taking Figure 2 the node b1 in [[]] as an example, the number of languages to which the other nodes connected to the node b1 belong is 2, and the number of reference edges is also 2. It is determined that the node a1 and the node c1 are translation pairs with each other.

[0076] Taking Figure 2 the node c3 in [[]] as another example, the number of languages to which the other nodes connected to the node c3 belong is 3, and the number of reference edges is 4. The number of languages and the number of reference edges are different, and the nodes a2, b2, d2 and d5 cannot be translation pairs with each other.

[0077] In other embodiments, partial determination can also be performed, that is, the number of languages to which some of the other nodes connected to the node belong and the corresponding number of reference edges are compared. If the number of languages and the number of reference edges are the same, some of the other nodes can be translation pairs with each other.

[0078] Similarly taking Figure 2 the node c3 in [[]] as an example, when some of the nodes connected to the node c3 are a2 and b2, the language to which a2 belongs is Chinese, the language to which b2 belongs is English, the number of languages is 2, and the number of reference edges of the nodes a2 and b2 connected to the node c3 is also 2, then the nodes a2 and b2 can be translation pairs with each other.

[0079] Furthermore, the extension of the translation pair can also be restricted by presetting a cross-edge threshold.

[0080] Optionally, when the number of languages and the number of reference edges are the same, and the number of reference edges spanned between other nodes is less than or equal to a preset cross-edge threshold, it is determined that the other nodes are translation pairs for each other.

[0081] Taking Figure 2 as an example, the ones that can become translation pairs are: a1-b1-c1, b1-c1-d3, b2-c2-d1, and a2-c3-b2. After merging, we can get: a1-b1-c1-d3, b2-c2-d1, and a2-c3-b2. Without restricting the number of cross-edges, node a1 and node d3 can be translated into each other. If the number of cross-edges is restricted, for example, when the preset cross-edge threshold is 2, the number of reference edges spanned between node a1 and node d3 is 3, and node a1 and node d3 cannot be translation pairs for each other. Additionally, when the preset cross-edge threshold is 3, node a1 and node d3 can be translation pairs for each other and be translated into each other.

[0082] In other embodiments, the value of the preset cross-edge threshold can be set according to the actual operation situation.

[0083] Step S104, connect the nodes that are translation pairs for each other to generate extended edges.

[0084] In the embodiment of the present application, the preset cross-edge threshold is 3. Connect the nodes that are translation pairs for each other to generate extended edges. For example, through a1-b1-c1, connect node a1 and node c1 to generate an extended edge.

[0085] It can be understood that the nodes connected by the extended edges are not two directly translated words, but are translated into each other through multiple dictionary conversions. Through the extended edges, preliminary translation comparison of the reference dictionaries not included in the comparison dictionary set can be performed, expanding the text similarity comparison range. For example, for a Chinese-meaning dictionary not included in the comparison dictionary set M, cross-language similarity comparison can be achieved through the method of the present application.

[0086] Furthermore, based on the translation weight of the reference edge and a preset attenuation coefficient, determine the translation weight of the extended edge.

[0087] Optionally, for the newly added translation pairs connected by the extended edges, relay translation is performed through multiple comparison dictionaries, and there may be a loss in translation accuracy. Therefore, an equivalent attenuation coefficient is set to re-determine the translation weight of the extended edge.

[0088] In the embodiment of the present application, the attenuation coefficient is set to , the number of reference edges spanned between the two nodes connected by the extended edge is , then the translation weight of the extended edge is , where 1 is the translation weight of the reference edge.

[0089] As an example, set the attenuation coefficient to 0.5, and the number of cross-edges between node a1 and node c1 is 2. Then the translation weight of the extended edge connecting node a1 and node c1 is 1 * 0.5 / 2 = 0.25.

[0090] Step S105, generate an undirected graph from nodes, reference edges, and extended edges.

[0091] Optionally, by adding extended edges to the undirected graph as Figure 2 shown, and then performing position adjustment to generate the final undirected graph, as Figure 4 shown.

[0092] As a possible implementation, please refer to Figure 5 , in step S40, determine the target mapping set corresponding to any term based on the undirected graph and the mapping dictionary, including the following steps: Step S401, select the node where any term is located in the undirected graph, and the edges connected by the node based on the languages of the first text and the second text, to form a subgraph.

[0093] Optionally, in the undirected graph as Figure 4 shown, which includes nodes in multiple languages, filter the node where any term is located after word segmentation in the undirected graph, and the edges that only include the languages to which the first text and the second text belong, and extract the subgraph as a tool for text similarity comparison.

[0094] As an example, if the language to which the first text belongs is Chinese and the language to which the second text belongs is English, then for Figure 4 the undirected graph shown, the subgraph obtained after extraction only retains nodes a1, node a2, node b1, and node b2, and the edges labeled 0, 2, and 14.

[0095] It should be noted that the embodiment of this application is a simple example. In actual use, the scope of the undirected graph includes all words in the standard dictionary, and professional terms and other vocabulary in different fields can also be added.

[0096] Step S402, based on the nodes in the subgraph and the identifiers of the edges connected by the nodes, generate the target mapping set.

[0097] Optionally, by searching the undirected graph, determine the target mapping set corresponding to any node in the subgraph.

[0098] For example, for the sub - graph in step S501, the target mapping set corresponding to node a1 is {0, 2}, the target mapping set corresponding to node a2 is {14}, the target mapping set corresponding to node b1 is {0}, and the target mapping set corresponding to node b2 is {2, 14}.

[0099] As a possible implementation, please refer to Figure 6 , the process of determining the candidate comparison unit of any text comparison unit in the first text in the second text in step S50 includes the following steps: Step S501: Taking the text comparison unit as the level, according to the identifiers in each target mapping set, establish an inverted index for the second text to determine the text comparison unit corresponding to any identifier in the second text.

[0100] Optionally, the inverted index is also called a reverse index. By changing the form of "document - word" to the form of "word - document", the retrieval efficiency is improved.

[0101] In the embodiment of the present application, the form of "text comparison unit - identifier" is changed to the form of "identifier - text comparison unit".

[0102] As an example, the second text R is divided into 4 text comparison units {R0, R1, R2, R3}, and the identifier sets corresponding to each text comparison unit generate a text comparison unit - identifier comparison table, as shown in Table 1: Table 1

[0103] By establishing an inverted index for each identifier in the identifier set, an identifier - text comparison unit comparison table is generated, as shown in Table 2: Table 2

[0104] Through the identifier - text comparison unit comparison table, any identifier in the text comparison unit can be quickly located.

[0105] Step S502: Number the text comparison units corresponding to each identifier in the second text and form a unit number set.

[0106] As an example, according to the identifier - text comparison unit comparison table shown in step S501, the unit number set corresponding to any identifier can be obtained. For example, the unit number set corresponding to identifier 0 is {R0}, the unit number set corresponding to identifier 1 is {R0}, and the unit number set corresponding to identifier 2 is {R0, R1}.

[0107] Step S503, count the unit numbers corresponding to any reference identifier, where the reference identifier is any identifier in the target mapping set corresponding to the term of the first text.

[0108] Optionally, the identifier represents the corresponding relationship between two languages, and through the identifier, the terms corresponding to each other in the first text and the second text can be determined.

[0109] As an example, if the reference identifiers in the first text are {2, 3, 6, 10}, then through the identifier-text comparison unit comparison table, the unit numbers corresponding to any reference identifier can be obtained: reference identifier 2: {R0 R1}, reference identifier 3: {R0 R1 R2}, reference identifier 6: {R1 R2}, reference identifier 10: {R3}.

[0110] By determining the text comparison unit corresponding to the reference identifier in the second text, the position where similar texts exist can be initially locked.

[0111] Step S504, determine the similarity ratio of any unit number based on the number of times the unit number appears during the statistics process and the corresponding translation weight.

[0112] Optionally, when the unit number appears, accumulate the corresponding translation weights to obtain the cumulative weight; calculate the ratio of the cumulative weight to the number of elements in the unit number set as the similarity ratio.

[0113] The calculation formula for the similarity ratio S of any unit number is as follows:

[0114] where H represents the number of times any unit number appears during the statistics process, represents the translation weight corresponding to the k-th appearance of any unit number, represents the cumulative weight of any unit number, represents the number of elements in the unit number set.

[0115] For example, for reference identifier 2: {R0 R1}, reference identifier 3: {R0 R1 R2}, reference identifier 6: {R1 R2}, reference identifier 10: {R3}, the number of elements in the unit number set {R0 R1 R2 R3} is 4, the number of times unit number R0 appears is 2, the number of times unit number R1 appears is 3, the number of times unit number R2 appears is 2, the number of times unit number R3 appears is 1. Assuming that the translation weight corresponding to any unit number each time it appears is 1, then the similarity ratio of unit number R0 , the similarity ratio of unit number R1 , the similarity ratio of unit number R2 , the similarity ratio of unit number R3 。

[0116] It should be noted that the similarity ratio represents the ratio of similarity between the text comparison unit corresponding to any unit number and any text comparison unit in the first text. When the similarity ratio is large, it indicates that the text comparison unit may be similar to any text comparison unit in the first text.

[0117] Step S505, select candidate comparison units according to the similarity ratio.

[0118] Optionally, when the similarity ratio of any unit number is greater than or equal to a preset similarity threshold, select the text comparison unit corresponding to the unit number as the candidate comparison unit.

[0119] As an example, the preset similarity threshold can be set to 0.7. In other embodiments, the value can also be set according to the actual situation.

[0120] Among them, the similarity ratio of unit number R1 is 0.75 > 0.7, then the text comparison unit corresponding to unit number R1 is listed as the candidate comparison unit.

[0121] As a possible implementation manner, please refer to Figure 7 , in step S70, the process of the first coverage ratio, the second coverage ratio, and the similarity comparison between any text comparison unit and the candidate comparison unit includes the following steps: Step S701, obtain the first union of the first hit words in any text comparison unit, and count the number of characters in the first union as the first hit coverage length of the first hit words in any text comparison unit.

[0122] Optionally, the first hit words may be repeated. For example, for the example text "He came to Shanghai Jiao Tong University", the first hit words may include both "Shanghai" and "Shanghai Jiao Tong University". Therefore, first obtain the first union of the first hit words in any text comparison unit as the hit character set, and count the number of characters in the first union as the first hit coverage length of the first hit words in any text comparison unit.

[0123] Among them, the first union is obtained by determining the original text positions corresponding to the hit identifiers.

[0124] As an example, the above text comparison unit "He came to Shanghai Jiao Tong University" has a total of 9 characters. The character positions in the original text are located by digital numbers, that is, each character is located by the numbers from 0 to 8.

[0125] Then a word-item-identifier-original text position record table can be generated, as shown in Table 3: Table 3

[0126] By obtaining the union of the original text positions, determine the number of characters hit, which is the first hit coverage length.

[0127] As shown in Table 3, if the hit identifiers are 0, 1, 2, 3, 5, then generate a position table that maps the hit identifiers to the original text, as shown in Table 4 below: Table 4

[0128] As shown in Table 4, the original text positions hit by the first hit word are [0,1][1,3][3,5][3,9][7,9]. By taking the union [0,9] of each original text position, it is the first union.

[0129] In other embodiments, a weight can also be assigned to any hit word. Using the translation weight as the weight of the hit word, perform a weighted sum of the lengths of each hit word to obtain the first hit coverage length.

[0130] Step S702, calculate the ratio of the first hit coverage length to the number of characters in any text comparison unit as the first coverage ratio.

[0131] Optionally, the first union represents the hit character length of the hit identifier in any text comparison unit. By calculating the ratio of the first hit coverage length to the number of characters in any text comparison unit, the ratio of the hit identifier hitting in any text comparison unit can be determined as the first coverage ratio.

[0132] As an example, as shown in Table 4, the first coverage ratio .

[0133] Step S703, obtain the second union of the second hit word in the candidate comparison unit, and count the number of characters in the second union as the second hit coverage length of the second hit word in the candidate comparison unit.

[0134] Optionally, the second hit word may be repeated. For example, for the example text "he Come to shanghaijiao tong University", the second hit word may be both "shanghai" and "shanghai jiao tongUniversity". Therefore, first obtain the second union of the second hit word in the candidate comparison unit as the hit character set, and count the number of characters in the second union as the second hit coverage length of the second hit word in the candidate comparison unit.

[0135] Among them, the second union is obtained by determining the original text positions corresponding to the hit identifiers.

[0136] As an example, the above text comparison unit "he come to shanghai jiao tongUniversity" has a total of 40 characters (including spaces). The character positions in the original text are located by digital numbers, that is, each character is located by numbers from 0 to 39.

[0137] Then a term-identifier-original text position record table can be generated, as shown in Table 5: Table 5

[0138] By obtaining the union of the original text positions, the number of hit characters is determined, that is, the second hit coverage length.

[0139] As shown in Table 5, the hit identifiers are 0, 1, 2, 3, 5, then a position table mapping the hit identifiers to the original text is generated, as shown in Table 6 below: Table 6

[0140] As shown in Table 6, the original text positions hit by the first hit term are [0, 2][3, 10][11, 19][11, 40][30, 40]. By taking the union of each original text position [0, 2][3, 10][11, 40], it is the second union, where the second union is represented by P in Table 6.

[0141] Step S704, calculate the ratio of the second hit coverage length to the number of characters in the candidate comparison unit as the second coverage ratio.

[0142] Optionally, the second union represents the hit character length of the hit identifier in the candidate comparison unit. By calculating the ratio of the second hit coverage length to the number of characters in the candidate comparison unit, the hit ratio of the hit identifier in the candidate comparison unit can be determined as the second coverage ratio.

[0143] As an example, as shown in Table 6, the second coverage ratio .

[0144] Step S705, when the first coverage ratio of the first hit term in any text comparison unit is greater than the first preset ratio and the second coverage ratio of the second hit term in the candidate comparison unit is greater than the second preset ratio, determine that any text comparison unit is similar to the candidate comparison unit.

[0145] Optionally, the more the hit term covers in the text, the more corresponding similar terms there are. By calculating the coverage length of the hit term in two text comparison units, the hit ratio of the hit term can be determined to be used to determine the similarity degree of the two text comparison units.

[0146] As an example, when both the first coverage ratio and the second coverage ratio are large, that is, when there are many similar terms in the two text comparison units, it is determined that any text comparison unit is similar to the candidate comparison unit.

[0147] For example, set the first similarity threshold and the second similarity threshold. When the first coverage ratio is greater than or equal to the first similarity threshold and the second coverage ratio is greater than or equal to the second similarity threshold, it is determined that any text comparison unit is similar to the candidate comparison unit.

[0148] In the embodiments of the present application, both the first similarity threshold and the second similarity threshold are set to 0.7, then and , then any text comparison unit is similar to the candidate comparison unit.

[0149] In the embodiments of the present application, first, an undirected graph constructed by obtaining a control dictionary between different languages and a mapping dictionary constructed by nodes and edges are used as tools for similarity comparison between different languages; then, the first text and the second text are split into small text comparison units to facilitate similarity comparison; further, the target mapping set corresponding to each term after word segmentation of the text comparison unit is determined through the undirected graph and the mapping dictionary. The identifier of each edge in the target mapping set represents the control relationship of the term between different languages, and the term can be located in the first text and the second text respectively through the identifier; then, the candidate comparison unit of any text comparison unit in the first text is determined through the target mapping set, realizing the preliminary positioning of the similar text segment; finally, similarity comparison is performed according to the coverage ratio of the hit identifiers in any text comparison unit and the candidate comparison unit, completing the cross-language text similarity comparison. Since the undirected graph includes multiple languages, the present application can directly perform similarity comparison on multi-language texts based on the undirected graph, and can realize complex text comparison across multiple languages without translation conversion, breaking language barriers in academic creation, avoiding cross-language malicious plagiarism, and maintaining the global academic ecosystem.

[0150] Although the above steps are described in the above order in the above embodiments, those skilled in the art can understand that in order to achieve the effects of this embodiment, different steps do not have to be executed in such an order, and they can be executed simultaneously (in parallel) or in reverse order, and these simple changes are all within the protection scope of the present application.

[0151] The multi-language text similarity comparison system of the second embodiment of the present application, as Figure 8 shown, includes: an undirected graph acquisition module 100, a mapping dictionary acquisition module 200, a text segmentation module 300, a word segmentation module 400, a candidate comparison unit determination module 500, a hit word determination module 600, and a similarity comparison module 700.

[0152] An undirected graph acquisition module 100, configured to acquire an undirected graph constructed based on a multilingual comparison dictionary set. Among them, any language in the multilingual comparison dictionary set corresponds to at least one comparison dictionary. The nodes of the undirected graph are words in any language, and the edges of the undirected graph are comparison relationships between nodes in different languages; A mapping dictionary acquisition module 200, configured to acquire a mapping dictionary constructed from a node and a mapping set of the node, where the mapping set includes the identifiers of each edge connected by the node; A text segmentation module 300, configured to perform text segmentation on the first text and the second text to be compared respectively, and generate a plurality of text comparison units; A word segmentation module 400, configured to perform word segmentation on each text comparison unit based on the nodes to obtain a plurality of terms, and determine a target mapping set corresponding to any term based on the undirected graph and the mapping dictionary; A candidate comparison unit determination module 500, configured to determine a candidate comparison unit of any text comparison unit in the first text in the second text based on the target mapping set, where the identifier that is hit in any text comparison unit and the candidate comparison unit is the hit identifier; A hit word determination module 600, configured to determine a first hit word of the hit identifier in any text comparison unit, and a second hit word of the hit identifier in the candidate comparison unit; A similarity comparison module 700, configured to perform a similarity comparison between any text comparison unit and the candidate comparison unit according to the first coverage ratio of the first hit word in any text comparison unit and the second coverage ratio of the second hit word in the candidate comparison unit.

[0153] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process and related descriptions of the above-described system can refer to the corresponding process in the foregoing method embodiment, and will not be repeated here.

[0154] It should be noted that the above-described multilingual text similarity comparison system provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the modules or steps in the embodiments of the present application can be further decomposed or combined. For example, the modules in the above embodiment can be combined into one module, or further split into multiple sub-modules to complete all or part of the functions described above. For the names of the modules and steps involved in the embodiments of the present application, they are only used to distinguish each module or step, and are not regarded as an improper limitation of the present application.

[0155] An electronic device according to the third embodiment of the present application includes: At least one processor; and A memory communicatively connected to at least one of the processors; wherein, the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned multi-lingual text similarity comparison method.

[0156] A computer-readable storage medium according to the fourth embodiment of the present application, the computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above-mentioned multi-lingual text similarity comparison method.

[0157] A computer program product according to the fifth embodiment of the present application, when the computer program product runs on an electronic device, it causes the electronic device to execute the above-mentioned multi-lingual text similarity comparison method.

[0158] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes and related descriptions of the above-described electronic devices and computer-readable storage media can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0159] Those skilled in the art should be able to realize that the modules and method steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. The programs corresponding to the software modules and method steps can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field. To clearly illustrate the interchangeability of electronic hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in the form of electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0160] Next, refer to Figure 9 , which shows a schematic structural diagram of a computer system of a server for implementing the method, system, and device embodiments of the present application. Figure 9 The server shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0161] As Figure 9As shown, the computer system includes a central processing unit (CPU), 901, which can perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 902 or a program loaded from a storage section 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for system operations are also stored. The CPU 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0162] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN (local area network) card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as required. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as required so that a computer program read therefrom is installed into the storage section 908 as required.

[0163] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 909 and / or installed from the removable medium 911. When the computer program is executed by the central processing unit (CPU) 901, the above-described functions defined in the method of the present application are performed. It should be noted that the computer-readable medium described above in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0164] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0165] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0166] The terms "first", "second", etc. are used to distinguish similar objects and not to describe or represent a particular order or sequence.

[0167] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or device / equipment that comprises a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent in these process, method, article, or device / equipment.

[0168] So far, the technical solution of the present application has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easily understood by those skilled in the art that the protection scope of the present application is obviously not limited to these specific embodiments. Without departing from the principle of the present application, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present application.

Claims

1. A multilingual text similarity comparison method, characterized in that: The method comprises: Step S10, obtaining an undirected graph constructed based on a multilingual comparison dictionary set, wherein any language in the multilingual comparison dictionary set corresponds to at least one comparison dictionary, the nodes of the undirected graph are words in any language, and the edges of the undirected graph are comparison relationships between nodes in different languages; Step S20, obtaining the node and a mapping dictionary constructed by a mapping set of the node, wherein the mapping set includes identifiers of each edge connected to the node; Step S30, performing text segmentation on the first text to be compared and the second text to be compared respectively, generating a plurality of text comparison units; Step S40, segmenting each text comparison unit based on the node to obtain multiple terms, and determining a target mapping set corresponding to any term based on the undirected graph and the mapping dictionary; Step S50, based on the target mapping set, determining a candidate comparison unit of any text comparison unit in the first text in the second text, wherein an identifier that is hit in both the any text comparison unit and the candidate comparison unit is a hit identifier; Step S60, determining a first hit word of the hit identifier in any text comparison unit, and a second hit word of the hit identifier in the candidate comparison unit; Step S70, performing a similarity comparison between the any text comparison unit and the candidate comparison unit according to a first coverage ratio of the first hit word in the any text comparison unit and a second coverage ratio of the second hit word in the candidate comparison unit.

2. A multilingual text similarity comparison method according to claim 1, characterized in that: The step of obtaining an undirected graph constructed based on a multilingual dictionary set includes: Taking words of any language as nodes, connecting two nodes that have a comparison relationship based on the comparison dictionary to generate a reference edge; Obtain the number of languages ​​of other nodes connected to any node, and the number of reference edges connected to any node; When the number of the languages ​​is the same as the number of the reference edges, determining that the other nodes are translation pairs; Connect the nodes that are translation pairs to each other to generate extended edges; The undirected graph is generated by the nodes, the base edges, and the extended edges.

3. A multilingual text similarity comparison method according to claim 2, characterized in that: The step of obtaining an undirected graph constructed based on a multilingual dictionary set further includes: determining a translation weight of the reference edge; The translation weight of the extended edge is determined based on the translation weight of the reference edge and a preset attenuation coefficient.

4. A multilingual text similarity comparison method according to claim 2, characterized in that: When the number of the languages ​​is the same as the number of the reference edges, determining that the other nodes are translation pairs includes: When the number of languages ​​is the same as the number of reference edges, and the number of reference edges crossing between the other nodes is less than or equal to a preset crossing-edge threshold, it is determined that the other nodes are translation pairs.

5. A multilingual text similarity comparison method according to claim 3, characterized in that: The determining, based on the target mapping set, a candidate comparison unit of any text comparison unit in the first text in the second text includes: Taking the text comparison unit as a hierarchy, according to the identifiers in each target mapping set, an inverted index is established for the second text, and a text comparison unit corresponding to any identifier in the second text is determined; Numbering the text comparison units corresponding to each identifier in the second text and forming a unit number set; Counting the unit number corresponding to any reference identifier, wherein the reference identifier is any identifier in the target mapping set corresponding to the term of the first text; Determine the similarity ratio of any unit number based on the number of occurrences of the unit number in the statistical process and the corresponding translation weight; The candidate comparison unit is selected according to the similarity ratio.

6. A multilingual text similarity comparison method according to claim 5, characterized in that: Based on the number of occurrences of the unit number in the statistical process and the corresponding translation weight, the similarity ratio of any unit number is determined, including: In the case where the unit number appears, the corresponding translation weights are accumulated to obtain a cumulative weight; The ratio of the cumulative weight to the number of elements in the unit number set is calculated as the similarity ratio.

7. A multilingual text similarity comparison method according to claim 5, characterized in that: The selecting the candidate comparison unit according to the similarity ratio includes: When the similarity ratio of any unit number is greater than or equal to a preset similarity threshold, the text comparison unit corresponding to any unit number is selected as the candidate comparison unit.

8. The multilingual text similarity comparison method according to claim 1, characterized in that: The determining a target mapping set corresponding to any term based on the undirected graph and the mapping dictionary includes: Selecting the nodes where any one of the terms is located in the undirected graph, and the edges connecting the nodes based on the languages ​​of the first text and the second text, to form a subgraph; The target mapping set is generated based on the nodes in the subgraph and the identifiers of the edges connecting the nodes.

9. The multilingual text similarity comparison method according to claim 1, characterized in that: Acquiring the first coverage ratio and the second coverage ratio includes: Obtaining a first union of the first hit word in any of the text comparison units, and counting the number of characters in the first union as a first hit coverage length of the first hit word in any of the text comparison units; Calculating the ratio of the first hit coverage length to the number of characters in any text comparison unit as the first coverage ratio; Obtaining a second union of the second hit word in the candidate comparison unit, and counting the number of characters in the second union as a second hit coverage length of the second hit word in the candidate comparison unit; The ratio of the second hit coverage length to the number of characters in the candidate alignment unit is calculated as the second coverage ratio.

10. The multilingual text similarity comparison method according to claim 1, characterized in that: The performing similarity comparison between any one of the text comparison units and the candidate comparison unit includes: When the first coverage ratio of the first hit word in any text comparison unit is greater than the first preset ratio, and the second coverage ratio of the second hit word in the candidate comparison unit is greater than the second preset ratio, it is determined that any text comparison unit is similar to the candidate comparison unit.

Citation Information

Patent Citations

  • Cyrillic Mongolian and traditional Mongolian bilingual knowledge map construction method

    CN109271529A

  • Generating desired discourse structure from an arbitrary text

    US20200184155A1

  • Detecting hypocrisy in text

    US20210150140A1