A method and system for constructing a multilingual comparable corpus based on similarity
By translating and embedding Uyghur and Tibetan texts into Mandarin and calculating similarities, the method addresses the inefficiencies in constructing comparable corpora for minority languages, resulting in a high-quality Uyghur-Mandarin-Tibetan corpus that enhances multi-lingual applications.
Patent Information
- Application Number
- CN202210694688.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-06-20
AI Technical Summary
It is difficult to build high-quality multilingual comparable corpus in the prior art, especially when it involves complex multilingual application scenarios. The distribution imbalance and cumbersomeness between multiple bilingual comparable corpus limit the timeliness and quality of the task. The corpus resources of ethnic minority languages are scarce and difficult to construct.
By obtaining Chinese, Uyghur and Tibetan corpus documents, using machine translation to generate corresponding Chinese texts, and performing semantic embedding processing, calculating the similarity between corpus, building a multilingual comparable corpus based on the set threshold, using cosine distance as a similarity measure, and automatically filtering corpus that meets the conditions using a comparable decision mechanism.
The construction of a comparable corpus of Chinese-Uighur-Tibetan has been realized, the quality and timeliness of the comparable corpus of multilingualism has been improved, the problems of resource scarcity and unbalanced distribution have been solved, and a high-quality data foundation is provided for cross-language information processing.
Smart Images

Figure CN115130482B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of corpus construction, and particularly to a method and system for constructing a multilingual comparable corpus based on similarity. Background Art
[0002] A corpus is an electronic text that has been carefully sampled and processed, and is an essential basic resource for theoretical research, applied research, and language engineering, especially natural language processing (NLP). In terms of language, corpora can be divided into monolingual corpora and multilingual corpora (bilingual corpora are a special case of multilingual corpora). In addition, according to the organization method of the corpus, multilingual corpora can be further divided into parallel corpora and comparable corpora. A parallel corpus is a set of text pairs consisting of source language texts in one or more target languages and their corresponding translated texts, and all language translations need to be aligned. As a type of corpus, large-scale parallel corpora are of great significance for modeling and learning in language research, statistical machine translation, lexicography, cross-language information retrieval, etc. However, due to the strict mutual translation relationship between the source language and the target language texts, it is difficult to obtain large-scale parallel corpora. In addition, the fields of the collected corpora are often unbalanced, and it is difficult to guarantee the alignment quality. These not only limit the rapid expansion of parallel corpora in terms of scale and field, but also make it difficult to meet the real-time requirements.
[0003] In view of the limitations of the above-mentioned parallel corpora, researchers have started to conduct research on comparable corpora. A comparable corpus refers to a set of two or more monolingual corpora that deal with the same topic. However, compared with parallel corpora, the source language and target language texts of comparable corpora are not strictly translatable and aligned. The acquisition of comparable corpora is flexible, the fields of the collected corpora are extensive, the construction means of the corpora are relatively convenient, and the scale and application fields of the corpora have expanded rapidly. In addition, as an important supplement to parallel corpora, comparable corpora have gradually become one of the indispensable research contents. So far, comparable corpora have been widely used in fields such as translation equivalence extraction, machine translation, cross-language information retrieval, and parallel sentence alignment.
[0004] Currently, comparable corpora mainly involve bilingual comparable corpora such as Chinese-English, Chinese-Russian, Chinese-Japanese, Chinese-French, and Chinese-Spanish, and there are relatively few comparable corpora involving three or more languages, and even fewer multilingual comparable corpora for low-resource ethnic minority languages. When it comes to relatively complex multilingual application scenarios, such as multilingual text translation, multilingual simultaneous interpretation, cross-language document retrieval, and cross-language interpretation of important policy documents, although the above tasks are completed by pairwise comparison based on multiple bilingual comparable corpora, the imbalance in the distribution between multiple bilingual comparable corpora and the cumbersome pairwise comparison greatly restrict the high-quality completion of the above tasks and the timeliness of completing the above tasks. Therefore, it is necessary to construct multilingual comparable corpora.
[0005] In addition, China is a unified multi-ethnic country. Different ethnic groups have different languages for communication. Therefore, language has become an important medium for communication between different ethnic groups. In other words, the barrier-free communication between one's own ethnic group and other ethnic groups has given rise to the emergence of cross-language information processing technology and the multilingual comparable corpus necessary for applying this technology. However, due to the relatively scarce resources and the great difficulty in constructing the ethnic minority language corpus, the scale and quality of the corpus have relatively large limitations. So far, no ethnic minority multilingual comparable corpus has been seen, and all are bilingual comparable corpora, such as the Chinese-Uyghur bilingual comparable corpus, the Chinese-Tibetan bilingual comparable corpus, and the Chinese-Mongolian bilingual comparable corpus. Therefore, it is imperative to construct an ethnic minority multilingual comparable corpus. Summary of the Invention
[0006] Based on this, the embodiments of the present invention provide a method and system for constructing a multilingual comparable corpus based on similarity to realize the construction of a Chinese-Uyghur-Tibetan comparable corpus.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] A method for constructing a multilingual comparable corpus based on similarity, comprising:
[0009] Obtaining a Chinese corpus document, a Uyghur corpus document, and a Tibetan corpus document;
[0010] Translating each Uyghur corpus in the Uyghur corpus document into a Chinese corpus text to obtain a Uyghur-to-Chinese translation corpus document, and translating each Tibetan corpus in the Tibetan corpus document into a Chinese corpus text to obtain a Tibetan-to-Chinese translation corpus document;
[0011] Performing semantic embedding processing on each corpus in the Chinese corpus document, the Uyghur-to-Chinese translation corpus document, and the Tibetan-to-Chinese translation corpus document to obtain a Chinese corpus semantic embedding word vector group, a Uyghur-to-Chinese translation corpus semantic embedding word vector group, and a Tibetan-to-Chinese translation corpus semantic embedding word vector group;
[0012] Calculating a first similarity, a second similarity, and a third similarity according to the Chinese corpus semantic embedding word vector group, the Uyghur-to-Chinese translation corpus semantic embedding word vector group, and the Tibetan-to-Chinese translation corpus semantic embedding word vector group; the first similarity is the similarity between each Chinese corpus in the Chinese corpus document and each Uyghur corpus in the Uyghur corpus document; the second similarity is the similarity between each Chinese corpus in the Chinese corpus document and each Tibetan corpus in the Tibetan corpus document; the third similarity is the similarity between each Uyghur corpus in the Uyghur corpus document and each Tibetan corpus in the Tibetan corpus document;
[0013] Determine a multilingual comparable corpus according to the first similarity, the second similarity, the third similarity and a set similarity threshold.
[0014] Optionally, the obtaining of the Chinese corpus document, the Uyghur corpus document and the Tibetan corpus document specifically includes:
[0015] Use a data scraping crawler software to search a set news website to obtain web page information;
[0016] Perform HTML parsing on the web page information, extract news titles, news contents and news times, and generate an initial corpus;
[0017] Preprocess the initial corpus to obtain a Chinese corpus document, a Uyghur corpus document and a Tibetan corpus document.
[0018] Optionally, the translating each Uyghur corpus in the Uyghur corpus document into a Chinese corpus text to obtain a Uyghur-to-Chinese translated corpus document, and translating each Tibetan corpus in the Tibetan corpus document into a Chinese corpus text to obtain a Tibetan-to-Chinese translated corpus document specifically includes:
[0019] Use a machine translation software to translate each Uyghur corpus in the Uyghur corpus document into a Chinese corpus text to obtain a Uyghur-to-Chinese translated corpus document;
[0020] Use a machine translation software to translate each Tibetan corpus in the Tibetan corpus document into a Chinese corpus text to obtain a Tibetan-to-Chinese translated corpus document.
[0021] Optionally, the calculating the first similarity, the second similarity and the third similarity according to the Chinese corpus semantic embedded word vector group, the Uyghur-to-Chinese translated corpus semantic embedded word vector group and the Tibetan-to-Chinese translated corpus semantic embedded word vector group specifically includes:
[0022] Calculate the word frequency vectors of every two semantic embedded word vector groups according to the Chinese corpus semantic embedded word vector group, the Uyghur-to-Chinese translated corpus semantic embedded word vector group and the Tibetan-to-Chinese translated corpus semantic embedded word vector group;
[0023] Calculate the first similarity, the second similarity and the third similarity according to the word frequency vectors.
[0024] Optionally, the determining a multilingual comparable corpus according to the first similarity, the second similarity, the third similarity and a set similarity threshold specifically includes:
[0025] For any Chinese corpus, Uyghur corpus, and Tibetan corpus, determine whether the intersection of the corresponding first similarity, the corresponding second similarity, and the corresponding third similarity is greater than the set similarity threshold;
[0026] If so, store the corresponding Chinese corpus, Uyghur corpus, and Tibetan corpus in the multilingual comparable corpus;
[0027] If not, delete the corresponding Chinese corpus from the Chinese corpus document, delete the corresponding Uyghur corpus from the Uyghur corpus document, and delete the corresponding Tibetan corpus from the Tibetan corpus document.
[0028] The present invention also provides a multilingual comparable corpus construction system based on similarity, including:
[0029] A corpus acquisition module for acquiring a Chinese corpus document, a Uyghur corpus document, and a Tibetan corpus document;
[0030] A corpus translation module for translating each Uyghur corpus in the Uyghur corpus document into a Chinese corpus text to obtain a Uyghur-to-Chinese translation corpus document, and translating each Tibetan corpus in the Tibetan corpus document into a Chinese corpus text to obtain a Tibetan-to-Chinese translation corpus document;
[0031] A semantic embedding module for performing semantic embedding processing on each corpus in the Chinese corpus document, the Uyghur-to-Chinese translation corpus document, and the Tibetan-to-Chinese translation corpus document to obtain a Chinese corpus semantic embedding word vector group, a Uyghur-to-Chinese translation corpus semantic embedding word vector group, and a Tibetan-to-Chinese translation corpus semantic embedding word vector group;
[0032] A similarity calculation module for calculating a first similarity, a second similarity, and a third similarity according to the Chinese corpus semantic embedding word vector group, the Uyghur-to-Chinese translation corpus semantic embedding word vector group, and the Tibetan-to-Chinese translation corpus semantic embedding word vector group; the first similarity is the similarity between each Chinese corpus in the Chinese corpus document and each Uyghur corpus in the Uyghur corpus document; the second similarity is the similarity between each Chinese corpus in the Chinese corpus document and each Tibetan corpus in the Tibetan corpus document; the third similarity is the similarity between each Uyghur corpus in the Uyghur corpus document and each Tibetan corpus in the Tibetan corpus document;
[0033] A corpus construction module for determining a multilingual comparable corpus according to the first similarity, the second similarity, the third similarity, and the set similarity threshold.
[0034] Optionally, the corpus acquisition module specifically includes:
[0035] A web information search unit for searching a set news website using data scraping crawler software to obtain web information;
[0036] A parsing unit for performing HTML parsing on the web information, extracting news titles, news contents, and news times, and generating initial corpus;
[0037] A preprocessing unit for preprocessing the initial corpus to obtain a Chinese corpus document, a Uyghur corpus document, and a Tibetan corpus document.
[0038] Optionally, the corpus translation module specifically includes:
[0039] A first translation unit for translating each Uyghur corpus in the Uyghur corpus document into a Chinese corpus text using machine translation software to obtain a Uyghur-to-Chinese translation corpus document;
[0040] A second translation unit for translating each Tibetan corpus in the Tibetan corpus document into a Chinese corpus text using machine translation software to obtain a Tibetan-to-Chinese translation corpus document.
[0041] Optionally, the similarity calculation module specifically includes:
[0042] A word frequency vector determination unit for calculating the word frequency vectors of every two semantic embedding word vector groups according to the Chinese corpus semantic embedding word vector group, the Uyghur-to-Chinese translation corpus semantic embedding word vector group, and the Tibetan-to-Chinese translation corpus semantic embedding word vector group;
[0043] A similarity calculation unit for calculating the first similarity, the second similarity, and the third similarity according to the word frequency vectors.
[0044] Optionally, the corpus construction module specifically includes:
[0045] A similarity determination unit for determining, for any Chinese corpus, Uyghur corpus, and Tibetan corpus, whether the intersection of the corresponding first similarity, the corresponding second similarity, and the corresponding third similarity is greater than a set similarity threshold;
[0046] A corpus construction unit for, if so, storing the corresponding Chinese corpus, the corresponding Uyghur corpus, and the corresponding Tibetan corpus into a multilingual comparable corpus;
[0047] A document update unit for, if not, deleting the corresponding Chinese corpus from the Chinese corpus document, deleting the corresponding Uyghur corpus from the Uyghur corpus document, and deleting the corresponding Tibetan corpus from the Tibetan corpus document.
[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0049] An embodiment of the present invention provides a method and system for constructing a multilingual comparable corpus based on similarity. The method includes obtaining Chinese-Uyghur-Tibetan corpus documents, translating the Uyghur and Tibetan corpus documents into corresponding Chinese corpus texts, solving the similarity of the Chinese-Uyghur-Tibetan trilingual texts in a unified Chinese mode to obtain the similarity between the corpora, and selecting comparable corpora according to the similarity between the corpora and a set similarity threshold, thereby constructing a Chinese-Uyghur-Tibetan comparable corpus. The present invention realizes the construction of a multilingual comparable corpus. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0051] Figure 1 It is a flowchart of the method for constructing a multilingual comparable corpus based on similarity provided by the embodiment of the present invention;
[0052] Figure 2 It is a flowchart of data scraping based on DCCS;
[0053] Figure 3 It is a flowchart of a comparable relationship decision mechanism based on text similarity;
[0054] Figure 4 It is an overall framework diagram of the method for constructing a multilingual comparable corpus based on similarity;
[0055] Figure 5 It is a schematic diagram of Chinese news corpus data;
[0056] Figure 6 It is a schematic diagram of Uyghur news corpus data;
[0057] Figure 7 It is a schematic diagram of Tibetan news corpus data. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0059] To make the above objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0060] Figure 1 It is a flowchart of a method for constructing a multilingual comparable corpus based on similarity provided by an embodiment of the present invention. Refer to Figure 1 , the method of this embodiment includes:
[0061] Step 101: Obtain Chinese corpus documents, Uyghur corpus documents, and Tibetan corpus documents.
[0062] Step 101 specifically includes:
[0063] 1) Use data scraping crawler software to search a set news website to obtain web page information. Specifically, as Figure 2 shown, first, based on the Scrapy framework, develop data scraping crawler software DCCS (data capture crawler software) for the news website; use DCCS to search the set news website and obtain all web page information on the web page that meets the requirements (such as including news keywords, time, location, etc.).
[0064] 2) Perform HTML parsing on the web page information, extract news titles, news contents, and news times, and generate initial corpora. Preprocess the initial corpora to obtain Chinese corpus documents, Uyghur corpus documents, and Tibetan corpus documents. Specifically:
[0065] Use BeautifulSoup in the Python library to perform HTML parsing and analysis on the obtained web page information, extract core information such as news titles, news contents, and news times, and form Chinese corpus documents, Uyghur corpus documents, and Tibetan corpus documents with the title as the title after preprocessing (such as cleaning) (each news corpus is saved in text format).
[0066] The Chinese corpus document, Uyghur corpus document, and Tibetan corpus document are respectively represented as C, U, and T, where:
[0067] C = {C1, C2,..., C M}, C i (i = 1,..., M) represents the i-th Chinese corpus in C, and M represents the number of Chinese corpora;
[0068] U = {U1, U2,..., U α}, U b (b = 1,..., α) represents the b-th Uyghur corpus in U, and α represents the number of Uyghur corpora;
[0069] T = {T1, T2,..., T β}, T τ (τ = 1,..., β) represents the τ-th Tibetan corpus in T, and β represents the number of Tibetan corpora.
[0070] Step 102: Translate each Uyghur corpus in the Uyghur corpus document into a Chinese corpus text to obtain a Uyghur-Chinese translation corpus document, and translate each Tibetan corpus in the Tibetan corpus document into a Chinese corpus text to obtain a Tibetan-Chinese translation corpus document. Specifically:
[0071] Use a machine translation software to translate each Uyghur corpus in the Uyghur corpus document into a Chinese corpus text to obtain a Uyghur-Chinese translation corpus document. Use a machine translation software to translate each Tibetan corpus in the Tibetan corpus document into a Chinese corpus text to obtain a Tibetan-Chinese translation corpus document. Specifically:
[0072] Translate the Uyghur corpus document U and the Tibetan corpus document T into corresponding Chinese corpus texts to obtain the Uyghur-Chinese translation corpus document U c and the Tibetan-Chinese translation corpus document T c , where:
[0073] U c = {U c1 , U c2 ,..., U cα},
[0074] U cb represents the Chinese corpus text corresponding to the b-th Uyghur corpus in U c ,
[0075] T c = {T c1 , T c2 ,..., T cβ}, T cτ represents the Chinese corpus text corresponding to the τ-th Tibetan corpus in T c .
[0076] Step 103: Perform semantic embedding processing on each corpus in the Chinese corpus document, the Uyghur-Chinese translation corpus document, and the Tibetan-Chinese translation corpus document to obtain a Chinese corpus semantic embedding word vector group, a Uyghur-Chinese translation corpus semantic embedding word vector group, and a Tibetan-Chinese translation corpus semantic embedding word vector group. This step realizes corpus semantic embedding processing. Specifically:
[0077] Denote the i-th Chinese corpus C in C as: i For: Where is the j-th word of the i-th Chinese corpus in C, and m represents the number of words in this corpus; denote U c the b-th Uyghur-to-Chinese translated corpus U in cb as: where is c the k-th word of the b-th Uyghur-to-Chinese translated corpus in U, and K represents the number of words in this corpus; denote T c the τ-th Tibetan-to-Chinese translated corpus T in cτ as: where is c the n-th word of the τ-th Tibetan-to-Chinese translated corpus in T, and N represents the number of words in this corpus.
[0078] For and perform embedding respectively to generate the corresponding semantic embedding word vector groups
[0079]
[0080] where e represents the word embedding vector. represents the word embedding vector of the c1-th word in the i-th Chinese corpus; represents c (Uyghur-to-Chinese translated corpus) the word embedding vector of the c1-th word in the b-th corpus of U; represents c (Tibetan-to-Chinese translated corpus) the word embedding vector of the c1-th word in the τ-th corpus of T. Correspondingly, cm represents the cm-th word in the i-th Chinese corpus, cK represents the cK-th word in the b-th corpus, and cN represents the cN-th word in the τ-th corpus.
[0081] Step 104: Calculate the first similarity, the second similarity, and the third similarity according to the Chinese corpus semantic embedding word vector group, the Uyghur-to-Chinese translated corpus semantic embedding word vector group, and the Tibetan-to-Chinese translated corpus semantic embedding word vector group. The first similarity is the similarity between each Chinese corpus in the Chinese corpus document and each Uyghur corpus in the Uyghur corpus document; the second similarity is the similarity between each Chinese corpus in the Chinese corpus document and each Tibetan corpus in the Tibetan corpus document; the third similarity is the similarity between each Uyghur corpus in the Uyghur corpus document and each Tibetan corpus in the Tibetan corpus document. What this step realizes is the similarity calculation between corpora. Specifically, this step includes:
[0082] 1) Calculate the word frequency vectors of every two semantic embedding word vector groups according to the Chinese corpus semantic embedding word vector group, the Uyghur-to-Chinese translation corpus semantic embedding word vector group, and the Tibetan-to-Chinese translation corpus semantic embedding word vector group. Specifically:
[0083] Let and be P1, P2, and P3 respectively, and P q (q ∈ (1, 2, 3)) includes m q words, then m1 = m, m2 = K, m3 = N. Define S uv as the union of P u and P v (u, v ∈ q, and u ≠ v). Assume the length L uv of S uv , where t y (y ∈ (1, 2,..., L uv )) represents the y-th element in S uv . Then the word frequency vectors of P u and P v are defined as:
[0084]
[0085] 2) Calculate the first similarity, the second similarity, and the third similarity according to the word frequency vectors. Specifically:
[0086]
[0087] Among them, represents the similarity between the i-th Chinese corpus and the b-th Uyghur corpus, that is, the first similarity. Specifically, calculate the co-occurring words in the i-th Chinese corpus and the b-th Uyghur-to-Chinese translation corpus, and then calculate (the word frequency of co-occurring words in the Chinese corpus) and (the word frequency of co-occurring words in the Uyghur-to-Chinese translation corpus) according to the first step, and then calculate the cosine similarity between these two word frequency vectors.
[0088] represents the similarity between the i-th Chinese corpus and the τ-th Tibetan corpus, that is, the second similarity. Specifically, calculate the co-occurring words in the i-th Chinese corpus and the τ-th Tibetan-to-Chinese translation corpus, and then calculate (the word frequency of co-occurring words in the Chinese corpus) and (the word frequency of co-occurring words in the Tibetan-to-Chinese translation corpus) according to the first step, and then calculate the cosine similarity between these two word frequency vectors.
[0089] Denote the similarity between the b-th Uyghur corpus and the τ-th Tibetan corpus, i.e., the third similarity. Specifically, calculate the co-occurring words in the b-th Uyghur-to-Chinese translated corpus and the τ-th Tibetan-to-Chinese translated corpus, and then calculate according to the first-step calculation (The word frequency of co-occurring words in the Uyghur-to-Chinese translated corpus) and (The word frequency of co-occurring words in the Tibetan-to-Chinese translated corpus), and then calculate the cosine similarity between these two word frequency vectors.
[0090] Then, according to the definitions of the first similarity, the second similarity, and the third similarity, calculate the similarities between all Chinese, Uyghur, and Tibetan corpora.
[0091] Step 105: Determine a multilingual comparable corpus according to the first similarity, the second similarity, the third similarity, and a set similarity threshold.
[0092] Step 105 specifically includes:
[0093] For any Chinese corpus, Uyghur corpus, and Tibetan corpus, determine whether the intersection of the corresponding first similarity, the corresponding second similarity, and the corresponding third similarity is greater than the set similarity threshold. If so, store the corresponding Chinese corpus, the corresponding Uyghur corpus, and the corresponding Tibetan corpus in the multilingual comparable corpus; if not, delete the corresponding Chinese corpus from the Chinese corpus document, delete the corresponding Uyghur corpus from the Uyghur corpus document, and delete the corresponding Tibetan corpus from the Tibetan corpus document. Among them, the set similarity threshold can be 0.5.
[0094] This step is implemented based on a decision-making mechanism of text similarity. As Figure 3 shown, in practical applications, a specific implementation process is as follows:
[0095] (1) Randomly select the i-th Chinese corpus from the Chinese corpus document C, and calculate the similarity between this corpus and all Uyghur corpora The maximum value of the similarities between the i-th Chinese corpus and all Uyghur corpora is the similarity between the i-th Chinese corpus and the j-th Uyghur corpus That is
[0096]
[0097] (2) For the j-th Uyghur news corpus in U c , calculate the similarity between this corpus and all Tibetan corpora
[0098] The similarity between the j-th Uyghur corpus and the l-th Tibetan corpus is
[0099]
[0100] (3) Calculate the similarity between the i-th Chinese corpus and the l-th Tibetan corpus
[0101] (4) Determine whether the condition is satisfied If the condition is not satisfied, jump to (6).
[0102] (5) If The i-th Chinese corpus in the Chinese corpus document C, the j-th Uyghur corpus in the Uyghur corpus document U, and the l-th Tibetan corpus in the Tibetan corpus document T form comparable corpora, and are entered into the multilingual comparable corpus. Delete the i-th Chinese corpus in the Chinese corpus document C, the j-th Uyghur corpus in the Uyghur corpus document U, and the l-th Tibetan corpus in the Tibetan corpus document T from the news corpus. (6) Traverse the Chinese corpus, extract the next corpus, and make a judgment on it.
[0103] (7) Repeat (1)-(6) until all the corpora in the Chinese corpus document C are comparable and end.
[0104] The above steps 101-104 complete the acquisition and basic processing of Chinese-Uyghur-Tibetan news comparable corpora. According to the comparability decision mechanism in step 105, the Chinese, Uyghur, and Tibetan (translation and original text) news corpora that meet the requirements are automatically selected to form comparable pairs and automatically stored in the corpus. Those that do not meet the requirements will be deleted, completing the construction of the Chinese-Uyghur-Tibetan comparable corpus.
[0105] The pseudo-code of the decision mechanism is as follows:
[0106]
[0107] In the method for constructing a multilingual comparable corpus based on similarity in this embodiment, the cosine distance is used as the similarity metric to solve the similarity of Chinese-Uyghur-Tibetan trilingual texts in a unified Chinese mode. According to the optimized decision mechanism, the Chinese-Uyghur-Tibetan trilingual texts with relatively high similarity are selected as comparable corpora to construct a Chinese-Uyghur-Tibetan news comparable corpus. A comparability decision mechanism is proposed so that the text similarity algorithm is applicable to discovering comparable corpora. The text corpora with a comparability less than or equal to 0.5 are removed using the minimum intersection value criterion of the similarity values, and the comparability decision mechanism of the text similarity algorithm can generally obtain good results of comparable corpora, thus greatly improving the overall quality of the comparable corpus.
[0108] The following gives a specific example to further illustrate the method for constructing a multilingual comparable corpus based on similarity in the above embodiment.
[0109] An embodiment of the present invention provides a method for obtaining Chinese-Uyghur-Tibetan news corpus by developing a data scraping crawler software DCCS based on the Scrapy framework. Refer to Figure 4 , the method includes:
[0110] 1. News corpus collection and processing
[0111] Step 1: First, based on the Scrapy framework, develop a data scraping crawler software DCCS for news websites.
[0112] In the embodiment of the present invention, the components such as "scheduler", "crawler", "downloader", "entity pipeline", and "scrapy engine" in the Scrapy framework are mainly referred to, and the crawling software that conforms to the present invention is developed in combination with its specific functions. The specific functions of the components are shown in the following table:
[0113] Table 1 Scrapy components and functions
[0114]
[0115]
[0116] Step 2: Use DCCS to search the preset news websites to obtain all web page information that meets the requirements (such as news keywords, time, location, etc.) in the web pages.
[0117] For the sample data described in the embodiment of the present invention, in principle, users can use the crawler software to scrape relevant news data from any news website, and after parsing, analyzing, and cleaning, form Chinese, Uyghur, and Tibetan original (raw) news corpus. As an example, the time interval of the news reports selected in the embodiment of the present invention is from January 1, 2019 to December 27, 2021 for two years. To ensure the authority, authenticity, and comprehensiveness of the corpus source in the embodiment of the present invention, the Chinese, Uyghur, and Tibetan online news published by People's Daily Online (http: / / www.people.com.cn / ; http: / / uyghur.people.com.cn / ; http: / / tibet.people.com.cn) are finally selected as the source of the original (raw) news corpus.
[0118] Step 3: Use BeautifulSoup in the Python library to perform HTML parsing and analysis on the obtained web pages, extract core information such as titles and texts, and after cleaning, form Chinese, Uyghur, and Tibetan original news corpus documents with the title as the title (each news is saved in text format).
[0119] From the numerous news corpora captured, considering the equivalence of the original news corpora in Chinese, Uyghur, and Tibetan, 2312 language texts each were selected for the Chinese, Uyghur, and Tibetan news texts, forming the original news corpora C, U, and T in Chinese, Uyghur, and Tibetan respectively, as shown in Figure 5 , Figure 6 and Figure 7 respectively.
[0120] 2. News Corpus Translation
[0121] Call the Niutrans machine translation plugin to translate the original Uyghur and Tibetan news corpus documents U and T into the corresponding Chinese news corpus texts, denoted as U c and T c respectively.
[0122] 3. Similarity Calculation between News Corpora
[0123] Use formulas (1)-(5) to calculate the similarity between the news corpora C, U c , T c pairwise. The results of the similarity calculation are shown in Tables 2, 3, and 4.
[0124] Table 2 Similarity Calculation Results between Uyghur-Chinese News Corpora (Partial)
[0125]
[0126] Table 3 Similarity Calculation Results between Tibetan-Chinese News Corpora (Partial)
[0127]
[0128] Table 4 Similarity Calculation Results between Uyghur-Tibetan News Corpora (Partial)
[0129]
[0130]
[0131] 4. Construction Results of News Comparable Corpora
[0132] Through the above processes and processing steps, in the corpus acquisition stage, the web news reports in Chinese, Uyghur, and Tibetan show an incomplete equivalence. Therefore, the news raw corpus dataset finally obtained in the embodiments of the present invention consists of 30,662 sample data, that is, the dataset contains 24,571 Chinese news corpora (111,290,526 bytes, data size 106 megabytes), 3,779 Uyghur news corpora (15,351,808 bytes, data size 14.6 megabytes), and 2,312 Tibetan news corpora (13,102,095 bytes, data size 12.4 megabytes).
[0133] According to the calculated similarity, use C and U c , T c The news comparable corpus decision-making mechanism is used to construct a comparable corpus, and a Chinese-Uyghur-Tibetan news comparable corpus with 53 Chinese news articles (76.8KB), 52 Uyghur news articles (320KB), and 52 Tibetan news articles (73.5KB) is obtained. Among them, C(3).txt, U(3).txt, and T(3).txt are a group of comparable corpus pairs.
[0134] The present invention uses a comparability decision-making mechanism to adjust the comparability threshold. In actual applications, different thresholds can be defined. By changing the threshold, poor comparable corpus results can be avoided. At the same time, the optimal comparable corpus is automatically found through the comparability decision-making mechanism, without the need for manual comparison and selection.
[0135] It should be noted that the preset comparability similarity threshold in the embodiments of the present invention has no maximum and minimum values, and the threshold can be changed according to the specific requirements of the corpus.
[0136] In order to further verify and analyze the constructed Chinese-Uyghur-Tibetan news comparable corpus, the text similarity calculation method and the comparability decision-making mechanism proposed by the present invention are used to calculate the similarity of the titles and contents of the Chinese-Uyghur-Tibetan comparable corpus, as shown in Table 5.
[0137] Table 5 Evaluation results of the Chinese-Uyghur-Tibetan news comparability corpus construction
[0138]
[0139]
[0140] *Note: In the process of calculating the similarity value, this experiment traverses the Chinese, Uyghur, and Tibetan folders and calculates the similarity pairwise in turn. Therefore, it is possible that one or more documents satisfy the conditions for establishing a comparable relationship with multiple documents and form comparable corpus pairs.
[0141] The following conclusions are drawn based on the similarity distribution: As can be seen from the above table, according to the content similarity screening, there are a total of 54 Chinese-Uyghur-Tibetan comparable corpus pairs with a similarity value greater than 0.5, among which 46 Chinese-Uyghur-Tibetan comparable corpus pairs have a similarity value greater than 0.8; according to the title similarity comparison and decision-making, a total of 545 Chinese-Uyghur-Tibetan comparable corpus pairs are obtained, among which 86 Chinese-Uyghur-Tibetan comparable pairs have a similarity value greater than 0.8. The average comparability between the texts of the comparable corpus constructed in the embodiments of the present invention is greater than 0.5 (content: 0.780995188; title: 0.658489125). The overall corpus has reached a relatively high comparability. The corpora in the range of 0.5-0.8 can be regarded as describing related events. The multilingual comparable corpora provide a good data basis for event detection, extraction, and further research. In addition, the corpora with a similarity greater than 0.8 can be used as potential multilingual parallel corpora and potential multilingual translation texts. Through further processing and research, the low-resource corpus can be expanded to a certain extent to achieve data augmentation. It can be seen that the comparability calculation and decision-making mechanism proposed based on the present invention has a good effect on the construction of the news comparable corpus. From the above construction results, it can be seen that the comparability between news corpora is relatively high, and the quality of the constructed comparable corpus is relatively high.
[0142] The method for constructing a multilingual comparable corpus based on similarity in this embodiment semi-automatically obtains Chinese-Uyghur-Tibetan news corpora by using web crawler technology, and translates Uyghur and Tibetan news into Chinese corpora with the help of machine translation; uses cosine similarity to evaluate the semantic similarity between Chinese-Uyghur-Tibetan news sentences, establishes a comparable relationship and makes a decision; adopts a decision-making mechanism that compares the similarity values in a pairwise loop and the minimum value of the intersection of the three satisfies being greater than a preset threshold, which well solves the method for constructing the comparability between multiple languages, provides an idea for the construction of a multilingual comparable corpus, obtains better corpus construction results, and better improves the quality of the comparable corpus.
[0143] The present invention also provides a multilingual comparable corpus construction system based on similarity, including:
[0144] A corpus acquisition module for acquiring Chinese corpus documents, Uyghur corpus documents, and Tibetan corpus documents.
[0145] A corpus translation module for translating each Uyghur corpus in the Uyghur corpus document into a Chinese corpus text to obtain a Uyghur-to-Chinese translation corpus document, and translating each Tibetan corpus in the Tibetan corpus document into a Chinese corpus text to obtain a Tibetan-to-Chinese translation corpus document.
[0146] A semantic embedding module for performing semantic embedding processing on each piece of corpus in the Chinese corpus document, the Uyghur-Chinese translated corpus document, and the Tibetan-Chinese translated corpus document, to obtain a Chinese corpus semantic embedding word vector group, a Uyghur-Chinese translated corpus semantic embedding word vector group, and a Tibetan-Chinese translated corpus semantic embedding word vector group.
[0147] A similarity calculation module for calculating a first similarity, a second similarity, and a third similarity according to the Chinese corpus semantic embedding word vector group, the Uyghur-Chinese translated corpus semantic embedding word vector group, and the Tibetan-Chinese translated corpus semantic embedding word vector group; the first similarity is the similarity between each piece of Chinese corpus in the Chinese corpus document and each piece of Uyghur corpus in the Uyghur corpus document; the second similarity is the similarity between each piece of Chinese corpus in the Chinese corpus document and each piece of Tibetan corpus in the Tibetan corpus document; the third similarity is the similarity between each piece of Uyghur corpus in the Uyghur corpus document and each piece of Tibetan corpus in the Tibetan corpus document.
[0148] A corpus construction module for determining a multilingual comparable corpus according to the first similarity, the second similarity, the third similarity, and a set similarity threshold.
[0149] In one example, the corpus acquisition module specifically includes:
[0150] A web information search unit for searching a set news website using data scraping crawler software to obtain web information.
[0151] An analysis unit for performing HTML analysis on the web information, extracting news titles, news contents, and news times, and generating initial corpus.
[0152] A preprocessing unit for preprocessing the initial corpus to obtain a Chinese corpus document, a Uyghur corpus document, and a Tibetan corpus document.
[0153] In one example, the corpus translation module specifically includes:
[0154] A first translation unit for translating each piece of Uyghur corpus in the Uyghur corpus document into Chinese corpus text using machine translation software to obtain a Uyghur-Chinese translated corpus document.
[0155] A second translation unit for translating each piece of Tibetan corpus in the Tibetan corpus document into Chinese corpus text using machine translation software to obtain a Tibetan-Chinese translated corpus document.
[0156] In one example, the similarity calculation module specifically includes:
[0157] A word frequency vector determination unit, configured to calculate the word frequency vectors of every two semantic embedded word vector groups according to the Chinese corpus semantic embedded word vector group, the Uyghur translated Chinese corpus semantic embedded word vector group, and the Tibetan translated Chinese corpus semantic embedded word vector group.
[0158] A similarity calculation unit, configured to calculate the first similarity, the second similarity, and the third similarity according to the word frequency vectors.
[0159] In one example, the corpus construction module specifically includes:
[0160] A similarity determination unit, configured to, for any Chinese corpus, Uyghur corpus, and Tibetan corpus, determine whether the intersection of the corresponding first similarity, the corresponding second similarity, and the corresponding third similarity is greater than a set similarity threshold.
[0161] A corpus construction unit, configured to, if so, store the corresponding Chinese corpus, the corresponding Uyghur corpus, and the corresponding Tibetan corpus into a multilingual comparable corpus.
[0162] A document update unit, configured to, if not, delete the corresponding Chinese corpus from the Chinese corpus document, delete the corresponding Uyghur corpus from the Uyghur corpus document, and delete the corresponding Tibetan corpus from the Tibetan corpus document.
[0163] The multi - language comparable corpus construction system based on similarity in this embodiment constructs a Chinese - Uyghur - Tibetan news comparable corpus based on a decision - making mechanism. The construction of the Chinese - Uyghur - Tibetan news comparable corpus is through the above - mentioned processes of corpus collection, pre - processing, machine translation of non - Chinese news corpora, similarity calculation between news corpora, and comparability decision - making. Based on the decision - making mechanism, news texts with a Chinese - Uyghur - Tibetan comparability higher than a preset similarity threshold are automatically screened out, obtaining Chinese - Uyghur - Tibetan news comparable corpora that meet the requirements of the present invention, and storing the corpus data to complete the construction of the Chinese - Uyghur - Tibetan news comparable corpus.
[0164] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0165] In this article, specific examples are used to illustrate the principles and implementation modes of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation modes and application scopes. To sum up, the content of this specification should not be construed as a limitation on the present invention.
Claims
1. A method for constructing a multilingual comparable corpus based on similarity, characterized in that Including: Obtain a Chinese corpus document, a Uyghur corpus document, and a Tibetan corpus document; Translate each Uyghur corpus in the Uyghur corpus document into a Chinese corpus text to obtain a Uyghur-Chinese translated corpus document, and translate each Tibetan corpus in the Tibetan corpus document into a Chinese corpus text to obtain a Tibetan-Chinese translated corpus document; Perform semantic embedding processing on each corpus in the Chinese corpus document, the Uyghur-Chinese translated corpus document, and the Tibetan-Chinese translated corpus document to obtain a Chinese corpus semantic embedding word vector group, a Uyghur-Chinese translated corpus semantic embedding word vector group, and a Tibetan-Chinese translated corpus semantic embedding word vector group; Calculate a first similarity, a second similarity, and a third similarity according to the Chinese corpus semantic embedding word vector group, the Uyghur-Chinese translated corpus semantic embedding word vector group, and the Tibetan-Chinese translated corpus semantic embedding word vector group; the first similarity is the similarity between each Chinese corpus in the Chinese corpus document and each Uyghur corpus in the Uyghur corpus document; the second similarity is the similarity between each Chinese corpus in the Chinese corpus document and each Tibetan corpus in the Tibetan corpus document; the third similarity is the similarity between each Uyghur corpus in the Uyghur corpus document and each Tibetan corpus in the Tibetan corpus document; Determine a multilingual comparable corpus according to the first similarity, the second similarity, the third similarity, and a set similarity threshold, specifically including: For any Chinese corpus, Uyghur corpus, and Tibetan corpus, determine whether the corresponding first similarity, the corresponding second similarity, and the corresponding third similarity are all greater than the set similarity threshold; If so, store the corresponding Chinese corpus, the corresponding Uyghur corpus, and the corresponding Tibetan corpus in the multilingual comparable corpus; If not, delete the corresponding Chinese corpus from the Chinese corpus document, delete the corresponding Uyghur corpus from the Uyghur corpus document, and delete the corresponding Tibetan corpus from the Tibetan corpus document.
2. The method for constructing a multi-lingual comparable corpus based on similarity according to claim 1, wherein The obtaining of the Chinese corpus document, the Uyghur corpus document, and the Tibetan corpus document specifically includes: Use data scraping crawler software to search a set news website to obtain web page information; Perform HTML parsing on the web page information, extract news titles, news contents, and news times, and generate initial corpus; Perform preprocessing on the initial corpus to obtain a Chinese corpus document, a Uyghur corpus document, and a Tibetan corpus document.
3. A method for constructing a multilingual comparable corpus based on similarity according to claim 1, characterized in that, The translating of each Uyghur corpus in the Uyghur corpus document into a Chinese corpus text to obtain a Uyghur-Chinese translated corpus document, and the translating of each Tibetan corpus in the Tibetan corpus document into a Chinese corpus text to obtain a Tibetan-Chinese translated corpus document specifically includes: Use machine translation software to translate each Uyghur corpus in the Uyghur corpus document into a Chinese corpus text to obtain a Uyghur-Chinese translated corpus document; Use machine translation software to translate each Tibetan corpus in the Tibetan corpus document into a Chinese corpus text to obtain a Tibetan-Chinese translated corpus document.
4. A method for constructing a multilingual comparable corpus based on similarity according to claim 1, characterized in that Calculating a first similarity, a second similarity, and a third similarity according to the semantic embedding word vector group of the Chinese corpus, the semantic embedding word vector group of the Uyghur-to-Chinese translation corpus, and the semantic embedding word vector group of the Tibetan-to-Chinese translation corpus specifically includes: Calculating the word frequency vectors of every two semantic embedding word vector groups according to the semantic embedding word vector group of the Chinese corpus, the semantic embedding word vector group of the Uyghur-to-Chinese translation corpus, and the semantic embedding word vector group of the Tibetan-to-Chinese translation corpus; Calculating the first similarity, the second similarity, and the third similarity according to the word frequency vectors.
5. A multi - language comparable corpus construction system based on similarity, characterized in that, Including: A corpus acquisition module for acquiring a Chinese corpus document, a Uyghur corpus document, and a Tibetan corpus document; A corpus translation module for translating each Uyghur corpus in the Uyghur corpus document into a Chinese corpus text to obtain a Uyghur-to-Chinese translation corpus document, and translating each Tibetan corpus in the Tibetan corpus document into a Chinese corpus text to obtain a Tibetan-to-Chinese translation corpus document; A semantic embedding module for performing semantic embedding processing on each corpus in the Chinese corpus document, the Uyghur-to-Chinese translation corpus document, and the Tibetan-to-Chinese translation corpus document to obtain a semantic embedding word vector group of the Chinese corpus, a semantic embedding word vector group of the Uyghur-to-Chinese translation corpus, and a semantic embedding word vector group of the Tibetan-to-Chinese translation corpus; A similarity calculation module for calculating a first similarity, a second similarity, and a third similarity according to the semantic embedding word vector group of the Chinese corpus, the semantic embedding word vector group of the Uyghur-to-Chinese translation corpus, and the semantic embedding word vector group of the Tibetan-to-Chinese translation corpus; the first similarity is the similarity between each Chinese corpus in the Chinese corpus document and each Uyghur corpus in the Uyghur corpus document; the second similarity is the similarity between each Chinese corpus in the Chinese corpus document and each Tibetan corpus in the Tibetan corpus document; the third similarity is the similarity between each Uyghur corpus in the Uyghur corpus document and each Tibetan corpus in the Tibetan corpus document; A corpus construction module for determining a multilingual comparable corpus according to the first similarity, the second similarity, the third similarity, and a set similarity threshold; The corpus construction module specifically includes: A similarity determination unit for determining, for any Chinese corpus, Uyghur corpus, and Tibetan corpus, whether the corresponding first similarity, the corresponding second similarity, and the corresponding third similarity are all greater than the set similarity threshold; A corpus construction unit for, if so, storing the corresponding Chinese corpus, the corresponding Uyghur corpus, and the corresponding Tibetan corpus into the multilingual comparable corpus; A document update unit for, if not, deleting the corresponding Chinese corpus from the Chinese corpus document, deleting the corresponding Uyghur corpus from the Uyghur corpus document, and deleting the corresponding Tibetan corpus from the Tibetan corpus document.
6. The system for constructing a multilingual comparable corpus based on similarity according to claim 5, wherein The corpus acquisition module specifically includes: A web information search unit for searching a set news website by using data scraping crawler software to obtain web information; The parsing unit is used to parse the web page information in HTML, extract the news title, news content and news time, and generate the initial corpus. The preprocessing unit is used to preprocess the initial corpus to obtain a Chinese corpus document, a Uyghur corpus document and a Tibetan corpus document.
7. A system for constructing a multilingual comparable corpus based on similarity according to claim 5, characterized in that, The corpus translation module specifically includes: The first translation unit is used to translate each Uyghur corpus in the Uyghur corpus document into a Chinese corpus text by using machine translation software to obtain a Uyghur-to-Chinese translation corpus document. The second translation unit is used to translate each Tibetan corpus in the Tibetan corpus document into a Chinese corpus text by using machine translation software to obtain a Tibetan-to-Chinese translation corpus document.
8. A system for constructing a multilingual comparable corpus based on similarity according to claim 5, characterized in that, The similarity calculation module specifically includes: The word frequency vector determination unit is used to calculate the word frequency vectors of every two semantic embedding word vector groups according to the semantic embedding word vector group of the Chinese corpus, the semantic embedding word vector group of the Uyghur-to-Chinese translation corpus and the semantic embedding word vector group of the Tibetan-to-Chinese translation corpus. The similarity calculation unit is used to calculate the first similarity, the second similarity and the third similarity according to the word frequency vectors.
Citation Information
Patent Citations
Bilingual text comparable corpus construction method
CN114118096A
Term synonym acquisition method and term synonym acquisition apparatus
US20150006157A1