A Method, Device, Equipment and Medium for Extracting Multilingual Text Terms
Through translation and alignment recognition technology, term extraction of text content in different languages is solved, and the problem of lack of multilingual term extraction in the prior art is achieved, and efficient term extraction of texts in different languages is achieved.
Patent Information
- Application Number
- CN202111615844.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-12-27
AI Technical Summary
There is a lack of technical solutions for extracting terminology of text content in different languages in the prior art.
By obtaining the original text of different languages corresponding to the same text content, the original text corresponding to each language is translated into a unified language, and the standard text is obtained, and then the standard text is aligned and identified, high-frequency noun vocabulary is determined, and terms are obtained through association matching.
It realizes term extraction of text content in different languages, solves the lack of multilingual term extraction in traditional technology, and improves the cross-lingual ability of text processing.
Smart Images

Figure CN114330380B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly to a method, apparatus, device and medium for extracting terms from multi-language texts. Background Art
[0002] A terminology is a set of appellations used to represent concepts in a specific discipline field, and is generally also called a scientific and technical term. In recent years, with the wide application of AI language processing technology, the automatic extraction of terms has become a research hotspot in the field of natural language. Traditional term extraction generally extracts terms from the same article or texts in the same language, lacking research on extracting terms from text contents in different languages. Summary of the Invention
[0003] The technical problem to be solved by the present invention is that there is a lack of a technical solution for extracting terms from text contents in different languages in the prior art. Therefore, the present invention provides a method, apparatus, device and medium for extracting terms from multi-language texts to achieve the technical effect of extracting terms from text contents in different languages.
[0004] The present invention is achieved by the following technical solutions:
[0005] A method for extracting terms from multi-language texts includes:
[0006] Obtaining original texts in different languages corresponding to the same text content;
[0007] Translating the text contents of the original texts corresponding to each language into a unified language to obtain standard texts;
[0008] Performing alignment recognition on different standard texts to obtain recognition results;
[0009] When the recognition result is alignment, taking the original texts corresponding to the aligned languages as term extraction texts, performing part-of-speech analysis and word frequency statistics on the words to be processed in the term extraction texts, and determining high-frequency noun vocabulary;
[0010] Performing association relationship matching on the high-frequency noun vocabulary to obtain terms.
[0011] Further, the performing part-of-speech analysis and word frequency statistics on the words to be processed in the term extraction texts and determining high-frequency noun vocabulary includes:
[0012] Performing sentence splitting and word segmentation preprocessing on the term extraction texts to obtain the words to be processed in each term extraction text;
[0013] Performing part-of-speech analysis on the words to be processed in each term extraction text through a part-of-speech analysis tool, and selecting the words to be processed with the part of speech of noun as valid nouns;
[0014] Count the word frequency of each valid noun in the corresponding term extraction text. When the word frequency of the valid noun meets the preset high-frequency judgment condition, define the corresponding valid noun as a high-frequency noun vocabulary.
[0015] Furthermore, the matching of the association relationship of the high-frequency noun vocabulary to obtain terms includes:
[0016] Select the corresponding bilingual dictionary according to the language in the term extraction text, and query the association relationship of the high-frequency noun vocabulary through the bilingual dictionary. When a matching relationship of the high-frequency noun vocabulary is found in the bilingual dictionary, it is considered that the high-frequency noun vocabulary is a term in the corresponding language;
[0017] When no matching relationship of the high-frequency noun vocabulary is found in the bilingual dictionary, obtain the sentences in the term extraction texts of different languages of the high-frequency noun vocabulary as term judgment sentences;
[0018] When the number of term judgment sentences in different languages is the same, and the number of times the high-frequency noun vocabulary appears in the term judgment sentences of each term extraction text is the same, it is considered that the high-frequency noun vocabulary is a term in the corresponding language.
[0019] Furthermore, the preprocessing of clause segmentation and word segmentation of the term extraction text to obtain the words to be processed in each term extraction text includes:
[0020] Perform clause segmentation processing on each term extraction text according to the sentence-breaking mark to obtain the sentences of each term extraction text;
[0021] Perform word segmentation processing on the sentences in different term extraction texts and remove stop words to obtain the words to be processed in each term extraction text.
[0022] Furthermore, the alignment recognition of each standard text to obtain the recognition result includes:
[0023] Determine the semantic relationship and positional relationship of the standard words in each standard text. If the semantic relationships of the standard words in different standard texts are the same and in the same position, it is considered that the standard words are aligned;
[0024] Count the number of aligned standard words. When the number of aligned standard words meets the preset condition, it is considered that the sentences where the standard words are located are aligned, and the aligned recognition result is obtained.
[0025] Furthermore, the determination of the semantic relationship and positional relationship of the words to be processed in each standard text includes:
[0026] Perform preprocessing of clause segmentation and word segmentation on different standard texts to obtain standard words; each sentence after clause segmentation carries a sentence sequence number;
[0027] Confirm the semantic relationships of all standard words numbered in the same order in a sentence. When the standard words numbered in the same order in different standard texts are the same or are synonyms or near-synonyms, it indicates that the semantic relationships of these standard words are consistent;
[0028] When the semantic relationships of the standard words are consistent, judge the positional relationships of the corresponding standard words;
[0029] When the standard words with consistent semantic relationships are in the same positions in their respective sentences, it indicates that the standard words with consistent semantic relationships are in the same positions in their respective sentences;
[0030] When the standard words with consistent semantic relationships are in different positions in their respective sentences, it indicates that the standard words with consistent semantic relationships are in different positions in their respective sentences.
[0031] A multi-language text term extraction device, comprising:
[0032] An original text acquisition module, configured to acquire original texts in different languages corresponding to the same text content;
[0033] An original text processing module, configured to translate the text content of the original texts corresponding to each language into a unified language to obtain standard texts;
[0034] A text alignment recognition module, configured to perform alignment recognition on different standard texts to obtain recognition results;
[0035] A high-frequency noun vocabulary determination module, configured to, when the recognition result is alignment, use the original texts corresponding to the aligned languages as term extraction texts, perform part-of-speech analysis and word frequency statistics on the words to be processed in the term extraction texts, and determine high-frequency noun vocabularies;
[0036] A term acquisition module, configured to perform association relationship matching on the high-frequency noun vocabularies to obtain terms.
[0037] Further, the high-frequency noun vocabulary determination module includes:
[0038] A term extraction text processing unit, configured to perform sentence splitting and word segmentation preprocessing on the term extraction texts to obtain the words to be processed in each term extraction text;
[0039] A part-of-speech analysis unit, configured to perform part-of-speech analysis on the words to be processed in each term extraction text through a part-of-speech analysis tool, and select the words to be processed with the part of speech of noun as valid nouns;
[0040] A word frequency analysis unit, configured to count the word frequencies of each valid noun in the corresponding term extraction texts, and when the word frequency of a valid noun meets a preset high-frequency judgment condition, define the corresponding valid noun as a high-frequency noun vocabulary.
[0041] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above multi-lingual text term extraction method is implemented.
[0042] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the above multi-lingual text term extraction method is implemented.
[0043] The present invention provides a multi-lingual text term extraction method, device, equipment, and medium. By obtaining the original texts in different languages corresponding to the same text content, translating the text content of the original texts corresponding to each language into a unified language to obtain standard texts, then performing sentence splitting and word segmentation preprocessing on each standard text to obtain standard words of different standard texts, and then performing alignment recognition on the standard words in different standard texts to confirm whether the text contents in different languages are aligned. When the recognition result is aligned, the original texts corresponding to the aligned languages are used as term extraction texts, performing part-of-speech analysis and word frequency statistics on the words to be processed in the term extraction texts to determine high-frequency noun vocabulary, and then performing correlation relationship matching on the high-frequency noun vocabulary to achieve the technical effect of extracting terms from the text contents in different languages. Description of the Drawings
[0044] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not constitute a limitation on the embodiments of the present invention. In the drawings:
[0045] Figure 1 is a flowchart of a multi-lingual text term extraction method of the present invention.
[0046] Figure 2 is Figure 1 a specific flowchart of step S50 in
[0047] Figure 3 is Figure 1 a specific flowchart of step S60 in
[0048] Figure 4 is Figure 1 a specific flowchart of step S40 in
[0049] Figure 5 is Figure 4 a specific flowchart of step S41 in
[0050] Figure 6 is a structural schematic diagram of a multi-lingual text term extraction device of the present invention.
[0051] Figure 7This is a schematic diagram of the computer device of the present invention. Detailed implementation manners
[0052] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with embodiments and drawings. The illustrative embodiments of the present invention and their descriptions are only used to explain the present invention and are not intended to limit the present invention.
[0053] The present invention provides a method for extracting multi - language text terms. This method can be applied to different electronic devices, including but not limited to various personal computers, laptop computers, smart phones and tablet computers.
[0054] In one embodiment, as Figure 1 shown, the present invention provides a method for extracting multi - language text terms, including:
[0055] S10: Obtain the original texts in different languages corresponding to the same text content.
[0056] S20: Translate the text content of the original texts corresponding to each language into a unified language to obtain standard texts.
[0057] S30: Perform alignment recognition on different standard texts to obtain recognition results.
[0058] S40: When the recognition result is alignment, use the original texts corresponding to the aligned languages as term - extraction texts, perform part - of - speech analysis and word - frequency statistics on the words to be processed in the term - extraction texts, and determine high - frequency noun vocabulary.
[0059] S50: Perform association - relationship matching on the high - frequency noun vocabulary to obtain terms.
[0060] As an example, the original texts in step S10 are texts with the same text content but different languages used for the text content, that is, texts with the same content but different languages.
[0061] As an example, in step S20, the language to be translated can be set according to actual needs, or a language can be randomly selected from the languages of the original texts as the standard language, and the text content of other languages in the original texts is translated into the standard language, so as to translate the original texts in different languages into texts in the same language to obtain standard texts.
[0062] As an example, in step S30, since the text content in the original text only uses different languages, but the specific text content is the same. In theory, after translating the texts in different languages with the same text content into the same language, the order of paragraphs in the text, the order of sentences in each paragraph, and the order of words in each sentence should be the same. However, differences may occur during the actual translation process. Therefore, in order to verify the accuracy of the translation and ensure the smooth progress of subsequent multi-language term extraction, after obtaining the words to be processed for each standard text, it is necessary to perform alignment recognition on the words to be processed in different standard texts to ensure the accuracy of the translation and provide a reliable data source for subsequent multi-language text term extraction.
[0063] As an example, in step S40, when the recognition result of determining whether the standard texts are aligned is aligned, it means that during the process of translating the original texts in different languages into the standard texts, the translation results are accurate, that is, when the recognition result is aligned, the original texts corresponding to the aligned languages are used as the text for term extraction, and part-of-speech analysis and word frequency statistics are performed on the words to be processed in the text for term extraction. When the part-of-speech of the words to be processed in the text for term extraction is a noun and the word frequency of the nouns to be processed is higher than the preset condition, it means that the importance of the word to be processed in the text for term extraction meets the requirements, and the word to be processed is considered a definite high-frequency noun vocabulary.
[0064] As an example, in step S50, since the high-frequency noun vocabulary in the text for term extraction does not necessarily mean that it is a term, after determining the high-frequency noun vocabulary in the text for term extraction, it is also necessary to judge the high-frequency noun vocabulary through the set term judgment criteria to obtain the required terms from the high-frequency noun vocabulary.
[0065] The term judgment criteria set in this embodiment are specifically as follows: Since some terms are not suitable for translation and are all literal translations, therefore, it is possible to query whether the expressions of the high-frequency noun vocabulary in the text for term extraction are consistent through a bilingual dictionary. When there is a matching relationship for the high-frequency noun vocabulary in the text for term extraction in different languages in the bilingual dictionary, such as "PDF", "excel", "MP3", etc., it means that the high-frequency noun vocabulary is a technical term.
[0066] In one embodiment, as Figure 2 shown, in step S50, performing part-of-speech analysis and word frequency statistics on the words to be processed in each standard text to determine the high-frequency noun vocabulary specifically includes the following steps:
[0067] S51: Perform sentence splitting and word segmentation preprocessing on the text for term extraction to obtain the words to be processed for each text for term extraction.
[0068] S52: Use a part-of-speech analysis tool to perform part-of-speech analysis on the words to be processed extracted from the text of each term, and select the words to be processed with the part-of-speech of noun as valid nouns.
[0069] S53: Count the word frequency of each valid noun in the corresponding term extraction text. When the word frequency of the valid noun meets the preset high-frequency judgment condition, define the corresponding valid noun as a high-frequency noun vocabulary.
[0070] Among them, the words to be processed refer to the words that need to be extracted for terms after clause segmentation, word segmentation, and removal of stop words from the term extraction text.
[0071] Word frequency refers to the number of times a word appears in a text file or a corpus, used to evaluate the repetition degree of a word for a file or a set of domain files in a corpus.
[0072] As an example, in step S51, after obtaining the term extraction text, perform clause segmentation on each term extraction text according to the sentence-breaking marks to obtain the sentences of each term extraction text; perform word segmentation on the sentences in different term extraction texts and remove the stop words to obtain the words to be processed for each term extraction text.
[0073] As an example, the preset high-frequency judgment condition in step S53 can be set according to the actual situation. It can be set as a preset high-frequency number, that is, the number of times preset to meet the high-frequency requirement judgment. When the word frequency of the valid noun is greater than the preset high-frequency number, it is considered that the valid noun is a high-frequency noun vocabulary; it can also be set as a preset high-frequency ratio, that is, the ratio preset to meet the high-frequency requirement judgment. When the ratio of the word frequency of the valid noun in its corresponding standard text is greater than the preset high-frequency ratio, it is considered that the valid noun is a high-frequency noun vocabulary.
[0074] In one embodiment, as Figure 3 shown, in step S60, perform association relationship matching on the high-frequency noun vocabulary to obtain terms, specifically including the following steps:
[0075] S61: Select the corresponding bilingual dictionary according to the language type in the term extraction text, and query the association relationship of the high-frequency noun vocabulary through the bilingual dictionary. When a matching relationship of the high-frequency noun vocabulary is found in the bilingual dictionary, it is considered that the high-frequency noun vocabulary is a term in the corresponding language type.
[0076] S62: When no matching relationship of the high-frequency noun vocabulary is found in the bilingual dictionary, obtain the sentences of the high-frequency noun vocabulary in the term extraction texts of different language types as term judgment sentences.
[0077] S63: When the number of term judgment sentences for different languages is the same, and the high-frequency noun vocabulary appears the same number of times in the term judgment sentences of each term extraction text, then this high-frequency noun vocabulary is considered a term in the corresponding language.
[0078] As an example, in step S61, since some terms are not suitable for translation and are all literal translations, therefore, the expressions of some terms in the bilingual dictionary are the same. When a high-frequency noun vocabulary has a matching relationship in two different languages in the bilingual dictionary, such as "PDF", "excel", "MP3", etc., it means that this high-frequency noun vocabulary is a technical term.
[0079] Since the term extraction texts in this embodiment are in different languages, to improve the accuracy of term matching, this embodiment needs to select the corresponding bilingual dictionary according to the language of the term extraction text. After determining the bilingual dictionary, select another term extraction text for term matching according to the other language in the bilingual dictionary to complete the matching of the high-frequency noun vocabulary, so as to obtain the terms in the two term extraction texts.
[0080] As an example, in steps S62 - S63, not all terms are not suitable for translation. For some terms that can be translated, no matching relationship can be found in the bilingual dictionary. For this situation, this embodiment obtains the sentences in the term extraction texts of different languages where the high-frequency noun vocabulary appears as term judgment sentences. When the number of term judgment sentences for different languages is the same, and the high-frequency noun vocabulary appears the same number of times in the term judgment sentences of each term extraction text, then this high-frequency noun vocabulary is considered a term in the corresponding language. If the number of term judgment sentences for different languages is inconsistent, or the high-frequency noun vocabulary appears a different number of times in the term judgment sentences of each term extraction text, it means that the high-frequency noun vocabulary is not a term in the corresponding language.
[0081] For example, the high-frequency noun vocabularies in the two languages of English and Chinese are China and porcelain. Select the sentences containing the word China from the term extraction text corresponding to the English language as the term judgment sentences for the English language, and select the sentences containing the word porcelain from the term extraction text corresponding to the Chinese language as the term judgment sentences for the Chinese language. If the number of term judgment sentences for both languages is 10, and China appears 11 times in the 10 term judgment sentences of the English language, and porcelain also appears 11 times in the 10 term judgment sentences of the Chinese language, then it means that China and porcelain are terms in the corresponding languages.
[0082] In one embodiment, as Figure 4 shown, step S40, perform alignment recognition on different standard texts to obtain the recognition result, which specifically includes the following steps:
[0083] S41: Determine the semantic relationship and positional relationship of the standard words in each standard text. If the semantic relationships of the standard words in different standard texts are the same and they are in the same position, it is considered that the standard words are aligned.
[0084] S42: Count the number of aligned standard words. When the number of aligned standard words meets the preset conditions, it is considered that the sentences where the standard words are located are aligned, and the aligned recognition result is obtained.
[0085] As an example, in step S41, since the original text only expresses the same text content in different languages, after the original text is translated into the standard text, theoretically, the order of paragraphs in the text, the order of sentences in each paragraph, and the order of words in each sentence should be the same. Therefore, in this embodiment, after obtaining the standard words corresponding to each standard text, the corpus alignment result of the original text is determined by determining whether the semantic relationship and positional relationship of the standard words in each standard text are the same. If the semantic relationships of the standard words in different standard texts are the same and they are in the same position, it is considered that the standard words are aligned. If the semantic relationships of the standard words in different standard texts are not the same and / or their positions are different, it is considered that the standard words are not aligned.
[0086] Because a word may appear multiple times in a sentence, and the position of each appearance is different, and the meaning it expresses will also be different. For example, the word "meaning" in the sentence "What I mean is not what I mean" is "meaning" in both cases, but the meanings expressed by "meaning" in different positions are not the same. Therefore, when determining whether the standard words in different standard texts are aligned, not only the semantic relationship needs to be the same, but also the positions in the sentence need to be the same to more accurately judge whether the standard words in different standard texts are aligned. By determining that the semantic relationship and positional relationship of the standard words in each standard text are the same, the accuracy of judging whether the standard words are aligned can be further improved.
[0087] As an example, in step S42, if the number of the to-be-processed aligned ones is small (such as only one or two), or the proportion is small (such as only 20%), it cannot be considered that the corpora in different standard texts match. Therefore, after determining that the standard words in different standard texts are aligned, it is also necessary to count the number of aligned standard words. Only when the number of aligned standard words meets the preset conditions is it considered that the corpora in different standard texts match. If the number of aligned standard words does not meet the preset conditions, it means that the corpora in different standard texts do not match.
[0088] The preset conditions in this embodiment refer to the conditions for judging whether the number of aligned standard words meets the requirements. The preset conditions in this embodiment can be limited according to the actual situation and are not limited here. For example, it can be set as the number of aligned standard words (10), or the proportion of the number of aligned standard words in all the standard words in the sentence where they are located (more than 70%).
[0089] In one embodiment, since the text content of the original text is the same, therefore, after being translated into the standard text, the sentence order in the text content corresponding to each language should theoretically be the same. Since a standard word may appear in multiple sentences, to improve the alignment accuracy, the position of each sentence in the text content is determined by setting a sentence order number for each sentence, and then it is determined whether the to-be-processed is aligned by determining whether the positions of the standard words in the sentences with the same sentence order number are the same. As Figure 5 shown, in step S41, determine the semantic relationship and positional relationship of the standard words in each standard text, which specifically includes the following steps:
[0090] S411: Perform sentence splitting and word segmentation preprocessing on different standard texts to obtain standard words; each sentence after sentence splitting carries a sentence order number.
[0091] S412: Confirm the semantic relationship of all standard words with the same sentence order number in different standard texts. When the standard words with the same sentence order number in different standard texts are the same or are synonyms or near-synonyms, it means that the semantic relationship of the standard words is the same.
[0092] S413: When the semantic relationship of the standard words is the same, judge the positional relationship of the corresponding standard words.
[0093] S414: When the standard words with the same semantic relationship are in the same position in their respective sentences, it means that the standard words with the same semantic relationship are in the same position in their respective sentences.
[0094] S415: When the standard words with the same semantic relationship are not in the same position in their respective sentences, it means that the standard words with the same semantic relationship are in different positions in their respective sentences.
[0095] As an example, in step S411, after obtaining the standard text, perform sentence splitting on each standard text according to the sentence breaking mark to obtain the sentences of each standard text, and then perform word segmentation on the sentences in different standard texts and remove the stop words to obtain the standard words of each standard text. Any method that can perform word segmentation can be used, and it is not limited here.
[0096] As an example, in step S412, when the standard words with the same sentence order number in different standard texts are the same or are synonyms or near-synonyms, it means that the semantic relationship of the standard words is the same; when the standard words with the same sentence order number in different standard texts are not synonyms or near-synonyms, it means that the semantic relationship of the standard words is different.
[0097] As an example, in steps S413-414, when the sentence sequence numbers of the sentences where the standard words with consistent semantic relationships are located are the same and their positions in their respective sentences are the same, it indicates that the standard words with consistent semantic relationships are in the same position in their respective sentences.
[0098] As an example, in step S415, when the sentence sequence numbers of the sentences where the standard words with consistent semantic relationships are located are different, or when the sentence sequence numbers of the sentences where the standard words with consistent semantic relationships are located are the same but their positions in their respective sentences are different, it indicates that the standard words with consistent semantic relationships are in different positions in their respective sentences.
[0099] Based on the semantic relationships and positional relationships of the standard words in different standard texts, the accuracy of determining whether the standard words are aligned can be improved, providing a reliable data source for subsequent term extraction from multilingual texts.
[0100] A method for extracting terms from multilingual texts provided by the present invention includes obtaining original texts in different languages corresponding to the same text content, translating the text content of the original texts corresponding to each language into a unified language to obtain standard texts, then performing sentence splitting and word segmentation preprocessing on each standard text to obtain standard words of different standard texts, and then performing alignment recognition on the standard words in different standard texts to confirm whether the text contents in different languages are aligned. When the recognition result is alignment, perform part-of-speech analysis and word frequency statistics on the words to be processed in each term extraction text to determine high-frequency noun vocabulary, and perform correlation relationship matching on the high-frequency noun vocabulary. When no matching relationship can be found in the bilingual dictionary, obtain the sentences where the high-frequency noun vocabulary appears in the term extraction texts in different languages as term judgment sentences. When the number of term judgment sentences in different languages is the same and the number of times the high-frequency noun vocabulary appears in the term judgment sentences in each term extraction text is the same, it is considered that the high-frequency noun vocabulary is a term in the corresponding language, so as to achieve the technical effect of extracting terms from the text contents in different languages.
[0101] In one embodiment, a device for extracting terms from multilingual texts is provided, and this device for extracting terms from multilingual texts corresponds one-to-one with a method for extracting terms from multilingual texts in the above embodiment. As Figure 6 shown, this device for extracting terms from multilingual texts includes an original text acquisition module 10, an original text processing module 20, a text alignment recognition module 30, a high-frequency noun vocabulary determination module 40, and a term acquisition module 50. The detailed description of each functional module is as follows:
[0102] The original text acquisition module 10 is used to obtain original texts in different languages corresponding to the same text content.
[0103] The original text processing module 20 is used to translate the text content of the original text corresponding to each language into a unified language to obtain a standard text.
[0104] The text alignment recognition module 30 is used to perform alignment recognition on different standard texts to obtain recognition results.
[0105] The high-frequency noun vocabulary determination module 40 is used to, when the recognition result is alignment, use the original text corresponding to each aligned language as the term extraction text, perform part-of-speech analysis and word frequency statistics on the words to be processed in the term extraction text, and determine the high-frequency noun vocabulary.
[0106] The term acquisition module 50 is used to perform association relationship matching on the high-frequency noun vocabulary to obtain terms.
[0107] Furthermore, the high-frequency noun vocabulary determination module 50 includes a term extraction text processing unit, a part-of-speech analysis unit, and a word frequency analysis unit.
[0108] The term extraction text processing unit is used to perform sentence splitting and word segmentation preprocessing on the term extraction text to obtain the words to be processed in each term extraction text.
[0109] The part-of-speech analysis unit is used to perform part-of-speech analysis on the words to be processed in each term extraction text through a part-of-speech analysis tool, and select the words to be processed with the part-of-speech of noun as valid nouns.
[0110] The word frequency analysis unit is used to count the word frequency of each valid noun in the corresponding term extraction text. When the word frequency of the valid noun meets the preset high-frequency judgment condition, the corresponding valid noun is defined as the high-frequency noun vocabulary.
[0111] For the specific limitations of the multi-language text term extraction device, reference can be made to the limitations of the multi-language text term extraction method in the above text, which will not be elaborated here. Each module in the above multi-language text term extraction device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in the form of hardware or independent of it, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0112] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 7As shown. The computer device includes a processor, a memory, a network interface, and a database connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a computer-readable storage medium and an internal memory. The computer-readable storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the computer-readable storage medium. The database of the computer device is used to store the data involved in the multi-lingual text term extraction method. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a multi-lingual text term extraction method.
[0113] A computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the multi-lingual text term extraction method in the above embodiments, such as Figure 1 the steps S10 - S50 shown, or Figures 2 to 5 the steps shown in, for the sake of avoiding repetition, they will not be elaborated here. Or, when the processor executes the computer program, it implements the functions of each module / unit of the multi-lingual text term extraction device in the above embodiments, such as Figure 6 the functions of module 10 to module 50 shown. For the sake of avoiding repetition, they will not be elaborated here.
[0114] In one embodiment, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements the steps of the multi-lingual text term extraction method in the above embodiments, such as Figure 1 the steps S10 - S50 shown, or Figures 2 to 5 the steps shown in, for the sake of avoiding repetition, they will not be elaborated here. Or, when the processor executes the computer program, it implements the functions of each module / unit in this embodiment of the multi-lingual text term extraction device, such as Figure 6 the functions of module 10 to module 50 shown. For the sake of avoiding repetition, they will not be elaborated here.
[0115] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0116] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0117] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for extracting terms from multilingual texts, characterized in that, it includes: obtaining the original texts in different languages corresponding to the same text content; translating the text content of the original texts corresponding to each language into a unified language to obtain standard texts; performing alignment recognition on different standard texts to obtain a recognition result; when the recognition result is alignment, taking the original texts corresponding to the aligned languages as term extraction texts, performing part-of-speech analysis and word frequency statistics on the words to be processed in the term extraction texts, and determining high-frequency noun vocabulary; performing association relationship matching on the high-frequency noun vocabulary to obtain terms; the performing part-of-speech analysis and word frequency statistics on the words to be processed in the term extraction texts and determining high-frequency noun vocabulary includes: performing sentence splitting and word segmentation preprocessing on the term extraction texts to obtain the words to be processed in each term extraction text; performing part-of-speech analysis on the words to be processed in each term extraction text through a part-of-speech analysis tool, and selecting the words to be processed with the part-of-speech of noun as valid nouns; counting the word frequency of each valid noun in the corresponding term extraction texts, and when the word frequency of the valid noun meets the preset high-frequency judgment condition, defining the corresponding valid noun as high-frequency noun vocabulary; the performing association relationship matching on the high-frequency noun vocabulary to obtain terms includes: selecting a corresponding bilingual dictionary according to the language in the term extraction text, querying the association relationship of the high-frequency noun vocabulary through the bilingual dictionary, and when a matching relationship of the high-frequency noun vocabulary is found in the bilingual dictionary, considering that the high-frequency noun vocabulary is a term in the corresponding language; when no matching relationship of the high-frequency noun vocabulary is found in the bilingual dictionary, obtaining the sentences in the term extraction texts of the high-frequency noun vocabulary in different languages as term judgment sentences; when the number of term judgment sentences in different languages is the same and the number of times the high-frequency noun vocabulary appears in the term judgment sentences in each term extraction text is the same, considering that the high-frequency noun vocabulary is a term in the corresponding language; the performing alignment recognition on each standard text to obtain a recognition result includes: determining the semantic relationship and position relationship of the standard words in each standard text, and if the semantic relationships of the standard words in different standard texts are the same and in the same position, considering that the standard words are aligned; counting the number of aligned standard words, and when the number of aligned standard words meets the preset condition, considering that the sentences where the standard words are located are aligned, and obtaining an aligned recognition result.
2. The method for extracting terms from multilingual texts according to claim 1, characterized in that, the performing sentence splitting and word segmentation preprocessing on the term extraction texts to obtain the words to be processed in each term extraction text includes: performing sentence splitting processing on each term extraction text according to the sentence-breaking mark to obtain the sentences of each term extraction text; performing word segmentation processing on the sentences in different term extraction texts and removing stop words to obtain the words to be processed in each term extraction text.
3. The method for extracting terms from multilingual texts according to claim 1, characterized in that, Determining the semantic relationship and positional relationship of the standard words in each standard text includes: performing sentence splitting and word segmentation preprocessing on different standard texts to obtain standard words; each sentence after sentence splitting carries a sentence sequence number; Confirming the semantic relationship of all standard words with the same sentence sequence number. When the standard words with the same sentence sequence number in different standard texts are the same or are synonyms or near-synonyms, it indicates that the semantic relationship of these standard words is consistent; When the semantic relationship of the standard words is consistent, judging the positional relationship of the corresponding standard words; When the positions of the standard words with consistent semantic relationships in their respective sentences are the same, it indicates that the standard words with consistent semantic relationships are in the same position in their respective sentences; When the positions of the standard words with consistent semantic relationships in their respective sentences are different, it indicates that the standard words with consistent semantic relationships are in different positions in their respective sentences.
4. A multi-lingual text term extraction device, characterized in that, it includes: An original text acquisition module for acquiring original texts in different languages corresponding to the same text content; An original text processing module for translating the text content of the original texts corresponding to each language into a unified language to obtain standard texts; A text alignment recognition module for performing alignment recognition on different standard texts to obtain recognition results; A high-frequency noun vocabulary determination module for, when the recognition result is alignment, using the original texts corresponding to the aligned languages as term extraction texts, performing part-of-speech analysis and word frequency statistics on the words to be processed in the term extraction texts, and determining high-frequency noun vocabulary; A term acquisition module for performing association relationship matching on the high-frequency noun vocabulary to obtain terms; Performing part-of-speech analysis and word frequency statistics on the words to be processed in the term extraction texts to determine high-frequency noun vocabulary, including: performing sentence splitting and word segmentation preprocessing on the term extraction texts to obtain the words to be processed in each term extraction text; Performing part-of-speech analysis on the words to be processed in each term extraction text through a part-of-speech analysis tool, and selecting the words to be processed with the part-of-speech of noun as valid nouns; Counting the word frequency of each valid noun in the corresponding term extraction texts. When the word frequency of the valid noun meets the preset high-frequency judgment condition, defining the corresponding valid noun as high-frequency noun vocabulary; Performing association relationship matching on the high-frequency noun vocabulary to obtain terms, including: selecting a corresponding bilingual dictionary according to the language in the term extraction text, querying the association relationship of the high-frequency noun vocabulary through the bilingual dictionary. When a matching relationship of the high-frequency noun vocabulary is found in the bilingual dictionary, it is considered that the high-frequency noun vocabulary is a term in the corresponding language; When no matching relationship of the high-frequency noun vocabulary is found in the bilingual dictionary, obtaining the sentences in the term extraction texts of the high-frequency noun vocabulary in different languages as term judgment sentences; When the number of term judgment sentences in different languages is the same, and the number of times the high-frequency noun vocabulary appears in the term judgment sentences in each term extraction text is the same, it is considered that the high-frequency noun vocabulary is a term in the corresponding language; Performing alignment recognition on different standard texts to obtain recognition results includes: determining the semantic relationship and positional relationship of standard words in each standard text, and if the semantic relationships of standard words in different standard texts are consistent and in the same position, it is considered that the standard words are aligned; Counting the number of aligned standard words, and when the number of aligned standard words meets a preset condition, it is considered that the sentences where the standard words are located are aligned, and an aligned recognition result is obtained.
5. A computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements the multi-lingual text term extraction method according to any one of claims 1 to 3.
6. A computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the multi-lingual text term extraction method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Bilingual corpus resource acquisition method and bilingual corpus resource acquisition system
CN102591857A