Method, device and storage medium for rapid construction of multilingual semantic understanding resource library
By utilizing Chinese resource migration technology, a multilingual semantic understanding resource library was constructed, which solved the problem of scarce multilingual text data, enabled the construction of a low-cost and efficient multilingual semantic understanding model, and improved the model's generalization ability.
Patent Information
- Application Number
- CN202211402192.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-09
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2042-11-09
AI Technical Summary
The development of multilingual language technology faces challenges such as the scarcity of multilingual text training data, high production difficulty and cost, resulting in high cost and insufficient diversity in the construction of data resources for multilingual semantic understanding modeling, which affects the generalization of the model.
By leveraging existing Chinese databases and real user data, and through multilingual translation engines and cross-lingual entity word retrieval, we construct multilingual core sentence structure libraries and entity word libraries. By combining deep learning language models and web retrieval proofreading, we reduce our reliance on high-cost multilingual resources and improve data diversity and model performance.
It enables the rapid construction of a multilingual semantic understanding resource library, reduces costs, decreases the need for manual proofreading, and ensures the performance and generalization of the multilingual semantic understanding model.
Smart Images

Figure CN115759114B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data resource construction technology for multilingual semantic understanding modeling, and in particular to a rapid construction method for a multilingual semantic understanding resource library. Background Technology
[0002] Currently, multilingual technology has gradually become a social necessity, and various translation software and software that supports multilingual input are springing up like mushrooms after rain.
[0003] However, the development of multilingual technologies faces significant challenges. One major challenge is the scarcity of multilingual text training data, which is crucial for the development process. This scarcity, coupled with the high difficulty and cost of production, makes it difficult to support the development of technologies in a large number of languages. The main reasons for this are: due to relevant constraints, some language data cannot be retrieved, necessitating manual collection, creation, and processing of related multilingual text data; and the vast diversity of languages, coupled with a shortage of linguistic experts for some languages.
[0004] Currently, the construction of text data resources for multilingual semantic understanding modeling still adopts the following methods: First, the method of manually collecting or creating multilingual text data and manually proofreading it is costly; Second, the translation route, that is, using multilingual translation technology, but the translation results need to be proofread by experts one by one, which is labor-intensive, and the translated data has insufficient diversity, which will have a negative impact on the generalization of subsequent multilingual semantic understanding models.
[0005] In view of this, the present invention is hereby proposed. Summary of the Invention
[0006] The purpose of this invention is to provide a method, device, and storage medium for rapidly constructing a multilingual semantic understanding resource library. This method can make full use of readily available Chinese interactive resources, and while ensuring the performance of the multilingual semantic understanding method, it can complete the migration of Chinese resources to the multilingual domain, thereby reducing the dependence on costly multilingual semantic understanding text resources and solving the aforementioned technical problems in the prior art.
[0007] The objective of this invention is achieved through the following technical solution:
[0008] A rapid method for constructing a multilingual semantic understanding resource library includes:
[0009] Step S1, Multilingual Text Translation:
[0010] Obtain Chinese text data with entity word annotations from existing Chinese databases and real user data as raw data;
[0011] The original data is translated using a multilingual translation engine to obtain initial multilingual text data with entity word annotation categories;
[0012] Step S2, Cross-language entity word retrieval:
[0013] The multilingual text data with entity word annotations obtained in step S1 is used to obtain multilingual text sentence structures and low-confidence multilingual entity words in the multilingual text through cross-language entity word retrieval;
[0014] Step S3, construct a multilingual semantic understanding resource library:
[0015] A multilingual core sentence structure library and a multilingual entity lexicon are constructed respectively, and the multilingual semantic understanding resource library is composed of the constructed multilingual core sentence structure library and multilingual entity lexicon;
[0016] (31) Construct a multilingual core sentence structure library:
[0017] The multilingual text sentence patterns obtained in step 2 are subjected to multilingual text normalization processing to obtain multilingual core sentence patterns, and a multilingual core sentence pattern library is constructed from the obtained multilingual core sentence patterns.
[0018] (32) Constructing a multilingual entity thesaurus:
[0019] The low-confidence multilingual entity words obtained in step 2 are processed through network retrieval and verification to obtain high-confidence multilingual entity words. The obtained high-confidence multilingual entity words are used to construct a multilingual entity thesaurus.
[0020] A processing apparatus, comprising:
[0021] At least one memory for storing one or more programs;
[0022] At least one processor is capable of executing one or more programs stored in the memory, such that when the processor executes one or more programs, the processor can implement the method of the present invention.
[0023] A readable storage medium storing a computer program that, when executed by a processor, enables the implementation of the methods described in this invention.
[0024] Compared with existing technologies, the method, apparatus, and storage medium for rapidly constructing a multilingual semantic understanding resource library provided by this invention have the following advantages:
[0025] Compared with manual collection and annotation methods, this invention utilizes existing and abundant Chinese resources to complete the resource migration in the multilingual domain, reducing dependence on costly multilingual resources and reducing costs. Compared with simple translation techniques, this method reduces the cost of manual proofreading and solves the impact of insufficient data diversity on subsequent multilingual semantic understanding models. Attached Figure Description
[0026] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 A flowchart illustrating a method for rapidly constructing a multilingual semantic understanding resource library as provided in an embodiment of the present invention.
[0028] Figure 2 A flowchart illustrating the rapid construction method for a multilingual semantic understanding resource library provided in this embodiment of the invention.
[0029] Figure 3 This is a flowchart illustrating the training process of the multilingual text normalization model in the method provided in this embodiment of the invention. Detailed Implementation
[0030] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the specific content of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments, which do not constitute a limitation of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0031] First, the following explanations are provided for the terms that may be used in this article:
[0032] The term "and / or" means that either or both can be achieved simultaneously. For example, X and / or Y means that it includes both "X" or "Y" as well as the three cases of "X and Y".
[0033] The terms “including,” “comprising,” “containing,” “having,” or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, “including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.)” should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.
[0034] The term "composed of" excludes any technical features not expressly listed. When used in a claim, it closes the claim to exclude all technical features other than those expressly listed, except for associated conventional impurities. If the term appears only in a clause of a claim, it limits the claim to the elements expressly listed in that clause; elements recited in other clauses are not excluded from the overall claim.
[0035] The terms “center,” “longitudinal,” “lateral,” “length,” “width,” “thickness,” “upper,” “lower,” “front,” “back,” “left,” “right,” “vertical,” “horizontal,” “top,” “bottom,” “inner,” “outer,” “clockwise,” and “counterclockwise” indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience and simplification of description and do not imply that the device or component referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this document.
[0036] The following is a detailed description of the rapid construction method for a multilingual semantic understanding resource library provided by this invention. Contents not described in detail in the embodiments of this invention are prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of this invention, they shall be performed according to conventional conditions in the art or conditions recommended by the manufacturer. Reagents or instruments used in the embodiments of this invention, unless otherwise specified by the manufacturer, are all commercially available conventional products.
[0037] like Figure 1 , Figure 2 As shown, this embodiment of the invention provides a method for rapidly constructing a multilingual semantic understanding resource library, including:
[0038] Step S1, Multilingual Text Translation:
[0039] Obtain Chinese text data with entity word annotations from existing Chinese databases and real user data as raw data;
[0040] The original data is translated using a multilingual translation engine to obtain initial multilingual text data with entity word annotation categories;
[0041] Step S2, Cross-language entity word retrieval:
[0042] The multilingual text data with entity word annotations obtained in step S1 is used to obtain multilingual text sentence structures and low-confidence multilingual entity words in the multilingual text through cross-language entity word retrieval;
[0043] Step S3, construct a multilingual semantic understanding resource library:
[0044] A multilingual core sentence structure library and a multilingual entity lexicon are constructed respectively, and the multilingual semantic understanding resource library is composed of the constructed multilingual core sentence structure library and multilingual entity lexicon;
[0045] (31) Construct a multilingual core sentence structure library:
[0046] The multilingual text sentence patterns obtained in step 2 are subjected to multilingual text normalization processing to obtain multilingual core sentence patterns, and a multilingual core sentence pattern library is constructed from the obtained multilingual core sentence patterns.
[0047] (32) Constructing a multilingual entity thesaurus:
[0048] The low-confidence multilingual entity words obtained in step 2 are processed through network retrieval and verification to obtain high-confidence multilingual entity words. The obtained high-confidence multilingual entity words are used to construct a multilingual entity thesaurus.
[0049] In step S2 of the above method, the multilingual text data with entity word annotations obtained in step S1 is processed through cross-language entity word retrieval to obtain multilingual entity words and multilingual text sentence structures, including:
[0050] Step S21: Using the entity words in the original Chinese text before translation, n candidate entity words for the language are obtained by translation;
[0051] Step S22: Search for each candidate entity word in the translated multilingual text data, and take the candidate entity words that can be found as low-confidence multilingual entity words;
[0052] Step S23: Replace low-confidence multilingual entity words in the translated multilingual text data with entity word annotations to obtain multilingual text sentence structures.
[0053] In step S3 of the above method, the multilingual text sentence patterns obtained in step 2 are processed by multilingual text normalization to obtain multilingual core sentence patterns in the following manner:
[0054] Multilingual text sentence patterns are input into a pre-trained multilingual text normalization model for text normalization processing to obtain multilingual core sentence patterns.
[0055] The aforementioned multilingual text normalization model consists of a deep learning language model with a binary classifier added to the last layer.
[0056] The deep learning language model described above can be either the BERT model or the ALBERT model. Other similar deep learning language models can also be used.
[0057] See Figure 3The aforementioned multilingual text normalization model is trained in the following manner:
[0058] Using existing Chinese databases and real user data as raw data, the raw data is translated by a multilingual translation engine to obtain multilingual text data;
[0059] The Chinese text is regularized using a trainable Chinese regularization model to obtain the core Chinese text.
[0060] A multilingual translation engine is used to translate the core Chinese text to obtain multilingual core text;
[0061] Multilingual text data and multilingual core text are used as the input and label of the multilingual text regularization model, respectively, to construct the training data of the multilingual text regularization model.
[0062] The multilingual text regularization model is trained using the constructed training data until the training termination condition is met. After training is completed, the parameters of the model with the best performance on the validation set are saved, thus completing the pre-training of the multilingual text regularization model.
[0063] The structure of the "trainable Chinese text normalization model" described above is similar to that of the multilingual text normalization model, consisting of a deep learning language model with a binary classifier added to the last layer. The deep learning language model used is either BERT or ALBERT. Other similar deep learning language models can also be used. The training data consists of existing Chinese text data used for training the Chinese normalization model, and the labels are existing labeled Chinese normalization data. The Chinese normalization model is trained using this data until the training termination condition is met.
[0064] In step S3 of the above method, the low-confidence multilingual entity words obtained in step 2 are processed through network retrieval and proofreading to obtain high-confidence multilingual entity words in the following manner:
[0065] Step S321: Use a search engine to search for the extracted low-confidence multilingual entity words; specifically, during the search process, enclose the multilingual entity words in quotation marks to use them as keywords for querying.
[0066] Step S322: The ratio of the number of occurrences M to N in the top N results (i.e., Top N) of the search results is used as the entity word credibility. If the entity word credibility exceeds the preset threshold α, then the entity word is marked as a high credibility word.
[0067] The M above indicates that the searched low-confidence multilingual entity word exists in M out of the first N results.
[0068] Step 323, construct a multilingual entity word library by using all the obtained high-confidence multilingual entity words.
[0069] In step S3 of the above method, the value range of the preset threshold α is 0 to 1. Preferably, the value range of the preset threshold α is 0.75 to 0.85. The preferred value of the preset threshold α in the present invention is 0.8. In practical applications, the value of the preset threshold α can be determined within the value range according to experience.
[0070] In summary, in the method of the embodiment of the present invention, since the existing Chinese database and the Chinese text data with entity word annotations in the real user data are used as the original data, the dependence on the costly multilingual semantic understanding text resources is reduced, and the easily obtained Chinese interaction resources are fully utilized. While ensuring the performance of the multilingual semantic understanding method, the migration of Chinese resources in the multilingual field is completed, and the text data resources for building the multilingual semantic understanding model can be constructed quickly and at low cost.
[0071] In order to more clearly show the technical solution provided by the present invention and the technical effects produced, the following uses specific embodiments to describe in detail the method for quickly constructing a multilingual semantic understanding resource library provided by the embodiments of the present invention.
[0072] Embodiment 1
[0073] As Figure 1 、 Figure 2 shown, the embodiment of the present invention provides a method for quickly constructing a multilingual semantic understanding resource library, which can quickly and at low cost construct text data resources for multilingual semantic understanding modeling. The method mainly includes: multilingual text translation, cross-lingual entity word retrieval, and construction of a multilingual semantic understanding resource library. Among them, in multilingual semantic understanding step S1, multilingual text translation:
[0074] The Chinese text data with entity word annotations from the existing Chinese database and real user data is translated by the existing translation engine to obtain the initial multilingual text data with entity word annotation types. Among them, multilingual text translation mainly uses the existing translation engine to achieve the purpose, which is the same as the multilingual translation technology mentioned in the background art. For example: the Chinese text data with entity word annotations "Could you please help me turn on the radio(device)", the translation engine is used to translate the Chinese text into "Would you please do me a favor to turn on the radio(device)". Among them, radio is the entity word, and device is the entity word annotation.
[0075] Step S2, cross-lingual entity word retrieval:
[0076] The multilingual text obtained through translation is processed by the cross-lingual entity word retrieval module to obtain the entity words and sentence patterns in the multilingual text. It should be noted that the text consists of two parts, namely sentence patterns and entities. The specific method can be as follows:
[0077] First, use the entity words in the Chinese text before translation to translate and obtain n candidate entity words in that language. Then, search for relevant entity words in the translated multilingual text and save them as low-confidence multilingual entity words. Finally, use entity word annotation to replace the entity words in this translated text to obtain the multilingual text sentence pattern. It should be noted that there are other ways of cross-lingual entity word retrieval technology, which will not be listed here. The example is as follows:
[0078] The multilingual text data input into this module is "Would you please do me a favor to turn on the radio" and the Chinese text entity word "收音机" before the translation of this text, and the entity word annotation of this text data is "device". Use the translation engine to translate "收音机" and get "radio" and "receiver". By searching in the translated text, the low-confidence entity word "radio" can be obtained. Through entity annotation replacement, the multilingual text sentence pattern "Would you please do me a favor to turn on the device" can be obtained.
[0079] Step S3, construct a multilingual semantic understanding resource library, including: constructing a multilingual core sentence pattern library and a multilingual entity word library. These two are parallel operations and there is no order; among them,
[0080] (31) Construction of the multilingual core sentence pattern library
[0081] Through the above steps, a large number of multilingual text sentence patterns can be obtained. However, it should be noted that there are the following problems in the initial multilingual text sentence patterns obtained through translation: the number of text sentence patterns is large, and it takes a long time and high cost for language experts to proofread one by one; after translation, the diversity of the text is reduced compared with the original text. If it is directly used for the construction of subsequent multilingual semantic understanding resources and the training of the multilingual semantic understanding model using this resource, it will have a negative impact on the generalization of the model. Therefore, this patent uses multilingual text regularization technology to solve the above problems here.
[0082] A large number of multilingual text sentence patterns are input into the multilingual text regularization module. Use the regularization technology to obtain the core sentence patterns, and thus construct a multilingual core sentence pattern library to form the training data for the subsequent multilingual semantic understanding model. So far, the first of the above problems has been solved. Since the scale of the core sentence pattern library is small, the number of texts that need to be proofread by experts is reduced, and the cost is reduced.
[0083] Then, the data constructed using core sentence patterns is used to train subsequent multilingual semantic understanding models, aiming to enable the models to learn relevant knowledge of core sentence patterns in each language. In practical application reasoning, when faced with complex multilingual input text, the regularization model regularizes it into core sentence patterns before inputting it into the subsequent semantic understanding model. This approach ensures both the small size of the multilingual core sentence pattern library and the performance of the semantic understanding model for complex inputs. This solves the second problem mentioned above: although the diversity of the constructed core sentence patterns is relatively reduced, it does not affect the performance of the subsequent multilingual semantic understanding model, while ensuring the model's generalization ability to complex inputs. To enhance understanding, this section provides an example: when the model reasones, it inputs complex text such as "Would you please do me a favor to turn on the radio," "Could you help me turn on the radio," and "Can you please...". After regularization, the core sentence pattern is always "turn on the radio," which is text (sentence pattern + entity words) that the subsequent semantic understanding model has learned and can be recognized by the model.
[0084] Therefore, the construction of a multilingual core database relies on a multilingual text regularization model. However, the training of this model depends on multilingual text regularization data. To truly eliminate the dependence on multilingual text resources throughout the process, this invention employs the following... Figure 2 The multilingual text normalization technique shown.
[0085] The process continues to use existing Chinese databases and real user data as the initial input for Chinese text, which is then translated by a translation engine to obtain complex multilingual text. Next, a trainable Chinese normalization model is used to normalize the complex Chinese text, yielding the core Chinese text.
[0086] The Chinese language normalization model used here is the BERT model, with a binary classifier added to the last layer of the BERT model to determine whether the current word in this text data constitutes a normalized sentence. The training data for this model comes from an existing training dataset. Other deep learning language models can also be used for this model structure.
[0087] Then, the translation engine is used again to translate the Chinese core text, obtaining multilingual core texts. Using the complex multilingual text and the multilingual core text as the input and label of the multilingual text regularization model respectively, the construction of the training data for the multilingual text regularization model is completed. Finally, a multilingual text regularization model is built. Here, the model structure selection is the same as the above form (Chinese regularization model) and will not be elaborated. Other deep learning language models such as ALBERT can also be used for the model structure here. For better understanding, an example of this process is given here. The Chinese data is "Hello, please help me turn on the air conditioner". After translation, the complex multilingual text for the training data of the multilingual text regularization model can be obtained, such as "Hello, please turn on the air conditioning for me". Then, the Chinese core text, that is, "Turn on the air conditioner", is obtained using the Chinese regularization model, and after translation, the multilingual core text "Turn on the air conditioning" for the above training data label is obtained. Thus, the training of the multilingual text regularization model is achieved.
[0088] The present invention utilizes the above text regularization technology to complete the migration of the text regularization function, achieving the function of multilingual text regularization while getting rid of the dependence on multilingual text resources.
[0089] (32) Construction of multilingual entity word library:
[0090] After step 2, a large number of low-confidence entity words can be obtained. Among them, the reason for naming them low-confidence entity words is mainly that there are limitations in the translation engine itself, and the entities after translation cannot guarantee their correctness. If manual quality inspection is used, there are disadvantages such as long cycle and high cost. To address this problem, the present invention adopts a network search and proofreading technology, using a search engine to retrieve the extracted multilingual entity words, thereby achieving automatic proofreading.
[0091] A large number of low-confidence entity words are obtained as high-confidence entity words through the network search and proofreading module, which are used to construct the multilingual entity word library. The specific method is as follows:
[0092] First, use a search engine to retrieve a large number of extracted low-confidence multilingual entity words. It should be noted that during the search process, quotation marks need to be added on both sides of the multilingual entity words to ensure that they are used as keywords for query. Then, it is measured according to the number of occurrences (M) in the search return result TopN. Among them, the value of M / N is the confidence of the entity word. A preset threshold α is set. If the confidence of the entity word exceeds the threshold α, the entity word is marked as a high-confidence word. Finally, use all the obtained high-confidence multilingual entity words to construct the multilingual entity word library.
[0093] At this point, the multilingual core sentence structure library and the multilingual entity lexicon have been completed, and together they form a multilingual semantic understanding resource library.
[0094] In summary, compared to the costly acquisition of multilingual data resources, Chinese data offers advantages such as abundant resources, ease of acquisition, and low cost. Therefore, the method in this invention utilizes existing Chinese databases with entity-annotated text and real user data to construct a multilingual semantic understanding modeling data resource library. In other words, the initial raw data for this invention consists of existing Chinese databases with entity-annotated text and real user data. Compared to methods that involve manual collection and annotation, this invention leverages existing and abundant Chinese resources, completing resource migration across multiple languages, reducing reliance on costly multilingual resources, and lowering costs. Compared to simple translation techniques, this method reduces the cost of manual proofreading and addresses the impact of insufficient data diversity on subsequent multilingual semantic understanding models.
[0095] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0096] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The information disclosed in the background section is intended only to enhance the understanding of the overall background technology of the present invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art.
Claims
1. A method for rapidly constructing a multilingual semantic understanding resource library, characterized in that, include: Step S1, Multilingual Text Translation: Obtain Chinese text data with entity word annotations from existing Chinese databases and real user data as raw data; The original data is translated using a multilingual translation engine to obtain initial multilingual text data with entity word annotation categories; Step S2, Cross-language entity word retrieval: The multilingual text data with entity word annotations obtained in step S1 is used to obtain multilingual text sentence structures and low-confidence multilingual entity words in the multilingual text through cross-language entity word retrieval; Step S3, construct a multilingual semantic understanding resource library: A multilingual core sentence structure library and a multilingual entity lexicon are constructed respectively, and the multilingual semantic understanding resource library is composed of the constructed multilingual core sentence structure library and multilingual entity lexicon; (31) Construct a multilingual core sentence structure library: The multilingual text sentence patterns obtained in step 2 are subjected to multilingual text normalization processing to obtain multilingual core sentence patterns, and a multilingual core sentence pattern library is constructed from the obtained multilingual core sentence patterns. (32) Constructing a multilingual entity thesaurus: The low-confidence multilingual entity words obtained in step 2 are processed through network retrieval and verification to obtain high-confidence multilingual entity words. The obtained high-confidence multilingual entity words are used to construct a multilingual entity thesaurus.
2. The method for rapidly constructing a multilingual semantic understanding resource library according to claim 1, characterized in that, In step S2, the multilingual text data with entity word annotations obtained in step S1 is processed through cross-language entity word retrieval to obtain multilingual entity words and multilingual text sentence structures, including: Step S21: Using the entity words in the original Chinese text before translation, n candidate entity words for the language are obtained by translation; Step S22: Search for each candidate entity word in the translated multilingual text data, and take the candidate entity words that can be found as low-confidence multilingual entity words; Step S23: Replace low-confidence multilingual entity words in the translated multilingual text data with entity word annotations to obtain multilingual text sentence structures.
3. The method for rapidly constructing a multilingual semantic understanding resource library according to claim 1 or 2, characterized in that, In step S3, the multilingual text sentence patterns obtained in step 2 are processed by multilingual text normalization to obtain multilingual core sentence patterns in the following manner: Multilingual text sentence patterns are input into a pre-trained multilingual text normalization model for text normalization processing to obtain multilingual core sentence patterns.
4. The method for rapidly constructing a multilingual semantic understanding resource library according to claim 3, characterized in that, The multilingual text normalization model consists of a deep learning language model with a binary classifier added to the last layer.
5. The method for rapidly constructing a multilingual semantic understanding resource library according to claim 4, characterized in that, The deep learning language model used is either the BERT model or the ALBERT model.
6. The method for rapidly constructing a multilingual semantic understanding resource library according to claim 3, characterized in that, The multilingual text normalization model is trained in the following manner, including: Using existing Chinese databases and real user data as raw data, the raw data is translated by a multilingual translation engine to obtain multilingual text data; The Chinese text is regularized using a trainable Chinese regularization model to obtain the core Chinese text. A multilingual translation engine is used to translate the core Chinese text to obtain multilingual core text; Multilingual text data and multilingual core text are used as the input and label of the multilingual text regularization model, respectively, to construct the training data of the multilingual text regularization model. The multilingual text regularization model is trained using the constructed training data until the training termination condition is met. After training is completed, the parameters of the model with the best performance on the validation set are saved, thus completing the pre-training of the multilingual text regularization model.
7. The method for rapidly constructing a multilingual semantic understanding resource library according to claim 1 or 2, characterized in that, In step S3, the low-confidence multilingual entity words obtained in step 2 are processed through network retrieval and proofreading to obtain high-confidence multilingual entity words in the following manner: Step S321: Use a search engine to search for the extracted low-confidence multilingual entity words; Step S322: The ratio of the number of occurrences M to N in the Top N search results is used as the entity word credibility. If the entity word credibility exceeds the preset threshold α, then the entity word is marked as a high credibility word. Step 323: Construct a multilingual entity thesaurus using all the obtained high-confidence multilingual entity words.
8. The method for rapidly constructing a multilingual semantic understanding resource library according to claim 7, characterized in that, In step S3, the preset threshold α is set to a value of 0.75 to 0.
85.
9. A processing device, characterized in that, include: At least one memory for storing one or more programs; At least one processor is capable of executing one or more programs stored in the memory, such that when the one or more programs are executed by the processor, the processor can perform the method according to any one of claims 1-8.
10. A readable storage medium, characterized in that, It contains a computer program that, when executed by a processor, can implement the method described in any one of claims 1-8.