Method, device and equipment for cross-language migration of language model and storage medium
By constructing pseudo-parallel corpus and adding feedforward neural networks to the language model, the problem of high cost of parallel corpus labeling in cross-language migration is solved, and a better migration effect is achieved.
Patent Information
- Application Number
- CN202510180316.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-10
AI Technical Summary
In the existing cross-language migration scheme, the labeling of parallel corpus is expensive, and the migration effect is not good compared to the model trained from the de novo after the model is migrated to the target language.
By constructing pseudo-parallel corpus, using text fragments of the first and second languages alternately arranged, the cost of data acquisition and labeling is significantly reduced, and feedforward neural networks are added to each layer of the language model to improve migration effect.
It realizes the improvement of the migration effect of the language model under the target language while reducing the cost of data acquisition and labeling, and solves the problem of high cost of parallel corpus labeling.
Smart Images

Figure CN120124643A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and particularly to a method, apparatus, device, and storage medium for cross - language transfer of a language model. Background Art
[0002] With the rapid development of artificial intelligence technology, language models play an increasingly important role in natural language processing tasks. Existing mainstream open - source large - language models are mainly developed by English - speaking countries and institutions, and these models have insufficient support for other languages in design and training. Therefore, cross - language transfer of language models is a research direction.
[0003] Existing cross - language transfer solutions usually start from a source model trained in a source language, construct a large - scale and high - quality target - language corpus, and bilingual parallel corpus (translation corpus) for training. By adjusting the data ratio of the source language / target language, careful parameter adjustment, etc., the source model is transferred to the target language to obtain a language model in the target language.
[0004] The problems with the above - mentioned method are that the annotation cost of parallel corpus is high, the scale is difficult to expand, and after the model is transferred to the target language, there is a large gap compared with a language model trained from scratch in the target language, that is, the transfer effect is not good. Summary of the Invention
[0005] The present disclosure provides a method, apparatus, device, and storage medium for cross - language transfer of a language model. By constructing a pseudo - parallel corpus with the vocabulary of the first language and the vocabulary of the second language, the cost of data acquisition and annotation is significantly reduced, and the problem of high annotation cost of parallel corpus is solved.
[0006] According to one aspect of the embodiments of the present disclosure, a method for cross - language transfer of a language model is provided. The method includes:
[0007] Obtaining a pseudo - parallel corpus, where the pseudo - parallel corpus includes text segments of the first language and text segments of the second language arranged alternately;
[0008] Adding a first feed - forward neural network to each layer of the first model to obtain a second model. The first model is a language model trained based on the first language, and each layer of the first model includes a second feed - forward neural network parallel to the first feed - forward neural network. The second feed - forward neural network is used to process data in the first language, and the first feed - forward neural network is used to process data in the second language;
[0009] Training the second model based on the pseudo - parallel corpus.
[0010] According to another aspect of the embodiments of the present disclosure, there is provided an apparatus for cross - language transfer of a language model, the apparatus comprising:
[0011] An acquisition unit configured to acquire pseudo - parallel corpora, the pseudo - parallel corpora including text segments of a first language and text segments of a second language arranged alternately;
[0012] A model adjustment unit configured to add a first feed - forward neural network to each layer of the first model to obtain a second model, the first model being a language model trained based on the first language, and each layer of the first model including a second feed - forward neural network juxtaposed with the first feed - forward neural network, the second feed - forward neural network being used to process data of the first language, and the first feed - forward neural network being used to process data of the second language;
[0013] A training unit configured to train the second model based on the pseudo - parallel corpora.
[0014] In some embodiments, the acquisition unit is configured to acquire a sample document, the sample document including text segments of the first language;
[0015] Replace at least one text segment in the sample document with a text segment of the second language to obtain the pseudo - parallel corpora.
[0016] In some embodiments, the acquisition unit is configured to randomly sample the sample document to obtain a plurality of first - language segments with an average length of a first length and a total length of a target random length, each first - language segment including at least one vocabulary of the first language; translate the plurality of first - language segments into a plurality of second - language segments, each second - language segment including at least one vocabulary of the second language; and replace the plurality of first - language segments in the sample document with the plurality of second - language segments to obtain the pseudo - parallel corpora.
[0017] In some embodiments, the acquisition unit is configured to segment the sample document to obtain a plurality of first - language vocabularies; translate some of the plurality of first - language vocabularies into second - language vocabularies; and replace the some of the plurality of first - language vocabularies with the translated second - language vocabularies to obtain the pseudo - parallel corpora.
[0018] In some embodiments, the training unit is configured to freeze the parameters of each layer of the first model except for the first feedforward neural network; input the pseudo-parallel corpus into the second model, and based on the first feedforward neural network in the second model, process the text segments in the pseudo-parallel corpus that belong to the second language; based on the second feedforward neural network in the second model, process the text segments in the pseudo-parallel corpus that belong to the first language; and train the second model according to the output result of the second model.
[0019] In some embodiments, the training unit is configured to train the first feedforward neural network in the second model according to the output result of the second model; unfreeze all the parameters in the second model; and use a target learning rate to adjust all the parameters in the second model based on the pseudo-parallel corpus.
[0020] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, which includes:
[0021] One or more processors;
[0022] A memory for storing executable program code of the processor;
[0023] Wherein, the processor is configured to execute the program code to implement the method for cross-lingual transfer language model as described above.
[0024] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method for cross-lingual transfer language model as described above.
[0025] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, which implements the method for cross-lingual transfer language model when executed by a processor.
[0026] The embodiments of the present disclosure provide a solution for cross-lingual transfer language model. By constructing a pseudo-parallel corpus with text segments in the first language and text segments in the second language, the cost of data acquisition and annotation is significantly reduced, and the problem of high annotation cost of parallel corpus is solved. Moreover, since the pseudo-parallel corpus includes text segments in the first language and text segments in the second language, and the first feedforward neural network added in the first model is used to process the second language, and the original second feedforward neural network in the first model is used to process the first language, the trained second model has good capabilities in the second language, that is, a good transfer effect is achieved.
[0027] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an undue limitation to the present disclosure.
[0029] Figure 1 is a schematic diagram of an implementation environment of a method for cross - language transfer language model shown according to an exemplary embodiment.
[0030] Figure 2 is a flowchart of a method for cross - language transfer language model shown according to an exemplary embodiment.
[0031] Figure 3 is a flowchart of another method for cross - language transfer language model shown according to an exemplary embodiment.
[0032] Figure 4 is a schematic diagram of a model structure provided according to an exemplary embodiment.
[0033] Figure 5 is a block diagram of a device for cross - language transfer language model shown according to an exemplary embodiment.
[0034] Figure 6 is a block diagram of an electronic device shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0036] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above - mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0037] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the vocabulary and documents involved in this disclosure are obtained under full authorization.
[0038] Figure 1 It is a schematic diagram of an implementation environment of a method for a cross - language transfer language model shown according to an exemplary embodiment. Refer to Figure 1 , and this implementation environment specifically includes: an electronic device 101 and a server 102. The electronic device 101 can be connected to the server 102 through a wireless network or a wired network.
[0039] The electronic device 101 can be at least one of devices such as a smart phone, a smart watch, a desktop computer, a laptop computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, and a laptop portable computer.
[0040] The electronic device 101 can generally refer to one of multiple electronic devices. In this embodiment, the electronic device 101 is used as an example for illustration. Those skilled in the art can know that the number of the above - mentioned electronic devices can be more or less. For example, the above - mentioned electronic devices can be several, or dozens or hundreds of the above - mentioned electronic devices, or more. The embodiments of the present disclosure do not limit the number and type of electronic devices.
[0041] The server 102 is at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Optionally, the number of the above - mentioned servers can be more or less, and the embodiments of the present disclosure do not limit this. Of course, the server 102 can also include other functional servers to provide more comprehensive and diverse services. In some embodiments, the server 102 undertakes the main computing work, and the electronic device 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing work, and the electronic device 101 undertakes the main computing work; or, the server 102 and the electronic device 101 adopt a distributed computing architecture for collaborative computing. The server 102 can be connected to the electronic device 101 and other electronic devices through a wireless network or a wired network. Optionally, the number of the above - mentioned servers can be more or less, and the embodiments of the present disclosure do not limit this.
[0042] Figure 2 is a flowchart of a method for cross - language transfer of a language model shown according to an exemplary embodiment. As Figure 2 shown, the method is executed by an electronic device and includes the following steps:
[0043] In step S201, pseudo - parallel corpora are obtained.
[0044] In the embodiments of the present disclosure, the pseudo - parallel corpora include text segments of a first language and text segments of a second language arranged alternately. Wherein, each text segment includes at least one vocabulary. That is to say, the pseudo - parallel corpora include both the vocabulary of the first language and the vocabulary of the second language, and the data in the parallel corpora is semantically complete and smooth, rather than randomly piled up with vocabulary.
[0045] In step S202, a first feed - forward neural network is added to each layer of the first model to obtain a second model.
[0046] In the embodiments of the present disclosure, each layer of the first model itself includes a second feed - forward neural network, and the first model is a language model trained based on the first language. By adding a first feed - forward neural network to each layer of the first model, a second model can be obtained. Correspondingly, each layer of the second model includes a second feed - forward neural network juxtaposed with the first feed - forward neural network. The second feed - forward neural network is used to process data in the first language, and the first feed - forward neural network is used to process data in the second language.
[0047] In step S203, the second model is trained based on the pseudo - parallel corpora.
[0048] In the embodiments of the present disclosure, by inputting the pseudo - parallel corpora into the second model, and then different feed - forward neural networks process text segments in different languages in the pseudo - parallel corpora, the first feed - forward neural network can learn the features of the second language, thereby obtaining the second model.
[0049] The embodiments of the present disclosure provide a solution for cross - language transfer of a language model. By constructing pseudo - parallel corpora with text segments of the first language and text segments of the second language, the cost of data acquisition and annotation is significantly reduced, and the problem of high annotation cost of parallel corpora is solved. And, since the pseudo - parallel corpora include text segments of the first language and text segments of the second language, and the first feed - forward neural network added to the first model is used to process the second language, and the original second feed - forward neural network in the first model is used to process the first language, the second model obtained by training has good capabilities in the second language, that is, a good transfer effect is achieved.
[0050] In some embodiments, obtaining the pseudo - parallel corpora includes:
[0051] Obtain a sample document, where the sample document includes text segments in a first language;
[0052] Replace at least one text segment in the sample document with a text segment in a second language to obtain pseudo-parallel corpus.
[0053] In the embodiments of the present disclosure, by obtaining a sample document and replacing some of the text segments therein with a second language, pseudo-parallel corpus can be obtained, which can simulate the corresponding relationship between two languages in a relatively simple manner, provide basic materials for preliminary exploration, simple training and testing in natural language processing, and improve the model training efficiency.
[0054] In some embodiments, replacing at least one text segment in the sample document with a text segment in a second language to obtain pseudo-parallel corpus includes:
[0055] Randomly sample the sample document to obtain multiple first-language segments with an average length of a first length and a total length of a target random length, and each first-language segment includes at least one vocabulary in the first language;
[0056] Translate the multiple first-language segments into multiple second-language segments, and each second-language segment includes at least one vocabulary in the second language;
[0057] Replace the multiple first-language segments in the sample document with the multiple second-language segments to obtain pseudo-parallel corpus.
[0058] In the embodiments of the present disclosure, by randomly sampling, translating and replacing the first-language segments of the sample document to generate pseudo-parallel corpus, it helps to enrich the corpus resources, and thus improves the efficiency of model training.
[0059] In some embodiments, replacing at least one vocabulary in the sample document with a vocabulary in a second language to obtain pseudo-parallel corpus includes:
[0060] Segment the sample document to obtain multiple first-language vocabularies;
[0061] Translate some of the multiple first-language vocabularies into second-language vocabularies;
[0062] Replace some of the multiple first-language vocabularies with the translated second-language vocabularies to obtain pseudo-parallel corpus.
[0063] In the embodiments of the present disclosure, by successively segmenting, translating and replacing some of the vocabularies of the sample document to obtain pseudo-parallel corpus, it can represent the corresponding relationship between different languages, provide a reference basis for model training, and improve the model training efficiency.
[0064] In some embodiments, training the second model based on pseudo-parallel corpora includes:
[0065] Freezing the parameters in each layer of the first model except for the first feed-forward neural network;
[0066] Inputting the pseudo-parallel corpora into the second model, and processing the text segments in the second language in the pseudo-parallel corpora based on the first feed-forward neural network in the second model;
[0067] Processing the text segments in the first language in the pseudo-parallel corpora based on the second feed-forward neural network in the second model;
[0068] Training the second model according to the output result of the second model.
[0069] In the embodiments of the present disclosure, by freezing some parameters, using the pseudo-parallel corpora to respectively process the text segments in the corresponding languages based on different feed-forward neural networks in the second model, and training the first feed-forward neural network according to the output result to obtain the second model, it helps to more accurately mine the associations between languages, improve the processing ability of the model for the second language, and can efficiently construct the second model to a certain extent using the existing model structure.
[0070] In some embodiments, training the second model according to the output result of the second model includes:
[0071] Training the first feed-forward neural network in the second model according to the output result of the second model;
[0072] Thawing all the parameters in the second model;
[0073] Using the target learning rate to adjust all the parameters in the second model based on the pseudo-parallel corpora.
[0074] In the embodiments of the present disclosure, by training all the parameters with a smaller learning rate in the second stage, it can not only prevent the model from having catastrophic forgetting in the first language, but also promote better alignment between the second language and the first language through fine-tuning of all the parameters. The process of first training the first feed-forward neural network, then thawing all the parameters and adjusting all the parameters based on the pseudo-parallel corpora to obtain the second model helps to improve the comprehensive processing ability of the model for the two languages and the language conversion effect, and improves the training efficiency of the model.
[0075] The above Figure 2 shows a flowchart of a method for cross-lingual transfer language model of the present disclosure. The following further elaborates on the cross-lingual transfer language model solution provided by the present disclosure. Figure 3 is a flowchart of another method for cross-lingual transfer language model shown according to an exemplary embodiment. Refer to Figure 3The method is performed by an electronic device and comprises the following steps:
[0076] In step S301, a sample document is obtained.
[0077] In the disclosed embodiments, a sample document refers to a text document written in a certain original language, such as a scientific paper originally written in English, a novel written in French, a business report drafted in Chinese, etc., where English, French, and Chinese are the first languages of the corresponding documents. The sample document includes vocabulary in the first language. Obtaining the sample document means obtaining these documents in the original language state through certain channels and methods.
[0078] It should be noted that the above sample documents are obtained with full authorization. The sample documents can be academic papers, classic works or news reports, etc.
[0079] In step S302, at least one text segment in the sample document is replaced with a text segment in a second language to obtain a pseudo-parallel corpus.
[0080] In the disclosed embodiment, an operation is performed based on a sample document, and at least one text segment is selected and replaced with an expression corresponding to a second language. The second language here refers to another language different from the first language. For example, if the sample document is in English, the second language may be Chinese; if the first language is in Japanese, the second language may be Korean, etc. Optionally, the second language may be a language with fewer corpus resources and it is difficult to construct a parallel corpus.
[0081] For example, taking a simple English sentence as an example, the sentence in the sample document is "The book is on the table." If the word "book" is replaced with the Chinese word "书", the sentence becomes "The book is on the table." This is a simple manifestation of vocabulary replacement.
[0082] It should be noted that in actual applications, when replacing words, it is necessary to consider many factors such as the grammatical properties of the words and the contextual adaptation. You cannot replace them arbitrarily, which will lead to incoherent sentences or semantic confusion. For example, for words with part of speech changes, singular and plural forms, tense expressions, etc., they should be replaced accurately according to the grammatical norms of the second language to ensure that the replaced sentences are as reasonable as possible in terms of structure and semantic understanding.
[0083] After replacing the text fragments in the sample document with text fragments in the second language, the resulting text is a pseudo-parallel corpus. The following is an introduction to pseudo-parallel corpus.
[0084] Parallel corpora usually refer to text content in two different languages, which correspond and match each other in terms of semantics, paragraph structure, etc., and can reflect the same expressive function. They are generally formed through formal translation, careful alignment and arrangement, etc. For example, a formally translated Chinese novel and its corresponding English translation can correspond to each other sentence by sentence and paragraph by paragraph, which is convenient for many natural language processing-related research and training work. For example, machine translation model training requires a large amount of high-quality parallel corpora to learn the conversion rules between the two languages.
[0085] Compared with parallel corpora, pseudo-parallel corpora are a combination of texts that are not as accurate as parallel corpora, but have certain similar meanings and correspondences. The text obtained by replacing some text fragments of sample documents with text fragments of the second language does not fully comply with the rigorous rules of language conversion and high-quality correspondence requirements like regular parallel corpora. It only simulates the appearance of parallel corpora to a certain extent, so it is called pseudo-parallel corpora. Correspondingly, the construction cost of pseudo-parallel corpora is low, and it is easy to obtain large-scale pseudo-parallel corpora.
[0086] The following introduces two methods of replacing at least one text segment in a sample document with a text segment in a second language to obtain a pseudo-parallel corpus.
[0087] Method 1: First, randomly sample sample documents to obtain multiple first language segments with an average length of a first length and a total length of a target random length, each of which includes at least one vocabulary in the first language. Then, translate the multiple first language segments into multiple second language segments, each of which includes at least one vocabulary in the second language. Finally, replace the multiple first language segments in the sample document with multiple second language segments to obtain a pseudo-parallel corpus. Generating a pseudo-parallel corpus by randomly sampling, translating, and replacing the first language segments of sample documents helps to enrich corpus resources and thereby improve the efficiency of model training.
[0088] Among them, the purpose of random sampling is to randomly select a part of the content from this complete sample document. The random method is adopted here to ensure that the selected content has a certain degree of randomness and universality, and to avoid the deviation caused by deliberate selection. The first language fragments obtained by random sampling will meet certain length requirements. That is, the average length of these fragments must reach the "first length", and the total length of all fragments added up is the "target random length". Optionally, the length of the sample document X is l, the first length is μ, and the target random length is (l / number of samplings)*ratio. Among them, ratio is a hyperparameter used to control the proportion of the first language fragments obtained by sampling, and the value is between 0-1.
[0089] For example, if the first length is 10 words and the target random length is 1000 words, then for the multiple first-language segments extracted finally, on average each segment contains approximately 10 words, and the total number of words in all segments is close to 1000. And each first-language segment must contain at least one vocabulary, that is, it is not allowed to extract blank or parts without actual semantic content, and it is necessary to ensure that the extracted segments are all segments with actual language semantic functions.
[0090] For each of the first-language segments obtained by random sampling previously, use translation means (such as a bilingual dictionary or manual translation) to convert its content into another language, that is, the second language. It should be noted that here it is required that each second-language segment must also contain at least one vocabulary to ensure that the content after conversion also has practical meaning and can express certain semantics, rather than being empty or invalid text.
[0091] For example, a sample document X = [x_0, x_1, x_2, x_3, \ldots, x_l], where x_i represents a first-language word or phrase. Sample to obtain multiple first-language segments, and for each segment, use a bilingual dictionary to translate it into the second language. Replace the corresponding positions in the sample document with the translated second-language segments, discard the corresponding first-language segments, and construct a pseudo-parallel corpus X' = [x_0, x_1, y_0, y_1, y_2, x_3, x_4, y_3, \ldots] with interleaved first and second languages. Optionally, as shown in Table 1 below, Table 1 exemplarily shows the process of replacing English with Chinese.
[0092] Table 1
[0093]
[0094] Method 2: First, segment the sample document to obtain multiple first-language vocabularies. Then, translate some of the first-language vocabularies among the multiple first-language vocabularies into second-language vocabularies. Finally, replace some of the first-language vocabularies among the multiple first-language vocabularies with the translated second-language vocabularies to obtain a pseudo-parallel corpus. By sequentially performing operations of segmenting, translating some vocabularies, and replacing on the sample document to obtain a pseudo-parallel corpus, the corresponding relationship between different languages can be represented, providing a reference basis for model training and improving the model training efficiency.
[0095] Among them, word segmentation is to split continuous text content into relatively independent and meaningful minimum language units, that is, the first language vocabulary, according to the language rules and characteristics of the first language itself. For example, in Chinese, for the sentence "The weather is nice today", through word segmentation, we can get words such as "today", "weather", "really", "nice"; in English, for the sentence "He is reading a book", after word segmentation, we will get words such as "He", "is", "reading", "a", "book", etc. Different languages have different word segmentation bases and tools. For example, some can rely on professional language processing software, and some can rely on manual splitting according to grammar and semantic rules. After obtaining a large number of first language vocabulary, select a part of them for translation. Optionally, we can focus on the conversion of some key, commonly used or representative vocabulary in the two languages to select vocabulary, or select vocabulary from the perspectives of workload, ease of operation, etc. Finally, take the translated second language vocabulary and go back to the original pile of first language vocabulary, and replace the previously selected and translated corresponding first language vocabulary with the corresponding second language vocabulary.
[0096] In step S303, a first feed-forward neural network is added to each layer of the first model to obtain a second model.
[0097] In the embodiments of the present disclosure, the first model is a language model trained based on the first language. The first model is a model constructed based on artificial intelligence technology and used to process tasks related to the first language. Usually, it is built based on architectures such as neural networks. The present disclosure improves the mainstream Transformer architecture, that is, improves the first model to obtain a second model. The design principle of the second model is that FFN (Feed-Forward Neural Network, feed-forward neural network) is used as a language-specific module to store language-specific knowledge and general world knowledge. The language-agnostic module is used to model the underlying logical relationships in the text that are independent of language. Correspondingly, each layer of the second model includes a second feed-forward neural network parallel to the first feed-forward neural network. The second feed-forward neural network is used to process data in the first language, and the first feed-forward neural network is used to process data in the second language.
[0098] Next, the model structure of the second model will be introduced.
[0099] First, in each Transformer Block, an additional FFN, called the Language Expert module, is introduced. This module is exclusive to the newly introduced second language and is parallel to and non-interfering with the FFN in the original model.
[0100] Then, the input to the second model is an interleaved input sequence of the first language and the second language, which is the pseudo-parallel corpus mentioned above. The second model can selectively choose specific FFNs for forward computation, that is, {x_i} selects the original FFN, and {y_i} selects the newly introduced FFN. That is, the tokens {x_i} in the first language segment select the original FFN. The tokens {y_i} in the second language segment select the newly introduced Language Expert module.
[0101] For example, see Figure 4 as shown Figure 4 is a schematic diagram of a model structure provided according to an exemplary embodiment. As Figure 4 shown, taking a Transformer Block as an example, this Transformer Block includes two parallel FFNs and a multi-head attention mechanism. The output of the upper layer is the input of this layer. The output of the upper layer first passes through a layer normalization module, then enters the multi-head attention module, and after passing through the layer normalization module respectively, it enters two different FFNs for processing. Among them, layer normalization is a normalization technique in neural networks. The purpose of layer normalization is to normalize the neuron inputs of a certain layer in the neural network so that the mean and variance of these inputs are within a stable range.
[0102] Optionally, the model processing process is as shown in the following formula.
[0103] Input: X' = [x_0, x_1, y_0, y_1, y_2, x_3,...].
[0104] Initialization: H^0 = [e_0, e_1, e_2, e_3, e_4, e_5,...]. The initialization indication is to perform normalization processing through the layer normalization module.
[0105] The multi-head attention module of the L-th block: A L = MHAL(H L-1 );
[0106] The FFN of the L-th block:
[0107] In step S304, the second model is trained based on the pseudo-parallel corpus.
[0108] In the embodiments of the present disclosure, a two-stage training strategy is adopted. In the first stage, only the Language Expert is trained, and in the second stage, full-scale training is performed for parameter fine-tuning.
[0109] In some embodiments, in the first stage, all parameters of the first model are frozen, and only the Language Expert module is trained. Through the above pseudo-parallel corpus for training, the Language Expert module learns the knowledge of the second language. For the convenience of description, the Language Expert module is referred to as the first feedforward neural network, and the original FFN is referred to as the second feedforward neural network. Correspondingly, first, freeze the parameters in the second model except for the first feedforward neural network. Then, input the pseudo-parallel corpus into the second model, and based on the first feedforward neural network in the second model, process the text segments belonging to the second language in the pseudo-parallel corpus. Based on the second feedforward neural network in the second model, process the text segments belonging to the first language in the pseudo-parallel corpus. Finally, train the second model according to the output result of the second model. By freezing some parameters, using the pseudo-parallel corpus to respectively process the text segments of the corresponding languages based on different feedforward neural networks in the second model, and training the first feedforward neural network according to the output result to obtain the second model, it helps to more accurately discover the correlations between languages, improve the model's processing ability for the second language, and can efficiently construct the second model to a certain extent using the existing model structure.
[0110] In some embodiments, in the second stage, a relatively small learning rate (e.g., 1e^{-6}) can be used to train all parameters to prevent catastrophic forgetting of the model in the first language during the training of all parameters. Through full-scale parameter fine-tuning, better alignment between the second language and the first language can be achieved. Correspondingly, first, train the first feedforward neural network in the second model according to the output result of the second model. Then, unfreeze all parameters in the second model. Finally, use the target learning rate to adjust all parameters in the second model based on the pseudo-parallel corpus. By using a relatively small learning rate to train all parameters in the second stage, it can not only prevent catastrophic forgetting of the model in the first language, but also promote better alignment between the second language and the first language through full-scale parameter fine-tuning. From the process of first training the first feedforward neural network, then unfreezing all parameters and adjusting all parameters based on the pseudo-parallel corpus to obtain the second model, it helps to improve the model's comprehensive processing ability for the two languages and the language conversion effect, and improves the training efficiency of the model.
[0111] Embodiments of the present disclosure provide a solution for cross - language transfer of a language model. By constructing pseudo - parallel corpus with text segments in a first language and text segments in a second language, the cost of data acquisition and annotation is significantly reduced, and the problem of high annotation cost of parallel corpus is solved. Moreover, since the pseudo - parallel corpus includes text segments in the first language and text segments in the second language, and the first feed - forward neural network added to the first model is used to process the second language, and the original second feed - forward neural network in the first model is used to process the first language, the trained second model has good capabilities in the second language, that is, a good transfer effect is achieved.
[0112] Figure 5 is a block diagram of an apparatus for cross - language transfer of a language model shown according to an exemplary embodiment. As Figure 5 shown, the apparatus includes: an acquisition unit 501, a model adjustment unit 502, and a training unit 503.
[0113] The acquisition unit 501 is configured to acquire a pseudo - parallel corpus, where the pseudo - parallel corpus includes text segments in a first language and text segments in a second language arranged alternately;
[0114] The model adjustment unit 502 is configured to add a first feed - forward neural network to each layer of the first model to obtain a second model. The first model is a language model trained based on the first language. Each layer of the second model includes a second feed - forward neural network juxtaposed with the first feed - forward neural network. The second feed - forward neural network is used to process data in the first language, and the first feed - forward neural network is used to process data in the second language;
[0115] The training unit 503 is configured to train the second model based on the pseudo - parallel corpus.
[0116] In some embodiments, the acquisition unit 501 is configured to acquire a sample document, where the sample document includes text segments in the first language; and replace at least one text segment in the sample document with a text segment in the second language to obtain the pseudo - parallel corpus.
[0117] In some embodiments, the acquisition unit 501 is configured to randomly sample the sample document to obtain a plurality of first - language segments with an average length of a first length and a total length of a target random length. Each first - language segment includes at least one vocabulary in the first language; translate the plurality of first - language segments into a plurality of second - language segments, where each second - language segment includes at least one vocabulary in the second language; and replace the plurality of first - language segments in the sample document with the plurality of second - language segments to obtain the pseudo - parallel corpus.
[0118] In some embodiments, an obtaining unit 501 is configured to perform word segmentation on the sample document to obtain a plurality of first-language words; translate some of the plurality of first-language words into second-language words; and replace the some of the plurality of first-language words with the translated second-language words to obtain the pseudo-parallel corpus.
[0119] In some embodiments, a training unit 503 is configured to freeze parameters in each layer of the second model except for the first feedforward neural network; input the pseudo-parallel corpus into the second model, and process text segments belonging to the second language in the pseudo-parallel corpus based on the first feedforward neural network in the second model; process text segments belonging to the first language in the pseudo-parallel corpus based on the second feedforward neural network in the second model; and train the second model according to the output result of the second model.
[0120] In some embodiments, the training unit 503 is configured to train the first feedforward neural network in the second model according to the output result of the second model; unfreeze all parameters in the second model; and adjust all the parameters in the second model based on the pseudo-parallel corpus using a target learning rate.
[0121] The embodiments of the present disclosure provide an apparatus for cross-lingual transfer of a language model. By constructing a pseudo-parallel corpus with text segments in a first language and text segments in a second language, the cost of data acquisition and annotation is significantly reduced, and the problem of high annotation cost of parallel corpus is solved. Moreover, since the pseudo-parallel corpus includes text segments in the first language and text segments in the second language, and the first feedforward neural network added in the first model is used to process the second language, and the original second feedforward neural network in the first model is used to process the first language, the trained second model has good capabilities in the second language, that is, a good transfer effect is achieved.
[0122] It should be noted that for the apparatus for cross-lingual transfer of a language model provided in the above embodiments, only the above division of each functional unit is used for illustration. In practical applications, the above functions can be assigned to different functional units according to needs, that is, the internal structure of the electronic device is divided into different functional units to complete all or part of the functions described above. In addition, the apparatus for cross-lingual transfer of a language model provided in the above embodiments and the method embodiments of cross-lingual transfer of a language model belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be elaborated here.
[0123] Regarding the cross - language transfer language model placement device in the above - mentioned embodiments, the specific ways in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0124] In the embodiments of the present disclosure, the electronic device can be a terminal or a server. When the electronic device is a terminal, the terminal serves as the execution entity to implement the technical solutions provided by the embodiments of the present disclosure; when the electronic device is a server, the server serves as the execution entity to implement the technical solutions provided by the embodiments of the present disclosure; or, the technical solutions provided by the present disclosure are implemented through the interaction between the terminal and the server. The embodiments of the present disclosure do not limit this.
[0125] Figure 6 It is a block diagram of an electronic device shown according to an exemplary embodiment. Generally, the electronic device 600 includes a processor 601 and a memory 602.
[0126] The processor 601 may include one or more processing cores, such as a 4 - core processor, an 8 - core processor, etc. The processor 601 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field - Programmable Gate Array), PLA (Programmable Logic Array). The processor 601 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low - power processor for processing data in the standby state. In some embodiments, the processor 601 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 601 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0127] The memory 602 may include one or more computer-readable storage media, which may be non-transitory. The memory 602 may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 602 is used to store at least one program code, and the at least one program code is used to be executed by the processor 601 to implement the method for cross-language migrating a language model provided in the method embodiments of the present disclosure.
[0128] In some embodiments, the electronic device 600 may further optionally include: a peripheral device interface 603 and at least one peripheral device. The processor 601, the memory 602, and the peripheral device interface 603 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 603 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 604, a display screen 605, a camera assembly 606, an audio circuit 607, and a power supply 608.
[0129] The peripheral device interface 603 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 601 and the memory 602. In some embodiments, the processor 601, the memory 602, and the peripheral device interface 603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 601, the memory 602, and the peripheral device interface 603 may be implemented on a separate chip or circuit board, and the present embodiment does not limit this.
[0130] The radio frequency circuit 604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 604 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 604 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like. The radio frequency circuit 604 may communicate with other electronic devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: a metropolitan area network, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 604 may further include a circuit related to NFC (Near Field Communication), and the present disclosure does not limit this.
[0131] The display screen 605 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 605 is a touch display screen, the display screen 605 also has the ability to collect touch signals on or above the surface of the display screen 605. The touch signal can be input to the processor 601 as a control signal for processing. At this time, the display screen 605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 605, which is provided on the front panel of the electronic device 600; in other embodiments, there may be at least two display screens 605, which are respectively provided on different surfaces of the electronic device 600 or are in a foldable design; in still other embodiments, the display screen 605 may be a flexible display screen, which is provided on a curved surface or a folding surface of the electronic device 600. Even, the display screen 605 can also be set to an irregular non-rectangular shape, that is, an irregular-shaped screen. The display screen 605 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0132] The camera module 606 is used to capture images or videos. Optionally, the camera module 606 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the electronic device, and the rear camera is provided on the back of the electronic device. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to realize the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fused shooting functions. In some embodiments, the camera module 606 may also include a flash. The flash can be a single-color-temperature flash or a two-color-temperature flash. A two-color-temperature flash refers to a combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.
[0133] The audio circuit 607 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 601 for processing, or input to the radio frequency circuit 604 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 600. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 601 or the radio frequency circuit 604 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 607 may further include a headphone jack.
[0134] The power supply 608 is used to supply power to each component in the electronic device 600. The power supply 608 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 608 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0135] Those skilled in the art can understand that Figure 6 the structure shown in
[0136] does not limit the electronic device 600, and may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.
[0137] A computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method for cross-language migration language model described above.
[0138] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0139] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for cross-language language model migration, characterized in that: The method comprises: Acquire a pseudo-parallel corpus, wherein the pseudo-parallel corpus includes text segments in a first language and text segments in a second language arranged alternately; Adding a first feedforward neural network to each layer of the first model to obtain a second model, wherein the first model is a language model trained based on the first language, and each layer of the second model includes a second feedforward neural network parallel to the first feedforward neural network, the second feedforward neural network is used to process data of the first language, and the first feedforward neural network is used to process data of the second language; The second model is trained based on the pseudo-parallel corpus.
2. The method for cross-language language model migration according to claim 1, characterized in that: The obtaining of pseudo-parallel corpus comprises: Obtaining a sample document, wherein the sample document includes a text segment in a first language; At least one text segment in the sample document is replaced with a text segment in the second language to obtain the pseudo-parallel corpus.
3. The method for cross-language language model migration according to claim 2, characterized in that: The step of replacing at least one text segment in the sample document with a text segment in the second language to obtain the pseudo-parallel corpus includes: Randomly sampling the sample document to obtain a plurality of first language segments with an average length of a first length and a total length of a target random length, each first language segment including at least one vocabulary of the first language; translating the plurality of first language segments into a plurality of second language segments, each second language segment comprising at least one vocabulary word in the second language; The multiple first language segments in the sample document are replaced with the multiple second language segments to obtain the pseudo-parallel corpus.
4. The method for cross-language language model migration according to claim 2, characterized in that: The step of replacing at least one text segment in the sample document with a text segment in the second language to obtain the pseudo-parallel corpus includes: Segmenting the sample document to obtain a plurality of first language words; translating some of the first language words in the plurality of first language words into second language words; The part of the first language words among the plurality of first language words is replaced with the translated second language words to obtain the pseudo parallel corpus.
5. The method for cross-language language model migration according to claim 1, characterized in that: The training of the second model based on the pseudo-parallel corpus includes: Freeze the parameters of each layer of the second model except the first feedforward neural network; Inputting the pseudo-parallel corpus into the second model, and processing text segments in the pseudo-parallel corpus that belong to the second language based on the first feedforward neural network in the second model; Based on the second feedforward neural network in the second model, processing the text segments belonging to the first language in the pseudo-parallel corpus; The second model is trained according to the output result of the second model.
6. The method for cross-language language model transfer according to claim 5, characterized in that: The step of training the second model according to the output result of the second model includes: Training the first feedforward neural network in the second model according to the output result of the second model; Unfreeze all parameters in the second model; The target learning rate is used to adjust the total parameters in the second model based on the pseudo-parallel corpus.
7. A device for transferring a language model across languages, characterized in that: The device comprises: An acquisition unit is configured to acquire a pseudo-parallel corpus, wherein the pseudo-parallel corpus includes text segments in a first language and text segments in a second language that are alternately arranged; a model adjustment unit configured to add a first feedforward neural network to each layer of a first model to obtain a second model, wherein the first model is a language model trained based on the first language, each layer of the first model includes a second feedforward neural network parallel to the first feedforward neural network, the second feedforward neural network is used to process data of the first language, and the first feedforward neural network is used to process data of the second language; A training unit is configured to train the second model based on the pseudo-parallel corpus.
8. An electronic device, characterized in that: The electronic device comprises: one or more processors; a memory for storing program code executable by the processor; The processor is configured to execute the program code to implement the method for cross-language language model transfer as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method for cross-language language model transfer as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, which, when executed by a processor, implements the method for cross-language language model transfer according to any one of claims 1 to 6.