Multi-language big language model training method and multi-language question and answer method and device
By obtaining question-and-answer combination samples and training task instructions for different language types, multilingual large language models are trained, and cross-language training is achieved using English as an intermediate hub. The problem of insufficient understanding of non-main languages in the existing technology is solved, and the effect of resource saving and application scope expansion is achieved.
Patent Information
- Application Number
- CN202411966033.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-06-03
AI Technical Summary
The existing multilingual model lacks high-quality data of different languages during training, resulting in limited understanding of non-main training languages. The existing methods require a large amount of hardware resources to generate simulated multilingual sample data, which limits the scope of application.
By obtaining question-answer combination samples and training task instructions for multiple different language types, training multilingual large language models until the output results meet the training task instructions requirements. This method uses English as an intermediate hub, and only needs to translate and transform the sample data to achieve cross-language training without introducing additional high-quality small language sample data.
It improves the understanding ability of multilingual large language models across multiple languages, saves the resources required for training, and allows the model to be deployed to hardware devices with smaller resources, expanding the scope of application.
Smart Images

Figure CN120087494A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method for training a multilingual large language model, a multilingual question answering method, and an apparatus therefor. Background Art
[0002] MLLMs (Multimodal Large Language Model) models are an important branch in the field of artificial intelligence, combining technologies in multiple fields such as natural language processing and computer vision, and capable of processing and generating various types of media content, such as text, images, audio, and video. MLLMs enhance the ability of large language models to learn rich language knowledge and knowledge of different countries around the world by using large language models to process and respond to queries in multiple languages.
[0003] However, due to the lack of high-quality training data in different languages during the training process of current multilingual large models, the understanding ability of the trained multilingual large models for languages other than the main training languages such as English is significantly limited. In addition, the existing technologies for enhancing the capabilities of multilingual large models require consuming a large amount of hardware resources to generate simulated multilingual sample data, which is likely to limit the application scope of multilingual large models. Summary of the Invention
[0004] In view of this, embodiments of this application provide a method for training a multilingual large language model, a multilingual question answering method, and an apparatus therefor, to enhance the capabilities of the multilingual large language model while expanding the application scope of the multilingual large language model.
[0005] In a first aspect, embodiments of this application provide a method for training a multilingual large language model, where the method includes:
[0006] Obtain training sample data and a training task instruction, and input the training sample data and the training task instruction into an initial multilingual large language model to train the multilingual large language model until the output result of the multilingual large language model meets the requirements of the training task instruction;
[0007] Wherein, the training sample data includes multiple question-and-answer combination samples of different language types, the question-and-answer combination sample includes a question sample and an answer sample, and the language types of the question sample and the answer sample in the same question-and-answer combination sample are the same. The training task instruction is used to instruct the multilingual large language model to convert non-English sample data into English sample data and then use the English sample data as a training intermediate sample to train the multilingual large language model;
[0008] Determine the multi - language large - language model whose output result meets the requirements of the training task instruction as the target multi - language large - language model.
[0009] In a second aspect, an embodiment of the present application provides a multi - language question - answering method. The method includes:
[0010] Obtain a target question text and a target language type, where the target language is the language type of the reply text corresponding to the target question text;
[0011] Input the target question text into the target multi - language large - language model to obtain a target reply text in the target language type output by the target multi - language large - language model;
[0012] Wherein, the target multi - language large - language model is a model trained according to the multi - language large - language model training method described in the first aspect.
[0013] In a third aspect, an embodiment of the present application provides a multi - language large - language model training device. The device includes:
[0014] An acquisition module, configured to acquire training sample data and a training task instruction, and input the training sample data and the training task instruction into an initial multi - language large - language model;
[0015] A model training module, configured to train the multi - language large - language model until the output result of the multi - language large - language model meets the requirements of the training task instruction;
[0016] Wherein, the training sample data includes multiple question - answer combination samples of different language types. The question - answer combination sample includes a question sample and an answer sample. The question sample and the answer sample in the same question - answer combination sample have the same language type. The training task instruction is used to instruct the multi - language large - language model to convert non - English sample data into English sample data and then use the English sample data as training intermediate samples to train the multi - language large - language model;
[0017] A model output module, configured to determine the multi - language large - language model whose output result meets the requirements of the training task instruction as the target multi - language large - language model.
[0018] In a fourth aspect, an embodiment of the present application provides a multi - language question - answering device. The device includes:
[0019] An input module, configured to acquire a target question text and a target language type, where the target language is the language type of the reply text corresponding to the target question text;
[0020] An output module, configured to input the target problem text into a target multilingual large language model to obtain a target response text in the target language type output by the target multilingual large language model;
[0021] Wherein, the target multilingual large language model is a model trained according to the multilingual large language model training method described in the first aspect.
[0022] In a fifth aspect, an embodiment of the present application provides an electronic device, where the electronic device includes:
[0023] A processor; and
[0024] A memory storing a program,
[0025] Wherein, the program includes instructions that, when executed by the processor, cause the processor to execute the multilingual large language model training method described in the first aspect and the multilingual question and answer method described in the second aspect.
[0026] In a sixth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium storing computer instructions, characterized in that the computer instructions are used to cause a computer to execute the multilingual large language model training method described in the first aspect and the multilingual question and answer method described in the second aspect.
[0027] Advantages of the present application:
[0028] The present application provides a multilingual large language model training method, a multilingual question and answer method and apparatus. By obtaining a plurality of question and answer combination samples of different language types and training task instructions, using the question samples and answer samples in different question and answer combination samples, and combining the training effect expected by the multilingual large language model indicated by the training task instructions, the multilingual large language model is trained until the result output by the multilingual large language model meets the requirements of the training task instructions. In this way, in the case of a shortage of training sample resources for minority languages, by means of the training task instructions, non-English sample data is converted into English sample data, and the English sample data is used as an intermediate training sample to train the multilingual large language model, so as to improve the understanding ability of the multilingual large language model across multiple languages.
[0029] In addition, the present application uses the good ability of the large language model in English as an intermediate hub, and only needs to perform translation conversion on the sample data to achieve cross-language training of the large language model. Without introducing additional high-quality minority language sample data, cross-language multilingual large language model training is carried out, effectively saving the resources required for training the multilingual large language model, which is beneficial to deploying the trained target multilingual large language model to hardware devices with less resources, and helps to expand the application scope of the multilingual large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In the following description of exemplary embodiments in conjunction with the accompanying drawings, more details, features, and advantages of the present application are disclosed. In the drawings:
[0031] Figure 1 A flowchart showing a process of a multi - language large language model training method provided by an embodiment of the present application is shown;
[0032] Figure 2 A change diagram showing the change of the capabilities of a multi - language large language model obtained by training with the multi - language large language model training method provided by an embodiment of the present application is shown;
[0033] Figure 3 A flowchart showing a process of a multi - language question - answering method provided by an embodiment of the present application is shown;
[0034] Figure 4 A schematic logical structure diagram of a multi - language large language model training device provided by an embodiment of the present application is shown;
[0035] Figure 5 A schematic logical structure diagram of a multi - language question - answering device provided by an embodiment of the present application is shown;
[0036] Figure 6 A block diagram showing the structure of an exemplary electronic device that can be used to implement the embodiments of the present application is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.
[0038] It should be understood that the steps recited in the method embodiments of the present application can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this regard.
[0039] As used herein, the term "including" and its variants are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first", "second", etc. mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0040] It should be noted that the modifications of "one" and "a plurality" mentioned in this application are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0041] LLM (Large Language Models) is an artificial intelligence model trained using a large amount of data. Multilingual large language models (MLLMs) refer to large language models that can process and respond to queries in multiple languages. Although MLLMs have made breakthroughs in multilingual environments, due to the large number of languages in the world, especially the large number of minority languages, and the fact that minority languages have fewer audiences and can provide fewer training samples for MLLM model training, this leads to the training difficulty of multilingual large language models and the required hardware resources. Especially for hardware devices with limited resources, the corresponding multilingual large language models cannot be deployed for training and use due to insufficient hardware resources. In view of this, the embodiments of this application provide a multilingual large language model training method, a multilingual question-answering method and device to improve the capabilities of multilingual large language models while expanding the application scope of multilingual large language models.
[0042] Among them, in the first aspect, this application provides a multilingual large language model training method, which is applied to any electronic device with multilingual large language model training capabilities, including but not limited to personal mobile terminals, computers or servers, etc. As Figure 1 shown, this method includes the following steps:
[0043] S11. Obtain training sample data and training task instructions, and input the training sample data and the training task instructions into an initial multilingual large language model;
[0044] S12. Train the multilingual large language model until the output result of the multilingual large language model meets the requirements of the training task instructions;
[0045] Among them, the training sample data includes multiple question-and-answer combination samples of different language types. Each question-and-answer combination sample includes a question sample and an answer sample. The language types of the question sample and the answer sample in the same question-and-answer combination sample are the same. The training task instruction is used to instruct the multilingual large language model to convert the non-English sample data into English sample data and then use the English sample data as the training intermediate sample to train the multilingual large language model;
[0046] S13. Determine the multilingual large language model corresponding to the output result that meets the requirements of the training task instruction as the target multilingual large language model.
[0047] Selecting the embodiment of the present application, leveraging the good capabilities of the large language model in English as an intermediate hub, only requires translation and conversion of the sample data to achieve cross-language training of the large language model. Without introducing additional high-quality small language sample data, cross-language multilingual large language model training can be carried out, effectively saving the resources required for multilingual large language model training. This is conducive to deploying the trained target multilingual large language model to hardware devices with smaller resources and helps to expand the application scope of the multilingual large language model.
[0048] The following will exemplarily illustrate the above steps S11 to S13 with specific examples:
[0049] In the embodiment of the present application, the multilingual large language model is the above-mentioned MLLMs model, specifically a natural language processing model based on the Transformer structure. The internal network structure of this multilingual large language model can refer to existing natural language processing models of various Transformer structures, which is not the improvement point of the present application and will not be introduced in detail here. As an implementation manner, this multilingual large language model may include: a word embedding module, an encoder-decoder module, a self-attention mechanism module, a feed-forward neural network, a position encoding module, a loss function, and an optimization algorithm module. Specifically:
[0050] The word embedding (Embeddings) module is used to convert words into continuous vectors so that the neural network can process them. The words represented by the continuous vectors contain semantic information, making similar words closer in the vector space.
[0051] The encoder-decoder module, where the encoder module is responsible for processing the input text and converting it into an internal representation; the decoder module is used to convert these internal representations into output text.
[0052] The Self-Attention Mechanism module is used when processing the input sequence, enabling the model to focus on different parts of the sequence and understand the dependencies between words. This mechanism can process all words in the sequence in parallel, improving computational efficiency.
[0053] Feedforward Neural Networks: In each layer of the Transformer, the feedforward neural network is used to further process and transform the encoded representations, which are typically fully connected layers with activation functions (such as ReLU).
[0054] Positional Encoding module: Since the Transformer architecture has no sequential information, the Positional Encoding module is used to add positional encodings to the word embeddings to provide the position information of each word in the sequence. This is typically achieved through fixed positional encodings generated by sine and cosine functions or trainable positional encodings.
[0055] The loss function and optimization algorithm module are used to measure the gap between the model output and the actual target, guiding the update of model parameters. Commonly used loss functions include the Cross-Entropy Loss, and optimization algorithms such as Adam, SGD (Stochastic Gradient Descent), etc. are used to adjust model parameters based on the feedback of the loss function.
[0056] In the embodiments of this application, the input data used in the training process of the multilingual large language model includes two categories in total: training sample data and training task instructions. Among them, the training sample data can be pre-generated data specifically used for training the multilingual large language model, and the training task instructions can be instruction information input by personnel specifically responsible for model training or subsequent model users to the model, used to indicate the specific training direction and training method of the model. Among them, the training direction specifies the specific low-resource language types that the large language model needs to learn, and the training method specifies to what extent the output result of the large language model needs to reach to stop the model training.
[0057] Exemplarily, for instance, in the model application stage, the customer requires that the multilingual large language model obtained through training can give answers corresponding to 10 small languages for the same question. At this time, the customer inputs an instruction message to the multilingual large language model: Please explain what gross national product is in 10 languages including Japanese, Korean, Hindi, Thai, Italian, Spanish, Burmese, Vietnamese, Portuguese, and Czech. At this time, the multilingual large language model will give explanations of gross national product in different language versions based on these 10 languages. In the training stage, it is necessary to improve the data processing ability of the large language model in dealing with such scenarios of multiple language questions and answers through the multilingual large language model training method provided by the embodiments of the present application. At this time, the training task instruction is used to instruct the large language model to learn the above 10 languages and stop the model training until it can accurately output answers corresponding to 10 languages.
[0058] In the embodiments of the present application, the training task instruction can be an instruction message instantaneously input by the user or obtained by reading a training task instruction list stored locally. The training sample data is specifically a question-and-answer combination sample of different language types. The question-and-answer combination sample is a combined data of a question sample and an answer sample that is pre-processed through web crawling or manually in advance, or can also be generated based on the question-and-answer results of an existing large language model.
[0059] Exemplarily, a question-and-answer combination sample includes a question sample and an answer sample. The question sample and the answer sample can be sample combination data generated during the use stage of the large language model. Exemplarily, the questions raised by the user during the use of the large language model and the answers output by the large language model according to the questions raised by the user can be used to construct a sample combination data, that is, a question-and-answer combination sample can be obtained. However, the quality of such data is relatively poor, and a high-quality question sample and the corresponding answer sample can be selected manually to construct a question-and-answer combination sample.
[0060] As another implementation manner, the training sample data can also be question-and-answer combination samples in different language versions obtained by translating existing question samples and answer samples of known language types. Exemplarily, if the language version of question 1 and the corresponding answer 1 is English, question 1 and answer 1 can be translated into question 1' and answer 1' in hundreds of language versions through translation to obtain question-and-answer combination samples in multiple languages.
[0061] In the embodiments of the present application, in order to distinguish question samples and answer samples in the same Q&A combination sample, as well as different Q&A combination samples, different data identifiers are used for distinction. Exemplarily, for example, a training sample data 1_QA = {Xq1_en, Xa1_en; Xq2_la1, Xa2_la1; Xq3_la2, Xa3_la2,..., Xqn_lan, Xan_lan}, where Xqn represents the nth question and Xan represents the answer corresponding to the nth question. "en" represents English, and in la1, la2 ······ lan, "la" is the abbreviation of "language", and the number after "la" specifically represents the language type. In the embodiments of the present application, different questions in the training sample data can correspond to their respective answers. Exemplarily, this training sample data 1_QA = {Xq1_en, Xa1_en; Xq2_la1, Xa2_la1; Xq3_la2, Xa3_la2,... Xqn_lan, Xan_lan}, that is, the English question Xq1 and the answer Xa1 corresponding to the English question 1, the Xq2_la1 in the la1 language and the answer Xa2_la1 corresponding to the question Xq2 in the la1 language ······.
[0062] As another implementation manner, the training sample data can be different language versions of the same question and the corresponding answer. Exemplarily, the training sample data 2_QA = {Xq1_en, Xa1_en; Xq1_la1, Xa1_la1; Xq1_la2, Xa1_la2,... Xqn_lan, Xan_lan}. Each Q&A combination sample in this training sample data 2_QA is the same question 1 corresponding to different language versions, as well as different language versions of the answer 1 corresponding to this question. The language types of the question sample and the answer sample in the same Q&A combination sample must be the same.
[0063] In this way, during the execution of step S11, by obtaining a training sample data, multiple Q&A combination samples can be obtained. While obtaining question samples in multiple languages, answer samples corresponding to the question samples of each language type can be obtained.
[0064] In the embodiments of the present application, when executing step S11 and step S12, the multi-language large language model can be trained in the following manner:
[0065] Use the multi-language large language model to parse the training task instruction to determine the task type to which the training task instruction belongs;
[0066] Based on the task type and combined with the language type of the problem sample, train the multilingual large language model until the target language type of the expected answer result output by the multilingual large language model meets the requirements of the training task instructions corresponding to the task type.
[0067] Among them, the process of the multilingual large language model parsing the training task instructions can be understood as the multilingual large language model trying to understand what processing the training task instructions input by the user require the multilingual large language model to perform, and what the expected effect is. Exemplarily, during the training process of the multilingual large language model, the user can instruct the multilingual large language model to give a corresponding English weather introduction result for "What's the weather like today". This "What's the weather like today, give a corresponding English weather introduction result" is both a question and a training task instruction. The multilingual large language model processes this training task instruction through natural language processing methods, extracts the keywords therein: today, weather, English version, weather introduction result, to determine the specific task type.
[0068] In the embodiments of the present application, the task types corresponding to the training task instructions are divided into three categories in total: monolingual Q&A tasks, non-English Q&A tasks with English as the hub, and bilingual Q&A tasks. Among them, the monolingual Q&A task can be simply understood as that the language type of the answer to be output is the same as the language type of the input problem sample. The non-English Q&A tasks with English as the hub and the bilingual Q&A tasks can be understood as converting the input problem sample or answer sample into an English version of the problem sample and answer sample, and then relying on the processing ability of the large language model on English data to determine the answer output results of different language types.
[0069] Specifically, as an implementation manner, if the task type is a monolingual Q&A task, the training of the multilingual large language model includes:
[0070] The multilingual large language model outputs a first expected answer result whose language type is the same as that of the problem sample according to the language type of the problem sample;
[0071] Adjust the model parameters of the multilingual large language model according to the first target data difference between the first expected answer result and the answer sample corresponding to the problem sample until the first target data difference is less than the preset data difference threshold.
[0072] Among them, the monolingual Q&A task aims to train the model through problem samples and answer samples of a single language type, so as to improve the model's ability to give accurate answers when dealing with problems of the same language type. It can be like Figure 2As shown in the topology diagram in the leftmost box, by simulating the questions proposed by Human, the corresponding question sample data Xq_la is input into the multilingual large language model, and then the multilingual large language model Assistant answers the corresponding Xa_la (in the model application stage, the result answered by the model is the expected answer result). The black arrows in the figure are the questions and answers of the language types that the multilingual large language model is good at processing, and the gray arrows are the questions and answers of the language types that the multilingual large language model is not good at processing. The arrow from Human to Assistant points to the input question to the multilingual large language model, and the arrow from Assistant to Human indicates that the multilingual large language model outputs the corresponding answer according to the input question.
[0073] In the embodiment of the present application, the training task instruction can be: giving the answer Xa_la according to Xq_la. At this time, the language types of the questions and answers are the same. By training the model through this monolingual Q&A task, the understanding ability of the model to handle different monolingual Q&A can be improved. If all the input question samples and answer samples are in EN (English), the ability of the model to give English answers when dealing with English Q&A can be continuously improved through this monolingual Q&A task.
[0074] If the target training language type (assumed to be la) mentioned in the input training task instruction is different from the language types of the input question samples and answer samples, at this time, a translation tool can be called to translate the question sample Xq_en and the answer sample Xa_en into the question sample Xq_la and the answer sample Xa_la corresponding to the target training language type la, and input the Xq_la and Xa_la into the multilingual large language model. The multilingual large language model outputs the expected answer result (Xa_la pre) according to the question sample Xq_la, and this expected answer result is the first expected answer result.
[0075] Then, by calculating the first target data difference between Xa_la pre and the answer sample Xa_la, and then adjusting the model parameters of the entire multilingual large language model, such as the weights of each neural network layer of the model, until the first target data difference between the first expected answer result output by the multilingual large language model and the answer sample is less than the preset data difference threshold. Specifically, through the loss function and the optimization algorithm model, the cross-entropy loss function can be used to determine whether the first target data difference is less than the preset data difference threshold. Among them, when the cross-entropy loss function converges, it can be determined that the first target data difference is less than the preset data difference threshold.
[0076] In the embodiments of the present application, the monolingual Q&A task with answer samples can be used in the pre-training stage and the instruction fine-tuning stage during the training process of the multilingual large language model, and is added to the training data according to a certain proportion to improve the ability of the multilingual large language model to give accurate answers when dealing with monolingual Q&A. Among them, the instruction fine-tuning stage refers to the process of further optimizing the pre-trained large language model according to specific tasks or domains. In the embodiments of the present application, the instruction fine-tuning stage specifically forms a Q&A combination sample with human instructions and expected outputs, so that the multilingual large language model can better understand human instructions and generate corresponding output answer results.
[0077] As an implementation manner, if the task type is a non-English Q&A task with English as the pivot, the training of the multilingual large language model includes:
[0078] If the language type of the initial question sample is non-English, the multilingual large language model translates the initial question sample into a corresponding English question sample;
[0079] The multilingual large language model outputs a second expected answer result of the same language type as the initial question sample according to the English question sample and in combination with the initial question sample;
[0080] According to the second target data difference between the second expected answer result and the answer sample corresponding to the question sample, the model parameters of the multilingual large language model are adjusted until the second target data difference is less than the preset data difference threshold.
[0081] In the embodiments of the present application, the non-English Q&A task with English as the pivot can enhance the cross-language ability of the multilingual large language model through the strong English understanding ability of the multilingual large language model. Among them, for the non-English question sample Xq_la, it is expected that the multilingual large language model outputs the answer result Xa_la of the corresponding language type la. At this time, the training task instructions can be: 1) Please translate Xq_la into the corresponding English question Xq_en. 2) Combine the English question Xa_en with the original question Xq_en to output the final second expected answer result Xa_la_pre. The non-English Q&A task with English as the pivot introduces the idea of the chain of thought, provides prompt information to the multilingual large language model in a step-by-step guiding manner, and enables the multilingual large language model to complete the output of the final second expected answer result.
[0082] When the multilingual large language model trained in this way subsequently receives the question Xq_la input by the user, it can first translate the question into the corresponding English question, and then combine the powerful English resource library to answer the question, so that the answers to English questions and non-English questions are consistent, thereby improving the cross-language ability of the model. Due to the particularity of the training task instructions, this non-English Q&A task with English as the pivot is more suitable for the instruction fine-tuning stage of the large language model.
[0083] Among them, referring to the relevant description of adjusting the model parameters according to the first target data difference, for the second target data difference between the second expected answer result and the corresponding answer sample, construct a cross-entropy loss function, and adjust the model parameters of the multilingual large language model according to the change of the function value of the cross-entropy loss function until the cross-entropy loss function converges.
[0084] As an implementation, if the task type is a bilingual Q&A task, the bilingual Q&A task stipulates that the language types of the question sample and the expected answer result output by the multilingual large language model are different, and one of the language types of the question sample and the expected answer result is English. Training the multilingual large language model includes:
[0085] If the language type of the initial question sample is non-English, the multilingual large language model translates the initial question sample into the corresponding English question sample, and translates the initial answer sample corresponding to the initial question sample into the corresponding English answer sample;
[0086] The multilingual large language model outputs a third expected answer result in English according to the English question sample;
[0087] According to the third target data difference between the third expected answer result and the English answer sample, adjust the model parameters of the multilingual large language model until the third target data difference is less than the preset data difference threshold.
[0088] As another implementation, on the basis that the above training task instruction is a bilingual Q&A task, the method further includes:
[0089] If the language type of the question sample is English and the language type of the expected answer result stipulated by the bilingual Q&A task is non-English, the multilingual large language model outputs a fourth expected answer result in English according to the question sample;
[0090] According to the fourth target data difference between the fourth expected answer result and the answer sample corresponding to the question sample, adjust the model parameters of the multilingual large language model until the fourth target data difference is less than the preset data difference threshold;
[0091] And translate the fourth expected answer result into the fifth expected answer result in the language type specified for the bilingual Q&A task.
[0092] In the embodiments of the present application, the process of adjusting the model parameters according to the third target data difference and the fourth target data difference is similar to the process of adjusting the model parameters according to the first target data difference described above, and will not be elaborated here. For the bilingual Q&A task, the purpose is to train the model using bilingual question samples and answer samples. During the process, the question samples and answer samples correspond to two different languages respectively, but it is necessary to ensure that one of the languages is English. If both are not English, but the training task instruction is in English.
[0093] Exemplarily, it can be simulated that the user inputs the training task instruction: Please answer the question Xq_la1 in la2 to the large language model. After receiving this training task instruction, the large language model can understand that the user knows English en and la2, but does not know la1. At this time, it can first translate the question Xq_la1 into the corresponding English question Xq_en, then query the expected answer result Xa_en_pre corresponding to the English question Xq_en in the English resources, and finally translate Xa_en into Xa_la2. The original answer sample Xa_la1 corresponding to Xq_la1 can also be translated into the answer sample Xa_en, and then the data difference between the expected answer result Xa_en_pre and the answer sample Xa_en is used to adjust the parameters of the model.
[0094] Selecting the embodiments of the present application to establish an alignment relationship between non-English and English through the bilingual Q&A task, and finally achieving feature alignment among various training languages. Similar to the aforementioned non-English Q&A task with English as the hub, the bilingual Q&A task is more suitable for single-round conversations in the instruction fine-tuning stage. Both the bilingual question task and the non-English Q&A task with English as the hub are to improve the ability of the multilingual large language model to output answer results that meet the requirements of the user's output task instruction when dealing with questions with different language requirements.
[0095] Specifically, it can be as Figure 2As shown in the topology diagrams in the three boxes on the right, by using English as the intermediate hub, the ability of the model Assistant to output cross-language results is enhanced. The single-language Q&A task can improve the ability of the multilingual large language model to output English answers to English questions. Then, by means of the Q&A task with English as the hub, the ability of the multilingual large language model to give non-English answers to non-English (Non-EN) questions can be improved. For the bilingual Q&A task, the ability of the multilingual large language model to give non-English answers to English questions, or the ability of the multilingual large language model to give English answers to non-English questions can be improved.
[0096] In the embodiments of the present application, the preset data difference threshold can be designed according to actual experience, and the present application does not make strict limitations.
[0097] In a second aspect, the present application provides a multilingual Q&A method, which is applied to any electronic device with a language Q&A function, including but not limited to a personal mobile terminal, a computer or a server. Among them, as Figure 3 shown, the method includes the following steps:
[0098] S31. Obtain a target question text and a target language type, where the target language is the language type of the reply text corresponding to the target question text;
[0099] S32. Input the target question text into a target multilingual large language model to obtain a target reply text of the target language type output by the target multilingual large language model;
[0100] Among them, the target multilingual large language model is a model trained according to the multilingual large language model training method described in the first aspect.
[0101] Using the embodiments of the present application, by using the multilingual large language model trained in the first aspect, the ability of the large language model to output answer results of corresponding different language types when dealing with questions of different language types can be improved.
[0102] In a third aspect, the present application provides a multilingual large language model training device. Among them, as Figure 4 shown, the device 40 includes:
[0103] An acquisition module 401, configured to acquire training sample data and a training task instruction, and input the training sample data and the training task instruction into an initial multilingual large language model;
[0104] A model training module 402, configured to train the multilingual large language model until the output result of the multilingual large language model meets the requirements of the training task instruction;
[0105] Among them, the training sample data includes multiple question-and-answer combination samples of different language types. Each question-and-answer combination sample includes a question sample and an answer sample. The language types of the question sample and the answer sample in the same question-and-answer combination sample are the same. The training task instruction is used to instruct the multilingual large language model to convert non-English sample data into English sample data and then use the English sample data as training intermediate samples to train the multilingual large language model;
[0106] The model output module 403 is used to determine the multilingual large language model corresponding to the output result that meets the requirements of the training task instruction as the target multilingual large language model.
[0107] In some possible embodiments, the model training module is specifically used for:
[0108] Use the multilingual large language model to parse the training task instruction to determine the task type to which the training task instruction belongs;
[0109] Based on the task type and combined with the language type of the question sample, train the multilingual large language model until the target language type of the expected answer result output by the multilingual large language model meets the requirements of the training task instruction corresponding to the task type.
[0110] In some possible embodiments, if the task type is a monolingual question-and-answer task, the training of the multilingual large language model includes:
[0111] The multilingual large language model outputs a first expected answer result of the same language type as the question sample according to the language type of the question sample;
[0112] According to the first target data difference between the first expected answer result and the answer sample corresponding to the question sample, adjust the model parameters of the multilingual large language model until the first target data difference is less than the preset data difference threshold.
[0113] In some possible embodiments, if the task type is a non-English question-and-answer task with English as the pivot, the training of the multilingual large language model includes:
[0114] If the language type of the initial question sample is non-English, the multilingual large language model translates the initial question sample into a corresponding English question sample;
[0115] The multilingual large language model outputs a second expected answer result of the same language type as the initial question sample according to the English question sample and in combination with the initial question sample;
[0116] Adjust the model parameters of the multilingual large language model according to the second target data difference between the second expected answer result and the answer sample corresponding to the question sample until the second target data difference is less than a preset data difference threshold.
[0117] In some possible embodiments, if the task type is a bilingual question-and-answer task, the bilingual question-and-answer task stipulates that the language types of the question sample and the expected answer result output by the multilingual large language model are different, and one of the language types of the question sample and the expected answer result is English. Training the multilingual large language model includes:
[0118] If the language type of the initial question sample is non-English, the multilingual large language model translates the initial question sample into a corresponding English question sample and translates the initial answer sample corresponding to the initial question sample into a corresponding English answer sample;
[0119] The multilingual large language model outputs an English third expected answer result according to the English question sample;
[0120] Adjust the model parameters of the multilingual large language model according to the third target data difference between the third expected answer result and the English answer sample until the third target data difference is less than a preset data difference threshold.
[0121] In some possible embodiments, if the language type of the question sample is English and the language type of the expected answer result stipulated by the bilingual question-and-answer task is non-English, the multilingual large language model outputs an English fourth expected answer result according to the question sample;
[0122] Adjust the model parameters of the multilingual large language model according to the fourth target data difference between the fourth expected answer result and the answer sample corresponding to the question sample until the fourth target data difference is less than a preset data difference threshold;
[0123] And translate the fourth expected answer result into a fifth expected answer result in the language type of the expected answer result stipulated by the bilingual question-and-answer task.
[0124] Wherein, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved in this application comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0125] Fourthly, this application provides a multilingual question-and-answer device, where, as Figure 5 shown, the device 50 includes:
[0126] An input module 501 is configured to obtain a target question text and a target language type, where the target language is the language type of the reply text corresponding to the target question text.
[0127] An output module 502 is configured to input the target question text into a target multilingual large language model to obtain a target reply text in the target language type output by the target multilingual large language model.
[0128] Wherein, the target multilingual large language model is a model trained according to the multilingual large language model training method described in the first aspect.
[0129] The names of the messages or information exchanged between multiple devices in the embodiments of the present application are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0130] In a fifth aspect, an exemplary embodiment of the present application further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, and the computer program, when executed by the at least one processor, is configured to cause the electronic device to execute the method according to the embodiments of the present application.
[0131] An exemplary embodiment of the present application further provides a non-transitory computer-readable storage medium storing a computer program, where the computer program, when executed by a processor of a computer, is configured to cause the computer to execute the method according to the embodiments of the present application.
[0132] An exemplary embodiment of the present application further provides a computer program product, including a computer program, where the computer program, when executed by a processor of a computer, is configured to cause the computer to execute the method according to the embodiments of the present application.
[0133] Reference Figure 6 , the following will describe the structural block diagram of an electronic device 600 that can be a server or a client of the present application, which is an example of a hardware device applicable to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present application described herein and / or claimed.
[0134] AsFigure 6 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0135] Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. The input unit 606 can be any type of device capable of inputting information into the electronic device 600. The input unit 606 can receive input digital or character information and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 607 can be any type of device capable of presenting information and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 608 can include, but is not limited to, a magnetic disk, an optical disk. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0136] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above. For example, in some embodiments, the foregoing multi-language large language model training method, multi-language question answering method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 600 via the ROM 602 and / or the communication unit 609. In some embodiments, the computing unit 601 can be configured to execute the foregoing multi-language large language model training method, multi-language question answering method in any other appropriate manner (for example, by means of firmware).
[0137] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0138] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0139] As used in the present application, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., a disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0140] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0141] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0142] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client - server relationship is generated by computer programs that run on the respective computers and have a client - server relationship with each other.
Claims
1. A multi-language large language model training method, characterized in that: The method comprises: Acquire training sample data and training task instructions, input the training sample data and the training task instructions into an initial multilingual large language model, and train the multilingual large language model until an output result of the multilingual large language model meets the requirements of the training task instructions; The training sample data includes a plurality of question-answer combination samples of different language types, the question-answer combination sample includes a question sample and an answer sample, the question sample and the answer sample in the same question-answer combination sample have the same language type, and the training task instruction is used to instruct the multilingual large language model to convert non-English sample data into English sample data, and then train the multilingual large language model using the English sample data as an intermediate training sample; The multilingual large language model corresponding to the output result that satisfies the requirements of the training task instruction is determined as the target multilingual large language model.
2. The method according to claim 1, characterized in that: The step of inputting the training sample data and the training task instruction into an initial multilingual large language model and training the multilingual large language model until an output result of the multilingual large language model meets the requirements of the training task instruction includes: Parsing the training task instruction using the multilingual large language model to determine the task type to which the training task instruction belongs; Based on the task type and the language type of the question sample, the multilingual large language model is trained until the target language type of the expected answer result output by the multilingual large language model meets the requirements of the training task instruction corresponding to the task type.
3. The method according to claim 2, characterized in that If the task type is a monolingual question answering task, the training of the multilingual large language model includes: The multilingual large language model outputs a first expected answer result having the same language type as the question sample according to the language type of the question sample; According to a first target data difference between the first expected answer result and the answer sample corresponding to the question sample, the model parameters of the multilingual large language model are adjusted until the first target data difference is less than a preset data difference threshold.
4. The method according to claim 2, characterized in that: If the task type is a non-English question-answering task with English as the hub, the training of the multilingual large language model includes: If the language type of the initial question sample is not English, the multilingual large language model translates the initial question sample into a corresponding English question sample; The multilingual large language model outputs a second expected answer result having the same language type as the initial question sample based on the English question sample and the initial question sample; According to a second target data difference between the second expected answer result and the answer sample corresponding to the question sample, the model parameters of the multilingual large language model are adjusted until the second target data difference is less than a preset data difference threshold.
5. The method according to claim 2, characterized in that: If the task type is a bilingual question-answering task, the bilingual question-answering task stipulates that the language types of the question sample and the expected answer result output by the multilingual large language model are different, and the language type of one item between the question sample and the expected answer result is English, the training of the multilingual large language model includes: If the language type of the initial question sample is not English, the multilingual large language model translates the initial question sample into a corresponding English question sample, and translates the initial answer sample corresponding to the initial question sample into a corresponding English answer sample; The multilingual large language model outputs a third expected answer result in English according to the English question sample; According to a third target data difference between the third expected answer result and the English answer sample, the model parameters of the multilingual large language model are adjusted until the third target data difference is less than a preset data difference threshold.
6. The method according to claim 5, characterized in that The method further comprises: If the language type of the question sample is English, and the language type of the expected answer result specified in the bilingual question-answering task is non-English, the multilingual large language model outputs a fourth expected answer result in English based on the question sample; According to a fourth target data difference between the fourth expected answer result and the answer sample corresponding to the question sample, adjusting the model parameters of the multilingual large language model until the fourth target data difference is less than a preset data difference threshold; And the fourth expected answer result is translated into a fifth expected answer result of the language type of the expected answer result specified by the bilingual question and answer task.
7. A multilingual question-answering method, characterized in that: The method comprises: Obtaining a target question text and a target language type, wherein the target language is the language type of a response text corresponding to the target question text; Inputting the target question text into a target multilingual large language model to obtain a target answer text of the target language type output by the target multilingual large language model; The target multilingual large language model is a model trained according to the method according to any one of claims 1 to 6.
8. A multi-language large language model training device, characterized in that: The device comprises: An acquisition module, used for acquiring training sample data and training task instructions, and inputting the training sample data and the training task instructions into an initial multilingual large language model; A model training module, used for training the multilingual large language model until the output result of the multilingual large language model meets the requirements of the training task instruction; The training sample data includes a plurality of question-answer combination samples of different language types, the question-answer combination sample includes a question sample and an answer sample, the question sample and the answer sample in the same question-answer combination sample have the same language type, and the training task instruction is used to instruct the multilingual large language model to convert non-English sample data into English sample data, and then train the multilingual large language model using the English sample data as an intermediate training sample; The model output module is used to determine the multilingual large language model corresponding to the output result that meets the requirements of the training task instruction as the target multilingual large language model.
9. An electronic device, characterized in that: The electronic device comprises: Processor; and Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1-6 or 7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to make a computer execute the method according to any one of claims 1-6 or 7.