Data processing method, data generating method and big language model training system
By introducing basic expert modules, expert routing modules and domain expert modules into the large language model, combined with basic answer data processing, the problem of insufficient accuracy in cross-domain problem processing in the existing technology is solved, and higher accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202411815386.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-05-13
AI Technical Summary
Existing large language models cannot accurately give answers when dealing with cross-domain problems, and the mixed expert model is too strong in data processing, which affects the accuracy of model prediction.
A large language model including basic expert modules, expert routing modules and domain expert modules is adopted to predict the basic answer data of the problem data through the basic expert module, and the expert routing module is used to select the matching target expert submodule in the domain expert module, and process it in combination with the basic answer data to obtain the answer data.
The accuracy and adaptability of the model in the processing of problem data in different fields is improved, the consumption of computing resources is reduced, and the accuracy of efficient model training and question-and-answer results is achieved.
Smart Images

Figure CN119988534A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and in particular to a data processing method, a generation method, and a large language model training system. Background Art
[0002] In recent years, the field of natural language processing has made significant progress. Open-source large-scale language models such as LLaMA, Qwen, and Yi have emerged and have achieved certain success in processing tasks in e-commerce, culture and art, psychological counseling, and scientific research. However, although these models have strong versatility and can answer questions in multiple fields, due to differences in logic, thinking, language, etc. between various fields, large language models cannot accurately give answers when dealing with cross-domain problems.
[0003] In the prior art, it is proposed to use a hybrid expert model to deal with problems in specific fields in a targeted manner. In the hybrid expert model, each expert network is a relatively independent model that is specifically responsible for processing a certain part or a certain pattern of data, while the gating network determines which expert network should be activated or the weight of each expert network in the calculation when processing the input data. However, since the hybrid expert model is too targeted during the data processing process, it will affect the accuracy of the model prediction, and the hybrid expert model has weak applicability to multi-field problems. Therefore, there is an urgent need for a more effective data processing method to solve the above problems. Summary of the invention
[0004] In view of this, an embodiment of this specification provides a data processing method. One or more embodiments of this specification also relate to a data processing system, a large language model training system, a data processing device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0005] According to a first aspect of an embodiment of this specification, a data processing method is provided, including: Inputting question data corresponding to the question-answering task into a large language model, wherein the large language model includes a basic expert module, an expert routing module, and a domain expert module; Predicting basic answer data corresponding to the question data by using the basic expert module, and selecting a target expert submodule matching the question-answering task in the domain expert module by using the expert routing module; The question data and the basic answer data serving as reference data for the target expert submodule are input into the target expert submodule for processing to obtain answer data corresponding to the question-answering task.
[0006] According to a second aspect of an embodiment of this specification, a generation method is provided, including: Inputting the to-be-processed data corresponding to the generated task into a large language model, wherein the large language model comprises a basic expert module, an expert routing module and a domain expert module; Predicting the basic generated data corresponding to the data to be processed by the basic expert module, and selecting the target expert submodule matching the generated task in the domain expert module by the expert routing module; The data to be processed and the basic generation data serving as reference data of the target expert submodule are input into the target expert submodule for processing to obtain the target generation data corresponding to the generation task.
[0007] According to a third aspect of an embodiment of this specification, there is provided a data processing system, including a client and a server; The client is used to generate a question-and-answer request based on the question-and-answer task, and send the question-and-answer request to the server; The server is used to input the question data corresponding to the question and answer task into the large language model in response to the question and answer request, and the large language model includes a basic expert module, an expert routing module and a domain expert module; predict the basic answer data corresponding to the question data through the basic expert module, and use the expert routing module to select a target expert sub-module matching the question and answer task in the domain expert module; input the question data and the basic answer data serving as reference data of the target expert sub-module into the target expert sub-module for processing, obtain the answer data corresponding to the question and answer task, and send the answer data to the client.
[0008] According to a fourth aspect of the embodiments of this specification, a large language model training system is provided, including a terminal side and a cloud side; The terminal side is used to send the initial large language model and the model training data set to the cloud side; The cloud side is used to determine the base large language model, the initial basic expert module, the initial domain expert module and the initial expert routing module included in the initial large language model; determine the basic expert training data, the expert routing training data and the domain expert training data in the model training data set; train the initial basic expert module based on the basic expert training data and the base large language model to obtain the basic expert module; train the initial domain expert module based on the domain expert training data, the base large language model and the basic expert module to obtain the domain expert module; train the initial expert routing module based on the expert routing training data, the base large language model, the basic expert module and the domain expert module to obtain the expert routing module; build a large language model based on the base large language model, the basic expert module, the expert routing module and the domain expert module, and send the large language model to the end side.
[0009] According to a fifth aspect of an embodiment of this specification, there is provided a data processing device, including: A question input module is configured to input question data corresponding to the question-answering task into a large language model, wherein the large language model includes a basic expert module, an expert routing module, and a domain expert module; An expert selection module is configured to predict basic answer data corresponding to the question data through the basic expert module, and select a target expert submodule matching the question-answering task in the domain expert module using the expert routing module; The data input module is configured to input the question data and the basic answer data serving as reference data of the target expert submodule into the target expert submodule for processing to obtain answer data corresponding to the question-answering task.
[0010] According to a sixth aspect of an embodiment of this specification, there is provided a generating device, including: An input module is configured to input the to-be-processed data corresponding to the generated task into a large language model, wherein the large language model comprises a basic expert module, an expert routing module and a domain expert module; A selection module is configured to predict the basic generated data corresponding to the to-be-processed data through the basic expert module, and to select a target expert submodule matching the generated task in the domain expert module using the expert routing module; The generation module is configured to input the data to be processed and the basic generation data serving as reference data of the target expert submodule into the target expert submodule for processing, so as to obtain the target generation data corresponding to the generation task.
[0011] According to a seventh aspect of an embodiment of this specification, there is provided a computing device, including: Memory and processor; The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions. When the computer executable instructions are executed by the processor, the steps of the above data processing method are implemented.
[0012] According to an eighth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned data processing method are implemented.
[0013] According to a ninth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program or instructions, which implement the steps of the above-mentioned data processing method when executed by a processor.
[0014] The data processing method provided by an embodiment of the present specification inputs the question data corresponding to the question-answering task into a large language model, and the large language model includes a basic expert module, an expert routing module, and a domain expert module. The basic answer data is obtained by predicting the question data through the basic expert module, so as to realize the prediction of the question data as a general task using a more general expert model. The target expert submodule matching the question-answering task is selected in the domain expert module by using the expert routing module included in the large language model, and the question data and the basic answer data as the reference data of the target expert submodule are input into the target expert submodule for processing to obtain the answer data corresponding to the question-answering task. The target expert submodule can use the basic answer data as a reference, and refer to the basic answer data when predicting the question data. The question data is predicted in combination with the specific domain data prediction ability of the expert module, so as to improve the accuracy of the answer data. The basic expert module, the expert routing module, and the domain expert module included in the large language model are used to collaboratively process the question data, so that the question data in different fields can be flexibly processed, and the large language model has high adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a processing diagram of a data processing method provided by an embodiment of this specification; Figure 2 is a flow chart of a data processing method provided by an embodiment of this specification; Figure 3 is a flow chart of a generation method provided by an embodiment of this specification; Figure 4 is a structural diagram of a data processing system provided by an embodiment of this specification; Figure 5It is a structural diagram of a large language model training system provided by an embodiment of this specification; Figure 6 It is a schematic diagram of the training phase of a large language model training in a large language model training system provided by one embodiment of this specification; Figure 7 is a structural schematic diagram of a data processing device provided by an embodiment of this specification; Figure 8 is a schematic diagram of the structure of a generating device provided by an embodiment of this specification; Fig. 9 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION
[0016] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.
[0017] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0018] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0019] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0020] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, which usually contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than 10 trillion model parameters. A large model can also be called a foundation model / foundation model. The large model is pre-trained with large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization ability, such as a large-scale language model (LLM), a multi-modal pre-training model, etc.
[0021] When the big model is used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. The big model can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, it can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of the big model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0022] First, the terms involved in one or more embodiments of this specification are explained.
[0023] Large Language Model: A language model built with a deep neural network containing tens of billions of parameters. It uses self-supervised learning methods and is trained on a large amount of unlabeled text to predict and generate text and other content.
[0024] Fine-tuning: Based on the pre-trained large language model, use a small amount of labeled data to further train and adjust the model to adapt it to a specific task or field.
[0025] LoRA fine-tuning: A parameter-efficient fine-tuning method. By training a low-dimensional matrix to construct an incremental matrix, the number of fine-tuning parameters can be significantly reduced, reducing the computational complexity and memory usage during model training.
[0026] MoDULA: A multi-task hybrid expert model based on general LoRA and domain-specific LoRA.
[0027] MoDULA-Res: A multi-task learning model for research in the e-commerce field. It can solve the problems of general experts and domain-specific experts, and improve the stability of computational cost and performance through residual connections. ICBU's self-developed base model TARS can provide support for business assistant services.
[0028] Mixed Expert Model (MoE): It consists of multiple "expert" networks and a "gating" network. Each expert network is a relatively independent model that is responsible for processing a certain part or a certain pattern of data, while the gating network determines which expert network should be activated or the weight of each expert network in the calculation when processing the input data.
[0029] With the continuous expansion of the scale of large language models (LLMs) and their increasing application in e-commerce scenarios, how to achieve efficient training with limited computing resources to meet the needs of multi-task learning has become an important challenge for the current application of LLMs. Traditional training methods often show the problem of insufficient stability in e-commerce multi-task learning, and have a large demand for training resources. Specifically, the most commonly used methods in the industry include full parameter fine-tuning, LoRA fine-tuning, etc. However, since the data ratio greatly affects the final performance of the model, the data ratio between different fields (mathematics, code, law, e-commerce, etc.) and different languages (Chinese, English, other multilingual) needs to be fully considered before fine-tuning. On the other hand, this type of method also faces the problem of high training cost: when incrementally learning professional knowledge for the base large model, it is often necessary to retrain on a new data set, which has a large demand for training hardware resources and time cost. Some existing new methods such as MoLoRA and SiRA integrate the efficient parameter fine-tuning (PEFT) method with the hybrid expert model (MoE) architecture, which can adapt the base large model to multiple professional fields while reducing the number of parameters and training costs. However, these methods still have certain defects.
[0030] To solve the above problems, the large language model used in the data processing method provided in one embodiment of this specification includes a basic expert module, an expert routing module and a domain expert module. It is a multi-task hybrid expert model based on general LoRA and domain-specific LoRA, which can realize parameter-efficient multi-task learning and new domain adaptation. By training general LoRA and domain-specific LoRA separately, the problem of data set fusion is avoided and the performance of the model in various fields is improved. LoRA technology is used to reduce computing resource consumption and achieve efficient model training. The residual connection structure is designed to maintain the original general capabilities of the base model while performing domain adaptation. A flexible training paradigm is provided that can easily add new tasks or domain experts without retraining the parameters of all experts. Using this large language model to process question and answer data can significantly improve the accuracy of question and answer results.
[0031] Figure 1 is a schematic diagram of a processing process of a data processing method provided by an embodiment of this specification, such as Figure 1 As shown, the question data corresponding to the question-answering task is input into the large language model, which includes a basic expert module, an expert routing module and a domain expert module. The basic expert module is used to predict the question data to obtain the basic answer data, so as to use a more general expert model to predict the question data as a general task. The expert routing module included in the large language model is used to select the target expert submodule in the domain expert module that matches the question-answering task, and the question data and the basic answer data as the reference data of the target expert submodule are input into the target expert submodule for processing to obtain the answer data corresponding to the question-answering task. The target expert submodule can use the basic answer data as a reference, and refer to the basic answer data when predicting the question data. The question data is predicted in combination with the specific domain data prediction ability of the expert module, so as to improve the accuracy of the answer data. The basic expert module, the expert routing module and the domain expert module included in the large language model are used to collaboratively process the question data, so as to flexibly process the question data in different fields. The basic expert module, the expert routing module and the domain expert module included in the large language model are trained in stages, so as to improve the adaptability of the large language model.
[0032] In this specification, a data processing method is provided. This specification also relates to a data processing system, a large language model training system, a data processing device, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.
[0033] See also Figure 2 , Figure 2 A flow chart of a data processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.
[0034] Step 202: Input the question data corresponding to the question-answering task into a large language model, wherein the large language model includes a basic expert module, an expert routing module, and a domain expert module.
[0035] Specifically, the question-answering task can be a question-answering task in the field of e-commerce, education, law, scientific research, mathematics, code generation, cultural and artistic content generation, content recommendation, psychological counseling, information search, and other fields. Users submit question-answering tasks and seek the execution results of question-answering tasks. In the case where the question-answering task is a question-answering task in the field of e-commerce, the product description can be input as question data into the large language model, and the large language model can determine the corresponding product based on the product description; in the case where the question-answering task is a question-answering task in the field of law, the legal problem can be input as question data into the large language model, and the large language model can give relevant legal interpretations; in the case where the question-answering task is a mathematical question-answering task, the mathematical problem can be input as question data into the large language model, and the large language model can give the solution ideas and answers to the mathematical problem. In the case where the question-answering task is a code generation task, the function description data can be input as question data into the large language model, and the large language model can give the function implementation code of the function description data.
[0036] The question data corresponding to the question-answering task is the question information of the questions to be answered contained in the question-answering task. It is a question that requires the use of a large language model to perform data analysis and output answers. In addition to the basic expert module, the expert routing module and the domain expert module, the large language model also includes a base model. Among them, the base model is generally trained on massive text and code data in a self-supervised training manner. It is a general model with basic capabilities and knowledge, also known as a pre-trained large model. The basic expert module contains a general expert model, which has the ability to capture general knowledge across domains. The domain expert module contains a domain expert model, which has the ability to handle specific domain tasks. The expert routing module contains routing, which is used to dynamically assign expert weights to expert sub-modules in different domains. The expert routing module selects the domain expert sub-module that matches the domain of the question data by calculating the expert weights.
[0037] Based on this, a question-answering task is received, and the question-answering task is parsed to obtain the question data corresponding to the question-answering task, and the question data is input into a large language model. The large language model including the basic expert module, the expert routing module and the domain expert module is used to collaboratively process the question data and output the answer data corresponding to the question data.
[0038] In practical applications, the general expert model in the basic expert module and the domain expert model in the domain expert module are obtained by combining LoRA fine-tuning training. The incremental matrix is constructed by training the low-dimensional matrix, which significantly reduces the number of fine-tuning parameters and reduces the computational complexity and memory usage during model training.
[0039] For example, when the question-answering task is a math problem solving task, the question data can be the title data of a math problem: "A carnival snack stand can earn 50 yuan a day by selling popcorn. It can earn three times as much by selling cotton candy. For a 5-day event, the snack stand must pay 30 yuan in rent and 75 yuan in raw material costs. After deducting the rent and raw material costs, how much money can the snack stand earn in 5 days?". By inputting the question data into the large language model, the basic expert module, expert routing module, and domain expert module contained in the large language model can be used to collaboratively solve the math problem.
[0040] Furthermore, in the training process of the large language model, a three-stage training paradigm can be adopted to train the base large language model, the initial basic expert module, and the initial domain expert module contained in the initial large language model separately. The specific implementation is as follows: Determine a base large language model, an initial basic expert module, an initial domain expert module, and an initial expert routing module included in the initial large language model; determine basic expert training data, expert routing training data, and domain expert training data in a model training data set associated with the initial large language model; The initial basic expert module is trained based on the basic expert training data and the base large language model to obtain the basic expert module; the initial domain expert module is trained based on the domain expert training data, the base large language model and the basic expert module to obtain the domain expert module; the initial expert routing module is trained based on the expert routing training data, the base large language model, the basic expert module and the domain expert module to obtain the expert routing module; the large language model is constructed based on the base large language model, the basic expert module, the expert routing module and the domain expert module.
[0041] Specifically, the base large language model is the base model, which is a pre-trained large language model, and is used as the basic model parameter training. The initial basic expert module includes a general expert model, the initial domain expert module includes a domain expert model, and the initial expert routing module includes the routing to be trained. The model training data in the model training data set is used to train the initial large language model. The basic expert training data is used to train the initial basic expert module, the expert routing training data is used to train the initial domain expert module, and the domain expert training data is used to train the initial expert routing module.
[0042] Based on this, determine the initial large language model that has not been trained, as well as the base large language model, initial basic expert module, initial domain expert module and initial expert routing module included in the initial large language model. Determine the model training data set provided for the initial large language model, and determine the basic expert training data for training the initial basic expert model, the expert routing training data for training the initial domain expert model, and the domain expert training data for training the initial expert routing module in the model training data set. In the first stage, the initial basic expert module is trained based on the basic expert training data and the base large language model to obtain the basic expert module. In the second stage, after the basic expert module is obtained through training, the initial domain expert module is trained based on the domain expert training data, the base large language model and the basic expert module to obtain the domain expert module. In the third stage, after the basic expert module and the domain expert module are obtained through training, the initial expert routing module is trained based on the expert routing training data, the base large language model, the basic expert module and the domain expert module to obtain the expert routing module. Construct a large language model based on the base large language model, the basic expert module, the expert routing module and the domain expert module. During the training of the initial large language model, residual connections were introduced to maintain the generalizability of the model.
[0043] In summary, the three-stage training paradigm is used to train a large language model, reduce the need for complex data processing, avoid tedious hyperparameter adjustments, and improve training stability. Reduce the cost of new tasks. When adding new tasks, only new domain expert models and routes need to be trained, without training the entire model, saving time and hardware resources, and reducing the cost of adding new tasks. In financial tasks, the demand for training resources can be greatly reduced, and scalability and adaptability can be improved.
[0044] Furthermore, when training the initial basic expert module, in order to avoid the influence of other modules in the initial large language model, the connection between the initial basic expert module and the initial domain expert module and the initial expert routing module can be cut off, and only the initial basic expert module can be trained independently, which is specifically implemented as follows: While fixing the model weights of the base large language model, interrupting the connection between the initial domain expert module and the initial expert routing module, and interrupting the connection between the initial domain expert module and the initial basic expert module, the initial basic expert module is trained based on the basic expert training data to obtain the basic expert module.
[0045] Based on this, when training the initial basic expert module, the model weight of the base large language model is fixed, the model weight of the base large language model is kept unchanged, and the connection between the initial domain expert module and the initial expert routing module, as well as the connection between the terminal initial domain expert module and the initial basic expert module are interrupted to freeze the initial expert routing module and the initial domain expert module, and it is prohibited to change the parameters of the initial expert routing module and the initial domain expert module when training the initial basic expert module, and the initial expert routing module and the initial domain expert module are prohibited from participating in the training of the initial basic expert module. The initial basic expert module is trained based on the basic expert training data to obtain the basic expert module. The basic expert module is enabled to quickly adapt to general tasks, so that it has a stronger ability to follow instructions and can better understand the received instructions. The domain experts and routing are disabled, which allows the basic experts to learn task-independent representations and provide basic understanding capabilities for subsequent tasks.
[0046] In summary, by cutting off the connection between the initial basic expert module, the initial domain expert module and the initial expert routing module, and only training the initial basic expert module independently, targeted training of the initial basic expert module can be achieved, and computational success can also be reduced.
[0047] Furthermore, when training the initial domain expert module, the connection between the initial expert routing module and the initial domain expert module is interrupted, and the initial domain expert module is independently trained using the domain expert training data, which is specifically implemented as follows: While fixing the model weight of the base large language model, fixing the module parameters of the basic expert module and the initial expert routing module, and interrupting the connection between the initial expert routing module and the initial domain expert module, the initial domain expert module is trained based on the domain expert training data to obtain the domain expert module.
[0048] Based on this, when training the initial domain expert module, the model weight of the base large language model is fixed, the model weight of the base large language model is kept unchanged, and the module parameters of the basic expert module and the initial expert routing module are fixed, the module parameters of the basic expert module and the initial expert routing module are kept unchanged, and the connection between the initial expert routing module and the initial domain expert module is interrupted to prevent the initial expert routing module from participating in the training of the initial domain expert module and avoid the parameter change of the initial expert routing module during the training of the initial domain expert module. The initial domain expert module is trained based on the domain expert training data to obtain the domain expert module. The domain expert module can focus on learning domain knowledge and improve the performance of the model in a specific field.
[0049] In actual applications, the domain expert module can include multiple domain experts, such as mathematics domain experts, code domain experts, medical domain experts, etc. Specific domain data is provided for each domain expert for training to improve the performance of each domain expert.
[0050] In summary, under the environment of fixing the module parameters of the basic expert module and the initial expert routing module and interrupting the connection between the initial expert routing module and the initial domain expert module, the initial domain expert module is trained in a targeted manner to reduce the complexity of data processing and avoid tedious hyperparameter adjustment.
[0051] Furthermore, when training the initial expert routing module, the module parameters of the basic expert module and the domain expert module are fixed to improve the training efficiency of the initial expert routing module, which is specifically implemented as follows: Under the condition that the model weight of the base large language model is fixed and the module parameters of the basic expert module and the domain expert module are fixed, the initial expert routing module is trained based on the expert routing training data to obtain the expert routing module.
[0052] Based on this, when training the initial expert routing module, the model weights of the base large language model are fixed, and the model weights of the base large language model are kept unchanged. At the same time, the module parameters of the basic expert module and the domain expert module are fixed, and the module parameters of the basic expert module and the domain expert module are kept unchanged. The initial expert routing module is trained based on the expert routing training data to obtain the expert routing module. The expert routing training data contains a small amount of data for training the initial basic expert module, and a small amount of data for training the initial domain expert module. When training the initial expert routing module, both the domain expert module and the basic expert module need to participate in the training, but the training of the initial expert routing module does not change the module parameters of the domain expert module and the basic expert module.
[0053] In summary, when the module parameters of the basic expert module and the domain expert module are fixed, the initial expert routing module is trained, and only the initial expert routing module is trained, which reduces the complexity of data processing and improves the training efficiency of the initial expert routing module.
[0054] Step 204: predicting basic answer data corresponding to the question data through the basic expert module, and selecting a target expert submodule matching the question-answering task in the domain expert module using the expert routing module.
[0055] Specifically, after the question data corresponding to the question-answering task is input into the large language model, which includes the basic expert module, the expert routing module and the domain expert module, the basic answer data corresponding to the question data can be predicted by the basic expert module, and the target expert submodule matching the question-answering task can be selected in the domain expert module using the expert routing module, wherein the basic answer data refers to the processing result output by the general expert model in the basic expert module after processing the question data, which is the intermediate answer data obtained by analyzing the question data based on general knowledge, and the basic answer data can be the expression form of the latent vector. The domain processing capability of the target expert submodule matches the problem domain to which the question data belongs.
[0056] Based on this, after the question data corresponding to the question-answering task is input into the large language model, which includes the basic expert module, the expert routing module and the domain expert module, the basic expert module included in the large language model is used to predict the question data in the general knowledge dimension to obtain the basic answer data corresponding to the question data. The expert routing module included in the large language model is used to calculate the matching degree between the question data in the domain dimension and each domain expert submodule in the domain expert module, and select the target expert submodule with a higher matching degree with the question data domain.
[0057] In actual applications, when the routing module calculates the matching degree between the problem data in the domain dimension and each domain expert sub-module in the domain expert module, an expert weight can be assigned to each domain expert sub-module. The expert weight can be used to select the target expert sub-module and assist in the subsequent answer data generation.
[0058] Furthermore, considering that the domain expert module contains at least two domain expert sub-modules, the target expert sub-module can be determined according to the domain to which the problem data belongs, ensuring that the target expert sub-module matches the problem data in the domain dimension, which is specifically implemented as follows: Generate a question vector corresponding to the question data; use the expert routing module to calculate the domain expert weights corresponding to at least two domain expert sub-modules in the domain expert module based on the question vector, the routing function and the weight matrix contained in the expert routing module; select a target expert sub-module that matches the question and answer task in the at least two domain expert sub-modules according to the domain expert weights corresponding to the at least two domain expert sub-modules.
[0059] Specifically, the question vector is a vector expression of the question data. The routing function can be a softmax function, and the weight matrix is a parameter contained in the expert routing module. The domain expert weight can represent the domain matching degree between the domain expert submodule and the question vector. The domain expert weight can also represent the influence of the prediction result on the answer data when the domain expert submodule predicts the question data in the future.
[0060] Based on this, the question data is converted into a vector expression form to obtain the question vector corresponding to the question data. The expert routing module is used to calculate the domain expert weights corresponding to at least two domain expert submodules in the domain expert module based on the question vector, the routing function contained in the expert routing module and the weight matrix. According to the domain expert weights corresponding to the at least two domain expert submodules, a target expert submodule with a higher domain expert weight is selected from the at least two domain expert submodules as the target expert submodule matching the question-answering task.
[0061] Using the above example, when there are three domain expert submodules and the problem data involves the field of mathematics, the domain expert weight of each domain expert submodule is calculated using the expert routing module. When the domain expert weight of domain expert submodule 1 is 0.8, the domain expert weight of domain expert submodule 2 is 0.5, and the domain expert weight of domain expert submodule 3 is 0.2, domain expert submodule 1 and domain expert submodule 2 can be selected as target expert submodules. The calculation of domain expert weights can be achieved using the following formula (3):
[0062] in, Represents the input question vector; Indicates the parameters of the routes included in the expert routing module. Represents the weight of the i-th domain expert sub-module. represents the weight matrix of the expert routing module, R represents Router, and res represents the large language model. The expert routing module Get the weight of each domain expert submodule. The expert routing module contains the weight matrix and the softmax function, the problem vector First pass , and then passes through the softmax function, the output of softmax is a vector.
[0063] In summary, according to the domain expert weights corresponding to the at least two domain expert submodules, a target expert submodule with a higher domain expert weight is selected from the at least two domain expert submodules as the target expert submodule matching the question-answering task. The target expert submodule is used to predict question data and improve the accuracy of answer data.
[0064] Step 206: Input the question data and the basic answer data serving as reference data of the target expert submodule into the target expert submodule for processing to obtain answer data corresponding to the question-answering task.
[0065] Specifically, after predicting the basic answer data corresponding to the question data through the basic expert module, and selecting the target expert submodule matching the question-answering task in the domain expert module using the expert routing module, the question data and the basic answer data as the reference data of the target expert submodule can be input into the target expert submodule for processing to obtain the answer data corresponding to the question-answering task, wherein the answer data is the answer to the question output by the processing of the question data by the large language model. The basic answer data plays an auxiliary role in the prediction when the target expert submodule predicts the question data.
[0066] Based on this, after predicting the basic answer data corresponding to the question data through the basic expert module and selecting the target expert submodule matching the question-answering task in the domain expert module using the expert routing module, the basic answer data is input into the target expert submodule as the reference data of the target expert submodule, and the question data is also input into the target expert submodule. The target expert submodule uses the basic answer data as a reference, combines the question data to predict the answer, and obtains the answer data corresponding to the question-answering task. The basic answer data is an influencing factor for the target expert submodule to predict the data based on the question data. In the process of predicting the answer data, the influence of the basic answer data, which is the data obtained based on the general knowledge prediction, is incorporated to improve the prediction accuracy of the answer data.
[0067] Furthermore, considering that the target expert submodule has the ability to predict domain-related data, when the basic expert module predicts and obtains the basic answer data, the basic answer data can be input into the target expert submodule as an influencing factor through residual linking to assist the target expert submodule in predicting the answer to the question data. The specific implementation is as follows: Determine the expert matrix contained in the target expert submodule and the target expert weight corresponding to the target expert submodule; make predictions based on the expert matrix, the target expert weight, the question data and the basic answer data serving as reference data for the target expert submodule to obtain answer data corresponding to the question-answering task.
[0068] Specifically, the expert matrix refers to the A matrix and B matrix in the domain-specific expert model in the target expert submodule, which is used to achieve parameter fine-tuning during the model training phase. The target expert weight represents the contribution of the target expert submodule in the prediction of the problem data.
[0069] Based on this, the expert matrix contained in the target expert submodule and the target expert weight corresponding to the target expert submodule are determined. The target expert submodule predicts the answer to the question data based on the expert matrix, target expert weight, and question data. At the same time, the basic answer data is used as the reference data in the answer prediction process of the target expert submodule through the residual connection method to predict the answer data corresponding to the question and answer task.
[0070] Using the above example, after selecting the target expert submodule from at least two domain expert submodules, the question data: "The carnival snack stand can make 50 yuan a day by selling popcorn. Selling cotton candy can make three times as much money. For the 5-day event, the snack stand must pay 30 yuan in rent and 75 yuan in raw material costs. After deducting the rent and raw material costs, how much money can the snack stand make in 5 days?" and the basic answer data (the latent vector of the question data) are input into the target expert submodule. The prediction of the target expert submodule can be expressed by the following formulas (2) and (3).
[0071]
[0072] in, represents the basic answer data, Represents the operations of domain experts in the target expert submodule. The weight matrix of the linear layer of the base language model that has been frozen; Represents the initial hidden state vector of the input target expert submodule; represents the domain expert weight assigned by the expert routing module in the target expert submodule; The B matrix represents the domain experts in the target expert submodule; Represents the A matrix of the domain expert in the target expert submodule (a domain expert contains two matrices A and B, and the vector passes through A first and then B during calculation).
[0073] After the collaborative processing of the basic expert module, expert routing module and domain expert module in the large language model, the question data "The carnival snack stand can make 50 yuan a day by selling popcorn. Selling marshmallows can make three times the money. For a 5-day event, the snack stand must pay 30 yuan in rent and 75 yuan in raw material costs. After deducting the rent and raw material costs, how much money can the snack stand make in 5 days?" The corresponding answer data is 895 yuan. The large language model can also synchronously output the problem-solving ideas and problem-solving process: "To determine how much the carnival snack stand earns in 5 days after paying the rent and ingredient costs, we need to follow the following steps: 1. Calculate the revenue from selling popcorn and marshmallows every day. 2. Calculate the total revenue for 5 days. 3. Subtract the rent and ingredient costs from the total revenue. First, let's calculate the revenue from selling popcorn and marshmallows every day: the stall earns 50 yuan a day from selling popcorn. Selling marshmallows, it earns 3 times a day (3×50=150 yuan). Next, Calculate the total daily revenue: 50 (popcorn) + 150 (marshmallow) = 200 (yuan / day); now calculate the total revenue for 5 days: 200*5=1000 yuan; next, we need to subtract the rental and ingredient costs from the total revenue: 30 yuan for 5 days of rent, and 75 yuan for raw materials; total expenses: 30 (rent) + 75 (ingredients) = 105 yuan; finally, subtract the total expenses from the total revenue: 1000 yuan - 105 yuan = 895 yuan; therefore, after paying the rent and ingredient costs, the booth earned 895 yuan in 5 days."
[0074] In addition, when the question-answering task is a question-answering task in the medical field, the question data can be the topic data of a math problem: "I have high blood pressure. My son-in-law came to me these two days and brought me some Codonopsis pilosula to drink. Hello, can you eat Codonopsis pilosula for high blood pressure?". Input the question data into the large language model, and you can use the basic expert module, expert routing module, and domain expert module contained in the large language model to collaboratively complete the math problem. After the collaborative processing of the basic expert module, expert routing module, and domain expert module in the large language model, the answer data corresponding to the question data is obtained: "Patients with high blood pressure need to be cautious when using Codonopsis pilosula. Codonopsis pilosula is a commonly used Chinese medicinal material with the effects of replenishing qi, nourishing blood, and calming the nerves. However, because it has a certain blood pressure-raising effect, it is not recommended for patients with high blood pressure to use it casually. Especially when blood pressure is not well controlled, you should avoid using medicinal materials that may cause blood pressure fluctuations. If you want to regulate your body with Chinese medicine, it is recommended to consult a professional Chinese medicine practitioner or doctor first and develop a suitable treatment plan based on your specific situation. At the same time, maintaining good living habits, such as moderate exercise, a balanced diet, and reducing salt intake, is also very important for controlling high blood pressure."
[0075] When the question-answering task is a code generation task, the input function description data is the question data. When the question data is input into the large language model, the basic expert module, expert routing module, and domain expert module contained in the large language model can be used to collaboratively complete the code generation of the specified function. The question data is as follows: fromtypingimportList defhas_close_elements(numbers:List[float],threshold:float)->bool: """Checks if, in a given list of numbers, any two numbers are closer than a given threshold.
[0076] >>>has_close_elements([1.0,2.0,3.0],0.5) False >>>has_close_elements([1.0,2.8,3.0,4.0,5.0,2.0],0.3) True """ Input the question data into the large language model, and then use the basic expert module, expert routing module and domain expert module contained in the large language model to collaboratively complete the code generation of the specified function. After the collaborative processing of the basic expert module, expert routing module and domain expert module in the large language model, the answer data corresponding to the question data is obtained: Here is a Python function that implements the logic of checking whether any two numbers in a given list are closer than a specified threshold: ```Python from typing import List def has_close_elements(numbers: List[float], threshold: float) ->bool: """ Check if in given list of numbers, are any two numbers closer to each other than given threshold. >>>has_close_elements([1.0, 2.0, 3.0], 0.5) False >>>has_close_elements([1.0, 2.8, 3.0, 4.0, 5.0, 2.0], 0.3) True """ for i in range(len(numbers)): for j in range(i + 1, len(numbers)): if abs(numbers[i] - numbers[j]) <threshold: return True return False ``` This function does the following: 1. We use two nested loops to compare each pair of numbers in the list.
[0077] 2. The outer loop 'for i in range(len(numbers))' iterates over each number in the list.
[0078] 3. For j in the range (i+1, len(numbers)), the inner loop starts with the next number after 'i' and goes to the end of the list. This ensures that we do not compare a number to itself and do not make redundant comparisons.
[0079] The data processing method provided by an embodiment of the present specification inputs the question data corresponding to the question-answering task into a large language model, and the large language model includes a basic expert module, an expert routing module, and a domain expert module. The basic answer data corresponding to the question data is predicted by the basic expert module, so that the question data can be predicted as a general task using a more general expert model. The expert routing module included in the large language model is used to select a target expert submodule that matches the question-answering task in the domain expert module, and the question data and the basic answer data as the reference data of the target expert submodule are input into the target expert submodule for processing to obtain the answer data corresponding to the question-answering task. The target expert submodule can use the basic answer data as a reference, and refer to the basic answer data when predicting the question data. The question data is predicted in combination with the specific domain data prediction ability of the expert module, thereby improving the accuracy of the answer data. The basic expert module, the expert routing module, and the domain expert module included in the large language model are used to collaboratively process the question data, so that the question data in different fields can be flexibly processed, and the large language model has high adaptability.
[0080] Corresponding to the above method embodiment, this specification also provides a generation method embodiment, Figure 3 A flow chart of a generation method provided by an embodiment of the present specification is shown, which specifically includes the following steps.
[0081] Step 302: inputting the to-be-processed data corresponding to the generated task into a large language model, wherein the large language model comprises a basic expert module, an expert routing module and a domain expert module; Step 304: predicting the basic generated data corresponding to the to-be-processed data through the basic expert module, and selecting a target expert submodule matching the generated task in the domain expert module using the expert routing module; Step 306: input the data to be processed and the basic generation data serving as reference data of the target expert submodule into the target expert submodule for processing to obtain the target generation data corresponding to the generation task.
[0082] Specifically, the generation task can be a generation task in multiple fields such as e-commerce, mathematics, code, culture and art. The generation task can be an image generation task, a text generation task, a title generation task, a code generation task, a summary generation task, and other tasks for generating text or image content. The data to be processed is the content generation prompt of the generation task. When the generation task is a summary generation task, the data to be processed is a long text data such as articles and papers. The basic generation data refers to the data obtained after the general expert model in the basic expert module makes a preliminary prediction on the data to be processed. The basic generation data can be the expression form of the latent vector.
[0083] In practical applications, in the field of e-commerce, the generation task can be a user portrait generation task, and products can be accurately recommended to users based on the generated user portrait. When generating a user portrait, the user's user behavior data is input into the large language model as the data to be processed. The basic expert module contained in the large language model is used to predict the basic user portrait data corresponding to the user behavior data. The expert routing module contained in the large language model is used to select the target expert submodule that matches the user portrait generation task in the domain expert module contained in the large language model. The basic user portrait data is input into the target expert submodule as the reference data of the target expert submodule, and the user behavior data is also input into the target expert submodule. The target expert submodule can obtain the user's user portrait data by processing the user behavior data and the basic user portrait data. Subsequently, the user can be recommended products that the user may be interested in based on the user portrait data to achieve accurate product recommendations.
[0084] The large language model can also generate descriptive text for images. When the generation task is a text generation task, the target image data input into the large language model can be regarded as data to be processed. The target image data is input into the large language model as data to be processed. The basic image description data corresponding to the target image data is generated using the basic expert module contained in the large language model. The expert routing module contained in the large language model is used to select the target expert submodule that matches the text generation task in the domain expert module contained in the large language model. The basic image description data is input into the target expert submodule as the reference data of the target expert submodule, and the target image data is also input into the target expert submodule. The target expert submodule can obtain the image description data by processing the target image data and the basic image description data. The image description data represents the image analysis result of the target image data.
[0085] In summary, the data to be processed corresponding to the generation task is input into the large language model, which includes a basic expert module, an expert routing module and a domain expert module. The basic generated data obtained by predicting the data to be processed through the basic expert module can be used to predict the data to be processed as a general task using a more general expert model. The expert routing module contained in the large language model is used to select the target expert submodule matching the generation task in the domain expert module, and the data to be processed and the basic generated data as the reference data of the target expert submodule are input into the target expert submodule for processing to obtain the target generated data corresponding to the generation task. The target expert submodule can use the basic generated data as a reference, and refer to the basic generated data when predicting the data to be processed. The data to be processed is predicted in combination with the specific domain data prediction ability of the expert module, thereby improving the accuracy of the target generated data. The basic expert module, the expert routing module and the domain expert module contained in the large language model are used to collaboratively process the data to be processed, which can flexibly process the data to be processed in different fields, and the large language model has high adaptability.
[0086] Corresponding to the above method embodiment, this specification also provides a data processing system embodiment. Figure 4 FIG. 1 shows a schematic diagram of a data processing system provided by an embodiment of the present specification. Figure 4As shown, the data processing system 400 includes a client 410 and a server 420; the client 410 is used to generate a question and answer request based on a question and answer task, and send the question and answer request to the server 420; the server 420 is used to input the question data corresponding to the question and answer task into a large language model in response to the question and answer request, and the large language model includes a basic expert module, an expert routing module and a domain expert module; predict the basic answer data corresponding to the question data through the basic expert module, and use the expert routing module to select a target expert sub-module matching the question and answer task in the domain expert module; input the question data and the basic answer data as reference data of the target expert sub-module into the target expert sub-module for processing, obtain the answer data corresponding to the question and answer task, and send the answer data to the client 410.
[0087] In practical applications, a data processing system provided by an embodiment of this specification includes a client and a server. When the client has a question-and-answer demand, it is constructed as a question-and-answer task based on the question-and-answer demand, and a question-and-answer request is generated based on the question-and-answer task. The question-and-answer request is sent to the server. The server inputs the question data corresponding to the question-and-answer task into the large language model, and the large language model includes a basic expert module, an expert routing module, and a domain expert module. The basic answer data corresponding to the question data is predicted by the basic expert module, so that the question data is predicted as a general task using a more general expert model. The expert routing module contained in the large language model is used to select a target expert submodule matching the question-and-answer task in the domain expert module, and the question data and the basic answer data as the reference data of the target expert submodule are input into the target expert submodule for processing to obtain the answer data corresponding to the question-and-answer task. The target expert submodule can use the basic answer data as a reference, and refer to the basic answer data when predicting the question data. The question data is predicted in combination with the specific domain data prediction ability of the expert module to improve the accuracy of the answer data. The basic expert module, expert routing module and domain expert module contained in the large language model are used to collaboratively process problem data. This allows for flexible processing of problem data in different fields, and the large language model has high adaptability.
[0088] The above is a schematic scheme of a data processing system of this embodiment. It should be noted that the technical scheme of the data processing system and the technical scheme of the above data processing method belong to the same concept, and the details not described in detail in the technical scheme of the data processing system can be referred to the description of the technical scheme of the above data processing method.
[0089] Corresponding to the above method embodiment, this specification also provides a large language model training system embodiment, Figure 5FIG. 2 shows a schematic diagram of a large language model training system provided by an embodiment of the present specification. Figure 5 As shown, the large language model training system 500 includes a terminal side 510 and a cloud side 520; the terminal side 510 is used to send the initial large language model and the model training data set to the cloud side 520; the cloud side 520 is used to determine the base large language model, the initial basic expert module, the initial domain expert module and the initial expert routing module contained in the initial large language model; determine the basic expert training data, the expert routing training data and the domain expert training data in the model training data set; train the initial basic expert module based on the basic expert training data and the base large language model to obtain the basic expert module; train the initial domain expert module based on the domain expert training data, the base large language model and the basic expert module to obtain the domain expert module; train the initial expert routing module based on the expert routing training data, the base large language model, the basic expert module and the domain expert module to obtain the expert routing module; build a large language model based on the base large language model, the basic expert module, the expert routing module and the domain expert module, and send the large language model to the terminal side 510.
[0090] In practical applications, large language models can be used to answer questions in the e-commerce field. Figure 6 As shown in the figure, the training of the large language model adopts a three-stage training process. The first stage is to train basic experts; the second stage is to train domain experts; and the third stage is to train expert routing.
[0091] In the first stage, all weights of the base language model are fixed. The expert routing module, basic expert module and domain expert module are initialized. During training, the connection between the expert routing module and the domain expert module is cut off (as shown by the dotted line in the figure) to prevent the expert routing module and the domain expert module from participating in the operation or parameter update. Only the basic expert module is trained to quickly adapt to general tasks, so that it has stronger instruction following ability and can better understand the received instructions. At the same time, the domain expert module and the expert routing module are disabled, so that the basic expert can learn task-independent representations and provide basic understanding capabilities for subsequent tasks.
[0092] In the second stage, all weights of the base large language model are fixed. The parameters of the basic expert module and the expert routing module are frozen, and the connection of the expert routing module is cut off to prevent the expert routing module from participating in the operation and prevent the parameters from being updated. Each domain expert in the domain expert module is trained separately to focus on its corresponding task, and the parameters of the basic expert module remain frozen. At this time, domain experts can focus on learning specific knowledge in their respective fields and improve the performance of the model in specific fields. For example, the current MoDULA-Res contains a general expert, a mathematics expert, a code expert and a medical expert. When training a mathematics domain expert, the corresponding domain data prepared only contains data in the mathematics field. During training, all weights of the base large language model and the weights of the basic experts are kept fixed. The gradient descent method is used to update only the parameters of the mathematics expert. The remaining domain-specific experts (code, medicine) neither participate in the operation nor update the parameters.
[0093] In the third stage, all weights of the base language model are fixed. The parameters of the basic expert module and the domain expert module are frozen. At this time, all modules must participate in the operation, but only the weight matrix in the expert routing module will be updated. Only the expert routing is trained (a certain amount of data is extracted from the data sets of the training basic expert module and the domain expert module, and the expert routing module is fine-tuned after mixing) to learn the optimal combination strategy for different tasks. By training the expert routing module, the model can dynamically assign weights to different experts according to the needs of different tasks, thereby improving model performance.
[0094] MoDULA-Res integrates basic expert modules and domain expert modules and introduces residual connections to balance the model's processing capabilities for general tasks and domain-specific tasks, improving performance and stability. In MoDULA-Res, the hidden vector is first calculated using the basic expert, then refined by the domain expert, and the output of the basic expert is directly integrated into the final result through the residual connection, ensuring that key information is retained and enhancing the robustness of the model.
[0095] In practical applications, after completing the training of the large language model and obtaining the large language model containing the basic expert module, expert routing module and domain expert module, the large language model can be used to answer questions in the e-commerce field. When the question data is "generate a user portrait of user A", the question data "generate a user portrait of user A" can be input into the large language model, and the basic expert module, expert routing module and domain expert module contained in the large language model can be used to collaboratively complete the code generation of the specified function. After the collaborative processing of the basic expert module, expert routing module and domain expert module in the large language model, the answer data corresponding to the question data is obtained, "User A is a primary school student who likes to buy pens and notebooks."
[0096] In summary, the three-stage training paradigm is adopted. Through three-stage training, the need for complex data processing is reduced, cumbersome hyperparameter adjustments are avoided, and training stability is improved. Basic experts are trained using general data sets, and then training is performed for specific fields to improve data utilization and training stability, significantly reducing computational costs. Residual connections are introduced to enhance multi-task performance and stability, so that specific experts receive the output of basic experts, integrate information, and perform better under multi-task and different-scale models. By training different types of experts separately, the impact of data inconsistency and interference on performance is reduced, and good performance can be maintained when processing tasks in different fields. From the perspective of model design, new domain experts can be easily added. Only new experts and routes need to be trained without retraining all parameters, which reduces the cost of new tasks, enhances the adaptability of the model to new tasks, and realizes pluggable design.
[0097] The above is a schematic scheme of a large language model training system of this embodiment. It should be noted that the technical scheme of the large language model training system and the technical scheme of the above data processing method belong to the same concept, and the details not described in detail in the technical scheme of the large language model training system can be found in the description of the technical scheme of the above data processing method.
[0098] Corresponding to the above method embodiment, this specification also provides a data processing device embodiment, Figure 7 FIG. 1 is a schematic diagram showing the structure of a data processing device provided by an embodiment of the present specification. Figure 7 As shown, the device comprises: A question input module 702 is configured to input question data corresponding to the question-answering task into a large language model, wherein the large language model includes a basic expert module, an expert routing module, and a domain expert module; The expert selection module 704 is configured to predict the basic answer data corresponding to the question data through the basic expert module, and select a target expert submodule matching the question-answering task in the domain expert module using the expert routing module; The data input module 706 is configured to input the question data and the basic answer data serving as reference data of the target expert submodule into the target expert submodule for processing to obtain answer data corresponding to the question-answering task.
[0099] In an optional embodiment, the expert selection module 704 is further configured to: Generate a question vector corresponding to the question data; Calculating the domain expert weights corresponding to at least two domain expert submodules in the domain expert module respectively using the expert routing module based on the problem vector, the routing function and the weight matrix included in the expert routing module; A target expert submodule matching the question-answering task is selected from the at least two domain expert submodules according to the domain expert weights respectively corresponding to the at least two domain expert submodules.
[0100] In an optional embodiment, the data input module 706 is further configured to: Determine the expert matrix included in the target expert submodule and the target expert weight corresponding to the target expert submodule; Prediction is performed based on the expert matrix, the target expert weight, the question data and the basic answer data serving as reference data for the target expert submodule to obtain answer data corresponding to the question-answering task.
[0101] In an optional embodiment, the question input module 702 is further configured to: Determine the base large language model, the initial basic expert module, the initial domain expert module and the initial expert routing module included in the initial large language model; Determining basic expert training data, expert routing training data, and domain expert training data in a model training data set associated with the initial large language model; Training the initial basic expert module based on the basic expert training data and the base large language model to obtain the basic expert module; Training the initial domain expert module based on the domain expert training data, the base large language model and the basic expert module to obtain the domain expert module; Training the initial expert routing module based on the expert routing training data, the base large language model, the basic expert module and the domain expert module to obtain the expert routing module; The large language model is constructed based on the base large language model, the basic expert module, the expert routing module and the domain expert module.
[0102] In an optional embodiment, the question input module 702 is further configured to: While fixing the model weights of the base large language model, interrupting the connection between the initial domain expert module and the initial expert routing module, and interrupting the connection between the initial domain expert module and the initial basic expert module, the initial basic expert module is trained based on the basic expert training data to obtain the basic expert module.
[0103] In an optional embodiment, the question input module 702 is further configured to: While fixing the model weight of the base large language model, fixing the module parameters of the basic expert module and the initial expert routing module, and interrupting the connection between the initial expert routing module and the initial domain expert module, the initial domain expert module is trained based on the domain expert training data to obtain the domain expert module.
[0104] In an optional embodiment, the question input module 702 is further configured to: Under the condition that the model weight of the base large language model is fixed and the module parameters of the basic expert module and the domain expert module are fixed, the initial expert routing module is trained based on the expert routing training data to obtain the expert routing module.
[0105] The data processing device provided by one embodiment of the present specification inputs the question data corresponding to the question-answering task into a large language model, and the large language model includes a basic expert module, an expert routing module, and a domain expert module. The basic answer data corresponding to the question data is predicted by the basic expert module, so that the question data can be predicted as a general task using a more general expert model. The expert routing module included in the large language model is used to select a target expert submodule matching the question-answering task in the domain expert module, and the question data and the basic answer data as the reference data of the target expert submodule are input into the target expert submodule for processing to obtain the answer data corresponding to the question-answering task. The target expert submodule can use the basic answer data as a reference, and refer to the basic answer data when predicting the question data. The question data is predicted in combination with the specific domain data prediction ability of the expert module, thereby improving the accuracy of the answer data. The basic expert module, the expert routing module, and the domain expert module included in the large language model are used to collaboratively process the question data, and the question data in different fields can be flexibly processed. The large language model has high adaptability.
[0106] The above is a schematic scheme of a data processing device of this embodiment. It should be noted that the technical scheme of the data processing device and the technical scheme of the above data processing method belong to the same concept, and the details of the technical scheme of the data processing device that are not described in detail can be referred to the description of the technical scheme of the above data processing method.
[0107] Corresponding to the above method embodiment, this specification also provides a generation device embodiment, Figure 8 FIG. 1 shows a schematic diagram of a generating device provided by an embodiment of the present specification. Figure 8 As shown, the device comprises: An input module 802 is configured to input the to-be-processed data corresponding to the generated task into a large language model, wherein the large language model comprises a basic expert module, an expert routing module and a domain expert module; The selection module 804 is configured to predict the basic generated data corresponding to the to-be-processed data through the basic expert module, and select a target expert submodule matching the generated task in the domain expert module using the expert routing module; The generation module 806 is configured to input the data to be processed and the basic generation data serving as reference data of the target expert submodule into the target expert submodule for processing, so as to obtain the target generation data corresponding to the generation task.
[0108] The generation device provided by one embodiment of the present specification inputs the data to be processed corresponding to the generation task into the large language model, and the large language model includes a basic expert module, an expert routing module and a domain expert module. The basic generated data obtained by predicting the data to be processed by the basic expert module can be used to predict the data to be processed as a general task using a more general expert model. The expert routing module included in the large language model is used to select the target expert submodule matching the generation task in the domain expert module, and the data to be processed and the basic generated data as the reference data of the target expert submodule are input into the target expert submodule for processing to obtain the target generated data corresponding to the generation task. The target expert submodule can use the basic generated data as a reference, and refer to the basic generated data when predicting the data to be processed. The data to be processed is predicted in combination with the specific domain data prediction ability of the expert module, thereby improving the accuracy of the target generated data. The basic expert module, the expert routing module and the domain expert module included in the large language model are used to collaboratively process the data to be processed, and the data to be processed in different fields can be flexibly processed. The large language model has high adaptability.
[0109] The above is a schematic scheme of a generating device of this embodiment. It should be noted that the technical scheme of the generating device and the technical scheme of the generating method described above are of the same concept, and the details not described in detail in the technical scheme of the generating device can be referred to the description of the technical scheme of the generating method described above.
[0110] Fig. 9 The structure block diagram of a computing device 900 provided according to an embodiment of the present specification is shown. The components of the computing device 900 include but are not limited to a memory 910 and a processor 920. The processor 920 is connected to the memory 910 via a bus 930, and a database 950 is used to store data.
[0111] The computing device 900 also includes an access device 940 that enables the computing device 900 to communicate via one or more networks 960. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of network interface (e.g., a network interface card (NIC)) that is wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, and a near field communication (NFC).
[0112] In one embodiment of the present specification, the above components of the computing device 900 and Fig. 9 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Fig. 9 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0113] The computing device 900 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 900 may also be a mobile or stationary server.
[0114] The processor 920 is used to execute the following computer executable instructions, which implement the steps of the above method when executed by the processor.
[0115] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above method belong to the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical scheme of the above method.
[0116] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which can implement the steps of the above method when executed by a processor.
[0117] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above method belong to the same concept, and the details not described in detail in the technical scheme of the storage medium can be referred to the description of the technical scheme of the above method.
[0118] An embodiment of the present specification also provides a computer program product, including a computer program or instructions, which implement the steps of the above method when executed by a processor.
[0119] The above is a schematic solution of a computer program product of this embodiment. It should be noted that the technical solution of the computer program product and the technical solution of the above method belong to the same concept, and the details not described in detail in the technical solution of the computer program product can be referred to the description of the technical solution of the above method.
[0120] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0121] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0122] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0123] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0124] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A data processing method, comprising: Inputting question data corresponding to the question-answering task into a large language model, wherein the large language model includes a basic expert module, an expert routing module, and a domain expert module; Predicting basic answer data corresponding to the question data by using the basic expert module, and selecting a target expert submodule matching the question-answering task in the domain expert module by using the expert routing module; The question data and the basic answer data serving as reference data for the target expert submodule are input into the target expert submodule for processing to obtain answer data corresponding to the question-answering task.
2. The data processing method according to claim 1, wherein the step of selecting a target expert submodule matching the question-answering task in the domain expert module by using the expert routing module comprises: Generate a question vector corresponding to the question data; Calculating the domain expert weights corresponding to at least two domain expert submodules in the domain expert module respectively using the expert routing module based on the problem vector, the routing function and the weight matrix included in the expert routing module; A target expert submodule matching the question-answering task is selected from the at least two domain expert submodules according to the domain expert weights respectively corresponding to the at least two domain expert submodules.
3. The data processing method according to claim 1, wherein the step of inputting the question data and the basic answer data as reference data of the target expert submodule into the target expert submodule for processing to obtain the answer data corresponding to the question-answering task comprises: Determine the expert matrix included in the target expert submodule and the target expert weight corresponding to the target expert submodule; Prediction is performed based on the expert matrix, the target expert weight, the question data and the basic answer data serving as reference data for the target expert submodule to obtain answer data corresponding to the question-answering task.
4. The data processing method according to claim 1, wherein the training of the large language model comprises: Determine the base large language model, the initial basic expert module, the initial domain expert module and the initial expert routing module included in the initial large language model; Determining basic expert training data, expert routing training data, and domain expert training data in a model training data set associated with the initial large language model; Training the initial basic expert module based on the basic expert training data and the base large language model to obtain the basic expert module; Training the initial domain expert module based on the domain expert training data, the base large language model and the basic expert module to obtain the domain expert module; Training the initial expert routing module based on the expert routing training data, the base large language model, the basic expert module and the domain expert module to obtain the expert routing module; The large language model is constructed based on the base large language model, the basic expert module, the expert routing module and the domain expert module.
5. The data processing method according to claim 4, wherein the initial expert training module is trained based on the basic expert training data and the base large language model to obtain the basic expert module, comprising: While fixing the model weights of the base large language model, interrupting the connection between the initial domain expert module and the initial expert routing module, and interrupting the connection between the initial domain expert module and the initial basic expert module, the initial basic expert module is trained based on the basic expert training data to obtain the basic expert module.
6. The data processing method according to claim 4, wherein the initial domain expert module is trained based on the domain expert training data, the base large language model and the basic expert module to obtain the domain expert module, comprising: While fixing the model weight of the base large language model, fixing the module parameters of the basic expert module and the initial expert routing module, and interrupting the connection between the initial expert routing module and the initial domain expert module, the initial domain expert module is trained based on the domain expert training data to obtain the domain expert module.
7. The data processing method according to claim 4, wherein the initial expert routing module is trained based on the expert routing training data, the base large language model, the basic expert module and the domain expert module to obtain the expert routing module, comprising: Under the condition that the model weight of the base large language model is fixed and the module parameters of the basic expert module and the domain expert module are fixed, the initial expert routing module is trained based on the expert routing training data to obtain the expert routing module.
8. A generation method comprising: Inputting the to-be-processed data corresponding to the generated task into a large language model, wherein the large language model comprises a basic expert module, an expert routing module and a domain expert module; Predicting the basic generated data corresponding to the data to be processed by the basic expert module, and selecting the target expert submodule matching the generated task in the domain expert module by the expert routing module; The data to be processed and the basic generation data serving as reference data of the target expert submodule are input into the target expert submodule for processing to obtain the target generation data corresponding to the generation task.
9. A data processing system, comprising a client and a server; The client is used to generate a question-and-answer request based on the question-and-answer task, and send the question-and-answer request to the server; The server is used to input the question data corresponding to the question-answering task into the large language model in response to the question-answering request, wherein the large language model includes a basic expert module, an expert routing module and a domain expert module; Predicting basic answer data corresponding to the question data by using the basic expert module, and selecting a target expert submodule matching the question-answering task in the domain expert module by using the expert routing module; The question data and the basic answer data serving as reference data of the target expert submodule are input into the target expert submodule for processing, answer data corresponding to the question-answering task is obtained, and the answer data is sent to the client.
10. A large language model training system, including a terminal side and a cloud side; The terminal side is used to send the initial large language model and the model training data set to the cloud side; The cloud side is used to determine the base large language model, the initial basic expert module, the initial domain expert module and the initial expert routing module included in the initial large language model; determine the basic expert training data, the expert routing training data and the domain expert training data in the model training data set; Training the initial basic expert module based on the basic expert training data and the base large language model to obtain a basic expert module; Training the initial domain expert module based on the domain expert training data, the base large language model and the basic expert module to obtain a domain expert module; Training the initial expert routing module based on the expert routing training data, the base large language model, the basic expert module and the domain expert module to obtain an expert routing module; A large language model is constructed based on the base large language model, the basic expert module, the expert routing module and the domain expert module, and the large language model is sent to the terminal side.
11. A computing device comprising: Memory and processor; The memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions. When the computer program or instructions are executed by the processor, the steps of the method described in any one of claims 1 to 8 are implemented.
12. A computer-readable storage medium storing a computer program or instruction, wherein the computer program or instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 8.
13. A computer program product, comprising a computer program or instructions, which implement the steps of the method according to any one of claims 1 to 8 when executed by a processor.
Citation Information
Cited By
Hierarchical text classification method and device based on hybrid experts
CN120541225A
Intelligent questioning and answering method and equipment for legal knowledge and medium
CN121029961A
A legal knowledge intelligent question answering method, device and medium
CN121029961B
Information consultation method and system based on artificial intelligence
CN121031796A