Vertical field large model construction method and device, equipment and storage medium
By constructing instruction data sets in vertical fields and introducing initialization parameter matrix into the original large model for iterative updates, the problems of poor application effect of large models in vertical fields and high training costs are solved, and efficient and safe improvement of large model construction and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510211816.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
The existing large models have poor application effects in vertical fields and are highly trained, making it difficult to effectively reduce them.
By constructing the instruction dataset of the target vertical domain and introducing the initial parameter matrix into the existing original large model for iterative updates, the vertical domain large model is obtained. The method includes constructing a rank reduction and rank increase matrix, updating parameters using the instruction data set, reducing training costs and improving model accuracy.
It realizes efficient and safe construction of large models in vertical fields, reducing training costs and time, and improving the generalization ability of the model on different tasks.
Smart Images

Figure CN120069079A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment, and storage medium for constructing a large model in a vertical domain. Background Art
[0002] Large Language Models (LLMs) are a breakthrough technology in the field of artificial intelligence, especially in Natural Language Processing (NLP). LLMs are usually based on deep learning architectures such as the Transformer model and are trained with a large amount of text data to capture the complex structures and semantic relationships of language. Among them, through self-supervised learning, LLMs can understand and generate natural language text, and perform well in various tasks such as language generation, translation, text summarization, and dialogue systems. However, due to the large computational resource consumption of LLMs and the direct use of open-source large models with relatively small parameter scales, the use effect of LLMs in specific application scenarios is generally not good.
[0003] Therefore, how to efficiently and securely implement large model applications remains an important research topic. To solve the problem of implementing large models in vertical domains, the current mainstream solutions in the industry mainly include prompt engineering and fine-tuning. Among them, prompt engineering has the advantages of not requiring retraining of the model, low cost, and strong flexibility, and is suitable for various application scenarios. However, this solution depends on the design quality of the prompts, and there may be a phenomenon of repeated experimentation and adjustment of the prompts, which in turn affects the quality of the large model and the generation efficiency of the large model. Compared with prompt engineering, fine-tuning can more deeply customize the capabilities of the model, making it perform better in professional tasks. However, fine-tuning has problems such as high implementation costs, the need for a large amount of computing resources and professional datasets, which is not conducive to reducing the model training cost. Summary of the Invention
[0004] In view of this, to better balance the quality and training cost of large models in vertical domains, the purpose of the present invention is to provide a method, device, equipment, and storage medium for constructing a large model in a vertical domain.
[0005] To achieve the above purpose, the technical solutions adopted in the embodiments of the present invention are as follows:
[0006] In the first aspect of the embodiments of the present invention, a method for constructing a large model in a vertical domain is provided, including:
[0007] Construct an instruction dataset corresponding to the target vertical domain; the instruction dataset includes multiple task instructions, as well as the input information and knowledge information corresponding to each task instruction; wherein, the input information represents the user intention, the knowledge information represents the user expectation, and the task instruction represents the task type to which the user intention belongs.
[0008] Construct multiple initialization parameter matrices corresponding one-to-one to the multiple task instructions; each initialization parameter matrix includes a rank-reducing matrix and a rank-increasing matrix connected in sequence.
[0009] Introduce the multiple initialization parameter matrices into the existing original large model in the target vertical domain, and iteratively update the parameters of the multiple initialization parameter matrices through the instruction dataset, and keep the original parameters of the original large model unchanged until the multiple initialization parameter matrices converge, to obtain a vertical domain large model; wherein, the dimensions of the rank-reducing matrix and the rank-increasing matrix are both smaller than the dimension of the original large model.
[0010] In an alternative embodiment, the step of constructing multiple initialization parameter matrices corresponding one-to-one to the multiple task instructions includes:
[0011] For each initialization parameter matrix, construct a rank-reducing matrix and a rank-increasing matrix; wherein, the number of rows of the rank-reducing matrix is the same as the number of columns of the rank-increasing matrix, and the number of columns of the rank-reducing matrix is the same as the number of rows of the rank-increasing matrix.
[0012] Assign values to the parameters of the rank-reducing matrices of the multiple initialization parameter matrices, and the parameters of the rank-reducing matrix after assignment satisfy the normal distribution and the mean is 0.
[0013] Initialize the parameters of the rank-increasing matrices of the multiple initialization parameter matrices to 0.
[0014] In an alternative embodiment, the step of iteratively updating the parameters of the multiple initialization parameter matrices through the instruction dataset includes:
[0015] Obtain sample data from the instruction dataset; the sample data includes task instructions, input information and knowledge information corresponding to the task instructions.
[0016] Process the input information in the sample data through the multiple initialization parameter matrices to obtain the predicted information output by each of the multiple initialization parameter matrices.
[0017] For each initialization parameter matrix, calculate the loss value corresponding to the initialization parameter matrix based on the prediction information output by the initialization parameter matrix, the knowledge information in the corresponding sample data, and the training weight; wherein, if the task instruction corresponding to the initialization parameter matrix is the same as the task instruction in the sample data, the training weight is 1; if the task instruction corresponding to the initialization parameter matrix is different from the task instruction in the sample data, the training weight is a positive number less than 1.
[0018] For each initialization parameter matrix, update the parameters of the initialization parameter matrix according to the loss value corresponding to the initialization parameter matrix.
[0019] Return to the step of obtaining sample data until the loss value corresponding to each currently calculated initialization parameter matrix meets the preset expected value.
[0020] In an alternative embodiment, after obtaining the vertical domain large model, the method further includes:
[0021] Construct a real-time knowledge processing module, a knowledge classifier, and a real-time knowledge model; the real-time knowledge processing module includes a scraping module, an instruction conversion module, and a storage module connected in sequence, and the output end of the knowledge classifier is respectively connected to the input end of the real-time knowledge model and the input end of the vertical domain large model; the knowledge classifier is used to classify existing knowledge and real-time knowledge, and the real-time knowledge model is used to process input information related to real-time knowledge to obtain corresponding output results.
[0022] Obtain real-time knowledge information not recorded in the instruction dataset through the scraping module.
[0023] Determine the input information and task instruction corresponding to the real-time knowledge information through the instruction conversion module to obtain a real-time knowledge instruction dataset, and store it in the storage module.
[0024] Construct a knowledge sample set according to the real-time knowledge instruction dataset and the instruction dataset.
[0025] Train the knowledge classifier through the knowledge sample set until the knowledge classifier converges.
[0026] Train the real-time knowledge model through the real-time knowledge instruction dataset until the real-time knowledge model converges.
[0027] When the knowledge classifier and the real-time knowledge model converge, enable the knowledge classifier and the real-time knowledge model to obtain an updated vertical domain large model; the updated vertical domain large model includes the knowledge classifier, the real-time knowledge model, and the vertical domain large model before the update; the knowledge classifier accesses the knowledge processing module.
[0028] In an alternative embodiment, during the application of the updated vertical domain large model, the method further includes:
[0029] When the updated vertical domain large model receives user input information, output corresponding target knowledge information through the vertical domain large model before the update;
[0030] Determine the knowledge type corresponding to the target knowledge information through the knowledge classifier;
[0031] When the knowledge type represents existing knowledge, output an output result corresponding to the target knowledge information;
[0032] When the knowledge type represents real-time knowledge, input the target knowledge information into the real-time knowledge model to obtain a corresponding output result;
[0033] Display the output result.
[0034] In an alternative embodiment, the construction process of the real-time knowledge model includes:
[0035] Construct a multi-layer perceptron with 3 layers;
[0036] Assign values to the network parameters of each layer in the multi-layer perceptron. After the assignment, the network parameters of each layer satisfy the normal distribution with a mean of 0, and the variance of the network parameters of each layer is the square root of the ratio of 2 to the input dimension of the network layer where it is located;
[0037] Use the multi-layer perceptron after the assignment as the real-time knowledge model.
[0038] In an alternative embodiment, the vertical domain large model is a recipe large model, and the recipe large model is used to output recommended recipes according to the information input by the user.
[0039] A second aspect of the embodiments of the present invention provides a vertical domain large model construction device, including:
[0040] A dataset construction module configured to construct an instruction dataset corresponding to the target vertical domain; the instruction dataset includes a plurality of task instructions, as well as the input information and knowledge information corresponding to each task instruction; wherein, the input information represents the user's intention, the knowledge information represents the user's expectation, and the task instruction represents the task type to which the user's intention belongs;
[0041] A matrix construction module, configured to: construct a plurality of initial parameter matrices corresponding one by one to the plurality of task instructions; each initial parameter matrix includes a rank-reducing matrix and a rank-increasing matrix connected in sequence;
[0042] A training module, configured to: introduce the plurality of initial parameter matrices into an existing original large model in the target vertical domain, and iteratively update the parameters of the plurality of initial parameter matrices through the instruction dataset, and keep the original parameters of the original large model unchanged until the plurality of initial parameter matrices converge, to obtain a vertical domain large model; wherein, the respective dimensions of the rank-reducing matrix and the rank-increasing matrix are smaller than the dimension of the original large model.
[0043] In a third aspect of the embodiments of the present invention, there is provided an electronic device, including a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor can execute the machine-executable instructions to implement the vertical domain large model construction method provided in the first aspect above.
[0044] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the vertical domain large model construction method provided in the first aspect above.
[0045] The vertical domain large model construction method, device, equipment and storage medium provided by the embodiments of the present invention introduce a plurality of initial parameter matrices into an existing original large model in the target vertical domain, and construct a structured instruction dataset based on the task types, user intents and user expectations related to the target vertical domain, so as to realize that only these instruction datasets are used to iteratively update the plurality of initial parameter matrices, without training the original large model, and a vertical domain large model more in line with the target vertical domain can be obtained, and the obtained vertical domain large model has a higher-precision output result. Among them, the respective dimensions of the rank-reducing matrix and the rank-increasing matrix included in each initial parameter matrix are smaller than the dimension of the original large model. Therefore, compared with training the original large model, only training the initial parameter matrices can reduce the training time and training cost. In addition, since the plurality of initial parameter matrices are constructed one by one based on a plurality of task instructions, the similarity of tasks corresponding to other task instructions can be combined during the training process of the plurality of initial parameter matrices, so that the plurality of initial parameter matrices can master cross-task shared knowledge and skills, and at the same time can maintain the differences and independence of the instruction tasks corresponding to the plurality of initial parameter matrices respectively, which is beneficial to improving the generalization ability of the vertical domain large model on different tasks.
[0046] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, details are described as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0048] Figure 1 Shows a structural block diagram of an electronic device provided by an embodiment of the present invention;
[0049] Figure 2 Shows a flowchart of a method for constructing a vertical domain large model provided by an embodiment of the present invention;
[0050] Figure 3 Shows a schematic diagram of the training process of a plurality of initial parameter matrices provided by an embodiment of the present invention;
[0051] Figure 4 Shows a structural block diagram of an updated vertical domain large model provided by an embodiment of the present invention;
[0052] Figure 5 Shows a schematic diagram of the training process of a knowledge classifier and a real-time knowledge model provided by an embodiment of the present invention;
[0053] Figure 6 Shows a functional module diagram of a device for constructing a vertical domain large model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0055] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0056] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0057] To solve traditional technical problems such as etc., the present invention provides a method for constructing a vertical domain large model. By introducing multiple initial parameter matrices into the existing original large model in the target vertical domain, and constructing a structured instruction data set based on the task types, user intents and user expectations related to the target vertical domain, it is realized that only by using these instruction data sets to iteratively update the multiple initial parameter matrices, without training the original large model, a vertical domain large model more in line with the target vertical domain can be obtained, and the obtained vertical domain large model has a higher-precision output result. Among them, the respective dimensions of the rank-reducing matrix and the rank-increasing matrix included in each initial parameter matrix are smaller than the dimension of the original large model. Therefore, compared with training the original large model, only training the initial parameter matrices can reduce the training time and training cost. In addition, since the multiple initial parameter matrices are constructed one by one based on multiple task instructions, during the training process of the multiple initial parameter matrices, the similarity of the tasks corresponding to other task instructions can be combined to enable the multiple initial parameter matrices to master cross-task shared knowledge and skills, while maintaining the differences and independence of the instruction tasks corresponding to the multiple initial parameter matrices respectively, which is beneficial to improving the generalization ability of the vertical domain large model on different tasks.
[0058] The method for constructing a vertical domain large model provided by the present invention can be applied to an electronic device. Please refer to Figure 1 , which is a structural block diagram of the electronic device. The electronic device 100 includes a memory 110, a processor 120 and a communication module 130. The elements of the memory 110, the processor 120 and the communication module 130 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines.
[0059] Among them, the memory is used to store programs or data. The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electric Erasable Programmable Read-Only Memory (EEPROM), etc.
[0060] The processor is used to read / write the data or programs stored in the memory and execute corresponding functions.
[0061] The communication module is used to establish a communication connection between the electronic device and other communication terminals through a network and is used to send and receive data through the network.
[0062] It should be understood that Figure 1 The structure shown is only a schematic diagram of the structure of the electronic device, and the electronic device may further include more or fewer components than those shown Figure 1 shown, or have a different configuration from that shown Figure 1 shown. Figure 1 Each component shown can be implemented by hardware, software, or a combination thereof.
[0063] In some embodiments, the electronic device can be applied to home appliances, industrial equipment, or public equipment. For example, assuming the target vertical domain is the recipe domain, the vertical domain large model can be used to provide recipes for users. In this way, the electronic device can be applied to kitchen equipment, and the input information entered by the user can be input into the vertical domain large model to output knowledge information related to the user input information. For example, assuming the input information is "What Cantonese cuisine can you recommend", the knowledge information output by the vertical domain large model can include relevant content such as Cantonese cuisine recipes and recipe recommendation orders, but is not limited thereto.
[0064] Based on the above example, it can be seen that in the case where the target vertical domain is the recipe domain, the vertical domain large model is a recipe large model, and the recipe large model can be used to output recommended recipes according to the information input by the user.
[0065] The following Figure 2 is used to illustrate the method for constructing a vertical domain large model provided by an embodiment of the present invention. Figure 2 is a flowchart of a method for constructing a vertical domain large model provided by an embodiment of the present invention. The method for constructing a vertical domain large model includes:
[0066] In step S100, an instruction dataset corresponding to the target vertical domain is constructed; the instruction dataset includes a plurality of task instructions, as well as input information and knowledge information corresponding to each task instruction; wherein, the input information represents the user intention, the knowledge information represents the user expectation, and the task instruction represents the task type to which the user intention belongs.
[0067] In step S200, a plurality of initial parameter matrices corresponding one-to-one to the plurality of task instructions are constructed; each initial parameter matrix includes a rank-reducing matrix and a rank-increasing matrix connected in sequence.
[0068] In step S300, the plurality of initial parameter matrices are introduced into the existing original large model in the target vertical domain, and the parameters of the plurality of initial parameter matrices are iteratively updated through the instruction dataset, and the original parameters of the original large model are kept unchanged until the plurality of initial parameter matrices converge, and a vertical domain large model is obtained; wherein, the dimensions of the rank-reducing matrix and the rank-increasing matrix are both smaller than the dimension of the original large model.
[0069] When it is necessary to obtain a vertical domain large model that is more suitable for the target vertical domain, so that the output of the vertical domain large model is more in line with the actual expectations of users and has a more accurate output result, the vertical domain model construction method provided by the embodiments of the present invention can be adopted to obtain the required vertical domain large model:
[0070] First, in the process of executing step S100, an instruction dataset can be constructed according to the existing data in the target vertical domain. The instruction dataset can be a triple-form instruction fine-tuning dataset, which includes a plurality of task instructions corresponding one-to-one to a plurality of downstream tasks, and there can be multiple input information and knowledge information corresponding to each task instruction. Based on this, a triple can represent a sample data, which includes a task instruction, an input information, and a knowledge information.
[0071] Taking the recipe domain as an example of the target vertical domain, the above-mentioned plurality of downstream tasks can be respectively: recipe hidden semantic mining task, user intention recognition task, recipe recommendation ranking task, and recommended copywriting generation task. Based on this, some examples of triples corresponding to each task can be:
[0072] Triple corresponding to the recipe hidden semantic mining task:
[0073] Task instruction: Please recommend recipes with heat-relieving effects;
[0074] Input information: Recipe data in the existing recipe database;
[0075] Knowledge information (equivalent to output information): Recipes with heat-relieving effects recommended.
[0076] Among the above, the recipe data includes but is not limited to: recipe ingredients and recipe cooking methods.
[0077] Triple corresponding to the user intention recognition task:
[0078] Task instruction: Please determine which of recipe recommendation, recipe generation, and chatting is the user's current intention based on the user's input information.
[0079] Input information: The information currently input by the user and the user information.
[0080] Knowledge information: The user's current intention is recipe recommendation.
[0081] Among the above, the user information may include the user's language style and the user's dish preference information, but is not limited to this.
[0082] Triple corresponding to the recipe recommendation sorting task:
[0083] Task instruction: Please output the recipe recommendation results and sorting based on the user information and the information currently input by the user.
[0084] Input information: The information currently input by the user and the user information.
[0085] Knowledge information: The recommended recipes and the recipe recommendation order.
[0086] Triple corresponding to the recommended copywriting generation task:
[0087] Task instruction: Please generate the most suitable recipe recommendation language for the user according to the user's language style.
[0088] Input information: The user behavior characteristics, the user's historical conversations, and the recipe recommendation results.
[0089] Knowledge information: The recipe recommendation copywriting including the recipe recommendation results.
[0090] Among the above, the recipe recommendation results include the recommended recipes and the recipe recommendation order.
[0091] After constructing the instruction data set corresponding to the target vertical domain through step S100, step S200 can be executed to construct multiple initial parameter matrices corresponding to multiple task instructions. Each initial parameter matrix includes a rank-reducing matrix and a rank-increasing matrix connected in sequence.
[0092] The purpose of using the initial parameter matrix with the above structure is: compared with the dimension W∈R of the original large model pre-training parameters d*d , the parameter dimension of the rank-reducing matrix A∈R d*r , the parameter dimension of the rank-increasing matrix B∈R r*d, where r << d. It can be seen that by adding multiple initial parameter matrices outside the original large model to optimize the output results, the amount of parameter training can be greatly reduced.
[0093] In addition, the subsequent training of multiple initial parameter matrices can achieve fine-tuning of the task instructions and output results based on the above instruction set, realize knowledge sharing of multiple tasks, and maintain the differences and independence of each instruction task.
[0094] In some embodiments, the embodiments of the present invention provide a method for constructing multiple initial parameter matrices, but not limited thereto. That is, in the above step S200, the step of constructing multiple initial parameter matrices corresponding to the multiple task instructions one by one may include:
[0095] In step S210, for each initial parameter matrix, construct a rank-reducing matrix and a rank-increasing matrix; wherein, the number of rows of the rank-reducing matrix is the same as the number of columns of the rank-increasing matrix, and the number of columns of the rank-reducing matrix is the same as the number of rows of the rank-increasing matrix;
[0096] In step S220, assign values to the parameters of the rank-reducing matrices of the multiple initial parameter matrices, and the parameters of the rank-reducing matrices after assignment satisfy the normal distribution and the mean is 0;
[0097] In step S230, initialize the parameters of the rank-increasing matrices of the multiple initial parameter matrices to 0.
[0098] Taking the construction of one initial parameter matrix as an example, the rank-reducing matrix and the rank-increasing matrix connected in sequence can be configured for this initial parameter matrix through step S210. Based on the above example, to achieve the mapping between the rank-reducing matrix and the rank-increasing matrix, the number of rows of the rank-reducing matrix is the same as the number of columns of the rank-increasing matrix, and the number of columns of the rank-reducing matrix is the same as the number of rows of the rank-increasing matrix, that is, the parameter dimension A of the rank-reducing matrix ∈ R d*r , and the parameter dimension B of the rank-increasing matrix ∈ R r*d , where the values of d and r can be set according to experience or actual needs, and the embodiments of the present invention do not limit this. Based on this, the matrix dimension mapping of the rank-reducing matrix and the rank-increasing matrix is d.
[0099] After constructing the rank-reducing matrix and the rank-increasing matrix of the initial parameter matrix through step S210, steps S220 and S230 can be executed to initialize the parameters of the rank-reducing matrix and the rank-increasing matrix respectively.
[0100] In the process of initializing the parameters of the reduced-rank matrix through step S220, the parameters in the reduced-rank matrix can be randomly assigned according to the condition that the parameters follow a normal distribution with a mean of 0. The variances of these parameters also satisfy the set shared variance value. This shared variance value can be set according to actual requirements or experience, and the embodiments of the present invention do not limit this. Among them, all the parameters of the initialized parameter matrices share the above-mentioned shared variance value.
[0101] In the process of initializing the parameters of the rank-increased matrix through step S230, the parameters of the rank-increased matrix can be uniformly initialized to 0.
[0102] In the above, step S220 and step S230 can be executed in parallel or serially. In the case of serial execution, the execution order of the two can be unrestricted.
[0103] After constructing multiple initialized parameter matrices through step S200, step S300 can be executed to train the multiple initialized parameter matrices. Continuing with the above recipe example, it can be known that there are four downstream tasks. Based on this, the number of initialized parameter matrices can also be set to four. Then the training process of the four initialized parameter matrices can be as Figure 3 shown, Figure 3 which is a schematic diagram of the training process of multiple initialized parameter matrices provided by the embodiments of the present invention.
[0104] Figure 3 Among them, there are four initialized parameter matrices, namely M1 to M4. In the training process of the multiple initialized parameter matrices, the original large model does not participate, so the parameters of the original large model are in a frozen state. In the process of iteratively updating the parameters of the multiple initialized parameter matrices using the instruction dataset, the data in the instruction dataset can be first converted into a numerical form that can be understood and processed by a computer to obtain the corresponding input representation. Then, the input representation can be input into the multiple initialized parameter matrices so that the multiple initialized parameter matrices output the corresponding output representation. Subsequently, the corresponding loss value can be calculated based on the output representation and the corresponding sample data in the instruction data, and the parameters of the initialized parameter matrices can be updated. The training principle and parameter update principle of the initialized parameters can be referred to the related technology and will not be elaborated here.
[0105] In some embodiments, to enable each initialized parameter matrix to share the knowledge of other tasks while increasing the attention to the knowledge corresponding to its own task to better enhance the model generalization ability, the embodiments of the present invention also provide a training scheme for the initialized parameter matrix, that is, in the above step S300, the step of iteratively updating the parameters of the multiple initialized parameter matrices through the instruction dataset may include:
[0106] In step S310, sample data is obtained from the instruction dataset; the sample data includes a task instruction, input information corresponding to the task instruction, and knowledge information;
[0107] In step S320, the input information in the sample data is processed by the multiple initial parameter matrices to obtain prediction information output by each of the multiple initial parameter matrices;
[0108] In step S330, for each initial parameter matrix, a loss value corresponding to the initial parameter matrix is calculated according to the prediction information output by the initial parameter matrix, the knowledge information in the corresponding sample data, and the training weight; wherein, if the task instruction corresponding to the initial parameter matrix is the same as the task instruction in the sample data, the training weight is 1; if the task instruction corresponding to the initial parameter matrix is different from the task instruction in the sample data, the training weight is a positive number less than 1;
[0109] In step S340, for each initial parameter matrix, the parameters of the initial parameter matrix are updated according to the loss value corresponding to the initial parameter matrix;
[0110] In step S350, return to the step of obtaining sample data until the loss value corresponding to each currently calculated initial parameter matrix meets a preset expected value.
[0111] Thus, by adding a training weight in the calculation process of the loss value corresponding to each initial parameter matrix, and setting the training weight to 1 when the task instruction corresponding to the initial parameter matrix is the same as the instruction of the currently input sample data, and setting the training weight to a positive number less than 1, for example, 0.5, when the task instruction corresponding to the initial parameter matrix is different from the instruction of the currently input sample data, it can be ensured that the initial parameter matrix can enhance the learning of its own task knowledge while learning knowledge related to other tasks, thereby enabling the model to have stronger generalization ability.
[0112] In some examples, the calculation principle of the loss value of each initial parameter matrix can be the same. For example, the corresponding loss value can be calculated by a loss function shown in the following calculation formula:
[0113]
[0114] In the above formula, n represents the subscript of each initial parameter matrix, used to distinguish different initial parameter matrices. For example, continuing with the above example, for the four initial parameter matrices M1 to M4, their corresponding n values can be 1, 2, 3, and 4 respectively. Correspondingly, m represents the total number of initial parameter matrices. MSE represents the mean square error. wn represents the training weight of the initialization parameter matrix with subscript n, w n ∈(0, 1]. In a certain task, the subscript of the task instruction in the corresponding sample data is the same as the subscript of the corresponding initialization parameter matrix. Therefore, it is possible to determine whether the task instruction corresponding to the initialization parameter matrix is the same as the task instruction corresponding to the currently input sample data based on this. A n represents the rank-reducing matrix in the initialization parameter matrix with subscript n, B n represents the rank-increasing matrix in the initialization parameter matrix with subscript n. X represents the input representation, W represents the parameter weight of the original large model, and y represents the knowledge information in the sample data, that is, the true value.
[0115] During the process of updating the parameters of the initialization parameter matrix, that is, during the process of gradient direction propagation, whether it is a rank-reducing matrix or an increment matrix, the calculation principle of parameter update can be the same. Taking the parameter update calculation of the rank-reducing matrix as an example below, the corresponding calculation formula is exemplified:
[0116] A′ n = A n - η n dA n
[0117] In the above formula, A′ n represents the updated parameter in the rank-reducing matrix; dA n represents the gradient, which can be obtained through the formula calculated, represents the gradient of the loss function Loss with respect to A n . η n represents the dynamic learning rate corresponding to the initialization parameter matrix with subscript n, and η n ∈[1e -5 , 3e -5 .
[0118] After multiple initialization parameter matrices converge through step S300, the corresponding vertical domain large model can be obtained.
[0119] In the application stage of the vertical domain large model, the information currently input by the user will be respectively input into the original large model and multiple initialization parameter matrices. The original large model and multiple initialization parameter matrices respectively output the corresponding output representations, and then fusion processing is performed through a layer of network to obtain the final output. The fusion processing therein can refer to related technologies and will not be elaborated here.
[0120] As time goes by, new knowledge information may emerge in the target vertical domain. Since this new knowledge information did not exist during the construction phase of the large model for the vertical domain, it is possible that the large model for the vertical domain cannot provide answers that meet user expectations when new knowledge information appears. Therefore, to solve this technical problem, in some embodiments, the method for constructing a large model for a vertical domain provided by the embodiments of the present invention may further include related solutions for configuring a real-time knowledge model for the large model of the vertical domain, so as to enable the large model of the vertical domain to keep up with the times. Based on this, after obtaining the large model for the vertical domain, the method for constructing a large model for a vertical domain provided by the embodiments of the present invention may further include:
[0121] In step S400, a real-time knowledge processing module, a knowledge classifier, and a real-time knowledge model are constructed; the real-time knowledge processing module includes a scraping module, an instruction conversion module, and a storage module connected in sequence, and the output end of the knowledge classifier is respectively connected to the input end of the real-time knowledge model and the input end of the large model for the vertical domain; the knowledge classifier is used to classify existing knowledge and real-time knowledge, and the real-time knowledge model is used to process input information related to real-time knowledge to obtain corresponding output results;
[0122] In step S500, the scraping module is used to obtain real-time knowledge information not recorded in the instruction dataset;
[0123] In step S600, the instruction conversion module is used to determine the input information and task instructions corresponding to the real-time knowledge information to obtain a real-time knowledge instruction dataset, which is stored in the storage module;
[0124] In step S700, a knowledge sample set is constructed according to the real-time knowledge instruction dataset and the instruction dataset;
[0125] In step S810, the knowledge classifier is trained with the knowledge sample set until the knowledge classifier converges;
[0126] In step S820, the real-time knowledge model is trained with the real-time knowledge instruction dataset until the real-time knowledge model converges;
[0127] In step S900, when the knowledge classifier and the real-time knowledge model converge, the knowledge classifier and the real-time knowledge model are enabled to obtain an updated large model for the vertical domain; the updated large model for the vertical domain includes the knowledge classifier, the real-time knowledge model, and the large model for the vertical domain before the update; the knowledge classifier is connected to the knowledge processing module.
[0128] After obtaining the vertical domain large model, a real-time knowledge processing module, a knowledge classifier, and a real-time knowledge model can be further added outside the vertical domain large model. The real-time knowledge processing module may include a scraping module, an instruction conversion module, and a storage module connected in sequence.
[0129] After configuring the above real-time knowledge processing module, knowledge classifier, and real-time knowledge model in the vertical domain large model, the real-time knowledge processing module, knowledge classifier, and real-time knowledge model can be periodically triggered to perform corresponding tasks or updates.
[0130] After the real-time knowledge processing module is triggered, step S500 can be executed to scrape real-time knowledge information not recorded in the instruction dataset through the scraping module. Since the real-time knowledge information is not recorded in the instruction dataset, it can be considered that the real-time knowledge information has emerged recently. For example, assuming that the European Cup in the summer of 2024 is held in China and users need to stay up all night, relevant food can be prepared in advance to liven up the atmosphere, such as fried chicken. This information can be used as real-time knowledge information, which can be crawled from the network by the scraping module or obtained from the database and preprocessed and filtered.
[0131] After scraping the real-time knowledge information through the scraping module, step S600 can be executed to use the instruction conversion module to process the real-time knowledge information to obtain the corresponding input information and task instructions. It can be understood as performing instruction conversion on the real-time knowledge information to better be compatible with the input and output of the large model and enhance the expression ability of the downstream real-time knowledge model at the same time.
[0132] The following continues to use the above example to illustrate the process of obtaining the corresponding triples based on the real-time knowledge information processing:
[0133] Assume the real-time knowledge information is: The European Cup in the summer of 2024 is held in China, users need to stay up all night, and fried chicken needs to be prepared in advance.
[0134] Based on this, the instruction conversion module will construct the following task instructions, input information, and knowledge information based on the real-time knowledge information:
[0135] Task instruction: Recommend recipes to users on specific holidays or dates;
[0136] Input information: The European Cup is coming soon. What is good to eat?
[0137] Knowledge information: Recommend fried chicken for the European Cup in the summer of 2024.
[0138] Thus, the instruction conversion module can construct triples corresponding to each real-time knowledge information, and then these triples can be stored in the real-time knowledge instruction dataset, and the real-time knowledge instruction dataset can be stored in the storage module.
[0139] The storage module can convert the real-time knowledge instruction dataset into corresponding real-time knowledge embedding representations for the downstream knowledge classifier to call this real-time knowledge.
[0140] As an example, Faiss can be used as the storage module. Faiss is a vector storage database that can accelerate the retrieval of real-time knowledge embedding representations.
[0141] After obtaining the real-time knowledge instruction dataset through step S600, step S700 can be executed to construct a knowledge sample set based on the real-time knowledge instruction dataset and the instruction dataset. It can be understood that the knowledge sample set includes new knowledge and old knowledge, where the new knowledge comes from the real-time knowledge instruction dataset and the old knowledge comes from the instruction dataset. The knowledge sample set is used as the training sample of the knowledge classifier.
[0142] In addition, since the amount of new knowledge in the real-time instruction dataset is much smaller than that of the old knowledge in the instruction dataset, all the data in the real-time instruction dataset can be incorporated into the knowledge sample set, which can better improve the classification accuracy of the knowledge classifier for new knowledge and old knowledge.
[0143] After obtaining the knowledge sample set, step S810 can be executed to train the knowledge classifier using the knowledge sample set until the knowledge classifier converges. The training principle can also be referred to the related technology. The converged knowledge classifier has the ability to distinguish the knowledge mastered by the vertical domain large model and the newly added knowledge.
[0144] In some examples, the knowledge classifier can use FastText, which can improve the training speed and knowledge classification performance of the knowledge classifier.
[0145] In some examples, the loss function applied by the knowledge classifier can be the cross-entropy loss function, and the expression of this function is as follows:
[0146]
[0147] In the above formula, D represents the knowledge sample set, E represents the mathematical expectation of the cross-entropy loss function, (x, y) ∼ D means the sample (x, y) belongs to the knowledge sample set D, represents the loss of a single sample for training parameter calculation, that is, during the training process of the knowledge classifier, the loss value calculated for a single sample, and this loss value is calculated based on the current parameters of the knowledge classifier. (y in , x on ) represents old knowledge and belongs to the instruction dataset; (y out , x out ) represents new knowledge and belongs to the real-time knowledge instruction dataset.
[0148] While or after performing step S810, step S820 may be performed to train the real-time knowledge model using the real-time knowledge instruction data set until the real-time knowledge model converges. The training principle therein may also be referred to the related art. The converged real-time knowledge model has the ability to handle problems related to new knowledge.
[0149] After both the knowledge classifier and the real-time knowledge model converge, step S900 may be performed to enable the knowledge classifier and the real-time knowledge model to obtain an updated vertical domain large model. Please refer to Figure 4 , Figure 4 FIG. is a structural block diagram of an updated vertical domain large model provided by an embodiment of the present invention. The updated vertical domain large model includes a knowledge classifier, a real-time knowledge model, and the vertical domain large model before update; the knowledge classifier is connected to the knowledge processing module. Among them, the vertical domain large model before update is the vertical domain large model obtained in the above step S300.
[0150] Based on the previous embodiment, to realize the simultaneous update of multiple new knowledge and further reduce the training cost of the knowledge classifier and the real-time knowledge model, in some embodiments, the present invention also provides another training scheme for the knowledge classifier and the real-time knowledge model. Please refer to Figure 5 , Figure 5 FIG. is a schematic diagram of the training process of a knowledge classifier and a real-time knowledge model provided by an embodiment of the present invention. It can be understood that the training processes of the knowledge classifier and the real-time knowledge model can be combined. In this training process, the original large model does not participate in the training, and its parameters are in a frozen state.
[0151] Based on this, during the training process or during the subsequent new knowledge update process, the training or update of the knowledge classifier and the real-time knowledge model can be realized through the training process shown in Figure 5 :
[0152] For the training stage, the knowledge sample set can be obtained through the above step S700, and then the knowledge sample set is input into the vertical domain large model before update. The original large model and each parameter matrix in the vertical domain large model before update do not participate in the training, and their parameters are all in a frozen state. Among them, the information included in the knowledge sample set can be representational information or non-representational information, and is converted into representational information before being input into the model. The real-time instruction data set and the instruction data set can be the same by analogy.
[0153] In addition to being input into the vertical domain large model before update, the real-time instruction data set in the knowledge sample set can also be input into the storage module for storage in the storage module.
[0154] Next, the storage module inputs the real-time instruction dataset into the knowledge classifier, and the vertical domain large model before update also inputs the output results processed from the knowledge sample set into the knowledge classifier.
[0155] Subsequently, the knowledge classifier will perform classification training based on the characterization information output by the storage module and the vertical domain large model before update, and output the corresponding classification results.
[0156] In the case where the classification result output by the knowledge classifier represents new knowledge, the knowledge classifier inputs the corresponding new knowledge characterization information into the real-time knowledge model, so that the real-time knowledge model is trained according to the new knowledge characterization information. On the contrary, in the case where the classification result output by the knowledge classifier represents old knowledge, the knowledge classifier directly outputs the old knowledge characterization information to the output layer.
[0157] Thus, in one training process, the training of both the knowledge classifier and the real-time knowledge model can be achieved simultaneously.
[0158] For the update stage, after the knowledge processing module grabs new knowledge and forms the corresponding instruction data according to the principle described above, it can be stored in the storage module, and the knowledge classifier and the real-time knowledge model are triggered to update. Or, to further ensure the effectiveness of the update trigger, a search operation can be performed on the real-time instruction set stored in the storage module according to the instruction data corresponding to the grabbed new knowledge. If no instruction data with a similarity not higher than the set similarity threshold to the instruction data corresponding to the grabbed new knowledge is found, it means that the grabbed new knowledge is indeed new knowledge, and then it is stored in the storage module, and the knowledge classifier and the real-time knowledge model are triggered to update. On the contrary, if instruction data with a similarity higher than the set similarity threshold to the instruction data corresponding to the grabbed new knowledge is found, it means that the grabbed new knowledge has been learned, and then there is no need to store it in the storage module, and the knowledge classifier and the real-time knowledge model are not triggered to update.
[0159] When there are multiple pieces of grabbed new knowledge, the simultaneous update of multiple pieces of knowledge can be achieved.
[0160] Thus, in one update process, the update of both the knowledge classifier and the real-time knowledge model can be achieved simultaneously.
[0161] It can be seen that whether in the training stage or the update stage, the vertical domain large model before update does not need to participate, which can greatly reduce the relevant parameter update amount and the required data volume; and, in one training or update process, the training or update of both the knowledge classifier and the real-time knowledge model can be achieved simultaneously, and the two can use the same new knowledge data without the need for more data preparation. Thus, the purpose of reducing the training cost and improving the training efficiency can be well achieved.
[0162] In some embodiments, to avoid overfitting of the real-time knowledge model and affect the output quality, the real-time knowledge model can be implemented using a multi-layer perceptron. Based on this, the embodiments of the present invention also provide a construction scheme for the real-time knowledge model, that is, the construction process of the real-time knowledge model includes:
[0163] In step S410, construct a multi-layer perceptron with 3 layers;
[0164] In step S420, assign values to the network parameters of each layer in the multi-layer perceptron. After the assignment, the network parameters of each layer satisfy the normal distribution with a mean of 0, and the variance of the network parameters of each layer is the square root of the ratio of 2 to the input dimension of the network layer where it is located;
[0165] In step S430, use the multi-layer perceptron with assigned values as the real-time knowledge model.
[0166] In the above step S410, configuring the number of layers of the multi-layer perceptron to 3 can ensure that the introduced number of parameters is appropriate and avoid overfitting and affecting the output quality.
[0167] In the above step S420, the corresponding scheme for assigning values to the network parameters of each layer can be expressed by the following formula:
[0168]
[0169] In the above formula, ω l represents the network parameter, n l-input represents the dimension of the input of the layer for which parameter assignment is currently required. If the information flow of each layer is kept with the same variance it will be more beneficial to optimize the output of the multi-layer perceptron. Moreover, in the embodiments of the present invention, during the construction process of the multi-layer perceptron, the He parameter initialization method is used, combined with the non-linear mapping function (Exponential Linear Unit, ELU) configured in each layer of the multi-layer perceptron, which is smoother than the Relu activation function and can avoid gradient explosion. On this basis, by configuring the network parameters of each layer to meet the requirements of Gaussian distribution, mean of 0, and variance normalization, etc., the convergence of the multi-layer perceptron can be accelerated, thereby improving the training efficiency.
[0170] Based on the previous embodiment, in some examples, the loss function of the multi-layer perceptron can be:
[0171]
[0172] In the above formula, D in represents the real-time knowledge instruction data set; E represents the mathematical expectation of the loss function of the multi-layer perceptron; P ψ represents the single-sample loss calculated by the training parameter; (yin , x in ) represents old knowledge and belongs to the instruction dataset.
[0173] It can be seen that the multi-layer perceptron can effectively learn the representation of new knowledge without interfering with the general knowledge output of the vertical domain large model before the update. In addition, by configuring the multi-layer perceptron in the output layer of the vertical large model before the update and editing the parameters of the multi-layer perceptron, it is possible to better achieve the simultaneous update of the multi-layer perceptron based on multiple pieces of knowledge and better reduce the training cost.
[0174] In the above, editing the parameters of the multi-layer perceptron includes technical means at two levels. One level of technical means is the technical solution shown in the above steps S410 to S430, and the other level of technical means is the training or update solution of the real-time knowledge model in any of the above embodiments.
[0175] After obtaining the updated vertical domain large model, in some embodiments, to ensure that the updated vertical domain large model can correctly process the input information related to new knowledge and the input information related to old knowledge, the vertical domain large model construction method provided by the embodiments of the present invention also provides an application solution related to the updated vertical domain large model, that is, in the application process of the updated vertical domain large model, the vertical domain large model construction method provided by the embodiments of the present invention may further include:
[0176] In step S1000, when the updated vertical domain large model receives user input information, the corresponding target knowledge information is output through the vertical domain large model before the update;
[0177] In step S2000, the knowledge type corresponding to the target knowledge information is determined through the knowledge classifier;
[0178] In step S3000, when the knowledge type represents existing knowledge, the output result corresponding to the target knowledge information is output;
[0179] In step S4000, when the knowledge type represents real-time knowledge, the target knowledge information is input into the real-time knowledge model to obtain the corresponding output result;
[0180] In step S5000, the output result is displayed.
[0181] Therefore, when the updated vertical domain large model receives the user input information, first use the pre-updated vertical domain large model to output the corresponding target knowledge information; then use the knowledge classifier to determine the knowledge type corresponding to the target knowledge information, where the knowledge type is existing knowledge or real-time knowledge. If the knowledge type determined by the knowledge classifier represents existing knowledge, it means that the pre-updated vertical domain large model can apply the knowledge it has mastered to process the user input information. Therefore, execute step S3000 and directly output the output result corresponding to the target knowledge information. On the contrary, if the knowledge type determined by the knowledge classifier represents real-time knowledge, it means that the pre-updated vertical domain large model does not master the relevant knowledge. Therefore, to obtain a more accurate output result through the real-time knowledge model, execute step S4000 to further process the target knowledge information through the real-time knowledge model to output the corresponding result.
[0182] After obtaining the output result, step S5000 can be executed to display the output result to the user.
[0183] It should be noted that the technical features or technical solutions in any of the above embodiments of the present invention can be combined with each other as long as there is no combination conflict.
[0184] To execute the corresponding steps in the above embodiments and various possible ways, an implementation manner of a vertical domain large model construction device is given below. Optionally, the vertical domain large model construction device can adopt the device structure of the electronic device shown above. Figure 1 Further, please refer to Figure 6 , Figure 6 FIG. is a functional module diagram of a vertical domain large model construction device provided by an embodiment of the present invention. It should be noted that the basic principle and the technical effects generated by the vertical domain large model construction device provided in this embodiment are the same as those in the above embodiments. For a brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above embodiments. The vertical domain large model construction device 600 includes:
[0185] A dataset construction module 610, configured to: construct an instruction dataset corresponding to the target vertical domain; the instruction dataset includes a plurality of task instructions, as well as the input information and knowledge information corresponding to each task instruction; wherein, the input information represents the user intention, the knowledge information represents the user expectation, and the task instruction represents the task type to which the user intention belongs;
[0186] A matrix construction module 620, configured to: construct a plurality of initialization parameter matrices corresponding one-to-one to the plurality of task instructions; each initialization parameter matrix includes a rank-reducing matrix and a rank-increasing matrix connected in sequence;
[0187] The training module 630 is configured to: introduce the plurality of initial parameter matrices into the existing original large model in the target vertical domain, iteratively update the parameters of the plurality of initial parameter matrices through the instruction dataset, and keep the original parameters of the original large model unchanged until the plurality of initial parameter matrices converge to obtain a vertical domain large model; wherein, the respective dimensions of the rank-reducing matrix and the rank-increasing matrix are smaller than the dimension of the original large model.
[0188] In some embodiments, the process of the matrix construction module 620 constructing the plurality of initial parameter matrices corresponding to the plurality of task instructions is configured to:
[0189] For each initial parameter matrix, construct a rank-reducing matrix and a rank-increasing matrix; wherein, the number of rows of the rank-reducing matrix is the same as the number of columns of the rank-increasing matrix, and the number of columns of the rank-reducing matrix is the same as the number of rows of the rank-increasing matrix;
[0190] Assign values to the parameters of the rank-reducing matrices of the plurality of initial parameter matrices, and the parameters of the rank-reducing matrices after assignment satisfy a normal distribution and have a mean of 0;
[0191] Initialize the parameters of the rank-increasing matrices of the plurality of initial parameter matrices to 0.
[0192] In some embodiments, the process of the training module 630 iteratively updating the parameters of the plurality of initial parameter matrices through the instruction dataset is configured to:
[0193] Obtain sample data from the instruction dataset; the sample data includes task instructions, input information corresponding to the task instructions, and knowledge information;
[0194] Process the input information in the sample data through the plurality of initial parameter matrices to obtain prediction information output by each of the plurality of initial parameter matrices;
[0195] For each initial parameter matrix, calculate a loss value corresponding to the initial parameter matrix according to the prediction information output by the initial parameter matrix, the knowledge information in the corresponding sample data, and the training weight; wherein, if the task instruction corresponding to the initial parameter matrix is the same as the task instruction in the sample data, the training weight is 1; if the task instruction corresponding to the initial parameter matrix is different from the task instruction in the sample data, the training weight is a positive number less than 1;
[0196] For each initial parameter matrix, update the parameters of the initial parameter matrix according to the loss value corresponding to the initial parameter matrix;
[0197] Return to the step of obtaining sample data until the loss value corresponding to each initialized parameter matrix calculated currently meets the preset expected value.
[0198] In some embodiments, the vertical domain large model construction device 600 may further include a knowledge construction module and an enabling module.
[0199] The knowledge construction module is configured to:
[0200] Construct a real-time knowledge processing module, a knowledge classifier, and a real-time knowledge model; the real-time knowledge processing module includes a grabbing module, an instruction conversion module, and a storage module connected in sequence, and the output end of the knowledge classifier is respectively connected to the input end of the real-time knowledge model and the input end of the vertical domain large model; the knowledge classifier is used to classify existing knowledge and real-time knowledge, and the real-time knowledge model is used to process input information related to real-time knowledge to obtain corresponding output results;
[0201] Obtain real-time knowledge information not recorded in the instruction dataset through the grabbing module;
[0202] Determine the input information and task instructions corresponding to the real-time knowledge information through the instruction conversion module to obtain a real-time knowledge instruction dataset, and store it in the storage module;
[0203] Construct a knowledge sample set according to the real-time knowledge instruction dataset and the instruction dataset;
[0204] Train the knowledge classifier through the knowledge sample set until the knowledge classifier converges;
[0205] Train the real-time knowledge model through the real-time knowledge instruction dataset until the real-time knowledge model converges.
[0206] The enabling module is configured to:
[0207] When the knowledge classifier and the real-time knowledge model converge, enable the knowledge classifier and the real-time knowledge model to obtain an updated vertical domain large model; the updated vertical domain large model includes the knowledge classifier, the real-time knowledge model, and the updated vertical domain large model before update; the knowledge classifier accesses the knowledge processing module.
[0208] In some embodiments, the process of the knowledge construction module constructing the real-time knowledge model is configured to:
[0209] Construct a multi-layer perceptron with 3 layers;
[0210] Assigning values to the network parameters of each layer in the multilayer perceptron, wherein the network parameters of each layer after the assignment satisfy a normal distribution and have a mean of 0, and the variance of the network parameters of each layer is the square root of the ratio of 2 to the input dimension of the network in the layer where the network is located;
[0211] The multilayer perceptron after the assignment is used as the real-time knowledge model.
[0212] In some embodiments, the vertical domain large model building device 600 may further include:
[0213] Application modules are configured as follows:
[0214] During the application of the updated vertical domain big model, when the updated vertical domain big model receives user input information, the corresponding target knowledge information is output through the vertical domain big model before the update;
[0215] Determining the knowledge type corresponding to the target knowledge information by the knowledge classifier;
[0216] When the knowledge type represents existing knowledge, outputting an output result corresponding to the target knowledge information;
[0217] When the knowledge type represents real-time knowledge, the target knowledge information is input into the real-time knowledge model to obtain a corresponding output result;
[0218] Display the output results.
[0219] In some embodiments, the vertical field big model is a recipe big model, and the recipe big model is used to output recommended recipes based on information input by the user.
[0220] Optionally, the above modules can be stored in the form of software or firmware. Figure 1 The memory shown in the figure or solidified in the operating system (OS) of the electronic device, and can be Figure 1 Meanwhile, the data and program codes required for executing the above modules can be stored in the memory.
[0221] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0222] In addition, each functional module in various embodiments of the present invention can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0223] If the above functions are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0224] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for constructing a large vertical model, characterized in that: include: Build a command dataset corresponding to the target vertical field; The instruction data set includes a plurality of task instructions, and input information and knowledge information corresponding to each task instruction; wherein the input information represents the user intention, the knowledge information represents the user expectation, and the task instruction represents the task type to which the user intention belongs; Constructing a plurality of initialization parameter matrices corresponding to the plurality of task instructions one by one; each initialization parameter matrix includes a reduced rank matrix and an increased rank matrix connected in sequence; The multiple initialization parameter matrices are introduced into the original large model already existing in the target vertical field, and the parameters of the multiple initialization parameter matrices are iteratively updated through the instruction data set, while the original parameters of the original large model are kept unchanged, until the multiple initialization parameter matrices converge to obtain a vertical field large model; wherein the respective dimensions of the reduced rank matrix and the increased rank matrix are smaller than the dimensions of the original large model.
2. The method according to claim 1, characterized in that The step of constructing a plurality of initialization parameter matrices corresponding one-to-one to the plurality of task instructions comprises: For each initialization parameter matrix, construct a reduced rank matrix and an increased rank matrix; wherein the number of rows of the reduced rank matrix is the same as the number of columns of the increased rank matrix, and the number of columns of the reduced rank matrix is the same as the number of rows of the increased rank matrix; Assigning values to the parameters of the reduced-rank matrix of the multiple initialization parameter matrices, wherein the parameters of the reduced-rank matrix after the assignment satisfy a normal distribution and have a mean of 0; Initialize the parameters of the rank-increasing matrices of the multiple initialization parameter matrices to 0.
3. The method according to claim 2, characterized in that The step of iteratively updating the parameters of the multiple initialization parameter matrices through the instruction data set includes: Acquire sample data from the instruction data set; the sample data includes task instructions, input information and knowledge information corresponding to the task instructions; Processing the input information in the sample data through the multiple initialization parameter matrices to obtain prediction information output by each of the multiple initialization parameter matrices; For each initialization parameter matrix, the loss value corresponding to the initialization parameter matrix is calculated according to the prediction information output by the initialization parameter matrix and the knowledge information in the corresponding sample data, as well as the training weight; wherein, if the task instruction corresponding to the initialization parameter matrix is the same as the task instruction in the sample data, the training weight is 1; if the task instruction corresponding to the initialization parameter matrix is different from the task instruction in the sample data, the training weight is a positive number less than 1; For each initialization parameter matrix, update the parameters of the initialization parameter matrix according to the loss value corresponding to the initialization parameter matrix; Return to the step of obtaining sample data until the loss value corresponding to each initialization parameter matrix currently calculated meets the preset expected value.
4. The method according to any one of claims 1 to 3, characterized in that: After obtaining the vertical domain macro model, the method further includes: Construct a real-time knowledge processing module, a knowledge classifier and a real-time knowledge model; the real-time knowledge processing module includes a capture module, an instruction module and a storage module connected in sequence, and the output end of the knowledge classifier is respectively connected to the input end of the real-time knowledge model and the input end of the vertical field large model; the knowledge classifier is used to classify existing knowledge and real-time knowledge, and the real-time knowledge model is used to process input information related to real-time knowledge to obtain corresponding output results; Acquire the real-time knowledge information not recorded in the instruction data set through the capture module; Determining input information and task instructions corresponding to the real-time knowledge information through the instruction module to obtain a real-time knowledge instruction data set, and storing it in the storage module; Constructing a knowledge sample set according to the real-time knowledge instruction data set and the instruction data set; Training the knowledge classifier using the knowledge sample set until the knowledge classifier converges; Training the real-time knowledge model using the real-time knowledge instruction data set until the real-time knowledge model converges; When the knowledge classifier and the real-time knowledge model converge, the knowledge classifier and the real-time knowledge model are enabled to obtain an updated vertical domain big model; the updated vertical domain big model includes the knowledge classifier, the real-time knowledge model and the vertical domain big model before update; the knowledge classifier is connected to the knowledge processing module.
5. The method according to claim 4, characterized in that During the application of the updated vertical domain macro model, the method further includes: When the updated vertical domain big model receives user input information, the corresponding target knowledge information is output through the vertical domain big model before the update; Determining the knowledge type corresponding to the target knowledge information by the knowledge classifier; When the knowledge type represents existing knowledge, outputting an output result corresponding to the target knowledge information; When the knowledge type represents real-time knowledge, the target knowledge information is input into the real-time knowledge model to obtain a corresponding output result; Display the output results.
6. The method according to claim 4, characterized in that The construction process of the real-time knowledge model includes: Construct a multi-layer perceptron with 3 layers; Assigning values to the network parameters of each layer in the multilayer perceptron, wherein the network parameters of each layer after the assignment satisfy a normal distribution and have a mean of 0, and the variance of the network parameters of each layer is the square root of the ratio of 2 to the input dimension of the network in the layer where the network is located; The multilayer perceptron after the assignment is used as the real-time knowledge model.
7. The method according to claim 1, characterized in that The vertical field big model is a recipe big model, and the recipe big model is used to output recommended recipes according to information input by the user.
8. A vertical field large model construction device, characterized in that: include: The data set construction module is configured to: construct an instruction data set corresponding to the target vertical field; The instruction data set includes a plurality of task instructions, and input information and knowledge information corresponding to each task instruction; wherein the input information represents the user intention, the knowledge information represents the user expectation, and the task instruction represents the task type to which the user intention belongs; A matrix construction module is configured to: construct a plurality of initialization parameter matrices corresponding to the plurality of task instructions one by one; each initialization parameter matrix includes a reduced rank matrix and an increased rank matrix connected in sequence; The training module is configured to: introduce the multiple initialization parameter matrices into the original large model already existing in the target vertical field, and iteratively update the parameters of the multiple initialization parameter matrices through the instruction data set, and keep the original parameters of the original large model unchanged, until the multiple initialization parameter matrices converge to obtain the vertical field large model; wherein the respective dimensions of the reduced rank matrix and the increased rank matrix are smaller than the dimensions of the original large model.
9. An electronic device, characterized in that: The invention comprises a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor can execute the machine executable instructions to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.