Training Method of Text Processing Model, Text Processing Method and Device
By introducing expert modules and routing modules to generate and select boot sequences, the problem of large amount of pre-trained models and limited training resources is solved, and the ability to improve the base model on composite tasks is realized.
Patent Information
- Application Number
- CN202510525695.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The existing pretrained models have large amounts of parameters, making it difficult to effectively stimulate performance in specific tasks, and are limited by training resources, so comprehensive parameter training cannot be carried out.
Multiple expert modules and a routing module are introduced to generate different boot sequences, and the final boot sequence is selected by the routing module and added to the base model for forward propagation, and only the parameters of the expert module and routing module are updated.
Without affecting the base model, training resources are saved, the capabilities of the base model in composite tasks are improved, and an effective way for training large-parameter model under low resource conditions is provided.
Smart Images

Figure CN120067697B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a training method for a text processing model, a text processing method, and an apparatus. Background Art
[0002] With the development of artificial intelligence and big data technologies, pre-training endows the model with more basic knowledge, and these pre-trained models can be effectively applied to various tasks, showing very excellent effects. In the prior art, after the model is pre-trained, it obtains a large amount of knowledge. In practical applications, the performance of the model in specific tasks can usually be stimulated by adjusting the input guiding sequence. At the same time, in the current model application process, the number of parameters of the model itself is getting larger and larger, and the data and knowledge solidified into the model through pre-training are also increasing. If it is necessary to stimulate the performance of the model in a specific task, it is often necessary to adjust a large number of guiding sequences in combination with the business.
[0003] However, due to the large number of parameters of the model itself, if more domain knowledge needs to be endowed to the model while ensuring that the learned knowledge is not lost, it is often necessary to find more general knowledge data than professional knowledge for training together. At the same time, limited by the training resources, it is often impossible to perform post-training on all the parameters of the model. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a training method for a text processing model, a text processing method, and an apparatus, so as to improve the ability of the base model in composite tasks and reduce the impact of model training on the base model.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0006] In the first aspect, the present invention provides a training method for a text processing model. The text processing model includes: a plurality of expert modules, a routing module, and a base model layer, including: obtaining a training sample set; wherein, the training sample set includes: a plurality of composite text data composed of text data of a single text task; the text data of the single text task at least includes: a text similarity recognition task, a text category recognition task, a text feature extraction task, and a text answering task; inputting each composite text data in the training sample set into the expert module and the routing module to obtain a guiding sequence for each composite text data; inputting each composite text data and the guiding sequence of each composite text data into the base model layer to obtain an inference result for each composite text data; determining a loss function value based on the inference result of each composite text data and the annotation result of the composite text data; and updating the model parameters of the expert module and the routing module based on the loss function value.
[0007] Optionally, each composite text data in the training sample set is input into the expert module and the routing module to obtain the guiding sequence of each composite text data, including: for each composite text data in the training sample set, the composite text data is input into the expert module to obtain the guiding sequence output by each expert module; each composite text data in the training sample set is input into the routing module to obtain the weight coefficients of each expert module, and the weight coefficients are sorted in descending order, and a preset number of weight coefficients of the expert modules are selected as the output result of the routing module according to the sorting result; wherein, the output result of the routing module includes: the numbers of a preset number of expert modules and the corresponding weight coefficients; based on the numbers of a preset number of expert modules output by the routing module and the corresponding weight coefficients, the weight coefficients are weighted and calculated with the guiding sequences output by the corresponding expert modules to obtain the guiding sequence of the composite text data.
[0008] Optionally, the expert module includes: a dynamic expert module with a multi-layer network structure and a static expert module with a single-layer network structure; the dynamic expert module includes: a first encoding structure, an upsampling structure, a gating structure, and a downsampling structure; the upsampling structure is a single-layer fully-connected neural network, and the gating structure is a single-layer fully-connected neural network and a single activation layer; the downsampling structure is a single-layer fully-connected neural network; the static expert module includes: a second encoding structure, and the second encoding structure is a single-layer fully-connected neural network; the routing module is a single-layer neural network structure.
[0009] Optionally, inputting the composite text data into the expert module to obtain the guiding sequence output by each expert module includes: inputting the composite text data into the dynamic expert module to obtain a first guiding sequence; inputting the composite text data into the static expert module, and generating a second guiding sequence with a preset length based on the fully-connected neural network of the second encoding structure.
[0010] Optionally, inputting the composite text data into the dynamic expert module to obtain a first guiding sequence includes: encoding the composite text data based on the first encoding structure to obtain an initial guiding sequence with a preset length; expanding the dimensions of the initial guiding sequence based on the upsampling structure to obtain an upsampling result; expanding the dimensions of the initial guiding sequence based on the fully-connected neural network of the gating structure to obtain an output result, and activating the output result based on the activation layer to obtain a gating output result; performing matrix multiplication on the upsampling result and the gating output result based on the downsampling structure to obtain a matrix multiplication result, and passing the matrix multiplication result through the fully-connected neural network of the downsampling structure to obtain the first guiding sequence output by the dynamic expert module.
[0011] Optionally, before obtaining the training sample set, it further includes: obtaining the hyperparameters of the expert module and the routing module; wherein, the hyperparameters at least include: the output dimension of the first encoding structure for generating the guiding sequence, the output dimension of the upsampling structure, the output dimension of the downsampling structure, the output dimension of the gating structure, the activation method of the gating structure, the output dimension of the second encoding structure for generating the guiding sequence, the number of expert modules of the routing module, and the number ratio of the dynamic expert module and the static expert module.
[0012] In a second aspect, the present invention provides a text processing method, including: obtaining text data of a text task to be processed; inputting the text data into a text processing model to obtain an inference result of the text task; wherein, the text processing model is obtained by using the training method of the text processing model according to any one of the above first aspects.
[0013] In a third aspect, the present invention provides a training device for a text processing model. The text processing model includes: a plurality of expert modules, a routing module, and a base model layer, including: a sample acquisition module, configured to obtain a training sample set; wherein, the training sample set includes: a plurality of composite text data composed of text data of a single text task; the text data of the single text task at least includes: a text similarity recognition task, a text category recognition task, a text feature extraction task, a text answering task; a guiding sequence generation module, configured to input each composite text data in the training sample set into the expert module and the routing module to obtain a guiding sequence of each composite text data; an inference module, configured to input each composite text data and the guiding sequence of each composite text data into the base model layer to obtain an inference result of each composite text data; a loss function value determination module, configured to determine a loss function value based on the inference result of each composite text data and the annotation result of the composite text data; a model parameter update module, configured to update the model parameters of the expert module and the routing module based on the loss function value.
[0014] In a fourth aspect, the present invention provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor. The processor executes the computer-executable instructions to implement the steps of the method according to any one of the above first aspects or the second aspect.
[0015] In a fifth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the method according to any one of the above first aspects or the second aspect.
[0016] The present invention brings the following beneficial effects:
[0017] The training method, text processing method and device for the above-mentioned text processing model provided by the present invention. The text processing model includes: a plurality of expert modules, a routing module and a base model layer. The method includes: First, obtain a training sample set (including: a plurality of composite text data composed of text data of a single text task; the text data of a single text task at least includes: text similarity recognition task, text category recognition task, text feature extraction task, text answering task); Then, input each composite text data in the training sample set into the expert module and the routing module to obtain a guiding sequence for each composite text data; Next, input each composite text data and the guiding sequence of each composite text data into the base model layer to obtain an inference result for each composite text data; After that, determine the loss function value based on the inference result of each composite text data and the annotation result of the composite text data; Finally, update the model parameters of the expert module and the routing module based on the loss function value. In the above method, on the basis of the existing keynote model, a plurality of expert modules and a routing module are introduced. The expert modules can generate different guiding sequences, and the routing module can select the final guiding sequence from the guiding sequences generated by the expert modules, and add the finally selected guiding sequence to the input prefix of each layer of the base model for forward propagation. During the training process, the model parameters of the expert module and the routing module are updated according to the loss function value of the keynote model. The above method can train only the expert module and the routing module without affecting the base model, saving training resources; at the same time, using composite text data for model training provides an effective training approach for post-training of large-parameter models under low-resource conditions, and improves the ability of the base model in composite tasks.
[0018] Other features and advantages of the present invention will be described in the following specification, and in part, will be obvious from the specification, or can be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims and drawings.
[0019] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0021] Figure 1Flowchart of a method for training a text processing model provided by an embodiment of the present invention;
[0022] Figure 2 Schematic diagram of the network structure of a dynamic expert module provided by an embodiment of the present invention;
[0023] Figure 3 Schematic diagram of the network structure of a static expert module provided by an embodiment of the present invention;
[0024] Figure 4 Flowchart of another method for training a text processing model provided by an embodiment of the present invention;
[0025] Figure 5 Flowchart of a text processing method provided by an embodiment of the present invention;
[0026] Figure 6 Schematic diagram of the composition structure of a training device for a text processing model provided by an embodiment of the present invention;
[0027] Figure 7 Schematic diagram of the composition structure of a text processing device provided by an embodiment of the present invention;
[0028] Figure 8 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0030] Currently, due to the large number of parameters of the model itself, if more domain knowledge needs to be incorporated into the model while ensuring that the learned knowledge is not lost, it is often necessary to find more general knowledge data than professional knowledge for joint training. At the same time, limited by training resources, it is often impossible to perform post-training on all parameters of the model.
[0031] Based on this, a method for training a text processing model, a text processing method, and a device provided by an embodiment of the present invention can improve the ability of the base model in composite tasks and reduce the impact of model training on the base model.
[0032] To facilitate the understanding of this embodiment, first, a detailed introduction is given to a method for training a text processing model disclosed in the embodiments of the present invention. The text processing model includes: multiple expert modules, a routing module, and a base model layer. The text processing can be applied to scenarios such as medical care and finance. This method can be executed by an electronic device, such as a smartphone, a computer, a tablet computer, etc. Refer to Figure 1 The flowchart of a method for training a text processing model as shown, which shows that this method mainly includes the following steps S101 to step S105:
[0033] Step S101: Obtain a training sample set.
[0034] In one implementation, the training sample set includes composite text data composed of text data of multiple types of single text tasks. The text data of single text tasks includes at least: text similarity recognition task, text category recognition task, text feature extraction task, text answering task. For example: For the text similarity task, the corresponding text data can be "Please identify the similarity between [sentence A] and [sentence B]". For the text category recognition task, the corresponding text data can be "Please determine the category to which the disease described in [sentence A] belongs". Among them, [sentence A] and [sentence B] can be replaced with specific sentences. Then, in this application, the text data of the text similarity task and the text category recognition task can be combined and used as training samples.
[0035] Step S102: Input each composite text data in the training sample set into the expert module and the routing module to obtain the guiding sequence of each composite text data.
[0036] In one implementation, when each composite text data in the training sample set is input into the expert module and the routing module, the expert module can generate different guiding sequences according to different types of text data; the routing module can select a suitable guiding sequence from the guiding sequences generated by the expert module according to the input different text data as the guiding sequence of each composite text data.
[0037] Step S103: Input each composite text data and the guiding sequence of each composite text data into the base model layer to obtain the inference result of each composite text data.
[0038] In one implementation, the guiding sequence generated by the expert module and the routing module is input into the base model layer, that is, the generated guiding sequence is added to each layer of the base model and participates in the forward propagation process of the model together, so as to obtain the inference result output by the base model. Among them, the base model can be an open-source base model such as Llama or Qwen. All model parameters of the base model are frozen and no model parameter updates are performed during the model training process.
[0039] Step S104: Determine the loss function value based on the inference result of each composite text data and the annotation result of the composite text data.
[0040] Step S105: Update the model parameters of the expert module and the routing module based on the loss function value.
[0041] In one implementation, according to the inference result output by the base model and the annotation result of the composite text data in the training sample set in advance, calculate the loss function value of the base model, and update the model parameters of the album module and the routing module according to the loss function value until the loss function value is minimized, and obtain the trained text processing module.
[0042] The training method of the above text processing model provided by the present invention introduces multiple expert modules and a routing module on the basis of the existing keynote model. The expert module can generate different guiding sequences, and the routing module can select the final guiding sequence from the guiding sequences generated by the expert module, and add the finally selected guiding sequence to the input prefix of each layer of the base model for forward propagation. During the training process, according to the loss function value of the keynote model, update the model parameters of the expert module and the routing module. The above method can train only the expert module and the routing module on the basis of ensuring that the base model is not affected, saving training resources; at the same time, using composite text data for model training provides an effective training method for post-training of large-parameter models under low-resource conditions, and improves the ability of the base model in composite tasks.
[0043] In one implementation, the expert module includes: a dynamic expert module with a multi-layer network structure and a static expert module with a one-layer network structure. See Figure 2 As shown, the dynamic expert module includes: a first encoding structure, an upsampling structure, a gating structure, and a downsampling structure.
[0044] In specific implementation, the first encoding structure (i.e., the Embedding structure) is used to receive input data (i.e., composite text data) and generate a guiding sequence with a fixed length.
[0045] The upsampling structure is a one-layer fully connected neural network, which is used to amplify the dimension of the output of the Embedding structure (i.e., the guiding sequence with a fixed length) through the fully connected neural network.
[0046] The gating structure is a one-layer fully connected neural network and a one-layer activation layer. The gating structure amplifies the dimension of the output of the Embedding structure through the fully connected neural network to obtain an output result, and then activates the output result through the activation layer to obtain a gated output.
[0047] The downsampling structure is a single-layer fully-connected neural network. The downsampling structure performs matrix multiplication on the output of the gating structure and the output of the upsampling structure, and then passes the result of the matrix multiplication through the fully-connected neural network to obtain the final result, which is the final output of the dynamic expert module, that is, the guiding sequence generated by the dynamic expert module.
[0048] See Figure 3 As shown, the static expert module includes: a second encoding structure, which is a single-layer fully-connected neural network. In a specific implementation, the static expert module only includes an Embedding structure (i.e., the second encoding structure) that generates a fixed-length guiding sequence, and this Embedding structure receives input data to generate a fixed-length guiding sequence.
[0049] The routing module is a single-layer neural network structure, which is used to receive input data and generate the weight coefficients of all expert modules.
[0050] Furthermore, before obtaining the training sample set, the above method further includes: obtaining the hyperparameters of the expert module and the routing module; where the hyperparameters at least include: the output dimension of the guiding sequence generated by the first encoding structure, the output dimension of the upsampling structure, the output dimension of the downsampling structure, the output dimension of the gating structure, the activation method of the gating structure, the output dimension of the guiding sequence generated by the second encoding structure, the number of expert modules of the routing module, and the number ratio of the dynamic expert module and the static expert module.
[0051] In a specific implementation, before model training, it is also necessary to pre-configure the hyperparameters of the expert module and the routing module. The hyperparameters affect the number of trainable parameters of the expert module and the routing module and their final performance on specific tasks, including:
[0052] Hyperparameters of the dynamic expert module: the output dimension of the guiding sequence generated by the first encoding structure (Embedding structure), the output dimension of the downsampling structure, and the changes in these two output dimensions can be controlled by the same parameter;
[0053] The output dimension of the upsampling structure, the output dimension of the gating structure, and the changes in these two output dimensions can be controlled by the same parameter;
[0054] The activation method of the gating structure. In the embodiments of the present invention, the activation method can be SiLU, or other activation methods, which are not limited herein.
[0055] Hyperparameters of the static expert module: the length of the output guiding sequence, that is, the output dimension of the guiding sequence generated by the second encoding structure, and this hyperparameter controls the number of its trainable parameters. It should be noted that in the embodiments of the present invention, the hyperparameters of the output dimensions of the Embedding structures of the dynamic expert module and the static expert module can be controlled by the same hyperparameter.
[0056] Hyperparameters of the routing module: the number of expert modules and the ratio of the number of dynamic expert modules to the number of static expert modules. Specifically, the total number of expert modules is controlled by a hyperparameter, which directly acts on the output dimension of the fully connected neural network layer; the number of expert modules includes the sum of the number of dynamic expert modules and the number of static expert modules, and the ratio of the number of experts in the two modules can be set separately, and this parameter is also a controllable hyperparameter.
[0057] Based on this, for the aforementioned step S102, that is, when each composite text data in the training sample set is input into the expert module and the routing module to obtain the guiding sequence of each composite text data, the following methods can be adopted, including but not limited to:
[0058] First, for each composite text data in the training sample set, the composite text data is input into the expert module to obtain the guiding sequence output by each expert module.
[0059] In specific implementation, inputting the composite text data into the expert module to obtain the guiding sequence output by each expert module includes:
[0060] (1) Input the composite text data into the dynamic expert module to obtain the first guiding sequence.
[0061] In practical applications, first encode the composite text data based on the first encoding structure to obtain an initial guiding sequence of a preset length; then amplify the dimension of the initial guiding sequence based on the upsampling structure to obtain an upsampling result; then amplify the dimension of the initial guiding sequence based on the fully connected neural network of the gating structure to obtain an output result, and activate the output result based on the activation layer to obtain a gating output result; finally, perform matrix multiplication on the upsampling result and the gating output result based on the downsampling structure to obtain a matrix multiplication result, and pass the matrix multiplication result through the fully connected neural network of the downsampling structure to obtain the first guiding sequence output by the dynamic expert module.
[0062] Specifically, input the composite text data into the dynamic expert module, and the first encoding structure generates an initial guiding sequence of a preset length (i.e., a fixed length, the output dimension set by setting the hyperparameter); then the upsampling structure amplifies the dimension of the initial guiding sequence output by the first encoding structure through the fully connected neural network and outputs an upsampling result; then the gating structure amplifies the dimension of the initial guiding sequence output by the first encoding structure through the fully connected neural network to obtain an output result, and then passes the output result through the activation layer to obtain a gating output result; finally, the downsampling structure performs matrix multiplication on the gating output result of the gating structure and the upsampling result of the upsampling structure, and passes the matrix multiplication result through the fully connected neural network to obtain the final result, that is, the first guiding sequence output by the dynamic expert module.
[0063] (2) Input the composite text data into the static expert module, and generate a second guiding sequence of a preset length based on the fully connected neural network of the second coding structure.
[0064] In specific implementation, input the composite text data into the static expert module, and the second coding structure receives the composite text data to generate a second guiding sequence of a fixed length.
[0065] Then, each composite text data in the training sample set is input into the routing module to obtain the weight coefficients of each expert module, and the weight coefficients are sorted in descending order. According to the sorting result, select the weight coefficients of a preset number of expert modules as the output result of the routing module.
[0066] In specific implementation, an expert selection method is set in the routing module: input each composite text data in the training sample set into the routing module, and the fully connected neural network calculates the weight coefficients of each expert module, sorts the weight coefficients from large to small, and then selects K (i.e., the preset number) of expert modules as the expert modules selected by the current routing module as the output of the routing module. The value of K is a controllable hyperparameter. The output result of the routing module includes: the numbers of a preset number of expert modules and the corresponding weight coefficients.
[0067] Finally, based on the numbers of a preset number of expert modules output by the routing module and the corresponding weight coefficients, perform a weighted calculation on the weight coefficients and the guiding sequences output by the corresponding expert modules to obtain the guiding sequence of the composite text data.
[0068] In specific implementation, perform a weighted summation on the weight coefficients of the K expert modules selected by the routing module and the guiding sequences output by the corresponding expert modules to obtain the final guiding sequence.
[0069] The finally obtained guiding sequence and the composite text data are input into the base model layer for forward calculation, and then the expert module and the routing module are updated backward until a better text processing model of the base model for complex tasks is obtained after superimposing these two modules.
[0070] The above-mentioned training method of the present processing model provided by the embodiments of the present invention can train only the expert module and the routing module without affecting the base model. During the training process, forward calculation is performed by generating a guiding sequence and input data together, and only the parameters of the expert module and the routing module are updated in the reverse process. In the inference stage, the expert module and the routing module are combined into the base model to improve the performance of the model on composite tasks. The above method uses the composite text data of multiple tasks in the professional field to train the model parameters of the introduced expert module and routing module, providing an effective training approach for post-training of large-parameter models under low-resource conditions, improving the ability of the base model on composite tasks, and also reducing the impact on the base model.
[0071] For ease of understanding, the embodiments of the present invention also provide a schematic diagram of the training of a specific text processing model, as shown in Figure 4 shown, including the following processes:
[0072] (1) Obtain a sample set, which is composite text sample data of multiple task types, that is, input data.
[0073] (2) Based on the sample set, pass through the expert module. The expert module will generate a guiding sequence according to the sample data. The expert module is divided into two types. One type is the dynamic expert module, which contains an Embedding layer and multiple neural network layers, and these network layers are all learnable. Its main function is to differentially generate a guiding sequence according to the input data. The static expert module only contains an Embedding structure and no other network layers. Its main function is residual connection to prevent the expert module from overly affecting the output of the base model. The number and proportion of the static expert module and the dynamic expert module can be manually controlled.
[0074] (3) Based on the sample set, pass through the routing module. The routing module is a single-layer network structure. Input the sample data through the routing module, output the weight words on each expert module, sort the weight words from largest to smallest, and select the weight coefficients of the top K expert modules for output. The routing module identifies the input data and selects K expert modules for the input data to perform guiding sequence fusion, achieving the purpose of outputting different guiding sequences for different types of tasks.
[0075] (4) Perform weighted summation through the guiding sequences of the generated K expert modules and the weight coefficients output by the routing module. The finally obtained guiding sequence is the final guiding sequence of the input data; splice the guiding sequence in front of the input data, pass it into the base model for forward calculation, and then update the model parameters of the expert module and the routing module in the reverse direction until the trained expert module and routing module are obtained. The trained expert module and routing module can be manually fused into the base model to improve the model performance of the base model on composite tasks.
[0076] The training method of the text processing model provided by the embodiment of the present invention has the following advantages:
[0077] First, by introducing additional training parameters and not training the base model, it ensures that the number of training parameters is smaller, providing a feasible solution for the post-training of the model in low-resource scenarios;
[0078] Second, using the decomposable expert module and routing module for parameter update without updating the parameters of the base model, it ensures the minimum interference to the base model during the training process;
[0079] Third, the expert module is split into a static expert module and a dynamic expert module, and different expert modules provide personalized support for different tasks, providing support for the training of complex tasks;
[0080] Fourth, the decomposability of the expert module and the routing module on the base model makes the model more flexible to use.
[0081] The embodiment of the present invention also provides a text processing method. Refer to Figure 5 the flowchart of a text processing method shown, which shows that the method mainly includes the following steps S501 to step S502:
[0082] Step S501: Obtain the text data of the text task to be processed.
[0083] In one implementation, the text data can be a question to be answered input by the user.
[0084] Step S502: Input the text data into the text processing model to obtain the inference result of the text task.
[0085] Among them, the text processing model is obtained by using the training method of the text processing model provided in the foregoing embodiment.
[0086] For the training method of the text processing model provided in the foregoing embodiment, the embodiment of the present invention also provides a training device for the text processing model. The text processing model includes: a plurality of expert modules, a routing module, and a base model layer. Refer to
[0087] Figure 6 Schematic diagram of the composition structure of a training device for a text processing model, showing that the device mainly includes the following parts:
[0088] A sample acquisition module 601, configured to acquire a training sample set; wherein, the training sample set includes: a plurality of composite text data formed by combining text data of a single text task; the text data of the single text task includes at least: a text similarity recognition task, a text category recognition task, a text feature extraction task, and a text answering task.
[0089] A guiding sequence generation module 602, configured to input each composite text data in the training sample set into an expert module and a routing module to obtain a guiding sequence for each composite text data.
[0090] An inference module 603, configured to input each composite text data and the guiding sequence of each composite text data into a base model layer to obtain an inference result for each composite text data.
[0091] A loss function value determination module 604, configured to determine a loss function value based on the inference result of each composite text data and the annotation result of the composite text data.
[0092] A model parameter update module 605, configured to update the model parameters of the expert module and the routing module based on the loss function value.
[0093] The above-mentioned training device for the text processing model provided by the present invention, on the basis of the existing keynote model, introduces a plurality of expert modules and a routing module. The expert module can generate different guiding sequences, and the routing module can select the final guiding sequence from the guiding sequences generated by the expert module, and add the finally selected guiding sequence to the input prefix of each layer of the base model for forward propagation. During the training process, according to the loss function value of the keynote model, the model parameters of the expert module and the routing module are updated. The above-mentioned device can train only the expert module and the routing module without affecting the base model, saving training resources; at the same time, using composite text data for model training provides an effective training approach for post-training of large-parameter models under low-resource conditions, and improves the ability of the base model in composite tasks.
[0094] For the text processing method provided in the foregoing embodiment, the embodiment of the present invention further provides a text processing device. Refer to Figure 7 Schematic diagram of the composition structure of a text processing device, showing that the device mainly includes the following parts:
[0095] A text data acquisition module 701, configured to acquire text data of a text task to be processed.
[0096] A text inference module 702 is configured to input text data into a text processing model to obtain an inference result of a text task. The text processing model is obtained by using the text processing model training method provided in the foregoing embodiments.
[0097] The text processing device provided in the embodiments of the present invention uses a pre-trained text processing model to perform inference to obtain an inference result. The text processing module includes multiple expert modules, a routing module, and a base model layer. The multiple expert modules are responsible for generating different guiding sequences. The single routing selection module is responsible for selecting different guiding sequences according to the input, and weighted summing the finally selected multiple guiding sequences to obtain a final guiding sequence, which is added to the input prefix of each layer of the base model for forward propagation to obtain an inference result of the text task, thereby improving the ability of the base model in composite tasks.
[0098] It should be noted that the device provided in the embodiments of the present invention has the same implementation principle and the same technical effects as those in the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding contents in the foregoing method embodiments.
[0099] The embodiments of the present invention further provide an electronic device. Specifically, the electronic device includes a processor and a storage device. A computer program is stored on the storage device, and the computer program, when run by the processor, executes the method according to any one of the above embodiments.
[0100] Figure 8 FIG. 13 is a schematic structural diagram of an electronic device provided in an embodiment of the present invention. The electronic device 100 includes: a processor 80, a memory 81, a bus 82, and a communication interface 83. The processor 80, the communication interface 83, and the memory 81 are connected through the bus 82. The processor 80 is configured to execute an executable module stored in the memory 81, such as a computer program.
[0101] Among them, the memory 81 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 83 (which may be wired or wireless), a communication connection between the system network element and at least one other network element is realized, and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used.
[0102] The bus 82 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 8 only a bidirectional arrow is used in FIG. 13, but it does not mean that there is only one bus or one type of bus.
[0103] Among them, the memory 81 is used to store a program. After receiving an execution instruction, the processor 80 executes the program. The method executed by the device defined by the flow process disclosed in any embodiment of the foregoing embodiments of the present invention can be applied to or implemented by the processor 80.
[0104] The processor 80 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 80 or by instructions in the form of software. The above-mentioned processor 80 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed by a hardware decoding processor or completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 81, and the processor 80 reads the information in the memory 81 and combines its hardware to complete the steps of the above method.
[0105] The computer program product of the readable storage medium provided by the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the foregoing method embodiments. For specific implementation, reference can be made to the foregoing method embodiments and will not be elaborated here.
[0106] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0107] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A training method for a text processing model, characterized in that, The text processing model includes: a plurality of expert modules, a routing module, and a base model layer. Among them, the expert modules include: a dynamic expert module with a multi-layer network structure and a static expert module with a single-layer network structure; the dynamic expert module includes: a first encoding structure, an upsampling structure, a gating structure, and a downsampling structure; the upsampling structure is a single-layer fully connected neural network, and the gating structure is a single-layer fully connected neural network and an activation layer; the downsampling structure is a single-layer fully connected neural network; the static expert module includes: a second encoding structure, and the second encoding structure is a single-layer fully connected neural network; the routing module is a single-layer neural network structure; the training method includes: Obtain a training sample set; among them, the training sample set includes: a plurality of composite text data composed of text data of a single text task; the text data of the single text task at least includes: a text similarity recognition task, a text category recognition task, a text feature extraction task, and a text answering task; Input each composite text data in the training sample set into the expert module and the routing module to obtain a guiding sequence for each composite text data; Input each composite text data and the guiding sequence of each composite text data into the base model layer to obtain an inference result for each composite text data; Based on the inference result of each composite text data and the annotation result of the composite text data, determine the loss function value; Based on the loss function value, update the model parameters of the expert module and the routing module; Input each composite text data in the training sample set into the expert module and the routing module to obtain a guiding sequence for each composite text data, including: for each composite text data in the training sample set, input the composite text data into the expert module to obtain a guiding sequence output by each expert module; input each composite text data in the training sample set into the routing module to obtain the weight coefficients of each expert module, and sort the weight coefficients in descending order, and select the weight coefficients of a preset number of expert modules as the output result of the routing module according to the sorting result; among them, the output result of the routing module includes: the numbers of the preset number of expert modules and the corresponding weight coefficients; based on the numbers of the preset number of expert modules output by the routing module and the corresponding weight coefficients, perform a weighted calculation on the weight coefficients and the guiding sequences output by the corresponding expert modules to obtain the guiding sequence of the composite text data.
2. The method according to claim 1, wherein Input the composite text data into the expert module to obtain a guiding sequence output by each expert module, including: Input the composite text data into the dynamic expert module to obtain a first guiding sequence; Input the composite text data into the static expert module, and generate a second guiding sequence with a preset length based on the fully connected neural network of the second encoding structure.
3. The method according to claim 2, characterized in that, Input the composite text data into the dynamic expert module to obtain a first guiding sequence, including: Encode the composite text data based on the first encoding structure to obtain an initial guiding sequence of a preset length; Perform dimensionality amplification on the initial guiding sequence based on the upsampling structure to obtain an upsampling result; Perform dimensionality amplification on the initial guiding sequence based on the fully connected neural network of the gating structure to obtain an output result, and activate the output result based on the activation layer to obtain a gated output result; Perform matrix multiplication on the upsampling result and the gated output result based on the downsampling structure to obtain a matrix multiplication result, and pass the matrix multiplication result through the fully connected neural network of the downsampling structure to obtain the first guiding sequence output by the dynamic expert module.
4. The method according to claim 1, wherein Before obtaining the training sample set, it further includes: Obtain the hyperparameters of the expert module and the routing module; wherein, the hyperparameters at least include: the output dimension of the guiding sequence generated by the first encoding structure, the output dimension of the upsampling structure, the output dimension of the downsampling structure, the output dimension of the gating structure, the activation method of the gating structure, the output dimension of the guiding sequence generated by the second encoding structure, the number of expert modules of the routing module, and the number ratio of the dynamic expert module and the static expert module.
5. A text processing method, characterized in that, It includes: Obtain the text data of the text task to be processed; Input the text data into the text processing model to obtain the inference result of the text task; wherein, the text processing model is obtained by using the training method of the text processing model according to any one of claims 1 to 4.
6. A training device for a text processing model, characterized in that, The text processing model includes: a plurality of expert modules, a routing module, and a base model layer, wherein the expert module includes: a dynamic expert module with a multi-layer network structure and a static expert module with a single-layer network structure; the dynamic expert module includes: a first encoding structure, an upsampling structure, a gating structure, and a downsampling structure; the upsampling structure is a single-layer fully connected neural network, the gating structure is a single-layer fully connected neural network and a single-layer activation layer; the downsampling structure is a single-layer fully connected neural network; the static expert module includes: a second encoding structure, and the second encoding structure is a single-layer fully connected neural network; the routing module is a single-layer neural network structure; the training device includes: A sample acquisition module for acquiring a training sample set; wherein, the training sample set includes: a plurality of composite text data composed of the text data of a single text task; the text data of the single text task at least includes: a text similarity recognition task, a text category recognition task, a text feature extraction task, a text answering task; A guiding sequence generation module for inputting each composite text data in the training sample set into the expert module and the routing module to obtain the guiding sequence of each composite text data; An inference module for inputting each composite text data and the guiding sequence of each composite text data into the base model layer to obtain the inference result of each composite text data. A loss function value determination module, configured to determine a loss function value based on the inference result of each of the composite text data and the annotation result of the composite text data; A model parameter update module, configured to update the model parameters of the expert module and the routing module based on the loss function value; The guiding sequence generation module is specifically configured to: for each composite text data in the training sample set, input the composite text data into the expert module to obtain a guiding sequence output by each expert module; input each composite text data in the training sample set into the routing module to obtain the weight coefficients of each expert module, sort the weight coefficients in descending order, and select the weight coefficients of a preset number of expert modules as the output result of the routing module according to the sorting result; wherein, the output result of the routing module includes: the numbers of the preset number of expert modules and the corresponding weight coefficients; based on the numbers of the preset number of expert modules output by the routing module and the corresponding weight coefficients, perform a weighted calculation on the weight coefficients and the guiding sequences output by the corresponding expert modules to obtain the guiding sequence of the composite text data.
7. An electronic device, characterized in that, It includes a processor and a memory, the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of the method according to any one of claims 1 to 4 or claim 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it executes the steps of the method according to any one of claims 1 to 4 or claim 5.
Citation Information
Patent Citations
Vehicle manufacturing quality defect detection method and device, electronic equipment and storage medium
CN118277838A
Deep learning text generation for upgrading machine learning systems
US20240220576A1