Text processing model training method and device and text processing method and device
By introducing expert modules and routing modules on the existing pedestal model, using composite text data for training, the problem of large number of model parameters and limited training resources is solved, and the model's performance ability on composite tasks is improved and training resources is saved.
Patent Information
- Application Number
- CN202510525695.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-25
AI Technical Summary
In the prior art, the model parameters are large, making it difficult to impart more domain knowledge in the model while ensuring that the knowledge learned is not lost. Due to the training resources, it is impossible to post-train the model after all parameters.
A training method for text processing models is adopted, multiple expert modules and a routing module are introduced, and model training is performed by compounding text data, and the model parameters of only expert modules and routing modules are updated to avoid the impact on the base model.
The ability of the base model in composite tasks is improved, training resources is saved, and an effective way to train large-parameter models under low resource conditions is provided.
Smart Images

Figure CN120067697A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a training method for a text processing model, a text processing method and a device. Background Art
[0002] With the development of artificial intelligence and big data technologies, pre-training endows the model with more basic knowledge, and these pre-trained models can be effectively applied to various tasks, showing very excellent effects. In the prior art, after the model is pre-trained, it obtains a large amount of knowledge. In practical applications, the performance of the model in specific tasks can usually be stimulated by adjusting the input guiding sequence. At the same time, in the current model application process, the number of parameters of the model itself is getting larger and larger, and the data and knowledge solidified into the model through pre-training are also increasing. If it is necessary to stimulate the performance of the model in a specific task, it is often necessary to adjust a large number of guiding sequences in combination with the business.
[0003] However, due to the large number of parameters of the model itself, if it is necessary to endow the model with more domain knowledge while ensuring that the learned knowledge is not lost, it is often necessary to find more general knowledge data than professional knowledge for training together. At the same time, due to the influence of training resources, it is often impossible to perform post-training on all parameters of the model. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a training method for a text processing model, a text processing method and a device, so as to improve the ability of the base model in composite tasks and reduce the impact of model training on the base model.
[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows: In the first aspect, the present invention provides a training method for a text processing model. The text processing model includes: a plurality of expert modules, a routing module and a base model layer, including: obtaining a training sample set; wherein, the training sample set includes: a plurality of composite text data composed of text data of a single text task; the text data of a single text task at least includes: a text similarity recognition task, a text category recognition task, a text feature extraction task, a text answering task; inputting each composite text data in the training sample set into the expert module and the routing module to obtain a guiding sequence of each composite text data; inputting each composite text data and the guiding sequence of each composite text data into the base model layer to obtain an inference result of each composite text data; determining a loss function value based on the inference result of each composite text data and the annotation result of the composite text data; updating the model parameters of the expert module and the routing module based on the loss function value.
[0006] Optionally, input each composite text data in the training sample set into the expert module and the routing module to obtain the guiding sequence of each composite text data, including: for each composite text data in the training sample set, input the composite text data into the expert module to obtain the guiding sequence output by each expert module; input each composite text data in the training sample set into the routing module to obtain the weight coefficients of each expert module, sort the weight coefficients in descending order, and select the weight coefficients of a preset number of expert modules as the output result of the routing module according to the sorting result; wherein, the output result of the routing module includes: the numbers of a preset number of expert modules and the corresponding weight coefficients; based on the numbers of a preset number of expert modules output by the routing module and the corresponding weight coefficients, perform weighted calculation on the weight coefficients and the guiding sequences output by the corresponding expert modules to obtain the guiding sequence of the composite text data.
[0007] Optionally, the expert module includes: a dynamic expert module with a multi-layer network structure and a static expert module with a single-layer network structure; the dynamic expert module includes: a first encoding structure, an upsampling structure, a gating structure, and a downsampling structure; the upsampling structure is a single-layer fully connected neural network, and the gating structure is a single-layer fully connected neural network and a single activation layer; the downsampling structure is a single-layer fully connected neural network; the static expert module includes: a second encoding structure, and the second encoding structure is a single-layer fully connected neural network; the routing module is a single-layer neural network structure.
[0008] Optionally, inputting the composite text data into the expert module to obtain the guiding sequence output by each expert module includes: inputting the composite text data into the dynamic expert module to obtain a first guiding sequence; inputting the composite text data into the static expert module, and generating a second guiding sequence with a preset length based on the fully connected neural network of the second encoding structure.
[0009] Optionally, inputting the composite text data into the dynamic expert module to obtain a first guiding sequence includes: encoding the composite text data based on the first encoding structure to obtain an initial guiding sequence with a preset length; performing dimension expansion on the initial guiding sequence based on the upsampling structure to obtain an upsampling result; performing dimension expansion on the initial guiding sequence based on the fully connected neural network of the gating structure to obtain an output result, and activating the output result based on the activation layer to obtain a gating output result; performing matrix multiplication on the upsampling result and the gating output result based on the downsampling structure to obtain a matrix multiplication result, and passing the matrix multiplication result through the fully connected neural network of the downsampling structure to obtain the first guiding sequence output by the dynamic expert module.
[0010] Optionally, before obtaining the training sample set, it further includes: obtaining the hyperparameters of the expert module and the routing module; wherein, the hyperparameters at least include: the output dimension of the first encoding structure for generating the guiding sequence, the output dimension of the upsampling structure, the output dimension of the downsampling structure, the output dimension of the gating structure, the activation method of the gating structure, the output dimension of the second encoding structure for generating the guiding sequence, the number of expert modules of the routing module, and the number ratio of the dynamic expert module and the static expert module.
[0011] In a second aspect, the present invention provides a text processing method, including: obtaining text data of a text task to be processed; inputting the text data into a text processing model to obtain an inference result of the text task; wherein, the text processing model is obtained by using the training method of the text processing model according to any one of the first aspects described above.
[0012] In a third aspect, the present invention provides a training device for a text processing model. The text processing model includes: a plurality of expert modules, a routing module, and a base model layer, including: a sample acquisition module, configured to obtain a training sample set; wherein, the training sample set includes: a plurality of composite text data composed of text data of a single text task; the text data of the single text task at least includes: a text similarity recognition task, a text category recognition task, a text feature extraction task, and a text answering task; a guiding sequence generation module, configured to input each composite text data in the training sample set into the expert module and the routing module to obtain a guiding sequence of each composite text data; an inference module, configured to input each composite text data and the guiding sequence of each composite text data into the base model layer to obtain an inference result of each composite text data; a loss function value determination module, configured to determine a loss function value based on the inference result of each composite text data and the annotation result of the composite text data; a model parameter update module, configured to update the model parameters of the expert module and the routing module based on the loss function value.
[0013] In a fourth aspect, the present invention provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of the method according to any one of the first aspects or the second aspect described above.
[0014] In a fifth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the method according to any one of the first aspects or the second aspect described above.
[0015] The present invention brings the following beneficial effects: The training method, text processing method and device for the above-mentioned text processing model provided by the present invention. The text processing model includes: a plurality of expert modules, a routing module and a base model layer. The method includes: First, obtain a training sample set (including: a plurality of composite text data composed of text data of a single text task; the text data of the single text task at least includes: text similarity recognition task, text category recognition task, text feature extraction task, text answering task); Then, input each composite text data in the training sample set into the expert module and the routing module to obtain a guiding sequence for each composite text data; Next, input each composite text data and the guiding sequence of each composite text data into the base model layer to obtain an inference result for each composite text data; After that, determine the loss function value based on the inference result of each composite text data and the annotation result of the composite text data; Finally, update the model parameters of the expert module and the routing module based on the loss function value. In the above method, on the basis of the existing keynote model, a plurality of expert modules and a routing module are introduced. The expert modules can generate different guiding sequences, and the routing module can select the final guiding sequence from the guiding sequences generated by the expert modules, and add the finally selected guiding sequence to the input prefix of each layer of the base model for forward propagation. During the training process, the model parameters of the expert module and the routing module are updated according to the loss function value of the keynote model. The above method can save training resources by only training the expert module and the routing module without affecting the base model; at the same time, using composite text data for model training provides an effective training approach for post-training of large-parameter models under low-resource conditions, and improves the ability of the base model in composite tasks.
[0016] Other features and advantages of the present invention will be described in the following specification, and in part will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures particularly pointed out in the specification, claims and drawings.
[0017] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1Flowchart of a training method for a text processing model provided by an embodiment of the present invention; Figure 2 Schematic diagram of the network structure of a dynamic expert module provided by an embodiment of the present invention; Figure 3 Schematic diagram of the network structure of a static expert module provided by an embodiment of the present invention; Figure 4 Flowchart of another training method for a text processing model provided by an embodiment of the present invention; Figure 5 Flowchart of a text processing method provided by an embodiment of the present invention; Figure 6 Schematic diagram of the composition structure of a training device for a text processing model provided by an embodiment of the present invention; Figure 7 Schematic diagram of the composition structure of a text processing device provided by an embodiment of the present invention; Figure 8 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] Currently, due to the large number of parameters of the model itself, if more domain knowledge needs to be incorporated into the model while ensuring that the already learned knowledge is not lost, it is often necessary to find more general knowledge data than professional knowledge for joint training. At the same time, limited by training resources, it is often impossible to perform post-training on all parameters of the model.
[0022] Based on this, a training method, a text processing method, and a device for a text processing model provided by an embodiment of the present invention can improve the ability of the base model in composite tasks and reduce the impact of model training on the base model.
[0023] For ease of understanding of this embodiment, a training method for a text processing model disclosed in an embodiment of the present invention will be introduced in detail first. The text processing model includes: a plurality of expert modules, a routing module, and a base model layer. The text processing can be applied to scenarios such as medical care and finance. This method can be executed by an electronic device, such as a smart phone, a computer, a tablet computer, etc. Refer to Figure 1Flowchart of a training method for a text processing model, showing that the method mainly includes the following steps S101 to S105: Step S101: Obtain a training sample set.
[0024] In one implementation, the training sample set includes composite text data composed of text data of multiple types of single text tasks. The text data of single text tasks at least includes: text similarity recognition task, text category recognition task, text feature extraction task, text answering task. For example: For the text similarity task, the corresponding text data can be "Please identify the similarity between [sentence A] and [sentence B]". For the text category recognition task, the corresponding text data can be "Please determine the category to which the disease described in [sentence A] belongs". Among them, [sentence A] and [sentence B] can be replaced with specific sentences. Then, in this application, the text data of the text similarity task and the text category recognition task can be combined and used as training samples.
[0025] Step S102: Input each composite text data in the training sample set into the expert module and the routing module to obtain the guiding sequence of each composite text data.
[0026] In one implementation, input each composite text data in the training sample set into the expert module and the routing module. The expert module can generate different guiding sequences according to different types of text data; the routing module can select a suitable guiding sequence from the guiding sequences generated by the expert module according to the input different text data as the guiding sequence of each composite text data.
[0027] Step S103: Input each composite text data and the guiding sequence of each composite text data into the base model layer to obtain the inference result of each composite text data.
[0028] In one implementation, input the guiding sequence generated by the expert module and the routing module into the base model layer, that is, add the generated guiding sequence to each layer of the base model to jointly participate in the forward propagation process of the model, so as to obtain the inference result output by the base model. Among them, the base model can be an open-source base model such as Llama or Qwen. Freeze all model parameters of the base model and do not update any model parameters during the model training process.
[0029] Step S104: Determine the loss function value based on the inference result of each composite text data and the annotation result of the composite text data.
[0030] Step S105: Update the model parameters of the expert module and the routing module based on the loss function value.
[0031] In one embodiment, according to the inference result output by the base model and the annotation result of the composite text data in the training sample set in advance, the loss function value of the base model is calculated, and the model parameters of the album module and the routing module are updated according to the loss function value until the loss function value is minimized, so as to obtain a trained text processing module.
[0032] The training method of the above-mentioned text processing model provided by the present invention introduces a plurality of expert modules and a routing module on the basis of the existing keynote model. The expert modules can generate different guiding sequences, and the routing module can select the final guiding sequence from the guiding sequences generated by the expert modules, and add the finally selected guiding sequence to the input prefix of each layer of the base model for forward propagation. During the training process, according to the loss function value of the keynote model, the model parameters of the expert modules and the routing module are updated. The above method can train only the expert modules and the routing module on the basis of ensuring that the base model is not affected, saving training resources; at the same time, using composite text data for model training provides an effective training approach for post-training of large-parameter models under low-resource conditions, and improves the ability of the base model in composite tasks.
[0033] In one embodiment, the expert module includes: a dynamic expert module with a multi-layer network structure and a static expert module with a single-layer network structure. See Figure 2 As shown, the dynamic expert module includes: a first encoding structure, an upsampling structure, a gating structure, and a downsampling structure.
[0034] In specific implementation, the first encoding structure (i.e., the Embedding structure) is used to receive input data (i.e., composite text data) and generate a guiding sequence with a fixed length.
[0035] The upsampling structure is a single-layer fully connected neural network, which is used to amplify the dimension of the output of the Embedding structure (i.e., the guiding sequence with a fixed length) through the fully connected neural network.
[0036] The gating structure is a single-layer fully connected neural network and a single-layer activation layer. The gating structure amplifies the dimension of the output of the Embedding structure through the fully connected neural network to obtain an output result, and then activates the output result through the activation layer to obtain a gated output.
[0037] The downsampling structure is a single-layer fully connected neural network. The downsampling structure multiplies the output of the gating structure and the output of the upsampling structure in matrix form, and then passes the matrix multiplication result through the fully connected neural network to obtain a final result, which is the final output of the dynamic expert module, that is, the guiding sequence generated by the dynamic expert module.
[0038] See Figure 3As shown in the figure, the static expert module includes: a second encoding structure, which is a one-layer fully connected neural network. In a specific implementation, the static expert module only includes an Embedding structure (i.e., the second encoding structure) that generates a fixed-length guiding sequence. This Embedding structure receives input data and generates a fixed-length guiding sequence.
[0039] The routing module is a single-layer neural network structure, which is used to receive input data and generate the weight coefficients of all expert modules.
[0040] Furthermore, before obtaining the training sample set, the above method further includes: obtaining the hyperparameters of the expert module and the routing module; where the hyperparameters at least include: the output dimension of the guiding sequence generated by the first encoding structure, the output dimension of the upsampling structure, the output dimension of the downsampling structure, the output dimension of the gating structure, the activation method of the gating structure, the output dimension of the guiding sequence generated by the second encoding structure, the number of expert modules of the routing module, and the number ratio of the dynamic expert module and the static expert module.
[0041] In a specific implementation, before model training, it is also necessary to pre-configure the hyperparameters of the expert module and the routing module. The hyperparameters affect the number of trainable parameters of the expert module and the routing module and their final performance on specific tasks, including: Hyperparameters of the dynamic expert module: the output dimension of the guiding sequence generated by the first encoding structure (Embedding structure), the output dimension of the downsampling structure, and the changes in these two output dimensions can be controlled by the same parameter; The output dimension of the upsampling structure, the output dimension of the gating structure, and the changes in these two output dimensions can be controlled by the same parameter; The activation method of the gating structure. In the embodiments of the present invention, the activation method can be SiLU or other activation methods, which are not limited herein.
[0042] Hyperparameters of the static expert module: the length of the output guiding sequence, that is, the output dimension of the guiding sequence generated by the second encoding structure. This hyperparameter controls the number of its trainable parameters. It should be noted that in the embodiments of the present invention, the hyperparameter of the output dimension of the Embedding structure of the dynamic expert module and the static expert module can be controlled by the same hyperparameter.
[0043] Hyperparameters of the routing module: the number of expert modules and the number ratio of the dynamic expert module and the static expert module. Specifically, the total number of expert modules is controlled by one hyperparameter, which directly acts on the output dimension of the fully connected neural network layer; the number of expert modules includes the total number of the dynamic expert module and the static expert module, and the number ratio of the two modules can be set separately. This parameter is also a controllable hyperparameter.
[0044] Based on this, for the aforementioned step S102, that is, when each composite text data in the training sample set is input into the expert module and the routing module to obtain the guiding sequence of each composite text data, the following methods can be adopted, including but not limited to: First, for each composite text data in the training sample set, the composite text data is input into the expert module to obtain the guiding sequence output by each expert module.
[0045] In specific implementation, inputting the composite text data into the expert module to obtain the guiding sequence output by each expert module includes: (1) Input the composite text data into the dynamic expert module to obtain the first guiding sequence.
[0046] In practical applications, first encode the composite text data based on the first encoding structure to obtain an initial guiding sequence of a preset length; then amplify the dimension of the initial guiding sequence based on the upsampling structure to obtain an upsampling result; then amplify the dimension of the initial guiding sequence based on the fully connected neural network of the gating structure to obtain an output result, and activate the output result based on the activation layer to obtain a gating output result; finally, perform matrix multiplication on the upsampling result and the gating output result based on the downsampling structure to obtain a matrix multiplication result, and pass the matrix multiplication result through the fully connected neural network of the downsampling structure to obtain the first guiding sequence output by the dynamic expert module.
[0047] Specifically, input the composite text data into the dynamic expert module, and the first encoding structure generates an initial guiding sequence of a preset length (i.e., a fixed length, the output dimension set by setting hyperparameters); then the upsampling structure amplifies the dimension of the initial guiding sequence output by the first encoding structure through the fully connected neural network and outputs an upsampling result; then the gating structure amplifies the dimension of the initial guiding sequence output by the first encoding structure through the fully connected neural network to obtain an output result, and then passes the output result through the activation layer to obtain a gating output result; finally, the downsampling structure performs matrix multiplication on the gating output result of the gating structure and the upsampling result of the upsampling structure, and passes the matrix multiplication result through the fully connected neural network to obtain the final result, that is, the first guiding sequence output by the dynamic expert module.
[0048] (2) Input the composite text data into the static expert module, and generate a second guiding sequence of a preset length based on the fully connected neural network of the second encoding structure.
[0049] In specific implementation, input the composite text data into the static expert module, and the second encoding structure receives the composite text data and generates a second guiding sequence of a fixed length.
[0050] Then, each composite text data in the training sample set is input into the routing module to obtain the weight coefficients of each expert module, and the weight coefficients are sorted in descending order. According to the sorting result, the weight coefficients of a preset number of expert modules are selected as the output result of the routing module.
[0051] In specific implementation, an expert selection method is set in the routing module: each composite text data in the training sample set is input into the routing module, and the fully connected neural network calculates the weight coefficients of each expert module, sorts the weight coefficients from large to small, and then selects K (i.e., the preset number) of expert modules as the expert modules selected by the current routing module as the output of the routing module. The K value is a controllable hyperparameter. The output result of the routing module includes: the numbers of a preset number of expert modules and the corresponding weight coefficients.
[0052] Finally, based on the numbers of a preset number of expert modules output by the routing module and the corresponding weight coefficients, the weight coefficients are weighted with the guiding sequences output by the corresponding expert modules to obtain the guiding sequence of the composite text data.
[0053] In specific implementation, the weight coefficients of the K expert modules selected by the routing module and the guiding sequences output by the corresponding expert modules are weighted and summed to obtain the final guiding sequence.
[0054] The finally obtained guiding sequence and the composite text data are input into the base model layer for forward calculation, and then the expert module and the routing module are updated backward until a better text processing model of the base model for complex tasks is obtained after superimposing these two modules.
[0055] The above training method of the text processing model provided by the embodiments of the present invention can train only the expert module and the routing module without affecting the base model. During the training process, forward calculation is performed by generating the guiding sequence and the input data together, and only the parameters of the expert module and the routing module are updated in the reverse process. In the inference stage, the expert module and the routing module are combined into the base model to improve the performance of the model in composite tasks. The above method uses the composite text data of multiple tasks in the professional field to train the model parameters of the introduced expert module and routing module, provides an effective training approach for post-training of large-parameter models under low-resource conditions, improves the ability of the base model in composite tasks, and also reduces the impact on the base model.
[0056] For easy understanding, the embodiments of the present invention also provide a schematic diagram of the training of a specific text processing model. Refer to Figure 4 as shown, which includes the following processes: (1) Obtain a sample set, and the sample set is composite text sample data of multiple task types, that is, input data.
[0057] (2) Based on the sample set, through the expert module, which will generate a guiding sequence according to the sample data. The expert module is divided into two types. One type is the dynamic expert module, which contains an Embedding layer and multiple neural network layers, all of which are learnable. Its main function is to differentially generate a guiding sequence according to the input data. The static expert module only contains an Embedding structure and no other network layers. Its main function is residual connection to prevent the expert module from having an excessive impact on the output of the base model. The number and proportion of the static expert module and the dynamic expert module can be manually controlled.
[0058] (3) Based on the sample set, through the routing module, which is a single-layer network structure. By inputting the sample data through the routing module, the weighted words on each expert module are output. The weighted words are sorted from largest to smallest, and the weight coefficients of the top K expert modules are selected and output. The routing module identifies the input data and selects K expert modules for the input data to fuse the guiding sequences, so as to achieve the purpose of outputting different guiding sequences for different types of tasks.
[0059] (4) Perform weighted summation through the guiding sequences of the generated K expert modules and the weight coefficients output by the routing module. The finally obtained guiding sequence is the final guiding sequence of the input data. The guiding sequence is spliced in front of the input data and passed into the base model for forward calculation. Then, the model parameters of the expert module and the routing module are updated backward until the trained expert module and routing module are obtained. The trained expert module and routing module can be manually fused into the base model to improve the model performance of the base model in composite tasks.
[0060] The training method of the above text processing model provided by the embodiments of the present invention has the following advantages: First, by introducing additional training parameters and not training the base model, it is ensured that the number of training parameters is less, providing a feasible solution for the post-training of low-resource scenario models; Second, the parameter update is performed using the decomposable expert module and routing module, without updating the parameters of the base model, ensuring the minimum interference to the base model during the training process; Third, the expert module is split into a static expert module and a dynamic expert module. Different expert modules provide personalized support for different tasks, providing support for complex task training; Fourth, the decomposability of the expert module and the routing module on the base model makes the model more flexible to use.
[0061] The embodiments of the present invention also provide a text processing method. Refer to Figure 5 The flowchart of a text processing method shown, which shows that the method mainly includes the following steps S501 to step S502: Step S501: Obtain the text data of the text task to be processed.
[0062] In one implementation, the text data may be a question to be answered input by the user.
[0063] Step S502: Input the text data into the text processing model to obtain the inference result of the text task.
[0064] Among them, the text processing model is obtained by using the training method of the text processing model provided in the foregoing embodiments.
[0065] The above text processing method provided by the embodiments of the present invention uses a pre-trained text processing model to perform inference to obtain an inference result. The text processing module includes multiple expert modules, a routing module, and a base model layer. The multiple expert modules are responsible for generating different guiding sequences. The single routing selection module is responsible for selecting different guiding sequences according to the input, and weighted summing the finally selected multiple guiding sequences to obtain a final guiding sequence and adding it to the input prefix of each layer of the base model for forward propagation to obtain the inference result of the text task, thereby improving the ability of the base model in composite tasks.
[0066] For the training method of the text processing model provided in the foregoing embodiments, the embodiments of the present invention also provide a training device for the text processing model. The text processing model includes: multiple expert modules, a routing module, and a base model layer. Refer to Figure 6 The composition structure diagram of a training device for a text processing model shown, which shows that the device mainly includes the following parts: A sample acquisition module 601, configured to acquire a training sample set; wherein, the training sample set includes: multiple composite text data composed of text data of a single text task; the text data of the single text task at least includes: a text similarity recognition task, a text category recognition task, a text feature extraction task, a text answering task; A guiding sequence generation module 602, configured to input each composite text data in the training sample set into the expert module and the routing module to obtain the guiding sequence of each composite text data; An inference module 603, configured to input each composite text data and the guiding sequence of each composite text data into the base model layer to obtain the inference result of each composite text data; A loss function value determination module 604, configured to determine the loss function value based on the inference result of each composite text data and the annotation result of the composite text data; A model parameter update module 605, configured to update the model parameters of the expert module and the routing module based on the loss function value.
[0067] The training device for the above-mentioned text processing model provided by the present invention, on the basis of the existing keynote model, introduces a plurality of expert modules and a routing module. The expert modules can generate different guiding sequences, and the routing module can select the final guiding sequence from the guiding sequences generated by the expert modules, and add the finally selected guiding sequence to the input prefix of each layer of the base model for forward propagation. During the training process, according to the loss function value of the keynote model, the model parameters of the expert modules and the routing module are updated. The above device can train only the expert modules and the routing module on the premise of not affecting the base model, saving training resources; at the same time, using composite text data for model training provides an effective training approach for post-training large-parameter models under low-resource conditions, and improves the ability of the base model in composite tasks.
[0068] For the text processing method provided in the foregoing embodiment, the embodiment of the present invention also provides a text processing device. Refer to Figure 7 the schematic composition structure diagram of a text processing device shown in The text data acquisition module 701 is used to acquire the text data of the text task to be processed.
[0069] The text inference module 702 is used to input the text data into the text processing model to obtain the inference result of the text task; wherein, the text processing model is obtained by using the text processing model training method provided in the foregoing embodiment.
[0070] The above-mentioned text processing device provided by the embodiment of the present invention uses a pre-trained text processing model to obtain an inference result. The text processing module includes a plurality of expert modules, a routing module, and a base model layer. The multi-expert modules are responsible for generating different guiding sequences, and the single routing selection module is responsible for selecting different guiding sequences according to the input, and adding the finally selected multiple guiding sequences after weighted summation to the input prefix of each layer of the base model for forward propagation to obtain the inference result of the text task, thereby improving the ability of the base model in composite tasks.
[0071] It should be noted that the device provided by the embodiment of the present invention has the same implementation principle and the same technical effects as those of the foregoing method embodiment. For the sake of brief description, for the parts not mentioned in the device embodiment, reference may be made to the corresponding content in the foregoing method embodiment.
[0072] The embodiment of the present invention also provides an electronic device. Specifically, the electronic device includes a processor and a storage device; a computer program is stored on the storage device, and the computer program executes the method according to any one of the above embodiments when being run by the processor.
[0073] Figure 8A schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device 100 includes: a processor 80, a memory 81, a bus 82, and a communication interface 83. The processor 80, the communication interface 83, and the memory 81 are connected through the bus 82. The processor 80 is configured to execute an executable module stored in the memory 81, such as a computer program.
[0074] Among them, the memory 81 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 83 (which can be wired or wireless), a communication connection is established between this system network element and at least one other network element. The Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0075] The bus 82 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 8 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0076] Among them, the memory 81 is used to store a program. After receiving an execution instruction, the processor 80 executes the program. The method executed by the device defined by the flow process disclosed in any one of the foregoing embodiments of the present invention can be applied to the processor 80 or implemented by the processor 80.
[0077] The processor 80 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 80 or the instructions in the form of software. The above-mentioned processor 80 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 81, and the processor 80 reads the information in the memory 81 and combines its hardware to complete the steps of the above method.
[0078] The computer program product of the readable storage medium provided by the embodiments of the present invention includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For the specific implementation, reference can be made to the foregoing method embodiments and will not be elaborated herein.
[0079] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0080] Finally, it should be noted that: the above-mentioned embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A training method for a text processing model, characterized in that: The text processing model includes: a plurality of expert modules, a routing module and a base model layer, including: Acquire a training sample set; wherein the training sample set includes: a plurality of composite text data formed by combining text data of a single text task; the text data of the single text task includes at least: a text similarity recognition task, a text category recognition task, a text feature extraction task, and a text answering task; Input each compound text data in the training sample set into the expert module and the routing module to obtain a guide sequence for each compound text data; Input each of the compound text data and the guide sequence of each of the compound text data into the base model layer to obtain the inference result of each of the compound text data; Determine a loss function value based on the inference result of each of the compound text data and the annotation result of the compound text data; Based on the loss function value, model parameters of the expert module and the routing module are updated.
2. The method according to claim 1, characterized in that Inputting each compound text data in the training sample set into the expert module and the routing module to obtain a guide sequence for each compound text data, including: For each compound text data in the training sample set, inputting the compound text data into the expert module to obtain a guide sequence output by each expert module; Input each composite text data in the training sample set into the routing module, obtain the weight coefficient of each expert module, and sort the weight coefficients in descending order, and select the weight coefficients of a preset number of expert modules as the output result of the routing module according to the sorting result; wherein the output result of the routing module includes: the numbers of the preset number of expert modules and the corresponding weight coefficients; Based on the numbers of the preset number of expert modules output by the routing module and the corresponding weight coefficients, the weight coefficients are weightedly calculated with the guide sequences output by the corresponding expert modules to obtain the guide sequence of the composite text data.
3. The method according to claim 2, characterized in that The expert module includes: a dynamic expert module with a multi-layer network structure and a static expert module with a single layer network structure; The dynamic expert module includes: a first encoding structure, an upsampling structure, a gating structure and a downsampling structure; the upsampling structure is a layer of fully connected neural network, the gating structure is a layer of fully connected neural network and an activation layer; the downsampling structure is a layer of fully connected neural network; The static expert module includes: a second coding structure, wherein the second coding structure is a layer of fully connected neural network; The routing module is a single-layer neural network structure.
4. The method according to claim 3, characterized in that Inputting the composite text data into the expert modules to obtain a guide sequence output by each of the expert modules, including: Inputting the compound text data into the dynamic expert module to obtain a first guide sequence; The compound text data is input into the static expert module, and a second guide sequence of a preset length is generated based on the fully connected neural network of the second coding structure.
5. The method according to claim 4, characterized in that Inputting the composite text data into the dynamic expert module to obtain a first guide sequence includes: Encoding the composite text data based on the first encoding structure to obtain an initial guide sequence of a preset length; Performing dimension expansion on the initial guide sequence based on the upsampling structure to obtain an upsampling result; Performing dimension expansion on the initial guide sequence based on the fully connected neural network of the gating structure to obtain an output result, and activating the output result based on the activation layer to obtain a gated output result; Based on the down-sampling structure, matrix multiplication is performed on the up-sampling result and the gated output result to obtain a matrix multiplication result, and the matrix multiplication result is passed through the fully connected neural network of the down-sampling structure to obtain a first guiding sequence output by the dynamic expert module.
6. The method according to claim 3, characterized in that Before obtaining the training sample set, it also includes: Obtain hyperparameters of the expert module and the routing module; wherein the hyperparameters include at least: the output dimension of the guide sequence generated by the first encoding structure, the output dimension of the upsampling structure, the output dimension of the downsampling structure, the output dimension of the gating structure, the activation method of the gating structure, the output dimension of the guide sequence generated by the second encoding structure, the number of expert modules of the routing module, and the ratio of the number of the dynamic expert modules to the number of the static expert modules.
7. A text processing method, characterized in that: include: Get the text data of the text task to be processed; The text data is input into a text processing model to obtain an inference result of the text task; wherein the text processing model is obtained by using the training method of the text processing model described in any one of claims 1 to 5.
8. A training device for a text processing model, characterized in that: The text processing model includes: a plurality of expert modules, a routing module and a base model layer, including: A sample acquisition module is used to acquire a training sample set; wherein the training sample set includes: a plurality of composite text data formed by combining text data of a single text task; the text data of the single text task at least includes: a text similarity recognition task, a text category recognition task, a text feature extraction task, and a text answer task; A guide sequence generating module, used for inputting each compound text data in the training sample set into the expert module and the routing module to obtain a guide sequence for each compound text data; An inference module, used for inputting each of the compound text data and the guide sequence of each of the compound text data into the base model layer to obtain an inference result of each of the compound text data; A loss function value determination module, used to determine a loss function value based on the inference result of each of the compound text data and the annotation result of the compound text data; A model parameter updating module is used to update the model parameters of the expert module and the routing module based on the loss function value.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the steps of the method according to any one of claims 1 to 6 or claim 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 6 or claim 7 are performed.
Citation Information
Patent Citations
Semantic text similarity calculation method based on attention
CN112101043A
Model training method, text classification method, system and device and storage medium
CN117421592A
Vehicle manufacturing quality defect detection method and device, electronic equipment and storage medium
CN118277838A
Heterogeneous multi-mode hybrid expert adapter
CN118708381A
Text classification method based on event tags
CN118733777A