Text processing method, text processing device and storage medium
By introducing a combination of discriminator and prefix encoder in the large language model, enabling or disabling prefix encoder according to the task type, the problem of large resource consumption of full-parameter fine-tuning and imbalanced fine-tuning effects of efficient model fine-tuning is solved, and resource saving and task effect balance is achieved.
Patent Information
- Application Number
- CN202410108157.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-25
AI Technical Summary
When using large language models for text processing in the prior art, full-parameter fine-tuning requires a large amount of data and computing resources, while efficient model fine-tuning saves resources but improves the effect of specific tasks but reduces the effect of other tasks.
The text processing task type is judged by pre-training the discriminator, and the prefix encoder is enabled under the target task type, and the prefix encoder is not enabled under the non-target task type. The text is processed using a large language model containing the prefix encoder.
While ensuring the good results of special tasks, it avoids the decline in the processing effects of other tasks and saves data and computing resources.
Smart Images

Figure CN120373448A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of natural language processing based on large language model technology, and particularly relates to a text processing method, a text processing device, and a storage medium. Background Art
[0002] When using a large language model to process text information, there is often a need to make it have better processing effects for certain fields. Since the number of parameters of the large language model is large (from billions to hundreds of billions), when further improving the large language model for tasks in certain fields, more training data and more computing resources are often required for fine-tuning.
[0003] In related technologies, there are technical means of full-parameter fine-tuning and efficient model fine-tuning (fine-tuning some parameters). However, full-parameter fine-tuning can only be performed on a base model and a model with a publicly available instruction set, and requires a large amount of data resources and computer hardware resources. And efficient model fine-tuning will make the model have better processing effects for certain fields, while the processing effects for other fields will be weaker than those of the model before fine-tuning. Summary of the Invention
[0004] To overcome the problems in related technologies, the present disclosure provides a text processing method, a text processing device, and a storage medium.
[0005] According to a first aspect of an embodiment of the present disclosure, a text processing method is provided, including: obtaining text description information, where the text description information includes a text to be processed and a task type of a text processing task, the text processing task is executed by a large language model, and the large language model is used to process the text to be processed; determining the task type of the text processing task by discriminating the text description information through a pre-trained discriminator, where the task type includes a target task type and a non-target task type; processing the text to be processed through the large language model based on the task type of the text processing task, determining the processed text to be processed as a target text, and outputting the target text; where the large language model includes a prefix encoder, the prefix encoder is trained based on a training data set of the target task type and is used to process text processing tasks of the target task type, the training data set includes text description information for training the prefix encoder, and the text description information for training the prefix encoder requires performing a text processing task of the target task type on the text to be processed.
[0006] In one implementation, processing the text to be processed by the large language model based on the task type of the text processing task includes: in response to the task type being the target task type, processing the text to be processed by the large language model on the premise of enabling the prefix encoder; in response to the task type being a non-target task type, processing the text to be processed by the large language model without enabling the prefix encoder.
[0007] In one implementation, processing the text to be processed by the large language model on the premise of enabling the prefix encoder includes: adding character information corresponding to the target task type before the embedding vector corresponding to each layer in the input layer of the large language model through the prefix encoder in the large language model; processing the embedding vector after adding the character information through the input layer of the large language model, and determining the text information output by the large language model as the target text.
[0008] In one implementation, the discriminator is obtained in the following manner: obtaining a positive sample data set and a negative sample data set, where the positive sample data set includes first training texts, the negative sample data set includes second training texts, the first training texts include the text to be processed for training the discriminator, and the task type corresponding to the first training texts is the target task type, the second training texts include the text to be processed for training the discriminator, and the task type corresponding to the second training texts is a non-target task type; training a discriminant model according to the positive sample data set and the negative sample data set, and determining the discriminant model after completion of training as the discriminator.
[0009] In one implementation, the large language model is obtained in the following manner: obtaining a training data set corresponding to the target task type, and training an original language model according to the training data set to obtain a prefix encoder corresponding to the target task type, where the original language model is used to process text description information of multiple task types; inserting the prefix encoder into the original language model to obtain the large language model.
[0010] According to the first aspect of the embodiments of the present disclosure, a text processing device is provided, including: an acquisition unit configured to acquire text description information, where the text description information includes the text to be processed and the task type of the text processing task, the text processing task is executed by a large language model, and the large language model is used to process the text to be processed; a determination unit configured to determine the task type of the text processing task by discriminating the text description information through a pre-trained discriminator, where the task type includes a target task type and a non-target task type; a processing unit configured to process the text to be processed through the large language model based on the task type of the text processing task, determine the processed text to be processed as the target text, and output the target text; wherein, the large language model includes a prefix encoder, the prefix encoder is trained based on a training data set of the target task type and is used to process the text processing task of the target task type, the training data set includes the text description information for training the prefix encoder, and the text description information for training the prefix encoder requires performing the text processing task of the target task type on the text to be processed.
[0011] In one implementation, the processing unit processes the text to be processed through the large language model based on the task type of the text processing task in the following manner, including: in response to the task type being the target task type, processing the text to be processed through the large language model with the prefix encoder enabled; in response to the task type being the non-target task type, processing the text to be processed through the large language model with the prefix encoder disabled.
[0012] In one implementation, the processing unit processes the text to be processed through the large language model with the prefix encoder enabled in the following manner, including: adding character information corresponding to the target task type in front of the embedding vector corresponding to each layer of the text description information in the input layer of the large language model through the prefix encoder in the large language model; processing the embedding vector after adding the character information through the input layer of the large language model, and determining the text information output by the large language model as the target text.
[0013] In one implementation, the discriminator is obtained by the processing unit in the following manner: obtaining a positive sample data set and a negative sample data set, where the positive sample data set includes first training texts, the negative sample data set includes second training texts, the first training texts contain the texts to be processed for training the discriminator, and the task type corresponding to the first training texts is the target task type, the second training texts contain the texts to be processed for training the discriminator, and the task type corresponding to the second training texts is a non-target task type; training a discriminant model according to the positive sample data set and the negative sample data set, and determining the discriminant model after completion of training as the discriminator.
[0014] In one implementation, the large language model is obtained by the processing unit in the following manner: obtaining a training data set corresponding to the target task type, and training an original language model according to the training data set to obtain a prefix encoder corresponding to the target task type, where the original language model is used to process text description information of multiple task types; inserting the prefix encoder into the original language model to obtain the large language model.
[0015] According to a third aspect of the embodiments of the present disclosure, there is provided a text processing device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to: execute the text processing method described in the first aspect or any implementation manner of the first aspect.
[0016] According to a fourth aspect of the embodiments of the present disclosure, there is provided a storage medium, in which instructions are stored, and when the instructions in the storage medium are executed by a processor, the processor is enabled to execute the text processing method described in the first aspect or any implementation manner of the first aspect.
[0017] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: obtaining text description information and determining the task type corresponding to the text description information in combination with a pre-trained discriminator. Processing the text to be processed included in the text description information through a large language model including a prefix encoder according to the task type. Among them, the prefix encoder is trained based on a training data set of the target task type, and the training data set includes training texts of the target task type. Through the present disclosure, after obtaining the text description information, the task type required to be executed by the text description information is discriminated by the discriminator, and the text to be processed is processed according to the task type through a large language model including a prefix encoder, while ensuring good effects for special text processing tasks, and the processing effects for other text processing tasks will not decline.
[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings
[0019] The drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0020] Figure 1 It is a flowchart of a text processing method shown according to an exemplary embodiment.
[0021] Figure 2 It is a flowchart of a method for processing a text to be processed based on a task type and a large language model shown according to an exemplary embodiment.
[0022] Figure 3 It is a flowchart of a text processing method shown according to an exemplary embodiment of the present disclosure.
[0023] Figure 4 It is a flowchart of a method for processing a text to be processed by a large language model on the premise of enabling a prefix encoder shown according to an exemplary embodiment.
[0024] Figure 5 It is a flowchart of a method for obtaining a discriminator shown according to an exemplary embodiment.
[0025] Figure 6 It is a flowchart of a method for obtaining a large language model shown according to an exemplary embodiment.
[0026] Figure 7 It is a schematic diagram of a method for fine-tuning a large language model shown according to an exemplary embodiment of the present disclosure.
[0027] Figure 8 It is a schematic diagram of a method for fine-tuning a large language model shown according to another exemplary embodiment of the present disclosure.
[0028] Figure 9 It is a block diagram of a text processing device shown according to an exemplary embodiment.
[0029] Figure 10 It is a block diagram of a device for text processing shown according to an exemplary embodiment. Detailed Description of the Embodiments
[0030] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure.
[0031] The text processing method provided by the embodiments of the present disclosure is applied to scenarios suitable for processing special tasks using large language models.
[0032] The text processing method provided by the embodiments of the present disclosure belongs to the field of large models (Large Model), and particularly relates to the field of large language models (Large Language Model) in natural language processing technology, and can also be extended to related artificial intelligence application fields such as computer vision.
[0033] When using a large language model to process text information, in order to meet the needs of different users and make the effect of the large language model processing text information fit the preferences and daily habits of users. Therefore, when using a large language model, there is often a need to process special text information, that is, the large language model has better processing effects for text information in certain fields than for text information in other fields. For the above processing requirements for special tasks, it is necessary to adjust the parameters of the large language model. In the actual adjustment process of the large language model, due to the large number of parameters of the large language model (from billions to trillions), when further improving and enhancing the large language model for tasks in certain fields, more training data and more computing resources are often required for fine-tuning.
[0034] In related technologies, there are technical means of full-parameter fine-tuning and efficient model fine-tuning (fine-tuning some parameters):
[0035] When fine-tuning the full-parameters of the model, since all model parameters need to be adjusted, more data resources and computer hardware resources are required. Moreover, full-parameter fine-tuning needs to add new instructions to the inherent instruction set of the language model to fine-tune the model base. Therefore, for models with unpublished model bases and unpublished instruction sets, the method of full-parameter fine-tuning cannot be used to improve them.
[0036] For the efficient model fine-tuning methods in related technologies (such as LoRA, Adapter, Prefix-tuning, P-tuning, P-tuning v2, etc.), by keeping most of the parameters of the large language model unchanged and only fine-tuning some parameters, the model fine-tuning on the new dataset with less computing resources is realized. The models trained by the efficient model fine-tuning method are all trained for specific task data (such as writing five-character quatrains). The prefix character embedding vectors and prefix encoders obtained by training are optimized for specific tasks. The processing effect of the model obtained after the efficient model fine-tuning method is trained on a specific task will be improved compared with the original model (i.e., the model before fine-tuning), but the text processing effect for other tasks (such as obtaining a summary of a piece of text information) will decline compared with the original model, showing a "one-sided" effect.
[0037] In summary, the full-parameter fine-tuning in the related art requires a large amount of data resources and computer hardware resources, and can only fine-tune large language models with publicly available instruction sets and model bases. The fine-tuned models obtained by the efficient model fine-tuning (fine-tuning some parameters) in the related art have better processing effects for specific tasks than the pre-fine-tuning models, but are weaker than the pre-fine-tuning models in terms of processing effects for other tasks.
[0038] In view of this, the present disclosure proposes a text processing method, which pre-trains a large language model for processing text and trains a discriminator for determining the task type of the input text. Among them, the large language model includes a prefix encoder dedicated to processing text tasks of the target type. In the actual text processing process, the task type of the text to be processed obtained by the discriminator is discriminated. For the text to be processed of the target task type, the prefix encoder is enabled, and the large language model is used to process the text to be processed; for the text to be processed of the non-target task type, the prefix encoder is not enabled, and the large language model is used to process the text to be processed. Through the present disclosure, after obtaining the text to be processed, the task type of the text to be processed is discriminated by the discriminator, and the text is processed by the large language model including the prefix encoder according to the task type, while ensuring good effects for special tasks, the processing effects for other tasks will not decline.
[0039] Figure 1 is a flowchart of a text processing method shown according to an exemplary embodiment. As Figure 1 shown, the method includes steps S101 to S103.
[0040] In step S101, text description information is obtained, and the text description information includes the text to be processed and the task type of the text processing task.
[0041] Among them, the text processing task is executed by the large language model, and the large language model is used to process the text to be processed.
[0042] In step S102, the text description information is discriminated by a pre-trained discriminator to determine the task type of the text processing task, and the task type includes a target task type and a non-target task type.
[0043] In step S103, based on the task type of the text processing task, the large language model is used to process the text to be processed, and the processed text to be processed is determined as the target text, and the target text is output.
[0044] Among them, the large language model includes a prefix encoder, which is trained based on a training data set of a target task type and is used to process text processing tasks of the target task type. The training data set includes text description information for training the prefix encoder, and the text description information for training the prefix encoder requires performing text processing tasks of the target task type on the text to be processed.
[0045] In the embodiments of the present disclosure, the text description information for indicating text processing can be obtained through various channels, including but not limited to the following categories: text obtained by recognizing and converting the user's language, text obtained by the user typing, handwriting, pasting and copying, text obtained by image semantic recognition to obtain image description text and combining with the user's instructions. The text to be processed in the present disclosure will include simple description text and text that directly or indirectly conveys task instructions. In one example, the text description information can be "Please briefly describe the following content,...", where "Please briefly describe the following content" corresponds to the task type of the text processing task, and "..." is the text to be processed. In another example, the text description information can be "... please output the above content in the form of a seven-character quatrain". Among them, "please output the above content in the form of a seven-character quatrain" corresponds to the task type of the text processing task, and "..." is the text to be processed. It can be understood that the text processing method proposed in the present disclosure is not only applicable to the text processing tasks in the above examples, but also applicable to any text processing tasks that can be performed by a large language model.
[0046] In the embodiments of the present disclosure, after obtaining the text description information, the discriminator is used to discriminate the task type of the text to be processed, and the large language model including the prefix encoder processes the text according to the task type, while ensuring good results for special tasks, the processing effect for other tasks will not decline.
[0047] In the embodiments of the present disclosure, the prefix encoder in the large language model is dedicated to processing the text to be processed of the target task type. Therefore, after determining the task type corresponding to the text description information through the discriminator, it is necessary to decide whether to enable the prefix encoder in the large language model based on the task type corresponding to the text description information, and then process the text to be processed. The embodiments of the present disclosure illustrate the method for processing the text to be processed.
[0048] Figure 2 It is a flowchart of a method for processing text to be processed based on a task type and a large language model shown according to an exemplary embodiment. As Figure 2 shown, the method includes step S201, step S202A, and step S202B.
[0049] In step S201, the discriminator trained in advance is used to discriminate the text description information to determine the task type of the text processing task, and the task type includes the target task type and the non-target task type.
[0050] In step S202A, in response to the task type being the target task type, the large language model processes the text to be processed on the premise of enabling the prefix encoder.
[0051] In step S202B, in response to the task type being a non-target task type, the large language model processes the text to be processed without enabling the prefix encoder.
[0052] In the embodiments of the present disclosure, the prefix encoder in the large language model is dedicated to processing text processing tasks of the target task type. When performing text processing tasks of the target task type, the processing effect of the large language model is improved through the prefix encoder. It can be understood that when performing text processing tasks of non-target task types, if the prefix encoder is enabled, the corresponding parameters for the text to be processed of the target type will be increased during the text processing of non-target type tasks, resulting in contamination by other data during the text processing process and a reduction in the processing effect when performing text processing tasks of non-target task types. Therefore, the present disclosure does not enable the prefix encoder when determining that the task type corresponding to the text description information is a non-target task type to ensure the processing effect for the text to be processed of non-target types.
[0053] In the embodiments of the present disclosure, for text processing tasks of the target task type, the large language model enables the prefix encoder to process them; for text processing tasks of non-target task types, the large language model does not enable the prefix encoder to process them. Through the present disclosure, while improving the text processing effect of performing text processing tasks of the target task type, the text processing effect when performing text processing tasks of non-target task types is ensured.
[0054] In an exemplary embodiment of the present disclosure, as Figure 3 shown in the flowchart of the text processing method, the present disclosure processes the text to be processed in the following manner: input task (i.e., input text description information); task discrimination, that is, judging the task type of the text processing task through a discriminator; when it is determined that the task type is the target task type, the prefix encoder becomes effective, enabling the prefix encoder to participate in the text processing process, and using the fine-tuned large language model to process the text and output the target text; when it is determined to be a non-target task type, the prefix encoder is invalid, the prefix encoder is not enabled, and the original large language model is used to process the text and output the target text.
[0055] In the embodiments of the present disclosure, when processing the text to be processed on the premise of enabling the prefix encoder, the prefix encoder will supplement the parameters corresponding to the target task type during the text processing process, thereby improving the text processing effect. The following embodiments of the present disclosure illustrate the method of processing the text to be processed by the large language model on the premise of enabling the prefix encoder.
[0056] Figure 4 It is a flowchart of a method for processing text to be processed by a large language model with a prefix encoder enabled according to an exemplary embodiment. As Figure 4 shown, the method includes steps S301 to S302.
[0057] In step S301, through the prefix encoder in the large language model, character information corresponding to the target task type is added before the embedding vector corresponding to the text description information in each layer of the input layer of the large language model.
[0058] In step S302, the embedding vector after adding the character information is processed through the input layer of the large language model, and the text information output by the large language model is determined as the target text.
[0059] In the embodiments of the present disclosure, the prefix encoder is trained based on a training data set of the target task type, and the training data set includes text description information (including the text to be processed) that requires performing the target task type. Therefore, after the prefix encoder is enabled, the prefix encoder supplements character information corresponding to the target task type during the process of the large language model processing the text to be processed and converting the text to be processed into an embedding vector, and then inputs the embedding vector and the supplemented character information into the input layer of the large language model together. Since the prefix encoder supplements character information corresponding to the target task type during the text processing of the large language model, the large language model is superior to the general large language model in processing the text to be processed of the target type.
[0060] In the embodiments of the present disclosure, when the task type of the text processing task is determined to be the target task type, the large language model enables the prefix encoder and processes the text to be processed. During the process of processing the text to be processed by the prefix encoder, character information corresponding to the target task type is supplemented to improve the text processing effect of the large language model for the special task (i.e., the text processing task of the target task type).
[0061] In the embodiments of the present disclosure, the discriminator for discriminating the task type corresponding to the text to be processed is used to distinguish the target task type (including a single task type) from the non-target task type (including multiple task types). During the process of training the discriminant model to obtain the discriminator, text description information that requires performing the text processing task of the target task type and text description information that requires performing the text processing task of the non-target task type need to be used as training data respectively. The following embodiments of the present disclosure illustrate the training method of the discriminator.
[0062] Figure 5 It is a flowchart of a method for obtaining a discriminator according to an exemplary embodiment. As Figure 5As shown, the method includes steps S401 to S402.
[0063] In step S401, a positive sample data set and a negative sample data set are obtained.
[0064] Among them, the positive sample data set includes first training texts, the negative sample data set includes second training texts, the first training texts contain the text to be processed for training the discriminator, and the task type corresponding to the first training texts is the target task type, the second training texts contain the text to be processed for training the discriminator, and the task type corresponding to the second training texts is the non-target task type.
[0065] In step S402, a discriminant model is trained according to the positive sample data set and the negative sample data set, and the discriminant model after completion of training is determined as the discriminator.
[0066] In the embodiments of the present disclosure, a training data set is obtained according to the task types supported by the large language model, that is, an equal amount of training texts are respectively obtained for each task type supported by the large language model as the training data set. The data set corresponding to the target task type among them is used as the positive sample data set, and the data set corresponding to the non-target task type among them is used as the negative sample data set. The positive sample data set and the negative sample data set are divided so that a part of them is used for training the discriminant model, and the other part is used for validating the discriminant model. After training and validating the discriminant model through the positive sample data set and the negative sample data set, a discriminator that can be used to discriminate the task type corresponding to the text to be processed is obtained.
[0067] In the embodiments of the present disclosure, by collecting texts corresponding to the task types supported by the large language model as the training data set, a discriminator is trained, and the task type required to be executed by the text description information is discriminated after obtaining the text description information. It is convenient to enable the large language model to select different processing strategies according to different task types in the subsequent process, while improving the text processing effect of the text processing task of the target task type and ensuring the text processing effect of the text processing task of the non-target task type.
[0068] In the embodiments of the present disclosure, the difference between the large language model and the general language model is that it includes a prefix encoder for processing special tasks, and the prefix encoder is obtained based on the text corresponding to the target task type and the general large language model after training. The following embodiments of the present disclosure illustrate the method for obtaining the large language model.
[0069] Figure 6 is a flowchart of a method for obtaining a large language model shown according to an exemplary embodiment. As Figure 6 shown, the method includes steps S501 to S502.
[0070] In step S501, a training data set corresponding to the target task type is obtained, and the original language model is trained based on the training data set to obtain a prefix encoder corresponding to the target task type.
[0071] Among them, the original language model is used to process text description information of multiple task types.
[0072] In step S502, the prefix encoder is inserted into the original language model to obtain a large language model.
[0073] In the embodiments of the present disclosure, the original large language model used to train and obtain the large language model supports text processing tasks of multiple task types. Based on conditions such as user preferences or usage habits, one task type among the multiple task types supported by the original language model can be selected as the target task type, and a predetermined number of target task type texts are obtained as the training data set. After obtaining the training data set, the original language model is trained based on the training data set to obtain a prefix encoder, and the prefix encoder is inserted into the original language model so that the prefix encoder can participate in the text processing of the model to obtain a large language model. Further, the large language model including the prefix encoder only modifies some model parameters compared with the original language model.
[0074] In the embodiments of the present disclosure, the original language model is trained according to the training data corresponding to the target task type to obtain a prefix encoder. The prefix encoder is inserted into the original language model to obtain a large language model that can be used for special tasks, that is, the fine-tuning of the original language model is realized. In the present disclosure, by inserting the prefix encoder into the original language model, the prefix encoder can participate in the text processing of the model, so that the processing effect of the trained large language model on the text to be processed of the target type is better than that of the original language model. And, the large language model including the prefix encoder only modifies some model parameters compared with the original language model, saving data resources and computer hardware resources compared with the full-parameter fine-tuning of the model. And, the method for obtaining the large language model in the present disclosure does not need to adjust the inherent instruction set and model base of the language model compared with the full-parameter fine-tuning of the model, and more large language models are applicable than those used in the full-parameter fine-tuning.
[0075] In an exemplary embodiment of the present disclosure, the P-tuning v2 efficient model fine-tuning method is used to obtain a large language model, such as Figure 7As shown in the schematic diagram of the large language model fine-tuning method, the principle of the P-tuning v2 efficient model fine-tuning method is as follows: special characters are added to the input layer of the model through a prefix encoder. Not only the vectors input to the Transformer layer are provided by the prefix encoder, but the vectors of each layer come from the prefix encoder (i.e., the first-layer prompt to the nth-layer prompt). Before the multi-layer downstream task embedding vectors corresponding to the downstream tasks [CLS], text content A, text content B, and text content C, namely e[CLS], word embedding vector A, namely e[A], word embedding vector B, namely e[B], and word embedding vector C, namely e[C], the prompt information (i.e., the first-layer prompt to the nth-layer prompt) is inserted, and the output class label [Class Label(with linear head)] and the optimized input are input into the parameterization module to achieve the fine-tuning of the large language model. It can be understood that P-tuning v2 adds more trainable parameters and has better training effects compared with general efficient model fine-tuning methods (such as P-tuning). And the number of trainable parameters is still smaller than that of the entire large language model and smaller than that to be adjusted by full-parameter fine-tuning. It still has the advantages of high-efficiency parameter fine-tuning.
[0076] In an exemplary embodiment of the present disclosure, an alternative solution without using a prefix encoder is proposed. The P-tuning efficient model fine-tuning method is used to obtain a large language model, such as Figure 8As shown in the schematic diagram of the large language model fine-tuning method, the principle of the P-tuning efficient model fine-tuning method is as follows: fine-tuning is performed based on a pure decoder language model. The pure decoder language model generates one by one in the order from front to back. If the input sentence is "A B C", the model will generate a reply after decoding the input sentence (the beginning of the reply is marked with [MASK]). The P-tuning method is to add i special marker characters before the [MASK] of the reply. When inputting the Transformers layer, the embedding vectors of the special characters (i.e., the first-layer prompts corresponding to the special character embedding vectors h0 to the special character embedding vector hi) and the word embedding vectors of the words in the original sentence are used as the input of the Transformers together. Among them, the word embedding vectors of the words in the original sentence are the multi-layer downstream task embedding vectors e[CLS], the word embedding vector A (i.e., e[A]), the word embedding vector B (i.e., e[B]), and the word embedding vector C corresponding to the downstream task and the text content A, B, and C respectively. Among them, the special character embedding vectors h0 to the special character embedding vector hi are the vectors that can be learned during the P-tuning training process. After the processing at the input layer, the processing result is subjected to answer space mapping [Verbalizer(with LMhead)] and optimization (Optimzation), and then input into the prompt encoder to achieve the fine-tuning of the large language model. It can be understood that compared with adjusting the parameters of the entire large language model LLM, P-tuning only needs to train the embedding vectors of dozens of special characters, reducing the training volume and memory requirements of the model. At the same time, because the number of training parameters is small, the required training data is also relatively small.
[0077] In an exemplary embodiment of the present disclosure, the discriminator is trained in the following manner, and the language model is used for text processing in combination with the discriminator: Training stage: Use P-tuning efficient model fine-tuning to train a large language model (such as chatglm) on a specific dataset (such as a summary dataset), and retain the trained prefix encoder. Train a discriminator to distinguish whether the input task is a task specifically trained by P-tuning. If it is, output 1; otherwise, output 0. Inference stage: First, use the discriminator to determine whether the input task belongs to a specific training task, and output the discrimination result 0 / 1. Then use the large language model to complete the input task. According to the discrimination result of the previous step, if the discrimination result is 1, indicating that the task is a specific task of P-tuning, the prefix encoder is set to be effective, and the current model is in a mode with better performance on the specific task after P-tuning. If the discrimination result is 0, indicating that the task is other tasks, the prefix encoder is set to be invalid, and the current model is the original model without P-tuning.
[0078] In the embodiments of the present disclosure, a large language model with a prefix encoder is pre-trained, and a discriminator for determining the task type corresponding to the text description information is trained. Among them, the large language model contains a prefix encoder dedicated to processing text processing tasks of the target type. In the actual text processing process, the discriminator is used to discriminate the task type (including the target task type and the non-target task type) required to be executed by the obtained text description information. For the text processing task of the target task type, the prefix encoder is enabled, and the large language model is used to process the text to be processed; for the text processing task of the non-target task type, the prefix encoder is not enabled, and the large language model is used to process the text to be processed. Through the present disclosure, after obtaining the text to be processed, the discriminator is used to discriminate the task type of the text processing task, and the large language model including the prefix encoder processes the text according to the task type, ensuring good results for special text processing tasks while not degrading the processing effect for other text processing tasks. Moreover, the model training method in the text processing method of the present disclosure saves data resources and computer hardware resources compared with the full-parameter fine-tuning method, does not require adjustment of the inherent instruction set and model base of the language model, and is applicable to more large language models than those used in full-parameter fine-tuning.
[0079] Based on the same concept, the embodiments of the present disclosure also provide a text processing device 100.
[0080] It can be understood that, in order to implement the above functions, the text processing device 100 provided by the embodiments of the present disclosure includes the corresponding hardware structure and / or software module for executing each function. Combining the units and algorithm steps of the various examples disclosed in the embodiments of the present disclosure, the embodiments of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiments of the present disclosure.
[0081] Figure 9 is a block diagram of a text processing device 100 shown according to an exemplary embodiment. Referring to Figure 8 , the device includes an acquisition unit 101, a determination unit 102, and a processing unit 103.
[0082] The acquisition unit 101 is configured to acquire text description information, where the text description information includes the text to be processed and the task type of the text processing task.
[0083] Among them, the text processing task is executed by a large language model, and the large language model is used to process the text to be processed.
[0084] A determination unit 102 is configured to determine the task type of a text processing task by discriminating text description information through a pre-trained discriminator. The task type includes a target task type and a non-target task type.
[0085] A processing unit 103 is configured to process the text to be processed through a large language model based on the task type of the text processing task, determine the processed text to be processed as the target text, and output the target text.
[0086] Among them, the large language model includes a prefix encoder, which is trained based on a training data set of the target task type and is used to process text processing tasks of the target task type. The training data set includes text description information for training the prefix encoder, and the text description information for training the prefix encoder requires performing text processing tasks of the target task type on the text to be processed.
[0087] In one implementation, the processing unit 103 processes the text to be processed through a large language model based on the task type of the text processing task in the following manner, including: in response to the task type being the target task type, processing the text to be processed through the large language model with the prefix encoder enabled; in response to the task type being the non-target task type, processing the text to be processed through the large language model with the prefix encoder disabled.
[0088] In one implementation, the processing unit 103 processes the text to be processed through the large language model with the prefix encoder enabled in the following manner, including: adding character information corresponding to the target task type in front of the embedding vector corresponding to each layer of the text description information in the input layer of the large language model through the prefix encoder in the large language model; processing the embedding vector with the added character information through the input layer of the large language model, and determining the text information output by the large language model as the target text.
[0089] In one implementation, the discriminator is obtained by the processing unit 103 in the following manner: obtaining a positive sample data set and a negative sample data set. The positive sample data set includes first training texts, and the negative sample data set includes second training texts. The first training texts include the text to be processed for training the discriminator, and the task type corresponding to the first training texts is the target task type. The second training texts include the text to be processed for training the discriminator, and the task type corresponding to the second training texts is the non-target task type. Training a discriminant model according to the positive sample data set and the negative sample data set, and determining the discriminant model after completion of training as the discriminator.
[0090] In one implementation, the large language model is obtained by the processing unit 103 in the following manner: Obtain a training data set corresponding to the target task type, and train the original language model according to the training data set to obtain a prefix encoder corresponding to the target task type. The original language model is used to process text description information of multiple task types. Insert the prefix encoder into the original language model to obtain the large language model.
[0091] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0092] Figure 10 FIG. 200 is a block diagram of a device 200 for text processing shown according to an exemplary embodiment. The device 200 may be provided as a terminal. For example, the device 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0093] Referring to Figure 10 , the device 200 may include one or more of the following components: a processing component 202, a memory 204, a power component 206, a multimedia component 208, an audio component 210, an input / output (I / O) interface 212, a sensor component 214, and a communication component 216.
[0094] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 202 may include one or more modules to facilitate the interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate the interaction between the multimedia component 208 and the processing component 202.
[0095] The memory 204 is configured to store various types of data to support the operation of the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, etc. The memory 204 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0096] The power component 206 provides power for various components of the device 200. The power component 206 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device 200.
[0097] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0098] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC) that is configured to receive external audio signals when the device 200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 further includes a speaker for outputting audio signals.
[0099] The I / O interface 212 provides an interface between the processing component 202 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0100] The sensor assembly 214 includes one or more sensors for providing a status assessment of various aspects of the device 200. For example, the sensor assembly 214 can detect the on / off state of the device 200, the relative positioning of components, such as the display and keypad of the device 200. The sensor assembly 214 can also detect a change in the position of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200, and the temperature change of the device 200. The sensor assembly 214 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 214 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 214 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0101] The communication component 216 is configured to facilitate communication between the device 200 and other devices in a wired or wireless manner. The device 200 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0102] In an exemplary embodiment, the device 200 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0103] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 204 including instructions, and the above instructions can be executed by the processor 220 of the device 200 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0104] It can be understood that in the present disclosure, "a plurality of" means two or more, and other quantifiers are similar thereto. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The singular forms of "a", "the", and "said" are also intended to include the plural forms unless the context clearly indicates otherwise.
[0105] Furthermore, it can be understood that the terms "first", "second", etc. are used to describe various information, but such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other and do not represent a specific order or degree of importance. In fact, the expressions such as "first" and "second" can be used interchangeably. For example, without departing from the scope of the present disclosure, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information.
[0106] Furthermore, it can be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "lateral", "front", "rear", "upper", "lower", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present embodiment and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation.
[0107] Furthermore, it can be understood that unless otherwise specified, "connection" includes direct connection between two parties without other components therebetween, and also includes indirect connection between two parties with other elements therebetween.
[0108] Furthermore, it can be understood that although the operations are described in a specific order in the drawings in the embodiments of the present disclosure, it should not be understood as requiring the operations to be performed in the specific order shown or in a serial order, or requiring all the operations shown to obtain the desired result. In a specific environment, multitasking and parallel processing may be advantageous.
[0109] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of this solution, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure.
[0110] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A text processing method, characterized in that, including: Obtain text description information, where the text description information includes the text to be processed and the task type of the text processing task, and the text processing task is executed by a large language model, and the large language model is used to process the text to be processed; Discriminate the text description information through a pre-trained discriminator to determine the task type of the text processing task, where the task type includes a target task type and a non-target task type; Based on the task type of the text processing task, process the text to be processed through the large language model, determine the processed text to be processed as the target text, and output the target text; Among them, the large language model includes a prefix encoder, and the prefix encoder is trained based on the training data set of the target task type and is used to process the text processing task of the target task type. The training data set includes the text description information for training the prefix encoder, and the text description information for training the prefix encoder requires performing the text processing task of the target task type on the text to be processed.
2. The method according to claim 1, wherein The processing the text to be processed through the large language model based on the task type of the text processing task includes: In response to the task type being the target task type, process the text to be processed through the large language model with the prefix encoder enabled; In response to the task type being the non-target task type, process the text to be processed through the large language model without the prefix encoder enabled.
3. The method according to claim 2, wherein The processing the text to be processed through the large language model with the prefix encoder enabled includes: Through the prefix encoder in the large language model, add character information corresponding to the target task type before the embedding vector corresponding to the text description information in each layer of the input layer of the large language model; Process the embedding vector after adding the character information through the input layer of the large language model, and determine the text information output by the large language model as the target text.
4. The method according to any one of claims 1 to 3, characterized in that, The discriminator is obtained in the following manner: Obtain a positive sample data set and a negative sample data set. The positive sample data set includes first training texts, and the negative sample data set includes second training texts. The first training texts include the text to be processed for training the discriminator, and the task type corresponding to the first training texts is the target task type. The second training texts include the text to be processed for training the discriminator, and the task type corresponding to the second training texts is the non-target task type; Train a discrimination model according to the positive sample data set and the negative sample data set, and determine the discrimination model after completion of training as the discriminator.
5. The method according to any one of claims 1 to 3, characterized in that, The large language model is obtained in the following manner: Obtain a training data set corresponding to the target task type, and train an original language model according to the training data set to obtain a prefix encoder corresponding to the target task type. The original language model is used to process text description information of multiple task types; Insert the prefix encoder into the original language model to obtain the large language model.
6. A text processing device, characterized in that, including: An acquisition unit, configured to acquire text description information, where the text description information includes the text to be processed and the task type of the text processing task, and the text processing task is executed by a large language model for processing the text to be processed; A determination unit, configured to determine the task type of the text processing task by discriminating the text description information through a pre-trained discriminator, where the task type includes a target task type and a non-target task type; A processing unit, configured to process the text to be processed through the large language model based on the task type of the text processing task, determine the processed text to be processed as the target text, and output the target text; Wherein, the large language model includes a prefix encoder, and the prefix encoder is trained based on a training data set of the target task type and is used to process text processing tasks of the target task type. The training data set includes text description information for training the prefix encoder, and the text description information for training the prefix encoder requires performing text processing tasks of the target task type on the text to be processed.
7. The device according to claim 6, characterized in that, The processing unit processes the text to be processed through the large language model based on the task type of the text processing task in the following manner, including: In response to the task type being the target task type, processing the text to be processed through the large language model with the prefix encoder enabled; In response to the task type being the non-target task type, processing the text to be processed through the large language model with the prefix encoder disabled.
8. The device according to claim 7, characterized in that, The processing unit processes the text to be processed through the large language model with the prefix encoder enabled in the following manner, including: Through the prefix encoder in the large language model, adding character information corresponding to the target task type before the embedding vector corresponding to the text description information in each layer of the input layer of the large language model; Processing the embedding vector after adding the character information through the input layer of the large language model, and determining the text information output by the large language model as the target text.
9. The device according to any one of claims 6 to 8, characterized in that The discriminator is obtained by the processing unit in the following manner: Acquire a positive sample data set and a negative sample data set. The positive sample data set includes first training texts, and the negative sample data set includes second training texts. The first training texts include the text to be processed for training the discriminator, and the task type corresponding to the first training texts is the target task type. The second training texts include the text to be processed for training the discriminator, and the task type corresponding to the second training texts is the non-target task type; Train a discriminant model according to the positive sample data set and the negative sample data set, and determine the discriminant model after completion of training as the discriminator.
10. The device according to any one of claims 6 to 8, characterized in that The large language model is obtained by the processing unit in the following manner: Obtain a training data set corresponding to the target task type, and train an original language model according to the training data set to obtain a prefix encoder corresponding to the target task type, where the original language model is used to process text description information of multiple task types; Insert the prefix encoder into the original language model to obtain the large language model.
11. A text processing device, characterized in that, It includes: A processor: A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the text processing method according to any one of claims 1 to 5.
12. A storage medium, characterized in that, Instructions are stored in the storage medium, and when the instructions in the storage medium are executed by the processor, the processor can execute the text processing method according to any one of claims 1 to 5.