Method, apparatus, device and storage medium for pre-training a language model

By constructing a hierarchical multi-template multi-task language dataset from unsupervised and supervised data, the method addresses continuous learning challenges, enhancing the robustness and transfer ability of language models.

JP7777059B2Active Publication Date: 2025-11-27BEIJING BAIDU NETCOM SCI & TECH CO LTD

Patent Information

Application Number
JP2022169714
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-05-20
Filing Date
2022-10-24
Publication Date
2025-11-27
Estimated Expiration
2042-10-24

AI Technical Summary

Technical Problem

Existing multitask-based fine-tuning and pre-training techniques for language models hinder continuous learning from unsupervised data and reduce model robustness due to single template design limitations.

Method used

A method involving the construction of a pre-training language dataset that includes unsupervised and supervised data, generating a hierarchical multi-template multi-task language dataset, and pre-training a language model using this dataset to enhance diversity and robustness.

Benefits of technology

The method improves the model's ability to learn general knowledge from unsupervised data, enhances task learning diversity, and increases transfer capabilities, especially in scenarios with few or no samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007777059000009
    Figure 0007777059000009
  • Figure 0007777059000010
    Figure 0007777059000010
  • Figure 0007777059000011
    Figure 0007777059000011
Patent Text Reader

Abstract

To provide a language model pre-training method by which a model can model multi-task data at the same time to increase varieties of the model, the robustness of task learning of the model can be improved and knowledge related to tasks and datasets can be learned more properly, and transfer performance of the model can be improved without or with a few samples, an apparatus, a device, a storage medium, and a program.SOLUTION: A method comprises: a step S101 of constructing a pre-training language data set including unsupervised language data and supervised language data; a step S102 of generating a hierarchical multi-template multi-task language data set according to the pre-training language data set; and a step S103 of pre-training a language model on the basis of the hierarchical multi-template multi-task language data set.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the field of computer technology, specifically to the field of artificial intelligence, in particular to the field of deep learning technology, and in particular to a method, apparatus, device, storage medium and program for pre-training a language model. [Background technology]

[0002] In related technologies, multitask-based fine-tuning and multitask-based pre-training techniques can give large-scale language models powerful general-purpose text generation capabilities.

[0003] However, in related technologies, the multi-task fine-tuning technique may prevent the model from learning general knowledge from unsupervised data, which may prevent the model from learning continuously. In the multi-task data pre-training technique, the single template design affects the robustness of the model.

[0004] Therefore, there is an urgent need for a "language model pre-training method" to solve the problem of continuous model training, increase template diversity, and improve the robustness of the model's multi-task learning. Summary of the Invention

[0005] The present disclosure provides a method, apparatus, device, storage medium, and program for pre-training a language model.

[0006] According to a first aspect of the present disclosure, there is provided a method for pre-training a language model, the method including: constructing a pre-training language dataset, wherein the pre-training language dataset includes unsupervised language data and supervised language data; generating a hierarchical multi-template multi-task language dataset based on the pre-training language dataset; and pre-training a language model based on the hierarchical multi-template multi-task language dataset.

[0007] According to a second aspect of the present disclosure, there is provided an apparatus for pre-training a language model, comprising: a construction module for constructing a pre-training language dataset, wherein the pre-training language dataset includes unsupervised language data and supervised language data; a generation module for generating a hierarchical multi-template multi-task language dataset based on the pre-training language dataset; and a pre-training module for pre-training a language model based on the hierarchical multi-template multi-task language dataset.

[0008] According to a third aspect of the present disclosure, there is provided an electronic device comprising at least one processor and a memory communicatively connected to the at least one processor, the memory storing instructions to be executed by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor performs a method according to any of the method embodiments described above.

[0009] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having computer instructions stored thereon, the computer instructions causing a computer to perform a method according to any of the method embodiments described above.

[0010] According to a fifth aspect of the present disclosure, there is provided a computer program, wherein instructions in the computer program, when executed by a processor, result in a method according to any of the method embodiments above. [Effects of the Invention]

[0011] In the language model pre-training method of the embodiment of the present disclosure, a hierarchical multi-template multi-task language dataset is generated based on the pre-training language dataset, and the multi-template multi-task language dataset is constructed. By collectively configuring tasks as templates, the model can simultaneously model multiple task data. By configuring multiple templates, the model can increase the diversity and improve the robustness of task learning. In addition, when pre-training the model, consecutive templates can better learn knowledge related to the task and dataset, thereby improving the model's transfer ability when there are no samples or few samples.

[0012] In addition, based on constructing a pre-training language dataset, the present disclosure proposes joint training with unsupervised general knowledge and supervised task knowledge based on a language model, in which the language model can not only model task knowledge (based on supervised task knowledge), but also continuously learn general knowledge from training unsupervised data, thereby avoiding knowledge forgetting.

[0013] It should be noted that the contents of this section are not intended to identify key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present application will become apparent from the following description. [Brief explanation of the drawings]

[0014] The drawings are for a better understanding of the present application and are not intended to limit the present disclosure. [Figure 1]FIG. 1 is a schematic diagram of a language model pre-training method provided by an embodiment of the present disclosure. [Figure 2] FIG. 1 is a schematic diagram of a pre-training dataset for a language model provided by an embodiment of the present disclosure. [Figure 3] FIG. 1 is a schematic diagram illustrating generating a layered multi-template multi-task dataset based on a pre-training dataset of a language model provided by an embodiment of the present disclosure. [Figure 4] 3A to 3C are schematic diagrams of sample language data of first to fourth granularities provided by the present embodiment. [Figure 5] FIG. 1 is a schematic diagram of pre-training a set language model based on layered multi-template multi-task data provided by an embodiment of the present disclosure. [Figure 6] FIG. 1 is a schematic configuration diagram of a language model pre-training device provided by an embodiment of the present disclosure. [Figure 7] FIG. 1 is a schematic configuration diagram of another language model pre-training device provided by an embodiment of the present disclosure. [Figure 8] FIG. 10 is a schematic configuration diagram of yet another language model pre-training device provided by an embodiment of the present disclosure. [Figure 9] FIG. 1 is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0015] Hereinafter, exemplary embodiments of the present disclosure will be described in conjunction with the drawings. For ease of understanding, various details of the embodiments of the present invention are included therein, and they should be considered as mere examples. Therefore, those skilled in the art can make various changes and modifications to the embodiments described herein without departing from the scope and spirit of the present invention. Also, for the sake of clarity and conciseness, the following description will omit descriptions of well-known functions and structures.

[0016] In the technical solution disclosed herein, the acquisition, storage, and application of personal information of involved users all comply with the provisions of relevant laws and regulations and do not violate public order and morals.

[0017] Hereinafter, an execution control method, device, and electronic device for training a model according to an embodiment of the present application will be described with reference to the drawings.

[0018] FIG. 1 is a schematic flowchart of a language model pre-training method provided by an embodiment of the present disclosure, which includes the following steps:

[0019] In step S101, a pre-training language dataset is constructed, and the pre-training language dataset includes unsupervised language data and supervised language data.

[0020] FIG. 2 is a schematic diagram of a pre-training dataset for a language model provided by an embodiment of the present disclosure.

[0021] Here, as shown in FIG. 2, in one embodiment of the present disclosure, the unsupervised language data may be a large amount of text data and a knowledge graph.

[0022] For example, in one embodiment of the present disclosure, the large amount of text data may be text data from web pages or other search engines, and the knowledge graph may be three-tuple data of a knowledge base with a directed graph structure.

[0023] As shown in FIG. 2, in one embodiment of the present disclosure, the supervised language data may be multi-task language data.

[0024] Specifically, in one embodiment of the present disclosure, the multi-task language data may include general natural language understanding and generation tasks.

[0025] By way of example, in one embodiment of the present disclosure, the supervised language dataset may include an open question answering dataset, a sentiment analysis dataset, a semantic matching dataset, a text classification dataset, a text summarization dataset, and the like.

[0026] In step S102, a hierarchical multi-template multi-task language dataset is generated based on the pre-training language dataset.

[0027] Here, in one embodiment of the present disclosure, the pre-training language dataset includes supervised language data, and the supervised language data includes a multi-task language dataset, and for the multi-task language dataset, corresponding task templates and at least one task sub-template corresponding to each task template are set.

[0028] In one embodiment of the present disclosure, each task language dataset is divided into at least one task category based on at least one task sub-template corresponding to each task language dataset, thereby generating a hierarchical multi-template multi-task language dataset.

[0029] For example, FIG. 3 is a schematic diagram illustrating the generation of a hierarchical multi-template multi-task language dataset based on a pre-training dataset for a language model provided by an embodiment of the present disclosure. In one embodiment of the present disclosure, FIG. 3 is divided into three sub-diagrams, of which the left sub-diagram can be a multi-task dataset. As shown in the diagram, the multi-task dataset includes corresponding task templates such as sentiment analysis, development question answering, question matching, and advertisement creation, each of which has corresponding sample text. The central sub-diagram can also illustrate the division of each task dataset into at least one task category based on at least one task sub-template corresponding to each task template. For example, a book review sub-template, a financial sentiment sub-template, etc., can be divided into the sentiment analysis task category. The right sub-diagram can also illustrate the generation of a hierarchical multi-template multi-task language dataset, i.e., a dataset of sample texts with a unified format.

[0030] In one embodiment of the present disclosure, the multi-task language dataset is structured supervised language data, and the method of dividing each task language dataset into at least one task category based on at least one task sub-template corresponding to each task language dataset can be based on experience and knowledge.

[0031] In step S103, a language model is pre-trained based on the hierarchical multi-template multi-task language dataset.

[0032] Here, in one embodiment of the present disclosure, a method for pre-training a language model based on the hierarchical multi-template multi-task language dataset includes a step of realizing hierarchical modeling by splicing a continuous template in front of a sample text, and the step of pre-training a language model based on the hierarchical multi-template multi-task language dataset specifically includes: step a) acquiring sample text from the language model; step b) acquiring a task template and a task sub-template corresponding to the sample text based on a task category to which the sample text belongs; step c) generating a continuous template based on the task template and the task sub-template corresponding to the sample text; and step d) inputting the sample text and the continuous template into the language model to pre-train the language model.

[0033] In one embodiment of the present disclosure, the language model may be generated by training it using multi-granularity unsupervised language data.

[0034] Specifically, in one embodiment of the present disclosure, the language model can be trained from a large amount of unsupervised language data, using sample language data of words, sentences, paragraphs, and chapters, from fine to coarse granularity and from a first granularity to a fourth granularity, and the training of the language model is bidirectional training.

[0035] For example, in one embodiment of the present disclosure, FIG. 4 is a schematic diagram of sample language data at the first to fourth granularities provided by this embodiment. As shown in FIG. 4, the input is divided into bidirectionally encoded input (bold) and one-way decoded input (standard font), and the input content is partially modeled in two directions. In the diagram, M, S, and E represent mask characters, generation start characters, and end characters, respectively. The difference between word, sentence, paragraph, and chapter granularities lies in the mask characters (M). For example, in a word-granular bidirectional generation task, if the two words Harbin and Bing Xue are replaced with mask characters, the model needs to learn how to restore the mask characters in the bidirectionally encoded input (bold) by modeling the input.

[0036] This allows for joint training with unsupervised general knowledge and supervised task knowledge through the generative branch of the language model.

number

[0037] Here, x is a sample text with a total length of n, and y is a supervised dataset with a total length of m. The first half of the loss is optimized with unsupervised generic data, and the second half of the loss is optimized with supervised data. The language model is modeled uniformly, and the i-th character can only see the information of the previous 0 to i-1 characters. The characters from 0 to s can be seen from both directions, but the characters from s to i can be seen in one direction.

[0038] 5 is a schematic diagram illustrating pre-training a language model based on a hierarchical multi-template multi-task language dataset provided by an embodiment of the present disclosure. In one embodiment of the present disclosure, as shown in FIG. 5, the continuous template may be a learned vector, which can be input to the model together with sample text. Based on the task template and task sub-template corresponding to the sample text, the continuous template is jointly trained with unsupervised general knowledge and supervised task knowledge in the generation branch of the pre-training model, and then optimized.

[0039] After hierarchical multi-template multi-task pre-training is complete, the pre-trained model can have stronger transfer capabilities. Because the task sequence templates are trained together with the multi-task data, they also have stronger transfer capabilities. This allows the new hierarchical multi-template multi-task dataset to have zero-sample and few-sample transfer capabilities for data of the same task type. Similarly, the task templates and task subtemplates corresponding to the sequence templates can better guide the model to complete tasks corresponding to a specific dataset. Examples include question-answering templates and open-question-answering subtemplates, as shown in Figure 5.

[0040] Furthermore, to introduce hierarchical artificial prior knowledge, we assign N trainable word vectors (i.e., continuous prompts) to each task type and language dataset, which are spliced ​​before the original text to help the model learn hierarchical multi-task knowledge. During the training phase, the supervised optimization goal of the objective function for pre-training the language model can be modified to depend on the continuous prompts of the task and dataset. The specific modifications are as follows:

number

[0041] In a pre-training method according to an embodiment of the present disclosure, a pre-training language dataset is constructed, the pre-training language dataset including unsupervised language data and supervised language data, a layered multi-template multi-task language dataset is generated based on the pre-training language dataset, and a language model is pre-trained based on the layered multi-template multi-task language dataset. In this way, the embodiment of the present disclosure constructs a multi-template multi-task language dataset, and by collectively creating templates for tasks, the model can simultaneously model multiple task data. By setting multiple templates, the diversity of the model can be increased, improving the robustness of the model's task learning. When pre-training the model, successive templates can better learn knowledge related to the task and dataset, thereby improving the model's transfer ability when there are no samples or few samples.

[0042] In addition, based on constructing a pre-training language dataset, the present disclosure proposes joint training with unsupervised general knowledge and supervised task knowledge based on a language model, in which the language model can not only model task knowledge (based on supervised task knowledge), but also continuously learn general knowledge from training unsupervised data, thereby avoiding knowledge forgetting.

[0043] To realize the above embodiment, the embodiment of the present disclosure further provides a pre-training device. Figure 6 is a schematic configuration diagram of a language model pre-training device provided by the embodiment of the present disclosure.

[0044] As shown in FIG. 6 , the apparatus includes: a construction module 61 that constructs a pre-training language dataset, where the pre-training language dataset includes unsupervised language data and supervised language data; a generation module 62 that generates a hierarchical multi-template multi-task language dataset based on the pre-training language dataset; and a pre-training module 63 that pre-trains a language model based on the hierarchical multi-template multi-task language dataset.

[0045] As shown in FIG. 7, the generation module 62 includes a template setting submodule 621 that sets, for each task dataset in the multi-task dataset, a corresponding task template and at least one task sub-template corresponding to each task template, and a first generation submodule 622 that divides each task dataset into at least one task category based on the at least one task sub-template corresponding to each task dataset, to generate the hierarchical multi-template multi-task dataset.

[0046] As shown in FIG. 7, this generation module 62 comprises a second generation sub-module 623 that is used to generate the language model by training it with multi-granularity unsupervised language data.

[0047] As shown in FIG. 7, the generation module 62 includes an extraction submodule 624 that extracts sample language data of first to fourth granularity levels from the unsupervised language data, and a third generation submodule 625 that trains an initial model based on each of the sample language data of the first to fourth granularity levels to generate the language model.

[0048] As shown in FIG. 8 , the pre-training module 63 includes a sample acquisition submodule 631 that acquires sample text from the language model; a template acquisition submodule 632 that acquires a task template and a task subtemplate corresponding to the sample text based on the task category to which the sample text belongs; a continuous template generation submodule 633 that generates a continuous template based on the task template and the task subtemplate corresponding to the sample text; and a first pre-training submodule 634 that inputs the sample text and the continuous template into the language model to pre-train the language model.

[0049] As shown in FIG. 8, the pre-training module 63 further comprises a splicing sub-module 635 for splicing the continuous template before the sample text.

[0050] As shown in FIG. 8, the pre-training module 63 further includes a second pre-training sub-module 636 that jointly pre-trains the language model using the unsupervised language data and the supervised language data.

[0051] In an embodiment, the apparatus is used to jointly pre-train the pre-training model using the unsupervised data and the supervised data.

[0052] The description of the method embodiment above is also applicable to the device of this embodiment, and the principles are the same, so they will not be repeated here.

[0053] In an embodiment of the present disclosure, a language model pre-training device constructs a pre-training language dataset, which includes unsupervised language data and supervised data; generates a layered multi-template multi-task language dataset based on the pre-training language dataset; and pre-trains a language model based on the re-layered multi-template multi-task language dataset. In this way, the embodiment of the present disclosure constructs a multi-template multi-task language dataset and collectively templates tasks, allowing the model to simultaneously model multiple task data, thereby increasing the diversity of the model and improving the robustness of task learning for the model. When pre-training a model, successive templates can better learn knowledge related to tasks and datasets, thereby improving the transfer ability of the model when there are no samples or few samples.

[0054] In addition, based on constructing a pre-training language dataset, the present disclosure proposes joint training with unsupervised general knowledge and supervised task knowledge based on a language model, in which the language model can not only model task knowledge (based on supervised task knowledge), but also continuously learn general knowledge from training unsupervised data, thereby avoiding knowledge forgetting.

[0055] To achieve the above embodiments, the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable by the processor, which, when executed by the processor by the program, achieves the method described in the above method embodiments.

[0056] To realize the above embodiments, the present disclosure further provides a non-transitory computer-readable storage medium having computer instructions stored thereon, the computer instructions causing a computer to perform the method according to the above method embodiments.

[0057] To realize the above embodiments, the present disclosure further provides a computer program, and when instructions in the computer program are executed by a processor, the method according to the above method embodiments is realized.

[0058] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program.

[0059] 9 is a schematic block diagram of an exemplary electronic device 900 for implementing embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing devices, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the present application as described and / or claimed herein.

[0060] 9, the electronic device 900 includes a computing unit 901 that can perform various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 902 or loaded from a storage unit 908 into a random access memory (RAM) 903. The RAM 903 can further store various programs and data necessary for the operation of the electronic device 900. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0061] Multiple components in the electronic device 900 are connected to an I / O interface 905, including an input unit 906 such as a keyboard or mouse, an output unit 907 such as various displays and speakers, a storage unit 908 such as a magnetic disk or optical disk, and a communication unit 909 such as a network card, modem, wireless communication transceiver, etc. The communication unit 909 enables the electronic device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various carrier networks.

[0062] The computing unit 901 may be various general-purpose and / or special-purpose processing components having processing and computational capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, computing units that execute various machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs each of the methods and processes described above, such as the pre-training method. For example, in some embodiments, the pre-training method may be implemented as a computer software program tangibly embodied in a machine-readable medium such as the storage unit 908. In some embodiments, some or all of the computer program may be loaded and / or installed into the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, it may perform one or more steps of the pre-training method described above. Alternatively, in other embodiments, the computing unit 901 may be configured to perform the pre-training method in any other suitable manner (eg, firmware).

[0063] Various embodiments of the systems or techniques described herein may be implemented using digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard packages (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. Each of these embodiments may involve execution by one or more computer programs executed and / or interpreted by a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor, capable of receiving data and instructions from, and transferring data and instructions to, a storage system, at least one input device, and at least one output device.

[0064] Program code for implementing the methods of the present disclosure may be written in one or a combination of programming languages. The program code may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable human image restoration device such that, when executed by the processor or controller, the functions and operations specified in the flowcharts and / or block diagrams are performed. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a separate software package, or entirely on a remote machine or server.

[0065] In the context of the present disclosure, a machine-readable medium may be a tangible medium that contains or can store a program for use by or in combination with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. Examples of machine-readable media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, devices, or any suitable combination thereof. More specific examples of machine-readable storage media include one or more line-based electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0066] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) by which a user can provide input to the computer. Other types of devices can also be used to provide for interaction with a user; for example, the feedback provided to the user can be any form of sensing feedback (e.g., visual feedback, auditory feedback, or haptic feedback) and can receive input from the user in any form (including acoustic, speech, and tactile input).

[0067] The systems and techniques described herein can be implemented in a computing system with a back-end component (e.g., a data server), or a computing system with a middleware component (e.g., an application server), or a computing system with a front-end component (e.g., a user computer having a graphical user interface or a web browser through which a user interacts with embodiments of the systems and techniques described herein), or any combination of such back-end, middleware, and front-end components. The components of the system can be connected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0068] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship is established by computer programs running on corresponding computers. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product within the cloud computing service system. It overcomes the drawbacks of traditional physical hosts and VPS services (Virtual Private Servers, also abbreviated as "VPS"), such as difficulty in management and poor business scalability. The server may be a distributed system server or a blockchain-connected server.

[0069] It should be understood that steps may be rearranged, added, or deleted using the various forms of flow described above. For example, the steps described in this disclosure may be performed in parallel, sequentially, or in a different order as long as the desired results of the technical solutions disclosed in this application are achieved. This specification is not limited thereto.

[0070] The above specific embodiments do not limit the scope of protection of the present disclosure. It is understood that those skilled in the art can make various modifications, combinations, subcombinations, and substitutions according to design requirements and other factors. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application should be included within the scope of protection of the present disclosure.

Claims

1. A method comprising: a construction module constructing a pre-training language dataset, the pre-training language dataset including unsupervised language data and supervised language data; A generation module generates a hierarchical multi-template multi-task language dataset based on the pre-training language dataset; a pre-training module pre-training a language model based on the layered multi-template multi-task language dataset; Including, the supervised language data comprises a multi-task language dataset; generating a hierarchical multi-template multi-task language dataset based on the pre-training language dataset, a template setting sub-module setting, for each task language dataset in the multi-task language dataset, a corresponding task template and at least one task sub-template corresponding to each task template; a first generation sub-module classifying each of the task language datasets into at least one task category based on at least one task sub-template corresponding to each of the task language datasets to generate the hierarchical multi-template multi-task language dataset; A method for pre-training a language model, including:

2. The method of claim 1 , further comprising jointly pre-training the language model using the unsupervised language data and the supervised language data.

3. The method for pre-training a language model according to claim 1 , wherein the language model is generated by training with multi-granularity unsupervised language data.

4. The language model: an extraction sub-module extracting sample language data of first to fourth granularities from the unsupervised language data; The method for pre-training a language model according to claim 3, wherein a third generation sub-module trains an initial model based on each of the sample language data of the first to fourth granularities to generate the language model.

5. The method for pre-training a language model according to claim 4 , wherein the first to fourth granularities are a word granularity, a sentence granularity, a paragraph granularity, and a chapter granularity.

6. The method of claim 3 , wherein the training is bidirectional training.

7. The objective function for pre-training the language model is: [Equation 3] where x is a sample text of total length n, y is a supervised dataset of total length m, [Equation 4] The loss value of is optimized on the unsupervised generic data, [Equation 5] The method of claim 1 , wherein a loss value of

8. a construction module that constructs a pre-training language data set, the pre-training language data set including unsupervised language data and supervised language data; a generation module for generating a hierarchical multi-template multi-task language dataset based on the pre-training language dataset; a pre-training module that pre-trains a language model based on the layered multi-template multi-task language dataset; Equipped with the supervised language data comprises a multi-task language dataset; The generation module: a template setting submodule that sets, for each task language dataset in the multi-task language dataset, a corresponding task template and at least one task subtemplate corresponding to each task template; a first generating sub-module that divides each of the task language datasets into at least one task category based on at least one task sub-template corresponding to each of the task language datasets to generate the hierarchical multi-template multi-task language dataset; 1. A language model pre-training device comprising:

9. the pre-training module: The apparatus for pre-training a language model of claim 8 , further comprising a second pre-training sub-module that jointly pre-trains the language model using the unsupervised language data and the supervised language data.

10. The generation module: The language model pre-training device of claim 8 , further comprising a second generation sub-module used to generate the language model by training it with multi-granularity unsupervised language data.

11. The generation module: an extraction submodule that extracts sample language data of first to fourth granularities from the unsupervised language data; a third generation sub-module that trains an initial model based on each of the sample language data of the first to fourth granularities to generate the language model; The language model pre-training device of claim 9 , comprising:

12. The apparatus for pre-training a language model according to claim 11 , wherein the first to fourth granularities are a word granularity, a sentence granularity, a paragraph granularity, and a chapter granularity.

13. The apparatus for pre-training a language model according to claim 10, wherein the training is bidirectional training.

14. The objective function for pre-training the language model is: [Equation 6] where x is a sample text of total length n, y is a supervised dataset of total length m, [Equation 7] The loss value of is optimized on the unsupervised generic data, [Equation 8] The apparatus for pre-training a language model according to claim 8 , wherein a loss value of

15. at least one processor; a memory communicatively coupled to the at least one processor; Equipped with 8. An electronic device, the memory storing instructions for execution by the at least one processor, the instructions causing the at least one processor to perform the language model pre-training method of any one of claims 1 to 7.

16. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: A non-transitory computer-readable storage medium having computer instructions that cause a computer to perform the method for pre-training a language model of any one of claims 1 to 7.

17. A computer program comprising: A computer program product, the instructions of which, when executed by a processor, result in the method for pre-training a language model according to any one of claims 1 to 7 being implemented.

Citation Information

Patent Citations

  • Language model training method and device, electronic equipment and storage medium

    CN114036300A

  • Language model training method, apparatus, and equipment

    JP2018502344A

  • Event series extraction apparatus, event series extraction method, and event extraction program

    JP2019046304A

  • Method, device and storage medium for constructing a speech decoding network for digit speech recognition

    JP2019504355A

  • Method, device, and electronic apparatus for updating parameters of multi-task model

    JP2022028871A

Cited By

  • Dialogue model training method

    US12585885B2

  • Dialogue model training method

    US20240412002A1