Pre-training method, device and equipment of language model and storage medium
Patent Information
- Application Number
- CN202310232945.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-02
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-03-02
AI Technical Summary
[0012]本公开的实施例所提供的方法相比于传统的语言模型的预训练方法而言,能够对输入文本进行领域自适应的掩码处理,以更好地学习到文本中特定于领域的部分词语的特征表示,使得所训练的语言模型能够更加充分地学习到该领域的知识。
Smart Images

Figure CN117216254B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of natural language processing, and more specifically, to a method, apparatus, device, and storage medium for pre-training a language model. Background Technology
[0002] In recent years, the rapid development of pre-trained language models has significantly improved the overall performance of many Natural Language Processing (NLP) tasks, and has also enabled many applications to enter the practical implementation stage. Pre-trained language models are neural network language models, characterized by the ability to be trained using large-scale unlabeled plain text corpora and applicability to various downstream NLP tasks. For example, in NLP, a large amount of unlabeled language text can be used to pre-train an initial language model to obtain a task-independent pre-trained language model. Subsequently, this pre-trained language model can be fine-tuned based on a specific task (e.g., reading comprehension or entity recognition) to obtain a target language model for performing that specific task.
[0003] When pre-training language models for a specific domain, there are often significant differences between texts in specific domains (e.g., e-commerce, healthcare) and texts in general domains. Texts in specific domains often contain a large number of technical terms (and these technical terms are often at the entity level), which are also the core semantic part of the text. Therefore, learning the representations of these entity words is the key to pre-training language models for a specific domain.
[0004] Therefore, an effective language model pre-training method is needed to enable the language model to better learn the data features related to the domain. Summary of the Invention
[0005] To address the aforementioned issues, this disclosure employs a domain-adaptive masking technique on the input text for a specific domain, thereby enabling better learning of the feature representations of core entity words specific to that domain, and thus achieving pre-training of a domain-adaptive language model.
[0006] Embodiments of this disclosure provide a method, apparatus, device, and computer-readable storage medium for pre-training a language model.
[0007] Embodiments of this disclosure provide a pre-training method for a language model applied to a specific domain. The method includes: acquiring input text, the input text comprising one or more words, the one or more words comprising entity words; performing component identification on the input text to determine multiple entity words and their entity types in the input text, wherein the entity type of each entity word in the multiple entity words is related to the specific domain; performing masking processing on the input text based on the multiple entity words and their entity types, wherein a first number of entity words in the multiple entity words are assigned a higher masking probability, and the entity types of the first number of entity words belong to the core entity types of the specific domain, wherein the first number includes zero, one, or more; and for at least one masked word in the masked input text, predicting the at least one word using the language model, and training the language model based on the prediction result.
[0008] Embodiments of this disclosure provide a pre-training apparatus for a language model applied to a specific domain. The apparatus includes: a text acquisition module configured to acquire input text, the input text including one or more words, the one or more words including entity words; a component recognition module configured to perform component recognition on the input text to determine multiple entity words and their entity types in the input text, the entity type of each entity word being related to the specific domain; a masking module configured to perform masking processing on the input text based on the multiple entity words and their entity types, wherein a first number of entity words among the multiple entity words are assigned a higher masking probability, and the entity types of the first number of entity words belong to the core entity types of the specific domain, wherein the first number includes zero, one, or more; and a model training module configured to predict at least one masked word in the masked input text using the language model, and train the language model based on the prediction result.
[0009] Embodiments of this disclosure provide a language model pre-training device, comprising: one or more processors; and one or more memories, wherein the one or more memories store a computer-executable program, which, when executed by the processor, performs the language model pre-training method as described above.
[0010] Embodiments of this disclosure provide a computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by a processor, are used to implement the pre-training method for the language model as described above.
[0011] Embodiments of this disclosure provide a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a pre-training method for a language model according to embodiments of this disclosure.
[0012] Compared to traditional language model pre-training methods, the method provided by the embodiments of this disclosure can perform domain-adaptive masking processing on the input text to better learn the feature representations of domain-specific words in the text, enabling the trained language model to learn the knowledge of that domain more fully.
[0013] The method provided in the embodiments of this disclosure identifies domain-specific entity words in the input text and performs masking processing on the input text based on these entity words and their entity types. Specifically, entity words belonging to core entity types specific to that domain are masked with a higher masking probability. This allows the language model to better learn the representations of these core entity words through pre-training, enabling the trained language model to more fully learn the characteristics of domain-specific data. The method in the embodiments of this disclosure enables domain-adaptive language model pre-training, thereby enhancing the language model's learning ability in a specific domain and allowing the trained language model to be better applied to various task scenarios within that specific domain. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some exemplary embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0015] Figure 1 This is a schematic diagram illustrating a scenario of generating a pre-trained language model based on input text, suitable for various task scenarios, according to an embodiment of the present disclosure;
[0016] Figure 2 This is a flowchart illustrating a pre-training method for a language model according to an embodiment of the present disclosure;
[0017] Figure 3 This is a schematic flowchart illustrating an example process of a pre-training method for a language model according to an embodiment of the present disclosure;
[0018] Figure 4 This is a schematic diagram illustrating an example structure of a component identification model according to an embodiment of the present disclosure;
[0019] Figure 5 This is a schematic diagram illustrating an example training process of a language model according to an embodiment of the present disclosure;
[0020] Figure 6 This is a schematic diagram illustrating an example process of language model pre-training according to an embodiment of the present disclosure;
[0021] Figure 7 This is a schematic diagram illustrating a pre-training apparatus for a language model according to an embodiment of the present disclosure;
[0022] Figure 8 A schematic diagram of a pre-training device for a language model according to an embodiment of the present disclosure is shown;
[0023] Figure 9 A schematic diagram of the architecture of an exemplary computing device according to embodiments of the present disclosure is shown; and
[0024] Figure 10 A schematic diagram of a storage medium according to an embodiment of the present disclosure is shown. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of this disclosure more apparent, exemplary embodiments according to this disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments of this disclosure. It should be understood that this disclosure is not limited to the exemplary embodiments described herein.
[0026] In this specification and accompanying drawings, steps and elements that are substantially the same or similar are indicated by the same or similar reference numerals, and repeated descriptions of these steps and elements are omitted. Furthermore, in the description of this disclosure, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance or order.
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used herein is for the purpose of describing embodiments of the invention only and is not intended to limit the invention.
[0028] To facilitate the description of this disclosure, the following concepts related to this disclosure are introduced.
[0029] The pre-training method for the language model disclosed herein can be based on Artificial Intelligence (AI). Artificial intelligence utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to obtain optimal results—theories, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine capable of reacting in a manner similar to human intelligence. For example, the pre-training method for the language model based on artificial intelligence can learn domain-specific data features from input text in a way similar to how humans acquire knowledge of a specific domain by reading and processing text within that domain. By studying the design principles and implementation methods of various intelligent machines, artificial intelligence enables the pre-training method for the language model disclosed herein to better learn the representations of core entity words in a specific domain from input text, thereby allowing the pre-trained language model to more fully learn the domain-specific data features.
[0030] The pre-training method for the language model disclosed herein can be based on Natural Language Processing (NLP). NLP is an important field within computer science and artificial intelligence. It studies various theories and methods that enable effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language, i.e., the language people use in daily life, and thus it is closely related to linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs. In the pre-training method for the language model disclosed herein, NLP techniques can be used to process the input text to enhance the learning ability of the pre-trained language model in a specific domain, enabling the pre-trained language model to be better applied to various task scenarios within that specific domain.
[0031] The pre-training method for the language model disclosed herein can be based on a pre-trained language model (PTM). A language model (LM) is a model used in the field of NLP for analyzing and processing language text, and can generally be divided into grammatical rule-based language models, statistical language models, and neural network language models. As a type of neural network language model, pre-trained language models have seen rapid development, and researchers are no longer satisfied with training them using simple paradigms and simple corpora. This has led to a series of extended pre-trained language models. For example, these can include knowledge-enriched pre-trained models, multilingual / cross-lingual pre-trained models, language-specific pre-trained models, multimodal (including video-text, image-text, sound-text, etc.) pre-trained models, domain-specific pre-trained models, and task-specific pre-trained models. Pre-training tasks for language models can be categorized into three types based on learning paradigms: supervised learning, unsupervised learning, and self-supervised learning. Self-supervised learning also involves input and output labels, thus employing training methods consistent with supervised learning. However, in self-supervised learning, such as in the Masked Language Modeling (MLM) task, the labels are automatically generated rather than manually labeled as in supervised learning.
[0032] Therefore, as mentioned above, the pre-training method for the language model disclosed herein can also be based on the Masked Language Model (MLM) task. The MLM task can be categorized as a self-supervised learning problem and is typically treated as a classification problem. Traditional language models predict the current word from left to right or right to left using the preceding or following context, a design based on the initial idea of probabilistic language models. However, in reality, the occurrence of the current word depends not only on the preceding or following context but also on the surrounding context. Therefore, the MLM task was proposed to better utilize contextual information to encode the current word. The MLM language model structure involves randomly selecting a subset of words in a sentence, replacing each word with a mask, and then predicting these words during model training. The encoder output is then used as the feature representation of the language.
[0033] In summary, the solutions provided by the embodiments of this disclosure involve technologies such as artificial intelligence, natural language processing, and masked language modeling. The embodiments of this disclosure will be further described below with reference to the accompanying drawings.
[0034] Figure 1 This is a schematic diagram illustrating a scenario of generating a pre-trained language model based on input text, applicable to various task scenarios, according to an embodiment of the present disclosure.
[0035] like Figure 1 As shown, the server can receive a large amount of unlabeled language text input from terminal devices via the network and generate a pre-trained language model based on the input text. The generated pre-trained language model can be widely applied to various task scenarios. For example, the pre-trained language model can be fine-tuned based on a specific task (e.g., matching task, content understanding task, and classification task) to obtain a target language model for performing that specific task.
[0036] Optionally, the terminal device may specifically include smartphones, tablets, laptops, desktop computers, in-vehicle terminals, wearable devices, etc., but is not limited to these. For example, the terminal device may also be a client with a browser or various applications (including system applications and third-party applications) installed. Optionally, the network may be an Internet of Things (IoT) based on the Internet and / or telecommunications networks. It can be a wired network or a wireless network. For example, it may be an electronic network capable of information exchange, such as a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), or a cellular data communication network. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, which is not limited in this application. The server may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0037] Pre-trained language models (e.g., those based on language model embeddings, generative pre-training (GPT) models, and bidirectional encoder representations from Transformers (BERT) models) have achieved significant breakthroughs in many natural language understanding (NLU) tasks, including text classification, reading comprehension, and natural language inference. These models are trained on large-scale text corpora using autoencoding or autoregressive methods. In embodiments of this disclosure, the aforementioned pre-trained language models can be domain-specific pre-trained models. In the case of domain-specific language model pre-training, since there are usually significant differences between domain-specific (e.g., e-commerce, medical, etc.) texts and general domain texts, domain-specific texts often contain a large number of technical terms (and these technical terms are often entity-level), which are also the core semantic part of the text. Therefore, learning representations of these entity words is crucial for domain-specific language model pre-training.
[0038] In existing domain-specific language model pre-training, some studies have focused on masking granularity (e.g., word-level masking, continuous segment masking, and entity masking) tailored to the characteristics of Chinese text. These methods can learn rich semantic patterns from raw text data through MLM tasks, and can also enhance the learning ability of pre-trained models in specific domains through techniques such as word-level masking, continuous segment masking, and entity masking.
[0039] However, as mentioned above, texts in specific domains often contain a large number of technical terms, which constitute the core semantic part of the text in that domain. In existing pre-training of language models for specific domains, the core entity words that constitute the core semantics of the specific domain in the input text are often not distinguished from other words. Instead, when randomly masking the text / words in the input text, all text / words in the input text are masked with the same masking probability. In other words, these pre-training methods treat core entity words and other words as words with the same importance to the semantics of the text, without considering that various technical terms in a specific domain have a greater impact on the semantics of the text in that domain than other text / words.
[0040] Based on this, this disclosure provides a pre-training method for a language model, which achieves pre-training of a domain-adaptive language model by performing domain-adaptive masking on the input text for a specific domain, in order to better learn the feature representations of core entity words specific to the domain.
[0041] Compared to traditional language model pre-training methods, the method provided by the embodiments of this disclosure can perform domain-adaptive masking processing on the input text to better learn the feature representations of domain-specific words in the text, enabling the trained language model to learn the knowledge of that domain more fully.
[0042] The method provided in the embodiments of this disclosure identifies domain-specific entity words in the input text and performs masking processing on the input text based on these entity words and their entity types. Specifically, entity words belonging to core entity types specific to that domain are masked with a higher masking probability. This allows the language model to better learn the representations of these core entity words through pre-training, enabling the trained language model to more fully learn the characteristics of domain-specific data. The method in the embodiments of this disclosure enables domain-adaptive language model pre-training, thereby enhancing the language model's learning ability in a specific domain and allowing the trained language model to be better applied to various task scenarios within that specific domain.
[0043] Figure 2 This is a flowchart illustrating a pre-training method 200 for a language model according to an embodiment of the present disclosure. Figure 3 This is a schematic flowchart illustrating an example process of a pre-training method for a language model according to an embodiment of the present disclosure.
[0044] In the pre-training method of the language model disclosed herein, the obtained pre-trained language model can be a domain-specific language model. That is, the language model can be applied to a specific domain and performs well in various downstream task scenarios within that specific domain. To pre-train a language model for a specific domain to obtain a pre-trained language model that performs better in various task scenarios within that specific domain, it is necessary to consider the differences between text in the specific domain and the general domain during the pre-training process. This involves learning the characteristics of various word data specific to the domain (e.g., domain-specific terminology) so that the generated pre-trained language model possesses more knowledge about that specific domain, thereby enabling it to be better applied to various task scenarios within that specific domain.
[0045] Therefore, as Figure 3As shown, in the pre-training method of the language model disclosed herein, component recognition can be performed on the input text based on entity granularity to determine each entity word in the input text and the entity type of each entity word. Then, based on domain-specific masking processing, the masked words in the masked text can be predicted through the MLM task, so that the language model can learn the data features of the domain more fully.
[0046] Below, we will refer to Figure 3 and Figures 4-6 The following describes in detail the specific processing steps of the pre-training method for the language model disclosed herein. To facilitate understanding of the pre-training process, the following description uses the e-commerce domain as a specific example; however, it should be understood that this should not be construed as a limitation on the language model pre-training method of this disclosure, which can also be applied to various other domains and their downstream task scenarios.
[0047] In step 201, input text can be obtained, which may include one or more words, and the one or more words may include entity words.
[0048] As mentioned above, text data for language model pre-training can be obtained from various terminal devices. This input text can consist of one or more words. Considering that various professional terms in domain-specific texts are often entity-level, these words can be divided into two main categories: entity words and non-entity words. Generally, words that are objectively distinguishable from each other can be considered entity words. Entity words can be concrete objects (e.g., products) or abstract concepts and relationships (e.g., product design styles). The key to distinguishing between entity words and non-entity words is that one entity word can be distinguished from another, while non-entity words cannot be objectively distinguished from each other.
[0049] Furthermore, each entity term can have a corresponding entity type, and entity terms with the same entity type can have the same characteristics and properties. For example, in the e-commerce field, under the entity type "product color," there can be multiple entity terms, such as "red," "yellow," and "blue." These entity terms are all used to describe the color attribute characteristics of the product, and these entity terms can be distinguished from each other based on the specific color attribute they describe.
[0050] Optionally, each domain may include multiple entity types. That is, for an input text, the entity type of each word included in the input text can be determined based on multiple entity types under a specific domain. The entity type of each determined word can be an entity type under that specific domain, or it can be another type that does not belong to any entity type under that specific domain. This will be explained below with reference to step 202.
[0051] Therefore, in order to mask the input text at the entity level to learn the data features of domain-specific terminology, it is first necessary to determine the word type of each word in the input text in order to determine the potential semantic impact of each word on the input text, so as to perform entity-granular masking of the input text based on the potential impact of each word on the input text.
[0052] In step 202, component recognition can be performed on the input text to determine multiple entity words and their entity types in the input text, wherein the entity type of each entity word is related to the specific domain.
[0053] Optionally, the word type of each word in the input text can be determined by performing component recognition on the input text, such as... Figure 3 As shown, component identification of input text can include three parts: word segmentation, feature representation, and word classification. Specifically, according to embodiments of this disclosure, component identification of the input text can include: performing word segmentation on the input text to determine one or more words in the input text; obtaining a vector representation for each of the determined one or more words; determining a feature representation for the word based on the vector representation; and determining the entity type of the word using a classifier based on the feature representation.
[0054] According to embodiments of this disclosure, the above-described component recognition of the input text can be performed using a pre-trained component recognition model, wherein the component recognition model can be pre-trained for the specific domain. Figure 4 This is a schematic diagram illustrating an example structure of a component identification model according to an embodiment of the present disclosure. Wherein, Figure 4The BERT model is used as an example encoder for a pre-trained language model. This BERT model may include word embedding layers, an encoder layer, and a classifier layer. The specific processing of the component recognition part is described below based on the BERT model. However, it should be understood that this disclosure uses the BERT model as an example only and not as a limitation; other models that can achieve the same purpose can also be applied to the language model pre-training method of this disclosure. Specifically, the BERT model is a language model trained in an unsupervised manner using a large amount of unlabeled text. Its architecture is the encoder part of the Transformer model. As an encoder model, the BERT model can achieve bidirectional reading through a self-attention mechanism. Its goal is to obtain a semantic representation of the input text containing rich semantic information by training on a large-scale unlabeled corpus.
[0055] Optionally, for input text used for language model pre-training, to determine the word type of each word in the input text, we can first determine each word in the input text, that is, perform word segmentation on the input text. The purpose of word segmentation is to convert the input text into data that the model can process, because the model can only process numbers, so it is necessary to convert the text input into numerical data. For example, the input text can be segmented by a tokenizer, splitting the original input text into individual words to determine the vocabulary corresponding to the input text, and determining a numerical representation for each word, that is, determining the mapping of each word to the corresponding identifier (ID) in the vocabulary.
[0056] like Figure 4 As shown, before inputting the text into the BERT model, a single sentence can be used as input text. Tokenization is then used to determine individual words (tokens), such as Tok1, Tok2, ..., TokN. The [CLS] flag indicates classification, which can be used for subsequent classification tasks.
[0057] Of course, it should be understood that although various processing methods are described at the word level in this paper, processing such as word segmentation of input text can also be performed at other levels of granularity, such as character level, without affecting the implementation of the language model pre-training of this disclosure. This disclosure does not limit this. Similarly, this disclosure does not limit the specific word segmentation method; various word segmenters that can achieve the same effect can be applied to the pre-training method of this disclosure.
[0058] Next, to learn the feature attributes of the input text, we can represent each word in the input text using fixed-dimensional vectors; that is, we convert each word in the input text into a fixed-dimensional vector form. Therefore, based on the vector representation of each word in the input text, a neural network can take the vector representation of each word as input and output the semantic representation corresponding to that vector representation, i.e., the feature representation of the word. Typically, it is desirable for characters / words with similar semantics to be relatively close in distance in the feature vector space; therefore, the feature representation derived from the vector representation of characters / words can contain more accurate semantic information.
[0059] Optionally, such as Figure 4 As shown, the BERT model can transform each word in the input text into a fixed-dimensional vector form (e.g., E) in the word embedding layer. [CLS] E1, E2, ..., E N Then, in the encoder layer, a fixed-length string representing the vector of each word is taken as input and computed from bottom to top through multiple Transformer layers. Each Transformer layer uses a self-attention mechanism, and its computation result is passed to the next encoder layer through a feedforward neural network (FFN). For example, the encoder layer of the BERT model can include 12 Transformer layers, where each Transformer layer can consist of a self-attention layer and a feedforward neural network layer. The self-attention layer can compute the attention weights between pairs of words (tokens) in the input text, while the feedforward neural network layer can contain two fully connected layers. In addition, each Transformer layer can also include residual connections and layer normalization operations. Of course, it should be understood that the above BERT model structure is only an example in this disclosure, and other structures that can achieve the same purpose can also be applied to the pre-training method of this disclosure, and this disclosure does not limit them.
[0060] Therefore, based on the feature representation of each word in the obtained input text, a classifier can predict the label (i.e., entity type) corresponding to each word. For example... Figure 4 As shown, the feature representation of each word in the input text is obtained through the word embedding layer and the encoder layer (e.g., C, T1, T2, ..., T). N After that, the entity type label of each word can be input through the classifier layer, where the classifier layer can be a neural network structure such as a multilayer perceptron (MLP), and this disclosure does not limit it.
[0061] Optionally, for pre-training a language model for a specific domain, the aforementioned component recognition model can be pre-trained for that specific domain. This component recognition model can possess information about entity types within that specific domain, which can, for example, be predetermined and stored on a server. In other words, for any specific domain, the entity types it includes can be predetermined, and words can be classified based on these predetermined instance types during subsequent pre-training of the language model for that specific domain.
[0062] Therefore, the pre-training method for the language model of this disclosure may further include pre-training the component recognition model for the specific domain. Specifically, according to embodiments of this disclosure, pre-training the component recognition model for the specific domain may include: acquiring training text related to the specific domain, the training text including multiple entity words with pre-labeled entity types; predicting the entity type of each entity word in the training text using the component recognition model; and training the component recognition model based on the prediction results of the entity type of each entity word in the training text and its pre-labeled entity type.
[0063] Optionally, the training data used to train the component recognition model for a specific domain can have known entity type labels, for example, through manual annotation. Therefore, the training of the component recognition model can be supervised using the pre-annotated entity type labels for each word, i.e., the component recognition model can be trained based on the entity type label prediction loss function of the component recognition model.
[0064] For example, the original input text used to train the component recognition model is: x = (x1, ..., x2) m , ..., x M After passing through the word segmenter, the IDs of each word in the input text can be determined: t = (t1, ..., t2). m , ..., t M Then, through the word embedding layer and encoder layer of the BERT model, the encoded feature vector corresponding to each word in the input text can be obtained: e = (e1, ..., e2) m , ..., e M The probability of correctly obtaining the entity type label for each word can be determined by the classifier layer as p. t Therefore, the entity type label prediction loss function L of this component recognition model... entity It can be represented as: L entity =-∑ t y t log(p t ), where y tThis represents the entity type label of the t-th word (token). By optimizing the prediction loss function of the component recognition model, the parameters in the component recognition model used for pre-training the language model can be determined.
[0065] Therefore, based on the above description, it can be seen that each word in the input text does not necessarily have a corresponding instance type label output in a specific domain. This is because not every word in the input text is necessarily related to a specific domain; there may even be words belonging to non-entity types, i.e., non-entity words. In other words, the one or more words may also include non-entity words. Optionally, if some words in the input text do not belong to any pre-determined entity type in a specific domain, the output labels of these words through the component recognition model can be other types of labels, and these other types of labels can include labels such as "other entity types" or "non-entity types" based on whether the word is an entity word.
[0066] Next, back to Figure 2 After identifying domain-specific entity words and their entity types (as well as other words and their corresponding type labels) in the input text through component recognition, the pre-training process of the language model can be described below for the MLM task of the BERT model. Specifically, the MLM task may involve randomly selecting a subset of words from the input text and replacing them with a mask (i.e., masking the words). Then, during model training, the masked words are predicted. Optionally, the BERT model can be used as the encoder, and this pre-training process may include two parts: data processing and model training. Specifically, the data processing part may involve masking the input text based on each word and its entity type label, while the model training part may involve training the language model based on the masked input text. These two parts will be described in detail in steps 203 and 204, respectively.
[0067] In step 203, the input text can be masked based on multiple entity words and their entity types in the input text. A first number of entity words among the multiple entity words can be set to have a higher masking probability, and the entity type of the first number of entity words belongs to the core entity type of the specific domain. The first number can include zero or one or more.
[0068] Optionally, as mentioned above, entity types in a specific domain can define certain entity words that have a significant semantic impact on the domain. These words constitute the core semantic part of the text in that domain. Therefore, in the language model pre-training method disclosed herein, entity words under these entity types can be treated differently from other words to train a pre-trained language model more suitable for performing tasks in that specific domain. Furthermore, among all entity types in that specific domain, these entity types can be further divided based on the importance of each entity type's semantic impact on the text in that domain. As an example, all entity types in a specific domain can also be divided into core entity types and non-core entity types, where core entity types can be considered entity types that have a greater semantic impact on the text in that domain compared to non-core entity types. For example, in the aforementioned e-commerce domain, the entity types that can be predetermined in this domain may include, but are not limited to, product brands, product styles, target customer genders, specific products, and product colors. Among these numerous entity types, some entity types, including product brands and specific products, have higher importance in this specific domain than other entity types, and they can be classified as core entity types in this e-commerce domain. Optionally, the entity types and core entity types in a specific domain can be determined, for example, by human setting or by various neural network models. This disclosure does not limit the methods for obtaining these entity types, and all methods that can achieve this purpose can be applied to the pre-training method of this disclosure.
[0069] According to embodiments of this disclosure, masking the input text based on multiple entity words and their entity types may include: setting mask probabilities for all words in the input text based on the multiple entity words and their entity types; and masking all words in the input text based on the set mask probabilities.
[0070] Optionally, the masking process for the input text can be performed based on the masking probability for each word in the input text. Therefore, before performing the masking process, the masking probability settings can be determined based on each entity word and its entity type label in the determined input text.
[0071] Based on the above description, in the masking processing of input text, unlike existing methods that apply the same mask probability (e.g., 15%) to all characters / words / entities, the pre-training method of the language model disclosed herein can employ different dynamic mask probability settings in MLM tasks. That is, it adaptively sets the mask probability based on the importance of the entity type of each word in the input text to the semantics of the text in a specific domain. For example, for the input text and a defined domain, a higher mask probability can be set for some characters / words / entities that may have a greater impact on the semantics of the text, so as to learn more data features from these characters / words / entities, thereby learning more knowledge in that specific domain.
[0072] According to embodiments of this disclosure, setting a mask probability for all words in the input text may include: for entity words whose entity type belongs to the core entity type of the specific domain, setting a higher mask probability compared to other words. Optionally, the core entity type of the specific domain may have the highest mask probability among all words included in the input text. Of course, if the first quantity is zero, it can be assumed that the current input text does not include entity words belonging to the core entity type of the specific domain. Therefore, in this case, the mask probability can be set differently for entity words belonging to the entity type of the specific domain from other words.
[0073] Furthermore, setting mask probabilities for all words in the input text may also include: setting lower mask probabilities for non-entity words and entity words unrelated to the specific domain in the input text compared to other words. Optionally, for entity words in the input text that do not belong to any entity type under the specific domain, and for non-entity words in the input text, lower mask probabilities may be set for these words.
[0074] According to embodiments of this disclosure, masking all words in the input text based on a set masking probability may include: determining multiple words in the input text to be masked based on the set masking probability; and performing masking processing on each of the determined multiple words to be masked using a predetermined masking configuration.
[0075] Optionally, after adaptively setting mask probabilities for a specific domain based on the semantic importance of the entity types of each word in the input text to the text in that specific domain, multiple words to be masked can be selected from the input text based on these mask probability settings. The entity types of these words can be one of the following: an entity type within the specific domain, an entity type not belonging to any entity type within the specific domain, or a non-entity type. Therefore, the selected multiple words to be masked can be masked according to a predetermined mask configuration.
[0076] Optionally, the predetermined mask configuration may include probability settings for each of a variety of masking processes. These various masking processes may include, but are not limited to, performing a mask, preserving the original word, or replacing it with another word. Therefore, according to embodiments of this disclosure, the predetermined mask configuration may include: a first probability for masking, a second probability for preserving the original word, and a third probability for replacing it with another word.
[0077] According to embodiments of this disclosure, for each of the determined plurality of words to be masked, masking with a predetermined masking configuration may include: for the word, determining a masking process for the word based on the predetermined masking configuration, wherein the probability that the masking process is masking is a first probability, the probability that the masking process is preserving the original word is a second probability, and the probability that the masking process is replacing the word with another word is a third probability; and applying the determined masking process to the word.
[0078] Optionally, the aforementioned pre-defined masking configuration can be applied to each of the selected words to be masked. That is, each selected word can further determine its specific masking processing based on a probability configuration, preventing the language model from pre-determining the words to be masked. Since every word in the input text may be masked or replaced, the trained language model, when encoding the current word, does not overly rely on the current word itself, but rather considers its context, and may even perform error correction based on its context.
[0079] Therefore, through the above data processing of the training data, the masked input text can be input into the language model to be trained in order to train the language model.
[0080] In step 204, for at least one masked word in the masked input text, the language model can be used to predict the at least one word, and the language model can be trained based on the prediction result.
[0081] Optionally, the input text used for language model pre-training can come from a large amount of unsupervised text data in a specific domain. Each input text in this text data can generate a corresponding masked input text through the data processing described in steps 201-203 above. The masked input text may include at least one masked word.
[0082] According to embodiments of this disclosure, predicting at least one masked word in the masked input text using the language model may include: determining the original word corresponding to the masked word based on the masked input text using the language model. Optionally, the language model can be used to predict the masked word, where the language model may be, for example, a BERT model. Therefore, the language model can output a predicted value for the original word corresponding to the masked word.
[0083] According to embodiments of this disclosure, training the language model based on the prediction results may include: optimizing the parameters in the language model based on the probability that the language model obtains the original word corresponding to at least one masked word based on the masked input text.
[0084] Optionally, in the MLM task, the masked words can be represented by [MASK] tags. These [MASK] tags, as part of the language model's input, can be encoded into corresponding word vectors through a word embedding layer. These word vectors, in turn, can be passed through a Transformer encoder layer to obtain corresponding feature vectors. Due to the Transformer's attention mechanism, these feature vectors can include all contextual information of the [MASK] tags. Therefore, a classifier (e.g., a softmax classifier) can output the probability of obtaining the original word corresponding to at least one masked word based on the masked input text. In the MLM task, to accurately predict the masked words, the loss between the predicted and actual values should be as small as possible; that is, the loss for calculating the original word corresponding to the masked word should be as small as possible.
[0085] For example, given the original input text: x = (x1, ..., x...) m , ..., x M After word segmentation, the resulting word ID can be represented as: t = (t1, ..., t2) m , ..., t M The masked text is: t ~ =(t1,...,[MASK],...,t M ), where tm For at least one word that is masked (corresponding to t) ~ The [MASK] marker part in the text, therefore, the loss function of all words pre-masked by the language model can be expressed as, for example: L MLM =logP(t) m |t ~ ), where P(t) m |t ~ ) represents the probability of obtaining the original word corresponding to at least one masked word based on the masked input text.
[0086] Figure 5 This is a schematic diagram illustrating an example training process of a language model according to an embodiment of the present disclosure.
[0087] like Figure 5 As shown, for the masked input text "[CLS]Anta...[mask][mask]...blue small code[SEP]", the language model can determine the original word prediction for the "[mask][mask]" part, for example, in Figure 5 The word "T-shirt" was predicted, therefore, it can be based on the loss function L. MLM =logP(t) m =T-shirt|t ~ =[CLS]Anta...[mask][mask]...blue small, code[SEP]) optimization to optimize the parameters of the language model. Among them, the input text is first processed by word segmentation before being input into the language model, and two special tags [CLS] and [SEP] can also be inserted into the word-segmented input text to indicate the beginning and end of the input text, respectively.
[0088] Below, we will refer to Figure 6 This paper introduces the pre-training method of the language model disclosed herein for example processing of input text in the e-commerce domain. Figure 6 This is a schematic diagram illustrating an example process of language model pre-training according to an embodiment of the present disclosure.
[0089] like Figure 6 As shown, it is similar to Figure 3 Correspondingly, for the original input text, for example, in the e-commerce field, the original input text is "Anta 2022 New Fashion Men's T-shirt Blue Small Size". Similarly, component recognition is performed on the input text based on entity granularity to output each entity word in the input text and its entity type. For example, for the above original input text, the corresponding component recognition output could be "Anta". mainBrand 2022 time New Fashion style malecrowdGender t-shirt mainProduct blue color Small size size The entity type of each entity term is indicated by a superscript. Specifically, for the entity term "Anta," its entity type is "mainBrand"; for the entity term "2022," its entity type is "time"; for the entity term "fashion," its entity type is "style"; for the entity term "men," its entity type is "crowdGender"; for the entity term "T-shirt," its entity type is "mainProduct"; for the entity term "blue," its entity type is "color"; and for the entity term "small size," its entity type is "size." Here, "mainBrand" can represent the brand, "time" can represent the product launch date, "style" can represent the product style, "crowdGender" can represent the target customer's gender, "mainProduct" can represent the specific product, "color" can represent the product color, and "size" can represent the product size. For the term "new model," which is not an entity, its output entity type can optionally be empty to indicate that it belongs to a non-entity type. Of course, the above entity types are only examples of entity types in a specific domain, exemplified by the e-commerce field, and are not limitations. The specific domain and the entity types included in this disclosure are not limited to these.
[0090] Next, domain-specific masking can be performed on the aforementioned entity words.
[0091] For example, in the e-commerce field mentioned above, it can include entity types such as product brand, product launch time, product style, target customer gender, specific product, product color, and product size. Considering that product brand and specific product often have a significant impact on semantics in various task scenarios within e-commerce, they can be considered part of the core entity types in e-commerce. Therefore, for entity words under core entity types such as product brand and specific product, a higher masking probability can be set for them in the masking process. For example, for the component recognition result "Anta" mentioned above... mainBrand 2022 time New Fashion style male crowdGender t-shirt mainProduct blue color Small size sizeSince the entity types of the words "Anta" and "T-shirt" are product brand and specific product, respectively (i.e., core entity types), the highest masking probability can be assigned to these two entity words for this input text. Optionally, other entity words besides those under the aforementioned core entity types (e.g., "2022" and "fashion") can have the second highest masking probability. Furthermore, non-entity words, such as "new style," can have the lowest masking probability among all words in the input text. As an example, for the component recognition result "Anta..." mainBrand 2022 time New Fashion styl e male crowdGender t-shirt mainProduct blue color Small size size The mask probability setting can be expressed as: [0.2, 0.1, 0.05, 0.1, 0.1, 0.2, 0.1, 0.1].
[0092] Therefore, based on the above component identification results and mask probability settings, masking processing of the input text can be performed according to a predetermined mask configuration. For example, for the above component identification result "Anta", mainBrand 2022 time New Fashion style male crowdGende r T-shirt mainProduct blue color Small size size Based on the masking probability settings [0.2, 0.1, 0.05, 0.1, 0.1, 0.2, 0.1, 0.1], the words selected for masking can be determined to be, for example, "fashion", "T-shirt", and "small size". Therefore, assuming the predetermined masking configuration is: a first probability of 80%, a second probability of 10%, and a third probability of 10%, the masked text generated according to this predetermined masking configuration could be, for example, "[cls]Anta 2022 New Fashion Men's [mask][mask]Blue Chasing Wind [sep]", where "fashion" is retained, "T-shirt" is masked, and "small size" is replaced with a random word. Finally, the masked words in the masked text can be predicted using an MLM task. For example, the language model can predict the masked word as "T-shirt", accurately predicting the masked word.
[0093] As described above, the language model pre-training method disclosed herein can generate a pre-trained language model for a specific domain, that is, a pre-trained language model that integrates knowledge of that specific domain. This model can be applied to various downstream task scenarios in that specific domain to enhance the task performance of these task scenarios.
[0094] Taking the e-commerce field as an example, pre-trained language models that integrate knowledge from the e-commerce field can be applied to various task scenarios such as product search (involving matching tasks), product recommendation (involving product representation learning), product knowledge base construction (involving product content understanding), and product page navigation (involving classification tasks). Among these, the fine-tuned language model based on the pre-trained language model can greatly enhance the performance of these tasks.
[0095] For example, the language model pre-training method disclosed herein can also be applied to the e-commerce field, specifically to the product search task and the Named Entity Recognition (NER) task in this field. The product search task can use the relevance (i.e. accuracy) between the recalled product results and the product search text as a performance metric, while the NER task is similar to the masked word prediction using the language model mentioned above, so its performance metric can be the probability of named entity recognition.
[0096] Specifically, Table 1 below shows the performance of the pre-trained language model of this disclosure and various existing language models in the two types of tasks mentioned above.
[0097] Table 1 shows the task performance in the e-commerce field.
[0098]
[0099] As shown in Table 1 above, compared with the Google BERT model, the RoBERTa model and the basic e-commerce BERT model, the pre-trained language model obtained by the pre-training method disclosed in this paper has achieved significant performance improvements in the above task scenarios. Among them, the basic e-commerce BERT model refers to the pre-trained language model in the e-commerce domain that is trained using e-commerce data only through the MLM task.
[0100] Therefore, as described above, the pre-training method of the language model disclosed herein, by performing domain-adaptive masking on the input text, can better learn the feature representations of domain-specific words in the text, enabling the trained language model to learn more fully the knowledge of that domain. Through the above pre-training method, domain-adaptive language model pre-training can be achieved, thereby enhancing the language model's learning ability in a specific domain, so that the trained language model can be better applied to various task scenarios within that specific domain.
[0101] Figure 7 This is a schematic diagram illustrating a pre-training apparatus 700 for a language model according to an embodiment of the present disclosure.
[0102] According to embodiments of this disclosure, the pre-training device 700 for the language model may include a text acquisition module 701, a component recognition module 702, a mask processing module 703, and a model training module 704.
[0103] In the embodiments of this disclosure, the obtained pre-trained language model can be a domain-specific language model, meaning that the language model can be applied to a specific domain and performs well in various downstream task scenarios within that specific domain. To pre-train a language model for a specific domain to obtain a pre-trained language model that performs better in various task scenarios within that specific domain, it is necessary to consider the differences between text in the specific domain and the general domain during the pre-training process, and to learn the characteristics of various word data specific to the domain (e.g., domain-specific terminology), so that the generated pre-trained language model has more knowledge about that specific domain, thereby better applying it to various task scenarios within that specific domain.
[0104] The text acquisition module 701 can be configured to acquire input text, which includes one or more words, and the one or more words include entity words. Optionally, the text acquisition module 701 can perform the operations described above with reference to step 201.
[0105] For example, text data for language model pre-training can be obtained from various terminal devices. This input text may consist of one or more words, and considering that various professional terms in domain-specific texts are often entity-level, these words can be divided into two main categories: entity words and non-entity words. Furthermore, each entity word can have a corresponding entity type, and entity words with the same entity type can have the same characteristics and properties. For example, in the e-commerce field, under the entity type "product color," multiple entity words can be included, such as "red," "yellow," and "blue." These entity words are all used to describe the color attribute characteristics of the product, and these entity words can be distinguished from each other based on the specific color attribute they describe.
[0106] Optionally, each domain can include multiple entity types. That is, for an input text, the entity type of each word included in the input text can be determined based on multiple entity types under a specific domain. The entity type of each determined word can be an entity type under that specific domain, or it can be another type that does not belong to any entity type under that specific domain.
[0107] Therefore, in order to mask the input text at the entity level to learn the data features of domain-specific terminology, it is first necessary to determine the word type of each word in the input text in order to determine the potential semantic impact of each word on the input text, so as to perform entity-granular masking of the input text based on the potential impact of each word on the input text.
[0108] The component recognition module 702 can be configured to perform component recognition on the input text to determine multiple entity words and their entity types in the input text, wherein the entity type of each entity word is related to the specific domain. Optionally, the component recognition module 702 can perform the operations described above with reference to step 202.
[0109] Optionally, the word type of each word in the input text can be determined by performing component recognition on the input text. Component recognition on the input text can include three parts: word segmentation, feature representation, and word classification. These processes can be performed using a pre-trained component recognition model, which can be pre-trained for a specific domain.
[0110] Optionally, the purpose of word segmentation includes converting the input text into data that the model can process, since the model can only process numbers, so it is necessary to convert the text input into numerical data. For example, the input text can be segmented by a word segmenter, splitting the original input text into individual words to determine the vocabulary corresponding to the input text, and determining a numerical representation for each word, that is, determining the mapping of each word to the corresponding identifier (ID) in the vocabulary.
[0111] Next, regarding feature representation learning, to learn the feature attributes of the input text, we can represent each word in the input text using fixed-dimensional vectors; that is, we convert each word in the input text into a fixed-dimensional vector form. Therefore, based on the vector representation of each word in the input text, a neural network can take the vector representation of each word as input and output the semantic representation corresponding to that vector representation, i.e., the feature representation of the word. Typically, it is desirable that characters / words with similar semantics are also relatively close in distance in the feature vector space; therefore, the feature representation derived from the vector representation of characters / words can contain more accurate semantic information.
[0112] Finally, for word classification, the label (i.e., entity type) corresponding to each word can be predicted by a classifier based on the feature representation of each word in the obtained input text. For example, after obtaining the feature representation of each word in the input text through the word embedding layer and the encoder layer, the entity type label of each word can be input through the classifier layer. The classifier layer can be a neural network structure such as a multilayer perceptron (MLP), which is not limited in this disclosure.
[0113] Therefore, in addition to modules 701-704, the pre-training device for the language model of this disclosure may also include a module for pre-training the component recognition model for the specific domain. Optionally, the training data for training the component recognition model for the specific domain can have known entity type labels, for example, through manual annotation. Therefore, the training of the component recognition model can be supervised using the pre-annotated entity type labels of each word, i.e., the component recognition model can be trained based on the entity type label prediction loss function of the component recognition model. Therefore, each word in the input text does not necessarily have a corresponding instance type label output in the specific domain, because not every word in the input text is necessarily related to the specific domain; there may even be words belonging to non-entity types, i.e., non-entity words. In other words, the one or more words may also include non-entity words. Optionally, if some words in the input text do not belong to any pre-determined entity type in the specific domain, the output labels of these words through the component recognition model can be other type labels, and these other type labels can include labels such as "other entity type" or "non-entity type" based on whether the word is an entity word.
[0114] Next, after the component recognition module 702 determines the domain-specific entity words and their entity types (as well as other words and their corresponding type labels) in the input text, the pre-training process of the language model can be introduced for the MLM task of the BERT model.
[0115] The masking module 703 can be configured to mask the input text based on multiple entity words and their entity types in the input text. A first number of entity words among the multiple entity words are assigned a higher masking probability, and the entity types of the first number of entity words belong to the core entity types of the specific domain. The first number may include zero, one, or more. Optionally, the masking module 703 can perform the operations described above with reference to step 203.
[0116] Optionally, the masking process for the input text can be performed based on the masking probability for each word in the input text. Therefore, before performing the masking process, the masking probability settings can be determined based on each entity word and its entity type label in the determined input text.
[0117] The aforementioned entity types in a specific domain can define certain entity words that have a significant semantic impact on that domain. These words constitute the core semantic part of the text in that domain. Therefore, in the language model pre-training method disclosed herein, entity words under these entity types can be treated differently from other words to train a pre-trained language model more suitable for performing tasks in that specific domain. Furthermore, among all entity types in that specific domain, these entity types can be further divided based on the importance of each entity type's influence on the semantics of the text in that domain. As an example, all entity types in a specific domain can also be divided into core entity types and non-core entity types, where core entity types can be considered entity types that have a greater influence on the semantics of the text in that domain compared to non-core entity types.
[0118] Therefore, based on the above description, in the masking processing of input text, the pre-training method of the language model disclosed herein can employ dynamic mask probability settings in MLM tasks. That is, it adaptively sets the mask probability for a specific domain based on the importance of the entity type of each word in the input text to the semantics of the text in that specific domain. For example, for the input text and a defined domain, a higher mask probability can be set for characters / words / entities that may have a greater impact on the semantics of the text, in order to learn more data features from these characters / words / entities, thereby learning more knowledge of that specific domain. For example, words belonging to the core entity type of a specific domain can have the highest mask probability among all words included in the input text. Furthermore, for entity words in the input text that do not belong to any entity type under that specific domain, and for non-entity words in the input text, a lower mask probability can be set for these words.
[0119] Therefore, after adaptively setting masking probabilities for a specific domain based on the semantic importance of the entity types of each word in the input text to the text within that domain, multiple words to be masked can be selected from the input text based on these masking probability settings. The entity types of these words can be one of the following: an entity type within the specific domain, an entity type not belonging to any entity type within the specific domain, or a non-entity type. Therefore, the selected multiple words to be masked can be masked according to a predetermined masking configuration. For example, the predetermined masking configuration can include probability settings for each of several masking processes. These multiple masking processes can include, but are not limited to, performing masking, preserving the original words, or replacing them with other words.
[0120] Specifically, the aforementioned pre-defined masking configuration can be applied to each of the selected words to be masked. That is, each selected word can further determine its specific masking process based on a probability configuration, preventing the language model from pre-determining which words will be masked. Since every word in the input text may be masked or replaced, the trained language model, when encoding the current word, does not overly rely on the word itself but considers its context, and may even perform error correction based on its context.
[0121] The model training module 704 can be configured to predict at least one masked word in the masked input text using the language model, and train the language model based on the prediction result. Optionally, the model training module 704 can perform the operations described above with reference to step 204. For example, through the above data processing of the training data, the masked input text can be input into the language model to be trained to train the language model.
[0122] According to another aspect of this disclosure, a pre-training device for a language model is also provided. Figure 8 A schematic diagram of a language model pre-training device 2000 according to an embodiment of the present disclosure is shown.
[0123] like Figure 8 As shown, the language model pre-training device 2000 may include one or more processors 2010 and one or more memories 2020. The memories 2020 store computer-readable code, which, when run by the one or more processors 2010, can execute the language model pre-training method described above.
[0124] The processor in the embodiments of this disclosure can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor, and can be based on an x86 architecture or an ARM architecture.
[0125] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0126] For example, the method or apparatus according to embodiments of this disclosure can also be used by means of Figure 9 The architecture of the computing device 3000 shown is used for implementation. For example... Figure 9 As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage devices in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the language model pre-training method provided in this disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 9 The architecture shown is merely exemplary and can be omitted as needed when implementing different devices. Figure 9 One or more components in the computing device shown.
[0127] According to another aspect of this disclosure, a computer-readable storage medium is also provided. Figure 10 A schematic diagram 4000 of a storage medium according to the present disclosure is shown.
[0128] like Figure 10As shown, computer-readable instructions 4010 are stored on the computer storage medium 4020. When the computer-readable instructions 4010 are executed by a processor, a pre-training method for a language model according to an embodiment of the present disclosure, as described with reference to the above figures, can be performed. The computer-readable storage medium in the embodiments of the present disclosure may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus random access memory (DRRAM). It should be noted that the memory used in the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0129] Embodiments of this disclosure also provide a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a pre-training method for a language model according to embodiments of this disclosure.
[0130] Embodiments of this disclosure provide a method, apparatus, device, and computer-readable storage medium for pre-training a language model.
[0131] Compared to traditional language model pre-training methods, the method provided by the embodiments of this disclosure can perform domain-adaptive masking processing on the input text to better learn the feature representations of domain-specific words in the text, enabling the trained language model to learn the knowledge of that domain more fully.
[0132] The method provided in the embodiments of this disclosure identifies domain-specific entity words in the input text and performs masking processing on the input text based on these entity words and their entity types. Specifically, entity words belonging to core entity types specific to that domain are masked with a higher masking probability. This allows the language model to better learn the representations of these core entity words through pre-training, enabling the trained language model to more fully learn the characteristics of domain-specific data. The method in the embodiments of this disclosure enables domain-adaptive language model pre-training, thereby enhancing the language model's learning ability in a specific domain and allowing the trained language model to be better applied to various task scenarios within that specific domain.
[0133] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing at least one executable instruction for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0134] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0135] The exemplary embodiments of this disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art will understand that various modifications and combinations of these embodiments or their features can be made without departing from the principles and spirit of this disclosure, and such modifications should fall within the scope of this disclosure.
Claims
1. A pre-training method for a language model, said language model being applied to a specific domain, said method comprising: Obtain input text, wherein the input text includes one or more words, and the one or more words include entity words; The input text is subjected to component recognition to determine multiple entity words and their entity types in the input text, wherein the entity type of each entity word is related to the specific domain. Based on multiple entity words and their entity types in the input text, the input text is masked. A first number of entity words among the multiple entity words are assigned a higher masking probability, and the entity types of the first number of entity words belong to the core entity types of the specific domain. The first number may include zero, one, or more entity words. For at least one masked word in the masked input text, the language model is used to predict the at least one word, and the language model is trained based on the prediction result.
2. The method as described in claim 1, wherein, Component recognition of the input text includes: The input text is segmented to identify one or more words in the input text; For each of the identified one or more words, obtain a vector representation of that word; Based on the vector representation of the words, determine the feature representation of the words; and Based on the feature representation of the words, the entity type of the words is determined by a classifier.
3. The method as described in claim 2, wherein, The component recognition of the input text is performed using a pre-trained component recognition model, wherein the component recognition model is pre-trained for the specific domain.
4. The method of claim 1, wherein, The method also includes pre-training a component recognition model for the specific domain; Specifically, pre-training the component recognition model for the specific domain includes: Obtain training text related to the specific domain, the training text including multiple entity words with pre-labeled entity types; The component recognition model is used to predict the entity type of each entity word in the training text; and The component recognition model is trained based on the prediction results of the entity type of each entity word in the training text and its pre-labeled entity type.
5. The method of claim 1, wherein, Masking the input text based on multiple entity words and their entity types includes: Based on the multiple entity words and their entity types in the input text, a mask probability is set for all words in the input text; and Based on the set masking probabilities, all words in the input text are masked.
6. The method of claim 5, wherein, Setting a mask probability for all words in the input text includes: For entity words whose entity type belongs to the core entity type of the specific domain among the multiple entity words, a higher mask probability is set compared to other words.
7. The method of claim 6, wherein, The one or more words mentioned also include non-entity words; The probability of setting a mask for all words in the input text also includes: For non-entity words and entity words that are not related to the specific domain in the input text, a lower mask probability is set compared to other words.
8. The method of claim 5, wherein, Masking all words in the input text based on the set mask probabilities includes: Based on the set masking probabilities, multiple words in the input text to be masked are determined; and For each of the determined words to be masked, masking is performed using a predetermined masking configuration.
9. The method of claim 8, wherein, The predetermined mask configuration includes: a first probability for masking, a second probability for preserving the original words, and a third probability for replacing them with other words; Specifically, for each of the determined words to be masked, masking processing with a predetermined mask configuration includes: For the word, a masking process is determined based on the predetermined masking configuration, wherein the probability of the masking process being a mask is a first probability, the probability of the masking process being to preserve the original word is a second probability, and the probability of the masking process being to replace the word with another word is a third probability; and The determined masking process is applied to the words.
10. The method of claim 1, wherein, For at least one masked word in the masked input text, predicting the at least one word using the language model includes: Based on the masked input text, the language model is used to determine the original word corresponding to at least one masked word.
11. The method of claim 1, wherein, Training the language model based on the prediction results includes: The parameters in the language model are optimized based on the probability that the original word corresponding to at least one masked word is obtained from the masked input text by the language model.
12. A pre-training device for a language model, said language model being applied to a specific domain, said device comprising: The text acquisition module is configured to acquire input text, which includes one or more words, and the one or more words include entity words. The component recognition module is configured to perform component recognition on the input text to determine multiple entity words and their entity types in the input text, wherein the entity type of each entity word is related to the specific domain. A masking module is configured to mask the input text based on multiple entity words and their entity types. Specifically, a first number of entity words among the multiple entity words are assigned a higher masking probability, and the entity types of the first number of entity words belong to the core entity types of the specific domain. The first number may include zero, one, or more entity words. The model training module is configured to predict at least one word masked in the masked input text using the language model, and train the language model based on the prediction results.
13. A pre-training device for a language model, comprising: One or more processors; as well as One or more memories storing a computer-executable program, which, when executed by the processor, performs the method of any one of claims 1-11.
14. A computer program product stored on a computer-readable storage medium and comprising computer instructions that, when executed by a processor, cause a computer device to perform the method of any one of claims 1-11.
15. A computer-readable storage medium having stored thereon computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-11.
Citation Information
Patent Citations
Language model training method, device and equipment and computer readable storage medium
CN113515938A
Text relation extraction method and device, equipment and storage medium
CN114610903A