Text sequence generation method, pre-training method, storage medium and program product

By obtaining and translating the description text of the language model entity from the knowledge graph and generating multilingual text sequences, the limitations of the existing single language model in multilingual applications are solved, and a multilingual supported language model is implemented, which is suitable for cross-language natural language processing.

CN114580438BActive Publication Date: 2025-07-01ALIBABA (CHINA) CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210205264.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-02
Publication Date
2025-07-01
Estimated Expiration
2042-03-02

AI Technical Summary

Technical Problem

The existing knowledge-based pre-trained language models are all single languages, limiting their widespread use in multilingual applications.

Method used

Multilingual text sequences are generated by obtaining the first description text of the entity pair described in the main language in the knowledge graph and replacing some or all elements with description texts of other languages ​​based on the annotation information, and for pre-training of the language model based on multilingual knowledge.

Benefits of technology

The multilingual supported language model is implemented, enabling it to be used for cross-language natural language processing, simplifying the steps of obtaining sample texts for pre-training, reducing dependence on other models, and ensuring the accuracy of knowledge in multilingual text sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114580438B_ABST
    Figure CN114580438B_ABST
Patent Text Reader

Abstract

The present application provides a text sequence generation method, a pre-training method for a language model, a storage medium, and a program product. The text sequence generation method includes: obtaining first description texts of a plurality of entity pairs described in a main language in a knowledge graph, where the elements in the entity pair include two entities and the entity relationship therebetween; obtaining annotation information corresponding to some or all of the elements in the plurality of entity pairs, and replacing at least some of the elements in the entity pair with texts described in another language different from the main language according to the annotation information to obtain second description texts of the entity pair; generating a multilingual text sequence based on the first description texts and the second description texts corresponding to the entity pair, where the multilingual text sequence is used for pre-training the language model based on multilingual knowledge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technologies, and in particular, to a method for generating a text sequence, a pre-training method, a storage medium, and a program product. Background Art

[0002] A language model can predict the probability of the occurrence of a text with a certain length. Fine-tuned large-scale pre-trained language models are widely used in natural language processing tasks, such as speech recognition, machine translation, part-of-speech tagging, syntactic analysis, and information retrieval.

[0003] However, existing knowledge-based pre-trained language models are all single-language, which greatly limits the application of pre-trained language models. Summary of the Invention

[0004] In view of this, embodiments of the present application provide a text sequence generation solution to at least partially solve the above problems.

[0005] According to a first aspect of the embodiments of the present application, there is provided a method for generating a text sequence, including: obtaining first description texts of a plurality of entity pairs described in a main language in a knowledge graph, where elements in the entity pair include two entities and the entity relationship therebetween; obtaining annotation information corresponding to some or all of the elements in the plurality of entity pairs, and replacing at least some of the elements in the entity pair with texts described in another language different from the main language according to the annotation information, to obtain second description texts of the entity pair; generating a multilingual text sequence based on the first description texts and the second description texts corresponding to the entity pair, where the multilingual text sequence is used for pre-training a language model based on multilingual knowledge.

[0006] According to a second aspect of the embodiments of the present application, there is provided a method for pre-training a language model, including: obtaining first description texts of a plurality of entity pairs described in a main language in a knowledge graph, where elements in the entity pairs include two entities and the entity relationship therebetween; obtaining annotation information corresponding to some or all of the elements in the plurality of entity pairs, and replacing at least some of the elements in the entity pairs with texts described in another language different from the main language according to the annotation information to obtain second description texts of the entity pairs; generating a multilingual text sequence based on the first description texts and the second description texts corresponding to the entity pairs; replacing the texts corresponding to entities or entity relationships in the multilingual text sequence with masks to generate a sample text sequence; inputting the sample text sequence into the language model, and predicting the texts corresponding to the masks through the language model; and adjusting the language model according to the difference between the predicted texts output by the language model and the texts replaced with masks, so as to perform pre-training of the language model based on multilingual knowledge.

[0007] According to a third aspect of the embodiments of the present application, there is provided a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned text sequence generation method is implemented.

[0008] According to a fourth aspect of the embodiments of the present application, there is provided a computer program product, including computer instructions, and the computer instructions direct a computing device to perform operations corresponding to the above-mentioned text sequence generation method.

[0009] According to the solution provided by the embodiments of the present application, by obtaining first description texts of a plurality of entity pairs described in a main language in a knowledge graph, where elements in the entity pairs include two entities and the entity relationship therebetween; obtaining annotation information corresponding to some or all of the elements in the plurality of entity pairs, and replacing at least some of the elements in the entity pairs with texts described in another language different from the main language according to the annotation information to obtain second description texts of the entity pairs; generating a multilingual text sequence based on the first description texts and the second description texts corresponding to the entity pairs, thus, a multilingual text sequence can be directly obtained according to the entity pairs and annotation information in the knowledge graph, and the multilingual text sequence is used for pre-training the language model based on multilingual knowledge, so that the pre-trained language model supports multiple languages and can be used for cross-language natural language processing; in addition, the solution provided by the embodiments of the present application does not need to obtain texts in corresponding languages through language generation models corresponding to different languages, simplifies the steps of obtaining sample texts for pre-training, and in this embodiment, generating a multilingual text sequence according to entity pairs does not require aligning entities or linking, reduces the dependence on other models, and ensures the accuracy of knowledge in the multilingual text sequence. Description of the Drawings

[0010] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can also be obtained based on these drawings.

[0011] Figure 1 It is a flowchart of the steps of a text sequence generation method according to an embodiment of the present application;

[0012] Figure 2 It is a schematic diagram of a scenario according to an embodiment of the present application;

[0013] Figure 3 It is a flowchart of the steps of another text sequence generation method according to an embodiment of the present application;

[0014] Figure 4 It is a schematic flowchart of a pre-training method of a language model according to an embodiment of the present application;

[0015] Figure 5 It is a schematic diagram of a language model according to an embodiment of the present application;

[0016] Figure 6 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0017] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the embodiments of the present application.

[0018] Before elaborating on the specific content of the present application in detail, the terms and nouns related to the present application will be described first.

[0019] 1) Pretrained Language Model: For a sequence composed of texts, the language model is used to calculate the probability distribution of the sequence, indicating the possibility of the existence of the text. After fine-tuning, large-scale pre-trained language models have wide applications in natural language processing tasks, such as speech recognition, machine translation, part-of-speech tagging, syntactic analysis, and information retrieval, etc.

[0020] 2) Transformer: A deep learning network structure based on the self-attention mechanism that can convert text into vectors. A language model is generally composed of multiple layers of Transformers.

[0021] 3) Knowledge graph: A knowledge database that describes entities and the relationships between entities. Generally, we can use a relational graph to represent a knowledge graph.

[0022] 4) Named Entity Recognition (NER): Also known as "Proper Name Recognition", it is used to identify entities with specific meanings in text, mainly including person names, place names, organization names, proper nouns, etc. For example, in product descriptions, extracting brand names, product categories, etc. all belong to the applications of named entity recognition.

[0023] The following further illustrates the specific implementation of the embodiments of the present application in conjunction with the accompanying drawings of the embodiments of the present application.

[0024] The large-scale pre-trained language model itself is a neural network language model. It can be trained using large-scale unlabeled pure text corpora; and can be used in downstream NIP tasks after fine-tuning, without the need to be specifically designed for downstream tasks, and thus can obtain good results. However, the pre-trained model lacks common sense knowledge, which hinders the large-scale promotion of the pre-trained model in actual application scenarios.

[0025] Therefore, in order to enhance the robustness of the pre-trained language model, generally, an attempt is made to inject knowledge into the pre-trained model, so as to better use the pre-trained model for tasks such as knowledge-driven named entity recognition and named entity relationship recognition. Such models are also called knowledge-enhanced pre-trained language models. For example, when inputting a text sequence, triples can be injected into the text sequence as domain knowledge, so as to convert the text sequence into a knowledge-rich tree. The top layer of the tree structure can be each text in the text sequence, and the lower layer of the text sequence can be the knowledge corresponding to the text added through the knowledge graph.

[0026] However, all existing knowledge-enhanced pre-trained language models are single language models, which limits their application in more languages.

[0027] For this reason, the embodiments of the present application provide a text sequence generation method, and the generated text sequence can be used for pre-training a language model based on multi-language knowledge.

[0028] Figure 1 The following is a schematic flowchart of a text sequence generation method provided by the embodiments of the present application. As shown in the figure, it includes:

[0029] S101. Obtain the first description texts of several entity pairs described in the main language in the knowledge graph.

[0030] Among them, the elements in the entity pair include two entities and the entity relationship between them.

[0031] The entity pairs obtained in this embodiment can specifically be triples. In a knowledge graph, knowledge content is generally stored in the form of triples (head entity, relationship, tail entity), which is a common representation of a triple knowledge graph. Of course, in other implementation manners of this application, entity pairs in other forms can also be obtained. For example, the entity pair can also include annotation information added to the above triples, etc. This embodiment does not limit this.

[0032] When the entity pair is a triple, the first description text corresponding to the entity pair described in the main language can also be a triple, and the triple includes the text of the elements described in the main language. By way of example, if the main language is English, the first description text can be, for example, (motor car, designed to carry, passenger), where motor car is the text describing the head entity, passenger is the text describing the tail entity, and designed to carry is the text describing the entity relationship.

[0033] It should be noted that the first description text corresponding to an entity pair can include one or more, and this embodiment does not limit this. By way of example, ( motor car , designed to carry, passenger) and ( autocar , designed to carry, passenger) are two first description texts of an entity pair, where the underscores are used to distinguish the differences between the two first description texts.

[0034] S102. Obtain the annotation information corresponding to some or all of the elements in a plurality of entity pairs, and replace at least some of the elements in the entity pair with the text described in a language different from the main language according to the annotation information, so as to obtain the second description text of the entity pair.

[0035] A knowledge graph generally includes annotation information, and the annotation information can include attribute information, aliases, timeliness information, etc. of entities or entity relationships. In the embodiments of this application, the annotation information also includes the text of entities or entity relationships described in other languages.

[0036] Each element (entity or entity relationship) corresponds to a unique identifier, and the annotation information corresponding to the element can be located through the unique identifier. By way of example, the identifiers of the elements in a plurality of entity pairs can be combined, and the annotation information of some or all of the elements can be obtained according to the identifiers of the elements.

[0037] Exemplarily, taking the entity "motorcar" as an example, its corresponding unique identifier can be Q1420, and its corresponding partial annotation information can be as follows:

[0038]

[0039] Among them, "language" represents the language that can describe the entity, "lable" represents the text describing the entity in the corresponding language, and "aliases" represents the alternative names (also known as synonyms) of the entity described in the corresponding language.

[0040] In this embodiment, English can be used as the main language, and Spanish, Hungarian, etc. can be used as other languages.

[0041] Exemplarily, if the other language is French, the second description text can be, for example, (motor car, pour transporter passenger), where "pour transporter" is the text describing the entity relationship in French, or the second description text can be, for example, (motor car, designed to carry, passager), where "passager" is the text describing the end entity in French.

[0042] It should be noted that the second description text corresponding to an entity pair can include one or more, and this embodiment does not limit this.

[0043] S103. Generate a multilingual text sequence based on the first description text and the second description text corresponding to the entity pair, where the multilingual text sequence is used for pre-training the language model based on multilingual knowledge.

[0044] Exemplarily, in this embodiment, the first description text and the second description text respectively include the texts corresponding to the multiple elements corresponding to the entity pair. Therefore, the texts corresponding to the multiple elements can be connected through text identifiers to obtain the multilingual text sequence corresponding to the first description text or the second description text.

[0045] Exemplarily, for the first description text (motor car, designed to carry, passenger), the texts corresponding to the multiple elements can be spliced using a mask to obtain the multilingual text sequence of this first description text: motor car[mask]designed to carry[mask]passenger; for the second description text (motor car. To transport passengers, multiple texts corresponding to respective elements can be spliced using a mask to obtain a corresponding multilingual text sequence, such as motor car [mask] To transport [mask] passengers, this multilingual text sequence is a mixed English-French text sequence; the second description text can be, for example, (motor car, designed to carry, passenger), and multiple texts corresponding to respective elements can be spliced using a mask to obtain a corresponding multilingual text sequence motor car [mask] designed to carry [mask] passenger.

[0046] The obtained multilingual text sequence can be used for pre-training the language model based on multilingual knowledge. For the specific method of performing pre-training based on multilingual knowledge, reference can be made to the subsequent embodiments and will not be elaborated here.

[0047] Of course, the above is only an example. In other implementation manners of the application, other text identifiers other than the mask can be used to connect the texts corresponding to multiple elements, which is not a limitation of the present application.

[0048] Next, through a specific implementation scenario, the solution of the present application will be exemplarily described.

[0049] See Figure 2 , multiple triples (h, r, t) can be extracted from the knowledge graph as entity pairs, where h is the head entity, r is the entity relationship, and t is the tail entity. Exemplarily, the extracted triples can include n groups.

[0050] For each triple, a first description text described in the main language can be obtained. If the main language is English, the obtained first description text can be (h en , r en , t en ), where the subscripts en of h, r, and t indicate that they are English texts.

[0051] For multiple triples, the identifiers of the elements included in the multiple triples are determined, and the annotation information corresponding to some or all of the elements is determined according to the element identifiers. According to the annotation information, at least some of the elements in the entity pair can be replaced with texts described in other languages different from the main language to obtain the second description text of the entity pair.

[0052] Exemplarily, the first description text corresponding to the triple can be (h en , r en , t en), one of the elements r can be en replaced with the text r described in Spanish es to obtain the second description text (h en , r es , t en ), where the subscript es of r es identifies it as the text in Spanish.

[0053] Based on the first description text and the second description text, a multilingual text sequence can be generated. Based on the multilingual text sequence, the language model can be pre-trained based on multilingual knowledge.

[0054] Exemplarily, the mask can be used to splice the text corresponding to the first description text and the second description text with the elements to obtain the multilingual text sequence corresponding to the first description text h en [mask]r en [mask]t en , and the multilingual text sequence corresponding to the second description text h en [mask]r es [mask]t en . Thus, the text sequence with multilingual text can be used to pre-train the language model based on multilingual knowledge.

[0055] The solution provided in this embodiment obtains the first description text of several entity pairs described in the main language in the knowledge graph, where the elements in the entity pair include two entities and the entity relationship between them; obtains the annotation information corresponding to some or all of the elements in several entity pairs, and replaces at least some of the elements in the entity pair with the text described in a language different from the main language according to the annotation information to obtain the second description text of the entity pair; generates a multilingual text sequence based on the first description text and the second description text corresponding to the entity pair. Thus, the multilingual text sequence can be directly obtained according to the entity pair and the annotation information in the knowledge graph. The multilingual text sequence is used to pre-train the language model based on multilingual knowledge, so that the pre-trained language model supports multiple languages and can be used for cross-language natural language processing. In addition, the solution provided in this embodiment of the present application does not need to obtain the text in the corresponding language through the language generation model corresponding to different languages, simplifies the steps of obtaining the sample text for pre-training, and in this embodiment, generating the multilingual text sequence according to the entity pair does not require aligning entities or linking, reduces the dependence on other models, and ensures the accuracy of the knowledge in the multilingual text sequence.

[0056] The text sequence generation method of this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to: servers, mobile terminals (such as mobile phones, PADs, etc.), and PC machines, etc.

[0057] Figure 3 It is a schematic flowchart of a text sequence generation method provided by an embodiment of this application. As shown in the figure, it includes:

[0058] S201. Obtain the first description text of a plurality of entity pairs described in the main language in the knowledge graph.

[0059] For the specific implementation manner of this step, refer to the above embodiment and will not be elaborated here.

[0060] S202. Obtain the annotation information corresponding to the elements in some or all of the entity pairs, and determine the target other language.

[0061] In this embodiment, for the specific method of obtaining the annotation information, refer to the above embodiment and will not be elaborated here.

[0062] In this embodiment, the main language and the target other language supported by the expected pre-trained language model can be determined. The main language is generally the default language for describing entities in the knowledge graph, such as English, etc. The target other language can include one or more, and this embodiment does not limit this.

[0063] S203. For any element of any entity pair, query in the annotation information according to the identifier of the element and the target other language;

[0064] In this embodiment, for any element of any entity pair, the partial annotation information corresponding to the element can be located from the annotation information according to the identifier of the element; then, query can be performed in the annotation information according to the target other language to determine whether there is text describing the element in the target other language in the annotation information.

[0065] S204. If it is determined according to the query result that the element corresponds to text described in the target other language, then replace the text corresponding to the element in the first description text of the entity pair with the text described in the target other language to obtain the second description text corresponding to the entity pair.

[0066] If it is determined whether there is text describing the element in the target other language in the annotation information, then the text corresponding to the element in the first description text can be replaced to obtain the second description text. The second description text is generally a text mixed with the main language and the target other language.

[0067] Optionally, in any embodiment of the present application, step S204 includes: generating a random number corresponding to the element, and determining whether to replace the element according to the random number; if replacement is required and it is determined according to the query result that the element corresponds to a target other-language description text, then replacing the text corresponding to the element in the first description text corresponding to the entity pair with the text described in the target other language, to obtain the second description text corresponding to the entity pair. By determining whether to replace the element through the generated random number, the diversity of the generated second description text can be ensured, and thus the randomness of the generated multilingual text sequence can be ensured.

[0068] For the specific method of generating a random number, reference may be made to related technologies and will not be elaborated here. After generating the random number, it can be determined whether the generated random number conforms to a preset replacement rule. If it conforms, the element is replaced; otherwise, it is not replaced. The preset replacement rule can be set by those skilled in the art, such as a numerical range, etc., and this embodiment does not limit it.

[0069] Optionally, in any embodiment of the present application, the method may further include: obtaining a plurality of language pairs including the main language and a target other language, where the target other languages included in different language pairs are different; then the generating a random number corresponding to the element and determining whether to replace the element according to the random number includes: for any language pair, generating a random number corresponding to the language pair and the element, and determining whether to replace the element according to the random number; the replacing the text corresponding to the element in the first description text corresponding to the entity pair with the text described in the target other language of the language pair to obtain the second description text corresponding to the language pair includes: replacing the text corresponding to the element in the first description text corresponding to the entity pair with the text describing the element in the target other language of the language pair, to obtain the second description text corresponding to the language pair.

[0070] In this embodiment, by obtaining a language pair including the main language and the target language, the corresponding second description text can be obtained based on the language pair. The second description text is a mixed text of the main language and the target language in the language pair. Thus, the number of second description texts corresponding to the target other language can be accurately determined through the language pair, and the languages mixed in the second description text can be identified through the language pair.

[0071] Exemplarily, the obtained language pair may include (en, es), where en corresponds to the main language English and es corresponds to the target other language Spanish.

[0072] Optionally, in the embodiments of the present application, the annotation information includes aliases of elements in the entity pair described in the main language. The method further includes: replacing at least some of the elements in the entity pair with aliases described based on the main language to obtain the first description text; or replacing at least some of the elements in the entity pair with aliases described based on the other language to obtain the second description text.

[0073] By replacing some of the elements in the entity pair with aliases described in the main language, the number of the first description texts can be increased; similarly, by replacing some of the elements in the entity pair with aliases described in the other language, the number of the second description texts can be increased, thereby increasing the number of multilingual text sequences for pre-training the language model and improving the pre-training effect.

[0074] Exemplarily, for the first description text (motor car, designed to carry, passenger), replacing the motercar therein with an alias described in English can obtain the first description text (autocar, designed to carry, passenger); for the first description text (motor car, designed to carry, passenger), replacing the motercar therein with an alias described in Spanish can obtain the second description text (carro, designed to carry, passenger).

[0075] S205. Use a mask to splice the text corresponding to the element in the first description text or the second description text corresponding to the entity pair to generate a multilingual text sequence.

[0076] Exemplarily, in the embodiments of the present application, for the first description text (motor car, designed to carry, passenger), a multilingual text sequence motor car [mask] designed to carry [mask] passenger can be obtained through splicing with a mask; similarly, for the second description text (carro, designed to carry, passenger), a multilingual text sequence carro [mask] designed to carry [mask] passenger passenger can be obtained through splicing with a mask.

[0077] The solution provided in this embodiment obtains the first description text of several entity pairs described in the main language in the knowledge graph, where the elements in the entity pair include two entities and the entity relationship between them; obtains the annotation information corresponding to some or all of the elements in the several entity pairs, and replaces at least some of the elements in the entity pair with text described in another language different from the main language according to the annotation information, to obtain the second description text of the entity pair; generates a multilingual text sequence based on the first description text and the second description text corresponding to the entity pair. Thus, a multilingual text sequence can be directly obtained according to the entity pair and the annotation information in the knowledge graph, and the multilingual text sequence is used for pre-training a language model based on multilingual knowledge, so that the pre-trained language model supports multiple languages and can be used for cross-lingual natural language processing. In addition, the solution provided in the embodiments of the present application does not need to obtain the text in the corresponding language through a language generation model corresponding to different languages, simplifies the steps of obtaining the sample text for pre-training, and in this embodiment, generating a multilingual text sequence according to the entity pair does not require aligning entities or links, reduces the dependence on other models, and ensures the accuracy of the knowledge in the multilingual text sequence.

[0078] The text sequence generation method in this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to: servers, mobile terminals (such as mobile phones, PADs, etc.) and PC machines, etc.

[0079] Figure 4 It is a schematic flowchart of a pre-training method for a language model provided in an embodiment of the present application. As shown in the figure, it includes:

[0080] S301. Obtain the first description text of several entity pairs described in the main language in the knowledge graph.

[0081] S302. Obtain the annotation information corresponding to some or all of the elements in the several entity pairs, and replace at least some of the elements in the entity pair with text described in another language different from the main language according to the annotation information, to obtain the second description text of the entity pair.

[0082] S303. Generate a multilingual text sequence based on the first description text and the second description text corresponding to the entity pair.

[0083] For the specific implementation manners of the above steps S301 - S303, reference can be made to the above embodiments, and details are not described herein again.

[0084] S304. Replace the text corresponding to the entity or the entity relationship in the multilingual text sequence with a mask to generate a sample text sequence.

[0085] When specifically performing training, for the text corresponding to the element in the multilingual text sequence, it is replaced with a mask to generate a sample text sequence input to the language model.

[0086] Exemplarily, for the multilingual text sequence motor car[mask]designed to carry[mask]passenger, the motor in it can be replaced with a mask to obtain the corresponding sample text sequence [mask]car[mask]designed to carry[mask]passenger.

[0087] S305. Input the sample text sequence into the language model, and the language model predicts the text corresponding to the mask.

[0088] After the sample text sequence is input into the language model, the language model can predict the text corresponding to the mask in the sample text sequence and input the prediction result.

[0089] The specific method for prediction can refer to the related technology and will not be elaborated here.

[0090] S306. According to the difference between the predicted text output by the language model and the text replaced with the mask, adjust the language model to perform pre-training of the language model based on multilingual knowledge.

[0091] For the text replaced with the mask in step S206, the loss function can be calculated according to the predicted text output by the language model and the text replaced with the mask, and the parameters in the language model are adjusted according to the calculation result to perform multilingual pre-training on the language model.

[0092] In addition, in the above embodiments of the present application, the main purpose is to enable the language model to learn entity knowledge based on multiple languages. In order to enable the language model to accurately process text, the multilingual text sequence can also be used to perform joint training with at least one of the pre-training tasks based on multilingual knowledge and the pre-training tasks based on cloze or logical reasoning to obtain the trained language model.

[0093] The pre-training task based on cloze (Masked Language Modeling, MLM), and the language model pre-trained through this task can also be called a masked language model. The specific training method can refer to the related technology and will not be elaborated here.

[0094] The pre-training task based on logical reasoning mainly pre-trains the language model through a text sequence with multiple short sentences, and there is a logical relationship between the multiple short sentences. Thus, the trained language model can have logical reasoning ability.

[0095] The sample text sequence in the pre-training task based on logical reasoning can be obtained from the knowledge graph or by other methods, which is not limited in this embodiment. Exemplarily, multiple short sentences can be generated according to multiple entity pairs with logical relationships included in the knowledge graph, so that there is a logical relationship between the multiple short sentences.

[0096] The language model obtained through joint training can not only learn multilingual knowledge, but also learn the distribution of words in natural sentences through the cloze task, or improve the logical reasoning ability of the language model through the pre-training task based on logical reasoning.

[0097] When performing joint training of the three tasks, the loss function can be:

[0098]

[0099] Where L is the total loss, LMLM is the loss of the pre-training task based on the cloze, LCS is the loss of the pre-training task based on multilingual knowledge, LL is the loss of the pre-training task of logical reasoning, and α is the weight used to adjust the weights of the latter two losses in the total loss.

[0100] Next, a specific implementation manner is used to illustrate the solution provided in the embodiment of the present application.

[0101] 1) Extract the entity pair (h, r, t) from the knowledge graph, obtain the original description text describing the entity pair in the main language, and the annotation information corresponding to the elements in the entity pair. The original description text belongs to the first description text.

[0102] Exemplarily, the obtained original description text Original(en) can be, for example: (motor car, designed to carry, passenger), where the en in the parentheses after original indicates that the main language is English.

[0103] Specifically, the corresponding annotation information can be obtained according to the identifier of the element.

[0104] 2) Determine multiple language pairs, and each language pair includes the main language and another language different from the main language.

[0105] Exemplarily, the language pair can be, for example, (en, es) corresponding to the English-Spanish language pair, and the language pair can also be (en, fr) corresponding to the English-French language pair.

[0106] 3) For any entity pair, generate corresponding random numbers for the three elements therein, and determine whether to replace the element according to the corresponding random number.

[0107] Exemplarily, for the three elements h, r, and t in an entity pair (h, r, t), three random numbers can be generated, and based on the random numbers, it is determined whether to replace the three elements h, r, and t.

[0108] 4) If it is determined to perform replacement, and it is determined according to the language pair that the annotation information includes text describing the element in another language, then replace the text of the element in the first description text corresponding to the entity pair with the text described in another language to obtain a second description text.

[0109] Exemplarily, for (motor car, designed to carry, passenger), if it is determined to perform replacement, then according to the language pair (en, es), in the partial annotation data corresponding to the element, search whether there is text corresponding to the Spanish language es for motor car. If it is included, then replace motor car in the first description text with the text describing motorcycle in Spanish to obtain a second description text.

[0110] It should be noted that an entity pair can correspond to multiple first description texts. Then, the text of the element in some or all of the first description texts can be replaced. This embodiment does not limit this.

[0111] After all replacements are completed, duplicate removal operations can be performed on the first description text and the second description text.

[0112] In addition, the text corresponding to the element can also be replaced with an alias of the element described in the main language or another language by a similar method.

[0113] The first description text and the second description text can also be identified by the triple (hen, 0, ren, 0 , ten, 0). The h, r, and t identify the elements of the triple, and the subscript en in the lower right corner is used to identify the language describing the element. en can also be replaced with fr, es, etc. In this embodiment, en corresponds to the main language. When the subscripts of all three elements are en, it indicates that it is the first description text; the digital subscript in the lower right corner is used to identify whether it is an alias. hen, 3 indicates that it is the 3rd English alias.

[0114] 5) Use a mask to splice the first description text and the second description text to obtain a multilingual text sequence.

[0115] Exemplarily, the obtained multilingual description text Code Switched (en - fr) can be, for example:

[0116] motorcar[mask ]concu pour transporter[mask]passenger.

[0117] motor car[mask]designed to carry [mask]passager.

[0118] The en-fr identification in the Code Switched post-bracket indicates a multilingual sequence of English-French mixture, with the bold part in French.

[0119] 5) Use multilingual text sequences to perform pre-training of the language model based on multilingual knowledge.

[0120] Exemplarily, for a multilingual text sequence, the text therein can be replaced with masks and input into the language model. The language model can output the predicted text corresponding to the masks. Subsequently, the loss function can be calculated based on the predicted text and the replaced text, and the language model can be adjusted accordingly until the pre-training of the language model is completed.

[0121] Exemplarily, referring to Figure 5 , taking the transformers as an example of the language model, the motor and designed in the multilingual text sequence motorcar[mask]designed to carry[mask]passager. can be replaced with masks to obtain a sample text sequence, and the sample text sequence can be input into the language model.

[0122] The language model can output the predicted text for motor and designed. Subsequently, the loss function can be calculated based on the difference between the predicted text and the text replaced with masks.

[0123] The pre-training method of the language model in this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to: servers, mobile terminals (such as mobile phones, PADs, etc.) and PC machines, etc.

[0124] Referring to Figure 6 , a schematic structural diagram of an electronic device provided by an embodiment of the present application is shown. The specific implementation of the electronic device is not limited in the specific embodiments of the present application.

[0125] As Figure 6 shown, the electronic device may include: a processor 502, a communication interface 504, a memory 506, and a communication bus 508.

[0126] Wherein:

[0127] The processor 502, the communication interface 504, and the memory 506 communicate with each other via the communication bus 508.

[0128] The communication interface 504 is used to communicate with other electronic devices or servers.

[0129] The processor 502 is used to execute the program 510, and specifically can execute the relevant steps in the above-described method embodiments for generating a text sequence.

[0130] Specifically, the program 510 may include program code, and this program code includes computer operation instructions.

[0131] The processor 502 may be a central processing unit (CPU), or a specific application integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0132] The memory 506 is used to store the program 510. The memory 506 may include high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory.

[0133] For the specific implementation of each step in the program 510, reference may be made to the corresponding steps and descriptions in the corresponding units in the above-described method embodiments for generating a text sequence, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be repeated here.

[0134] The embodiments of the present application also provide a computer program product, including computer instructions, and these computer instructions direct a computing device to execute the operations corresponding to any one of the above-described method embodiments for generating a text sequence.

[0135] It should be noted that according to the needs of implementation, each component / step described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of a component / step can be combined into a new component / step to achieve the objectives of the embodiments of the present application.

[0136] The method according to the embodiments of the present application can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code that is originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and will be stored in a local recording medium, so that the method described herein can be stored on such a software process on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as RAM, ROM, flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the text sequence generation method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the text sequence generation method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the text sequence generation method shown herein.

[0137] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0138] The above embodiments are only used to illustrate the embodiments of the present application, rather than to limit the embodiments of the present application. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application. The patent protection scope of the embodiments of the present application shall be defined by the claims.

Claims

1. A method for generating a text sequence, comprising: Obtaining first description texts of a plurality of entity pairs described in a main language in a knowledge graph, wherein the elements in the entity pair include two entities and the entity relationship therebetween; Obtaining annotation information corresponding to some or all of the elements in the plurality of entity pairs, and replacing at least some of the elements in the entity pair with texts described in another language different from the main language according to the annotation information, to obtain second description texts of the entity pair; Generating a multilingual text sequence based on the first description text and the second description text corresponding to the entity pair, where the multilingual text sequence is used for pre-training a language model based on multilingual knowledge; The obtaining annotation information corresponding to the elements in the plurality of entity pairs, and replacing at least some of the elements in the entity pair with texts described in another language different from the main language according to the annotation information, to obtain second description texts corresponding to the entity pair, includes: Obtaining annotation information corresponding to some or all of the elements in the entity pair, and determining a target other language; Querying in the annotation information for any element of any entity pair according to the identifier of the element and the target other language; If it is determined according to the query result that the element corresponds to a text described in the target other language, then replacing the text corresponding to the element in the first description text corresponding to the entity pair with the text described in the target other language, to obtain the second description text corresponding to the entity pair.

2. The method according to claim 1, wherein, If it is determined according to the query result that the element corresponds to a text described in the target other language, then replacing the text corresponding to the element in the first description text corresponding to the entity pair with the text described in the target other language, to obtain the second description text corresponding to the entity pair, includes: Generating a random number corresponding to the element, and determining whether to replace the element according to the random number; If replacement is required and it is determined according to the query result that the element corresponds to a text described in the target other language, then replacing the text corresponding to the element in the first description text corresponding to the entity pair with the text described in the target other language, to obtain the second description text corresponding to the entity pair.

3. The method according to claim 2, wherein The method further includes: Obtaining a plurality of language pairs including the main language and the target other language, where the target other languages included in different language pairs are different; The generating a random number corresponding to the element, and determining whether to replace the element according to the random number, includes: For any language pair, generating a random number corresponding to the language pair and the element, and determining whether to replace the element according to the random number; The replacing the text corresponding to the element in the first description text corresponding to the entity pair with the text described in the target other language, to obtain the second description text corresponding to the entity pair, includes: Replacing the text corresponding to the element in the first description text corresponding to the entity pair with the text describing the element in the target other language in the language pair, to obtain the second description text corresponding to the language pair.

4. The method according to claim 1, wherein, The annotation information includes aliases of elements in the entity pair described in the main language, and the method further includes: Replacing at least some of the elements in the entity pair with aliases described based on the main language according to the annotation information to obtain the first description text; Alternatively, replacing at least some of the elements in the entity pair with aliases described based on the other language according to the annotation information to obtain the second description text.

5. The method according to claim 1, wherein, Generating a multilingual text sequence based on the first description text and the second description text corresponding to the entity pair, including: Using masking to splice the text corresponding to the elements in the first description text or the second description text corresponding to the entity pair to generate a multilingual text sequence.

6. A method for pre-training a language model, including: Obtaining a first description text of a plurality of entity pairs described in the main language in a knowledge graph, where the elements in the entity pair include two entities and the entity relationship between them; Obtaining annotation information corresponding to some or all of the elements in the plurality of entity pairs, and replacing at least some of the elements in the entity pair with text described in another language different from the main language according to the annotation information to obtain a second description text of the entity pair; Generating a multilingual text sequence based on the first description text and the second description text corresponding to the entity pair; Replacing the text corresponding to the entity or the entity relationship in the multilingual text sequence with a mask to generate a sample text sequence; Inputting the sample text sequence into the language model, and predicting the text corresponding to the mask through the language model; Adjusting the language model according to the difference between the predicted text output by the language model and the text replaced with the mask, so as to perform pre-training of the language model based on multilingual knowledge; The obtaining annotation information corresponding to some or all of the elements in the plurality of entity pairs, and replacing at least some of the elements in the entity pair with text described in another language different from the main language according to the annotation information to obtain a second description text of the entity pair includes: Obtaining annotation information corresponding to some or all of the elements in the entity pairs, and determining the target other language; Querying in the annotation information for any element of any entity pair according to the identifier of the element and the target other language; If it is determined according to the query result that the element corresponds to a text described in the target other language, replacing the text corresponding to the element in the first description text corresponding to the entity pair with the text described in the target other language to obtain the second description text corresponding to the entity pair.

7. The method according to claim 6, wherein, The method further includes: Using the multilingual text sequence to perform joint training with at least one of a pre-training task based on multilingual knowledge and a pre-training task based on cloze or a pre-training task based on logical reasoning to obtain a trained language model.

8. A computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the text sequence generation method described in any one of claims 1-5 or the pre-training method of the language model described in any one of claims 6-7.

9. A computer program product comprising computer instructions that direct a computing device to perform operations corresponding to the text sequence generation method according to any one of claims 1-5 or the pre-training method of the language model according to any one of claims 6-7.

Citation Information

Patent Citations

  • Text entity relationship extraction method and model training method

    CN111339774A

  • Knowledge representation learning method, device and equipment and storage medium

    CN111680145A

  • Translation model training method, device and equipment and storage medium

    CN112560510A