Model training method, device, computer equipment, storage medium and program product

By designing the tasks of word recovery, speaker prediction and speech order determination in social conversation scenarios, training the language model solves the problem of language feature learning in social conversation scenarios and achieves better language model performance.

CN114330701BActive Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111200203.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-14
Publication Date
2025-07-11
Estimated Expiration
2041-10-14

AI Technical Summary

Technical Problem

现有的语言模型在社交会话场景中难以有效学习语言特征,因为社交会话内容与独立内容存在较大区别,导致预训练任务不适用。

Method used

Pre-training tasks are designed for social conversation scenarios, including word recovery tasks, speaker prediction tasks and speech order judgment tasks. By performing specific transformation processing on social conversation data, language models are trained to learn language features in social conversations.

Benefits of technology

The trained language model can better understand and encode language features in social conversations, improving language processing capabilities in social conversation scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114330701B_ABST
    Figure CN114330701B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a model training method, apparatus, computer device, storage medium, and program product, which can be applied to the field of natural language processing in artificial intelligence. The method includes: obtaining training data for a language model, where the training data includes conversation data, and the conversation data includes speech content generated by multiple rounds of speech in a social conversation; obtaining a pre-training task for the language model, where the pre-training task includes at least one of the following: a word recovery task, a speaker prediction task, and a speech order determination task; performing transformation processing on the training data according to the pre-training task to obtain training samples for the pre-training task; and based on the training samples, invoking the language model to execute the pre-training task to obtain a trained language model. By using the embodiments of the present application, pre-training tasks for a language model can be proposed for social conversation scenarios, so that the trained language model can better learn the language features in social conversation scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, in particular to the field of artificial intelligence technology, and specifically relates to a model training method, device, computer device, storage medium, and program product. Background Art

[0002] With the continuous development of computer technology, scholars have begun to study theories and methods that enable communication between humans and machines through natural language, and the NPL (Natural Language Processing) technology has emerged as the times require. A core task of NPL technology is to pre-train a language model so that the trained language model can fully learn the language features of natural language.

[0003] Currently, most of the pre-training tasks of language models are proposed for some independent content. The language models trained in this way are not suitable for learning the language features in social conversation scenarios. This is because independent content is generated independently by the author. For example, the content published in news reports, the content recorded in books, etc. are all independent content, while the content in social conversation scenarios is generated by multiple speakers taking turns to speak. There are significant differences between the content in social conversation scenarios and independent content; therefore, there is an urgent need to propose corresponding pre-training tasks for social conversation scenarios to train language models. Summary of the Invention

[0004] Embodiments of this application provide a model training method, device, computer device, storage medium, and program product, which can propose pre-training tasks for language models for social conversation scenarios, so that the trained language models can better learn the language features in social conversation scenarios.

[0005] On the one hand, embodiments of this application provide a model training method, which includes:

[0006] Obtain the training data of the language model. The training data includes conversation data, and the conversation data includes the speech content generated by multiple rounds of speeches in a social conversation. Each round of speech is initiated by a speaker participating in the social conversation;

[0007] Obtain the pre-training tasks of the language model. The pre-training tasks include at least one of the following: word recovery task, speaker prediction task, and speech order determination task;

[0008] Perform transformation processing on the training data according to the task requirements of the pre-training tasks to obtain the training samples of the pre-training tasks;

[0009] Based on training samples, a language model is called to perform a pre-training task to obtain a trained language model; the trained language model is used to encode the conversation data in a social conversation.

[0010] Correspondingly, an embodiment of the present application provides a model training device, and the model training device includes:

[0011] An acquisition unit, configured to acquire training data of a language model, where the training data includes conversation data, and the conversation data includes speech contents generated by multiple rounds of speeches in a social conversation, and each round of speech is initiated by a speaker participating in the social conversation; and acquire a pre-training task of the language model, where the pre-training task includes at least one of the following: a word restoration task, a speaker prediction task, and a speech order determination task;

[0012] A processing unit, configured to perform transformation processing on the training data according to the task requirements of the pre-training task to obtain training samples of the pre-training task; and based on the training samples, call a language model to perform the pre-training task to obtain a trained language model; the trained language model is used to encode the conversation data in a social conversation.

[0013] In one implementation, the pre-training task includes a word restoration task, and the speech content is composed of words; when the processing unit performs transformation processing on the training data according to the task requirements of the pre-training task to obtain training samples of the pre-training task, it is specifically configured to perform the following steps:

[0014] According to the speech order of each round of speech in the social conversation, perform splicing processing on the speaker identifier and speech content of each round of speech in the social conversation in sequence to obtain the reference content of the social conversation;

[0015] Use a replacement identifier to replace the target word in the reference content to obtain training samples of the word restoration task.

[0016] In one implementation, when the processing unit is configured to, based on the training samples, call a language model to perform the pre-training task to obtain a trained language model, it is specifically configured to perform the following steps:

[0017] Obtain a representation vector of the training samples of the word restoration task;

[0018] Use the language model to encode the representation vector of the training samples of the word restoration task to obtain an encoding of the replacement identifier in the training samples of the word restoration task;

[0019] Perform word prediction based on the encoding of the replacement identifier to obtain the probability that the replacement identifier is correctly predicted as the target word;

[0020] Determine the loss information of the word restoration task according to the probability that the replacement identifier is correctly predicted as the target word;

[0021] Update the model parameters of the language model according to the loss information of the word restoration task to obtain a trained language model.

[0022] In one implementation, the representation vector of the training sample of the word restoration task includes: the representation vector of the speaker identifier in the training sample of the word restoration task, the representation vector of the word, and the representation vector of the replacement identifier; any word in the training sample of the word restoration task is represented as a reference word; when the processing unit is used to obtain the representation vector of the reference word, it is specifically used to perform the following steps:

[0023] Obtain the word vector of the reference word;

[0024] Determine the position vector of the reference word according to the arrangement position of the reference word in the speech content to which the reference word belongs;

[0025] Determine the speech content marker vector of the reference word according to the speech content to which the reference word belongs;

[0026] Determine the speaker representation vector of the reference word according to the speaker to which the reference word belongs;

[0027] Determine the representation vector of the reference word according to the word vector of the reference word, the position vector of the reference word, the speech content marker vector of the reference word, and the speaker representation vector of the reference word.

[0028] In one implementation, the pre-training task includes a speaker prediction task; when the processing unit is used to perform transformation processing on the training data according to the task requirements of the pre-training task to obtain the training sample of the pre-training task, it is specifically used to perform the following steps:

[0029] Sequentially splice the speaker identifier and the speech content of each round of speech in the social conversation according to the speech order of each round of speech in the social conversation to obtain the reference content of the social conversation;

[0030] Replace the target speaker identifier in the reference content with a replacement identifier to obtain the training sample of the speaker prediction task.

[0031] In one implementation, when the processing unit is used to call the language model to perform the pre-training task based on the training sample to obtain a trained language model, it is specifically used to perform the following steps:

[0032] Obtain the representation vector of the training sample of the speaker prediction task;

[0033] Encode the representation vector of the training sample of the speaker prediction task by using the language model to obtain the encoding of the replacement identifier in the training sample of the speaker prediction task;

[0034] Perform speaker prediction based on replacement identifier coding to obtain the probability that the replacement identifier is correctly predicted as the target speaker identifier;

[0035] Determine the loss information of the speaker prediction task according to the probability that the replacement identifier is correctly predicted as the target speaker identifier;

[0036] Update the model parameters of the language model according to the loss information of the speaker prediction task to obtain the trained language model.

[0037] In one implementation, the pre-training task includes a speaking order determination task; when the processing unit is used to transform the training data according to the task requirements of the pre-training task to obtain the training samples of the pre-training task, it is specifically used to perform the following steps:

[0038] Concatenate the classification symbol with the speaker identifier and the speech content of each turn in the social conversation to obtain the concatenated content of each turn in the social conversation;

[0039] Perform multiple random-order concatenation processes on the concatenated content of each turn in the social conversation to obtain the training samples of the speaking order determination task;

[0040] Among them, the arrangement order of the concatenated content of each turn in each random-order concatenation process is different, and each random-order concatenation process obtains a training sample of the speaking order determination task.

[0041] In one implementation, when the processing unit is used to call the language model to perform the pre-training task based on the training samples to obtain the trained language model, it is specifically used to perform the following steps:

[0042] Obtain the representation vector of the training samples of the speaking order determination task;

[0043] Use the language model to perform encoding processing on the representation vector of the training samples of the speaking order determination task to obtain the encoding of each classification symbol in the training samples of the speaking order determination task;

[0044] Perform speaking order prediction based on the encoding of each classification symbol to obtain the prediction probability that the predicted order of each turn in the training samples of the speaking order determination task is consistent with the actual order;

[0045] Determine the loss information of the speaking order determination task according to the prediction probability of the training samples of the speaking order determination task;

[0046] Update the model parameters of the language model according to the loss information of the speaking order determination task to obtain the trained language model.

[0047] In one implementation, the pre-training tasks include a word recovery task, a speaker prediction task, and a speaking order determination task; the processing unit is used to call a language model to execute the pre-training tasks based on training samples, and when obtaining a trained language model, it is specifically used to perform the following steps:

[0048] Based on the training samples of the word recovery task, call the language model to execute the word recovery task to obtain the loss information of the word recovery task;

[0049] Based on the training samples of the speaker prediction task, call the language model to execute the speaker prediction task to obtain the loss information of the speaker prediction task;

[0050] Based on the training samples of the speaking order determination task, call the language model to execute the speaking order determination task to obtain the loss information of the speaking order determination task;

[0051] According to the loss information of the word recovery task, the loss information of the speaker prediction task, and the loss information of the speaking order determination task, update the model parameters of the language model to obtain a trained language model.

[0052] In one implementation, the acquisition unit is further used to perform the following steps: acquire a decoding model for the language processing task;

[0053] The processing unit is further used to perform the following steps: encode the conversation data in the social conversation using the trained language model to obtain a conversation encoding;

[0054] According to the task requirements of the language processing task, train the decoding model based on the conversation encoding.

[0055] In one implementation, the language processing task is a conversation summary extraction task; the training data further includes the marked summary of the social conversation; when the processing unit is used to train the decoding model according to the task requirements of the language processing task based on the conversation encoding, it is specifically used to perform the following steps:

[0056] Use the decoding model of the conversation summary extraction task to decode the conversation encoding to obtain the predicted summary of the social conversation;

[0057] Train the decoding model based on the difference between the marked summary and the predicted summary.

[0058] In one implementation, the language processing task is a conversation prediction task; the social conversation includes the speech content generated by N rounds of speeches, where N is an integer greater than 1; the conversation encoding includes the encoding of the speech content of each round in the N rounds of speeches;

[0059] A processing unit, when training the decoding model according to the session encoding in accordance with the task requirements of a language processing task, is specifically configured to perform the following steps:

[0060] Use the decoding model of the session prediction task to decode the encoding of the speech content of the first M rounds of speech in N rounds of speech to obtain the predicted content of the subsequent N - M rounds of speech, where M is a positive integer less than N;

[0061] Train the decoding model according to the difference between the predicted content of the subsequent N - M rounds of speech and the speech content of the subsequent N - M rounds of speech.

[0062] In one implementation, the language processing task is a session retrieval task; the training data further includes retrieval questions for social conversations and marked answers to the retrieval questions;

[0063] The processing unit is further configured to perform the following steps: Encode the retrieval question using the trained language model to obtain the encoding of the retrieval question; wherein, the session encoding includes the encoding of the speech content of each round of speech in the social conversation;

[0064] A processing unit, when training the decoding model according to the session encoding in accordance with the task requirements of a language processing task, is specifically configured to perform the following steps:

[0065] Use the decoding model of the session retrieval task to calculate the similarity between the encoding of the retrieval question and the encoding of the speech content of each round of speech in the social conversation, and perform decoding processing on the encoding of the speech content corresponding to the maximum similarity among the calculated similarities to obtain the predicted answer to the retrieval question;

[0066] Train the decoding model based on the difference between the marked answer and the predicted answer.

[0067] Correspondingly, an embodiment of the present application provides a computer device, which includes a processor and a computer-readable storage medium. Among them, the processor is adapted to implement a computer program, and the computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the above model training method.

[0068] Correspondingly, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is read and executed by the processor of the computer device, the computer device is enabled to perform the above model training method.

[0069] Accordingly, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above model training method.

[0070] An embodiment of the present application proposes a pre-training task for a language model for a social conversation scenario, and trains the language model in the social conversation scenario by invoking the language model to execute the pre-training task; the pre-training task may include at least one of the following: a word recovery task, a speaker prediction task, and a speaking order determination task; the word recovery task can be used to train the ability of the language model to learn word-level features in the social conversation scenario, the speaker prediction task can be used to train the ability of the language model to learn speaker features in the social conversation scenario, and the speaking order prediction task can be used to train the ability of the language model to learn the logical features of the speaking order in the social conversation scenario. The trained language model can be used to encode the conversation data in the social conversation, so that the trained language model can better learn the language features in the social conversation scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0072] Figure 1 is a schematic flowchart of a natural language processing technology provided by an embodiment of the present application;

[0073] Figure 2 is a schematic flowchart of a model training method provided by an embodiment of the present application;

[0074] Figure 3a is a schematic diagram of the model processing logic of a language model provided by an embodiment of the present application;

[0075] Figure 3b is a schematic diagram of the execution process of a word recovery task provided by an embodiment of the present application;

[0076] Figure 3c is a schematic diagram of the execution process of a speaker prediction task provided by an embodiment of the present application;

[0077] Figure 3dIt is a schematic diagram of the execution process of a speaking order determination task provided by an embodiment of the present application;

[0078] Figure 4 It is a schematic flowchart of another model training method provided by an embodiment of the present application;

[0079] Figure 5 It is a schematic diagram of the interface of a model application scenario provided by an embodiment of the present application;

[0080] Figure 6 It is a schematic diagram of the structure of a model training device provided by an embodiment of the present application;

[0081] Figure 7 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0082] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0083] To more clearly understand the technical solutions proposed in the embodiments of the present application, the following first introduces the key terms involved in the embodiments of the present application:

[0084] (1) Artificial intelligence technology. Artificial Intelligence (AI) technology refers to the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.

[0085] (2) Natural language processing technology. Natural language processing technology is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language used in daily life by people, so it has a close connection with the research of linguistics. Natural language processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graph.

[0086] As Figure 1 shown, natural language processing technology can generally be divided into an upstream stage and a downstream stage. The main task of the upstream stage is to pre-train (Pre-Train) a language model through a large-scale training corpus sample, so that the trained language model can fully learn the language features of natural language; the language model can be understood as an encoding model, and the process of the language model learning the language features of natural language can be understood as the process of encoding natural language. That is to say, the language model can be used to encode natural language, and the encoding of natural language is the representation of the language features of natural language. The downstream stage can also be called the fine-tuning (Fine-Tune) stage. The main task of the downstream stage is to train different decoding models for various language processing tasks, so that the trained decoding models can achieve various language processing tasks by decoding the encoding obtained from the language model; the decoding model in the downstream stage can be a very lightweight output layer and can converge quickly based on the small-scale corpus samples of various language processing tasks. It should be noted that the language model mentioned in the embodiments of this application may include, but is not limited to, at least one of the following: BERT (a language model), RoBERTa (a language model), GPT2 (a language model). In the embodiments of this application, the language model is taken as BERT for introduction. When the language model is other situations (such as the language model is RoBERTa, or GPT2, etc.), the relevant descriptions when the language model is BERT can be referred to.

[0087] (3) Social conversation. A social conversation, also known as a multi-party dialog, refers to a conversation or chat among multiple speakers; medical consultations, business negotiations, work regular meetings, etc. belong to specific social conversation scenarios. A social conversation can have the following four conversation characteristics: ① The expression in a social conversation is more colloquial; for example, the idiom "Constant dripping wears away the stone" is often expressed in a social conversation as colloquial expressions like "Success will come with perseverance" or "You have to persevere in finishing this thing"; another example is that colloquial modal particles such as "ah", "ba", "ha" are often used in social conversations. ② Each speaker participating in a social conversation takes turns speaking, and the speakers have different personality traits, conversation purposes, and language habits. ③ There are complex interaction relationships in a social conversation; specifically, there are generally multiple conversation topics in a social conversation, and the speakers speak on each conversation topic. The interaction relationship refers to the topic correlation relationship between the speech contents of each speaker. The topic correlation relationship can include two types. One is that the speech contents of each speaker belong to the same conversation topic, and the other is that the speech contents of each speaker belong to different conversation topics; for example, the speech content of speaker A on conversation topic A is "What to eat this afternoon", the speech content of speaker B on conversation topic B is "Where to eat this afternoon", and the speech content of speaker C on conversation topics A and B is "Order takeout in the company this afternoon". The speech content of speaker C belongs to the same conversation topic (i.e., conversation topic A) as the speech content of speaker A, and the speech content of speaker C also belongs to the same conversation topic (i.e., conversation topic B) as the speech content of speaker B; generally speaking, the more speakers there are and the more conversation topics there are in a social conversation, the more complex the interaction relationship tends to be, and the fewer speakers there are and the fewer conversation topics there are in a social conversation, the simpler the interaction relationship tends to be. ④ There is a certain sequential logic between the speech contents in a social conversation. For example, the speech order of the speech content "Order takeout this afternoon" generally comes after the speech content "What to eat this afternoon".

[0088] Based on the above description of the language model and social conversations, an embodiment of the present application proposes a model training solution. This model training solution proposes three pre-training tasks for the language model according to the conversation characteristics of social conversations, and trains the language model by calling the language model to execute these three pre-training tasks. These three pre-training tasks are the word recovery task, the speaker prediction task, and the speaking order determination task. Among them, the word recovery task can be used to train the language model's ability to learn the characteristics of the constituent words of the speech content in social conversations, the speaker prediction task can be used to train the language model's ability to learn the characteristics of each speaker in social conversations, and the speaking order determination task can be used to train the language model's ability to learn the characteristics of the speaking order logic in social conversations. The language model trained through these three pre-training tasks can better learn the language characteristics in social conversations.

[0089] It should be noted that the model training solution provided by the embodiment of the present application can be executed by a computer device, which can be an intelligent terminal or a server. The intelligent terminal mentioned here can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, a smart TV, etc., but is not limited thereto. The server mentioned here can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0090] In addition, it is hereby stated that the "multiple" mentioned in the embodiment of the present application indicates a quantity of two or more, and the "multiple rounds" mentioned in the embodiment of the present application indicates a quantity of two rounds or more.

[0091] Next, in combination with Figures 2 to 5 the embodiments shown, the model training solution provided by the embodiment of the present application will be introduced in more detail.

[0092] An embodiment of the present application proposes a model training method. This model training method mainly introduces the transformation process of the training data by three pre-training tasks of the language model (that is, the word recovery task, the speaker prediction task, and the speaking order determination task mentioned above), and the specific training process of the three pre-training tasks on the language model. This model training method can be executed by the computer device mentioned above. As Figure 2 shown, this model training method can include the following steps S201 to step S204:

[0093] S201, Obtain the training data of the language model.

[0094] The training data of the language model may include conversation data, and the number of conversation data may be one or more; when the number of conversation data is multiple, the multiple conversation data may come from different social conversations. For example, the first conversation data comes from the first social conversation, and the second conversation data comes from the second social conversation; the multiple conversation data may also come from the same social conversation. Specifically, the multiple conversation data may be from different time periods of the same social conversation. For example, the first conversation data comes from the first time period of the social conversation, and the second conversation data comes from the second time period of the social conversation. Conversation data refers to the data generated in a social conversation. Conversation data may be text data or voice data. That is to say, conversation data may be the text data generated by multiple rounds of speech in text form in a social conversation, or may be the language data generated by multiple rounds of speech in voice form in a social conversation; in the embodiments of the present application, the case where conversation data is text data is taken as an example for introduction. When conversation data is voice data, it can be first converted into text data through speech recognition technology, and then refer to the description in the embodiments of the present application where conversation data is text data.

[0095] The conversation data may include the speaker identifier of multiple rounds of speech in the social conversation and the speech content generated by multiple rounds of speech. The speech content generated by multiple rounds of speech together constitutes the conversation content of the social conversation. Each round of speech is initiated by a speaker participating in the social conversation. For the sake of easy understanding, in the embodiments of the present application, the number of speech rounds in the social conversation may be represented as N rounds, and the conversation content of the social conversation may be represented as D = {U1, U2, …, U N}, D represents the conversation content, {U1, U2, …, U N} represents the speech content of N rounds of speech included in the conversation content; any round of speech in the N rounds of speech may be represented as the i-th round of speech, the speech content of the i-th round of speech may be represented as U i , and the speaker of the i-th round of speech may be represented as s i ; where N is an integer greater than 1, and i is a positive integer less than or equal to N. For example, an exemplary conversation data is as follows:

[0096] A: Where shall we have dinner?

[0097] B: Let's go to a fast food restaurant.

[0098] C: Shall we leave at six?

[0099] The conversation data shown above includes the speaker identification and speech content of three rounds of speeches in the social conversation. The speaker identification of the first round of speeches is "A", that is, "s1=A", and the speech content of the first round of speeches is "Where to eat dinner?", that is, "U1=Where to eat dinner?"; the speaker identification of the second round of speeches is "B", that is, "s2=B", and the speech content of the second round of speeches is "Fast food restaurant.", that is, "U2=Fast food restaurant."; the speaker identification of the third round of speeches is "C", that is, "s3=C", and the speech content of the third round of speeches is "Leave at six?", that is, "U3=Leave at six?".

[0100] The speech content in the conversation data may be composed of words (tokens). When the speech content is expressed in Chinese, a word may refer to a single character (a character may include a Chinese character or a punctuation mark), for example, the speech content "Where are we going to have dinner?" is composed of the seven characters "晚", "饭", "去", "哪", "儿", "吃", and "?"; or, a word may be a split word obtained by semantically splitting the speech content, for example, the speech content "Where are we going to have dinner?" is composed of the five split words "晚饭", "去", "哪", "吃", and "?". When the speech content is expressed in English, a word may refer to an English word. For ease of understanding, in the embodiment of the present application, the number of words contained in the speech content of the i-th round of speeches may be expressed as li, and the speech content of the i-th round of speeches may be expressed as U. i ={w i1 ,w i2 ,…,w ili}, where {w i1 ,w i2 ,…,w ili} represents the li words that make up the content of the i-th round of speech, and li is a positive integer.

[0101] S202, obtaining a pre-training task for a language model.

[0102] Before introducing the pre-training task of the language model, Figure 3a This section introduces the model processing logic of the language model. The model processing logic of the language model can roughly include three steps: content splicing, embedding layer representation, and model encoding. The details are as follows:

[0103] (1) Content splicing, that is, splicing the speaker identification and the speech content. Specifically, according to the speaking order of each turn of speech in the social conversation, the speaker identification and speech content of each turn of speech in the social conversation are spliced in sequence, that is, in the order of the speaker identification of the first turn of speech, the speech content of the first turn of speech, the speaker identification of the second turn of speech, and the speech content of the second turn of speech, the speaker identification and speech content of each turn of speech are spliced. Generally speaking, a classification identifier [CLS] can also be spliced before the speaker identification of the first turn of speech, and a segmentation identifier [SEP] can be spliced at the end of the speech content of the last turn of speech; or, a classification identifier [CLS] can also be spliced before the speaker identification of each turn of speech, and a segmentation identifier [SEP] can be spliced at the end of the speech content of each turn of speech. The embodiments of the present application do not limit this. As Figure 3a shown, the speaker identifications {s1, s2, …, s N} and speech contents {U1, U2, …, U N} of N turns of speech can be spliced in sequence to obtain "[CLS]s1, U1, s2, U2, …, s N , U N ". Applied to the conversation data exemplified in the above step S201, the following splicing result can be obtained: "[CLS][A] Where to have dinner? [B] Fast food restaurant. [C] Leave at six?". Another example is that the speaker identifications {s1, s2, …, s N} and speech contents {U1, U2, …, U N} of N turns of speech can be spliced in sequence to obtain "s1, U1, s2, U2, …, s N , U N [SEP]".

[0104] (2) Embedding layer representation, that is, representing the content obtained by splicing the speaker identification and speech content of each turn of speech in the social conversation in the form of a vector. Specifically, the speaker identification, words, classification identifier, and segmentation identifier in the spliced content are represented in the form of vectors. For the convenience of description, the embodiments of the present application collectively refer to the speaker identification, words, classification identifier, and segmentation identifier as content objects. The embedding layer may include but is not limited to at least one of the following: Token Embedding Layer, Segment Embedding Layer, Soft-Position Embedding Layer, and Speaker Embedding Layer. Among them:

[0105] ① The word embedding layer can map the content object to a word vector. When the content object is a word, this mapping process can be achieved by looking up a word vector table, which includes multiple words with determined word vectors and their corresponding word vectors. By looking up the word vector table, the word vector of the word can be obtained. Moreover, the dimension of the word vector can be determined based on the number of words in the word vector table. For example, the dimension of the word vector is equal to the number of words in the word vector table. When the content object is various identifiers (i.e., speaker identifier, classification identifier, segmentation identifier), their corresponding word vectors can be specified word vectors. And the specified word vectors of various identifiers can be the same or different. For example, the word vectors of the speaker identifier, classification identifier, and segmentation identifier are all the same specified word vector. Another example is that the specified word vectors of the speaker identifier, classification identifier, and segmentation identifier are different from each other.

[0106] ② The position embedding layer can be used to determine the position vector of the content object, and the position vector of the content object can be used to mark the arrangement position of the content object in the speech content. When the content object is a word, the position vector of the word can be determined according to the arrangement position of the word in the speech content to which the word belongs. For example, the word "late" is arranged in the first position in its speech content "Where to have dinner?", and its position vector can be determined as [1]; the word "?" is arranged in the seventh position in its speech content "Where to have dinner?", and its position vector can be determined as [7]. When the content object is various identifiers (i.e., speaker identifier, classification identifier, segmentation identifier), the position vectors are all empty.

[0107] ③ The segment embedding layer can assign different marks to the content objects belonging to different speech contents to distinguish the content objects belonging to different speech contents. This mark is the speech content mark vector of the content object. That is to say, the speech content mark vectors of the content objects belonging to the same speech content are the same, and the speech content mark vectors of the content objects belonging to different speech contents are different. For example, each word in {w 11 ,w 12 ,…,w 1l1} belongs to the speech content of the first round of speech, then the speech content mark vectors of each word in {w 11 ,w 12 ,…,w 1l1} are all [A]; each word in {w 21 ,w 22 ,…,w 2l2} belongs to the speech content of the second round of speech, then the speech content mark vectors of each word in {w 21 ,w 22 ,…,w 2l2The speech content marker vector for each word in {} is [B]. The segment embedding layer does not distinguish between words and various identifiers during marking, that is, the segment embedding layer marks various identifiers in the same way as words.

[0108] ④ The marking method of the speaker representation layer is similar to that of the segment embedding layer. It can mark content objects belonging to different speakers with different marks to distinguish content objects belonging to different speakers. This mark is the speaker representation vector of the content object; that is to say, the speaker representation vectors of content objects belonging to the same speaker are the same, and the speaker representation vectors of content objects belonging to different speakers are different. For example, if each word in {w 11 , w 12 , …, w 1l1} and each word in {w 21 , w 22 , …, w 2l2} belongs to the same speaker, then each word in {w 11 , w 12 , …, w 1l1} and each word in {w 21 , w 22 , …, w 2l2} has the same speaker representation vector (for example, they are all [α]). Another example is that if each word in {w 11 , w 12 , …, w 1l1} belongs to the first speaker and each word in {w 21 , w 22 , …, w 2l2} belongs to the second speaker, then each word in {w 11 , w 12 , …, w 1l1} has a speaker representation vector of [α], and each word in {w 21 , w 22 , …, w 2l2} has a speaker representation vector of [β]. The speaker representation layer also does not distinguish between words and various identifiers during marking, that is, the speaker representation layer marks various identifiers in the same way as words.

[0109] It should also be noted that for any word or identifier, the vectors of the above four embedding layers can be concatenated to obtain the final representation vector of the word or identifier. That is, the word vector, position vector, utterance content marker vector, and speaker representation vector of any word can be concatenated to obtain the representation vector of the word; the word vector, position vector, utterance content marker vector, and speaker representation vector of any identifier (i.e., the speaker identifier, classification identifier, or segmentation identifier mentioned above) can be concatenated to obtain the representation vector of the identifier.

[0110] (3) Model encoding, that is, using a language model to encode the representation vectors of each word to obtain the encoding of each word, and encoding the representation vectors of each identifier to obtain the encoding of each identifier. As Figure 3a shown, the content "[CLS]s1,U1,s2,U2,…,s N ,U N " concatenated in (1) above can obtain the following encoding result after being encoded by the language model "e c ,e s1 ,E U1 ,e s2 ,E U2 ,…,e sN ,E UN ", where e c represents the encoding of the classification symbol [CLS], e sN represents the encoding of the speaker identifier of the Nth utterance, and E UN represents the encoding of the content of the Nth utterance.

[0111] After introducing the model processing logic of the language model, the pre-training tasks of the language model will be introduced below. As described above, the pre-training tasks of the language model can include but are not limited to at least one of the following: word recovery task, speaker prediction task, and speech order determination task. Among them, the word recovery task can be used to train the language model to learn the ability of the characteristics of the constituent words of the utterance content in a social conversation, the speaker prediction task can be used to train the language model to learn the ability of the characteristics of each speaker in a social conversation, and the speech order determination task can be used to train the language model to learn the ability of the characteristics of the speech order logic in a social conversation.

[0112] S203, transform the training data according to the task requirements of the pre-training task to obtain the training samples of the pre-training task.

[0113] The ways in which the word recovery task, speaker prediction task, and speech order determination task transform the training data are different, as follows:

[0114] (1) When the pre-training task includes a word recovery task, the process of transforming the training data according to the task requirements of the word recovery task to obtain the training samples of the word recovery task may include: sequentially splicing the speaker identifiers and speech contents of each turn of speech in the social conversation included in the conversation data according to the order of speech of each turn of speech in the social conversation to obtain the reference content of the social conversation; using a replacement identifier to replace the target words in the reference content to obtain the training samples of the word recovery task. It should be noted that the number of target words can be one or more, that is, part of the words in the reference content can be replaced by a replacement identifier (for example, it can be [mask]); the number of target words can be calculated according to the first replacement ratio and the total number of words in the reference text. For example, the number of target words is equal to the first replacement ratio (for example, it can be 10%, 20%, etc.) multiplied by the total number of words in the reference text. The reference text obtained by splicing the conversation data as exemplified in step S201 above is "[CLS][A] Where to have dinner?[B] Fast food restaurant.[C] Leave at six?", and the training samples obtained by transforming according to the task requirements of the word recovery task can be "[CLS][A] Where to have [mask] for dinner?[B] Fast [mask] restaurant.[C] Leave at six?", and the two target words "eat" and "restaurant" in the reference content are replaced by the replacement identifier [mask].

[0115] (2) When the pre-training task includes a speaker prediction task, the process of transforming the training data according to the task requirements of the speaker prediction task to obtain the training samples of the speaker prediction task may include: sequentially splicing the speaker identifiers and speech contents of each turn of speech in the social conversation included in the conversation data according to the order of speech of each turn of speech in the social conversation to obtain the reference content of the social conversation; using a replacement identifier to replace the target speaker identifiers in the reference content to obtain the training samples required for the speaker prediction task. It should be noted that the number of target speaker identifiers can be one or more, that is, part of the speaker identifiers in the reference content can be replaced by a replacement identifier (for example, it can be [mask]); the number of target speaker identifiers can be calculated according to the second replacement ratio and the total number of speaker identifiers in the reference text. For example, the number of target speaker identifiers is equal to the second replacement ratio (for example, it can be 10%, 20%, etc.) multiplied by the total number of speaker identifiers in the reference text. The reference text obtained by splicing the conversation data as exemplified in step S201 above is "[CLS][A] Where to have dinner?[B] Fast food restaurant.[C] Leave at six?", and the training samples obtained by transforming according to the task requirements of the speaker prediction task can be "[CLS][A] Where to have dinner?[mask] Fast food restaurant.[C] Leave at six?", and one target speaker identifier "[B]" in the reference content is replaced by the replacement identifier [mask].

[0116] (3) When the pre-training task includes a speaking order determination task, the process of transforming the training data according to the task requirements of the speaking order determination task to obtain the training samples of the speaking order determination task may include: First, the classification symbol can be concatenated with the speaker identifier and the speech content of each turn of speech in the social conversation included in the conversation data to obtain the concatenated content of each turn of speech in the social conversation; that is to say, after the speaker identifier and the speech content of each turn of speech are concatenated, a classification identifier [CLS] is added before the speaker identifier of each turn of speech; in the conversation data exemplified in the above step S201, the concatenated content of the first turn of speech is "[CLS][A] Where shall we have dinner?", the concatenated content of the second turn of speech is "[CLS][B] How about a fast food restaurant.", and the concatenated content of the third turn of speech is "[CLS][C] Leave at six?". Then, the concatenated content of each turn of speech in the social conversation can be concatenated multiple times in a random order to obtain the training samples of the speaking order determination task, where the arrangement order of the concatenated content of each turn of speech in each random order concatenation process is different, and each random order concatenation process obtains a training sample of the speaking order determination task; that is to say, when the concatenated content of each turn of speech is concatenated again, the speaking order is not considered, and it can be concatenated in a random order. For the conversation data exemplified in the above step S201, the training sample obtained by transforming according to the task requirements of the speaking order determination task can be "[CLS][A] Where shall we have dinner?[CLS][B] How about a fast food restaurant.[CLS][C] Leave at six?", and the speaking order of each turn of speech in this training sample is not reversed; it can also be "[CLS][B] How about a fast food restaurant.[CLS][C] Leave at six?[CLS][A] Where shall we have dinner?", and the speaking order of each turn of speech in this training sample is reversed; there can be many other situations for the conversation data exemplified in step S201 to obtain the training sample by transforming according to the task requirements of the speaking order determination task, and no further examples will be given here.

[0117] The transformation processing scheme for the speech order determination task described above for the conversation data is a relatively general scheme. The speech order determination task also commonly uses the following transformation processing scheme: After splicing the classification symbol with the speaker identifier and the speech content of each turn of the social conversation included in the conversation data to obtain the spliced content of each turn of the social conversation, the spliced content of each turn can be divided into two parts D1 and D2, and then the spliced content of each turn is spliced again in the order of D1⊕D2 or D2⊕D1 to obtain the training sample for the speech order determination task. For the conversation data exemplified in step S201 above, the spliced content of each turn can be divided into the following two parts: D1 = "[CLS][A] Where to have dinner?[CLS][B] Fast food restaurant.[CLS][C] Leave at six?", D2 = "[CLS][C] Leave at six?", D1⊕D2 = "[CLS][A] Where to have dinner?[CLS][B] Fast food restaurant.[CLS][C] Leave at six?", D2⊕D1 = "[CLS][C] Leave at six?[CLS][A] Where to have dinner?[CLS][B] Fast food restaurant.". Another example, for the conversation data exemplified in step S201 above, the spliced content of each turn can be divided into the following two parts: D1 = "[CLS][A] Where to have dinner?", D2 = "[CLS][B] Fast food restaurant.[CLS][C] Leave at six?", D1⊕D2 = "[CLS][A] Where to have dinner?[CLS][B] Fast food restaurant.[CLS][C] Leave at six?", D2⊕D1 = "[CLS][B] Fast food restaurant.[CLS][C] Leave at six?[CLS][A] Where to have dinner?".

[0118] It should be noted that the conversation data for transformation processing in the word recovery task can refer to the first conversation data in the training data, the conversation data for transformation processing in the speaker prediction task can refer to the second conversation data in the training data, and the conversation data for transformation processing in the speech order determination task can refer to the third conversation data in the training data; the first conversation data, the second conversation data, and the third conversation data can be the same conversation data or can be mutually different conversation data, and the embodiments of the present application do not limit this.

[0119] S204, Based on the training samples, call the language model to perform a pre-training task to obtain a trained language model.

[0120] As described in step S203 above, the ways of transforming the training data in the word recovery task, the speaker prediction task, and the speech order determination task are different, and the processes of calling the language model to perform the word recovery task, the speaker prediction task, and the speech order determination task are also different, specifically as follows:

[0121] (1) For the training samples of the word recovery task, the process of calling the language model to execute the word recovery task and obtaining the trained language model can be referred to Figure 3b , Figure 3b where w ij a type of symbol represents a word, s i a type of symbol represents a speaker identifier, p ij a type of symbol represents a replacement identifier (the replacement identifier for the target word), e c a type of symbol represents the encoding of the classification identifier [CLS], e si a type of symbol represents the encoding of the speaker identifier, e wij a type of symbol represents the encoding of the word, e pij a type of symbol represents the encoding of the replacement identifier; Figure 3b Taking the example of using the replacement identifier to replace a target word in the first session data of the word recovery task for illustration. The process of calling the language model to execute the word recovery task and obtaining the trained language model can specifically include: ① The representation vector of the training samples of the word recovery task can be obtained. The representation vector of the training samples of the word recovery task can include: the representation vector of the speaker identifier in the training samples of the word recovery task, the representation vector of the word, and the representation vector of the replacement identifier. ② The language model can be used to encode the representation vector of the training samples of the word recovery task to obtain the encoding of the training samples of the word recovery task; specifically, that is, to encode the representation vector of the speaker identifier, the representation vector of the word, and the representation vector of the replacement identifier in the training samples of the word recovery task to obtain the encoding of the speaker identifier, the encoding of the word, and the encoding of the replacement identifier in the training samples of the word recovery task. ③ Word prediction can be performed based on the encoding of the replacement identifier to obtain the probability that the replacement identifier is correctly predicted as the target word. Specifically, the word prediction layer can be used to perform word prediction on the encoding of the replacement identifier to obtain the scores of the replacement identifier being predicted as each preset word in the model vocabulary, and the model vocabulary contains the target word; then the score that the replacement identifier is correctly predicted as the target word it replaces can be input into the activation layer, and the activation layer can map the score to a value between the interval [0, 1], thereby obtaining the probability that the replacement identifier is correctly predicted as the target word it replaces. ④ The loss information of the word recovery task can be determined according to the probability that the replacement identifier is correctly predicted as the target word it replaces; when using the replacement identifier to replace multiple target words in the first session data of the word recovery task, the loss information of the word recovery task is determined according to the probability that each replacement identifier is correctly predicted as the corresponding target word it replaces; further, the model parameters of the language model can be updated according to the loss information of the word recovery task to obtain the trained language model.

[0122] Among them, the calculation process of the loss information of the word recovery task can be referred to the following formula 1:

[0123]

[0124] As shown in the above formula 1, L w represents the loss information of the word restoration task; z represents any replaced target word; Z represents the set formed by the replaced target words; N w represents the model vocabulary; represents the labels of each preset word in the model vocabulary. The label of the target word replaced by the replacement identifier in the model vocabulary is 1, and the labels of other preset words are 0; represents the probability that the replacement identifier is correctly predicted as the target word it replaces.

[0125] The loss information of the word restoration task may refer to the loss value calculated by the above formula 1; the model parameters of the language model are updated according to the loss information of the word restoration task to obtain a trained language model. Specifically, it may refer to: updating the model parameters of the language model in the direction of reducing the loss value to obtain a trained language model; the "in the direction of reducing the loss value" mentioned in the embodiments of the present application refers to: the model optimization direction with the goal of minimizing the loss value; through this direction of model optimization, the loss value generated by the language model again after each optimization needs to be less than the loss value generated by the language model before optimization. For example, if the loss value of the language model calculated this time is 0.85, then after optimizing the model parameters of the language model in the direction of reducing the loss value, the loss value generated by optimizing the language model should be less than 0.85. The process of updating the model parameters of the language model based on the loss information mentioned in the embodiments of the present application can be referred to the above description.

[0126] In addition, to obtain the representation vectors of the training samples for the word restoration task, that is, to obtain the representation vectors of the speaker identifiers, words, and replacement identifiers in the training samples for the word restoration task. The processes of obtaining the representation vectors of the speaker identifiers and words can refer to the specific descriptions in step S202 above. Taking any word in the speech content (for example, it can be represented as a reference word) as an example, the process of obtaining the representation vector of the reference word may include: obtaining the word vector of the reference word; determining the position vector of the reference word according to the arrangement position of the reference word in the speech content to which the reference word belongs; determining the speech content marker vector of the reference word according to the speech content to which the reference word belongs; determining the speaker representation vector of the reference word according to the speaker to which the reference word belongs; determining the representation vector of the target word according to the word vector of the reference word, the position vector of the reference word, the speech content marker vector of the reference word, and the speaker representation vector of the reference word; that is, performing vector concatenation on the word vector of the target word, the position vector of the target word, the speech content marker vector of the reference word, and the speaker representation vector of the reference word to obtain the representation vector of the reference word. The manner of obtaining the representation vector of the replacement identifier is the same as the manner of obtaining various identifiers (i.e., the speaker identifier, classification identifier, and segmentation identifier mentioned above) in step S202.

[0127] It should be noted that the word prediction function of the word prediction layer can be implemented by, for example, an MLP (Multi-Layer Perceptron, multi-layer perceptron). An MLP is a forward-structured artificial neural network that maps a set of input vectors to a set of output vectors. An MLP can be regarded as a directed graph composed of multiple node layers, and each layer is fully connected to the next layer. Except for the input nodes, each node is a neuron (or called a processing unit) with a non-linear activation function. The mapping function of the activation layer can be implemented by an activation function, and the activation function can be, for example, softmax, sigmoid, etc.

[0128] (2) Based on the training samples for the speaker prediction task, the process of calling the language model to perform the speaker prediction task and obtaining the trained language model can be referred to Figure 3c , Figure 3c in which w ij A type of symbol represents a word, s i A type of symbol represents a speaker identifier, q i A type of symbol represents a replacement identifier (the replacement identifier for replacing the target speaker identifier), e c A type of symbol represents the encoding of the classification identifier [CLS], e si A type of symbol represents the encoding of the speaker identifier, e wij A type of symbol represents the encoding of a word, e qiA class of symbol representation replacement identifier encoding; Figure 3c Taking the example of replacing a target speaker identifier in the second session data of the speaker prediction task with a replacement identifier for illustration. The process of calling a language model to perform the speaker prediction task and obtaining a trained language model can specifically include: ① The representation vectors of the training samples of the speaker prediction task can be obtained. The representation vectors of the training samples of the speaker prediction task can include: the representation vectors of the speaker identifiers in the training samples of the speaker prediction task, the representation vectors of words, and the representation vectors of replacement identifiers. ② The language model can be used to encode the representation vectors of the training samples of the speaker prediction task to obtain the encoded training samples of the speaker prediction task; specifically, that is, to encode the representation vectors of the speaker identifiers, the representation vectors of words, and the representation vectors of replacement identifiers in the training samples of the speaker prediction task to obtain the encoded speaker identifiers, the encoded words, and the encoded replacement identifiers in the training samples of the speaker prediction task. ③ Speaker prediction can be performed based on the encoded replacement identifier to obtain the probability that the replacement identifier is correctly predicted as the target speaker identifier. Specifically, the speaker prediction layer can be used to perform speaker prediction on the encoded replacement identifier to obtain the scores of the replacement identifier being predicted as all speaker identifiers in the second session data; then the score of the replacement identifier being correctly predicted as the target speaker identifier it replaces can be input into the activation layer, and the activation layer can map the score to a value between the interval [0, 1], thereby obtaining the probability that the replacement identifier is correctly predicted as the target speaker identifier it replaces. ④ The loss information of the speaker prediction task can be determined according to the probability that the replacement identifier is correctly predicted as the target speaker identifier it replaces; when multiple target speaker identifiers in the second session data of the speaker prediction task are replaced with replacement identifiers, the loss information of the speaker prediction task is determined according to the probabilities that each replacement identifier is correctly predicted as the corresponding target speaker identifier it replaces; further, the model parameters of the language model can be updated according to the loss information of the speaker prediction task to obtain a trained language model.

[0129] Among them, the calculation process of the loss information of the speaker prediction task can be seen in the following formula 2:

[0130]

[0131] As shown in the above formula 2, L r represents the loss information of the speaker prediction task; g represents any replaced target speaker identifier; G represents the set formed by the replaced target speaker identifiers; N r represents the set formed by all speaker identifiers in the second session data; Labels for each speaker ID in the set formed by all speaker IDs in the second session data. The label of the target speaker ID replaced by the replacement ID in the speaker ID set is 1, and the labels of other speaker IDs are 0; Indicates the probability that the replacement ID is correctly predicted as the target speaker ID it replaces.

[0132] Similarly, the speaker prediction function of the speaker prediction layer can be implemented by, for example, an MLP (Multi-Layer Perceptron), and the mapping function of the activation layer can be implemented by an activation function, such as softmax, sigmoid, etc.

[0133] (3) Based on the training samples of the speech order determination task, the process of calling the language model to perform the speech order determination task and obtaining the trained language model can be referred to Figure 3d , Figure 3d where w ij One type of symbol represents a word, s i One type of symbol represents a speaker ID, e ci One type of symbol represents the encoding of the classification identifier [CLS], e si One type of symbol represents the encoding of the speaker ID, e wij One type of symbol represents the encoding of a word, t ci One type of symbol represents the interactive encoding of the classification identifier [CLS]. Based on the training samples of the speech order determination task, the process of calling the language model to perform the speech order determination task and obtaining the trained language model can specifically include: ① The representation vector of the training samples of the speech order determination task can be obtained. The representation vector of the training samples of the speech order determination task can include: the representation vector of the speaker ID, the representation vector of the word, and the representation vector of the classification identifier in the training samples of the speech order determination task. ② The language model can be used to encode the representation vector of the training samples of the speech order determination task to obtain the encoding of the training samples of the speech order determination task; specifically, that is, to encode the representation vector of the speaker ID, the representation vector of the word, and the representation vector of the classification identifier in the training samples of the speech order determination task to obtain the encoding of the speaker ID, the encoding of the word, and the encoding of the classification identifier in the training samples of the speech order determination task. ③ Based on the encodings of each classification symbol in the training samples of the speech order determination task, the speech order can be predicted to obtain the prediction probability that the predicted order of each turn of speech in the training samples of the speech order determination task is consistent with the actual order. Specifically, an interactive encoding layer (such as Figure 3dThe interaction coding layer 1 and the interaction coding layer 2 as shown) further encode the encodings of each classification symbol to obtain the interaction encoding of each classification symbol. Further, a speaking order determination layer can be used to determine the speaking order of the interaction encodings of each classification symbol, and obtain the prediction score that the predicted order of each turn of speech in the training sample of the speaking order determination task is consistent with the actual order. Then, the prediction score that the predicted order of each turn of speech in the training sample of the speaking order determination task is consistent with the actual order can be input into the activation layer, and the activation layer can map the prediction score to a value between the interval [0, 1], so as to obtain the prediction probability that the predicted order of each turn of speech in the training sample of the speaking order determination task is consistent with the actual order. ④ The loss information of the speaking order determination task can be determined according to the prediction probability of the training sample of the speaking order determination task; when there are multiple training samples of the speaking order determination task, the loss information of the speaking order determination task can be determined according to the prediction probabilities of each training sample; and then the model parameters of the language model can be further updated according to the loss information of the speaking order determination task to obtain a trained language model.

[0134] Among them, the calculation process of the loss information of the speaking order determination task can be seen in the following formula 3:

[0135]

[0136] As shown in the above formula 3, L t represents the loss information of the speaking order determination task; t represents any training sample of the speaking order determination task; T represents all training samples of the speaking order determination task. N T represents the relationship between the speaking order in the training sample of the speaking order determination task and the speaking order in the third session data, including two cases of inversion (i.e., inconsistent order) and non-inversion (i.e., consistent order). y tk represents the labels in two cases of inversion (i.e., inconsistent order) and non-inversion (i.e., consistent order); when the speaking order in the training sample of the speaking order determination task is consistent with the speaking order in the third session data, the non-inversion label is 1 and the inversion label is 0; when the speaking order in the training sample of the speaking order determination task is inconsistent with the speaking order in the third session data, the non-inversion label is 0 and the inversion label is 1. p tk represents the prediction probability that the predicted order of each turn of speech in the training sample of the speaking order determination task is consistent with the actual order.

[0137] Similarly, the speech order determination function of the speech order determination layer can be implemented by, for example, an MLP (Multi-Layer Perceptron), and the mapping function of the activation layer can be implemented by an activation function, such as softmax, sigmoid, etc. The interactive encoding layer (TransformerLayer, TL) is a neural network model constructed based on the self-attention mechanism. For a set of vectors, the encoder of the TL deeply encodes this set of vectors through a multi-head self-attention interaction layer (Multi-HeadSelfAttentionLayer) and a neural network in sequence. In the specific application of the TL encoder, multiple multi-head self-attention interaction layers can be stacked.

[0138] It should be noted that the content described in the above (1)-(3) separately trains the language model using the word recovery task, the speaker prediction task, and the speech order determination task. In the actual usage scenario, in order to enable the language model to more fully learn the language features in the social conversation scenario, the word recovery task, the speaker prediction task, and the speech order determination task are often used to jointly train the language model; that is, the pre-training task can include the word recovery task, the speaker prediction task, and the speech order determination task. Based on the training samples, the process of calling the language model to execute the pre-training task and obtaining the trained language model can include: based on the training samples of the word recovery task, calling the language model to execute the word recovery task to obtain the loss information of the word recovery task; based on the training samples of the speaker prediction task, calling the language model to execute the speaker prediction task to obtain the loss information of the speaker prediction task; based on the training samples of the speech order determination task, calling the language model to execute the speech order determination task to obtain the loss information of the speech order determination task; and updating the model parameters of the language model according to the loss information of the word recovery task, the loss information of the speaker prediction task, and the loss information of the speech order determination task to obtain the trained language model.

[0139] Among them, the process of updating the model parameters of the language model according to the loss information of the word recovery task, the loss information of the speaker prediction task, and the loss information of the speaking order determination task to obtain the trained language model may include: determining the loss information of the language model according to the loss information of the word recovery task, the loss information of the speaker prediction task, and the loss information of the speaking order determination task. The loss information of the language model may be equal to the sum of the loss information of the word recovery task, the loss information of the speaker prediction task, and the loss information of the speaking order determination task. Then, the language model may be trained according to the loss information of the language model to obtain the trained language model. The process of determining the loss information of the language model according to the loss information of the word recovery task, the loss information of the speaker prediction task, and the loss information of the speaking order determination task can be seen in the following formula 4:

[0140] L total =L w +L r +L t Formula 4

[0141] As shown in the above formula 4, L total represents the loss information of the language model; L w represents the loss information of the word recovery task, and the specific calculation process can be seen in the above formula 1; L r represents the loss information of the speaker prediction task, and the specific calculation process can be seen in the above formula 2; L t represents the loss information of the speaking order determination task, and the specific calculation process can be seen in the above formula 3. It should also be noted that the model training scheme described in the embodiments of the present application only introduces the process of one training of the language model. In the actual training scenario, the language model needs to be trained iteratively multiple times until the loss information of the language model meets the convergence condition (for example, the loss value indicated by the loss information of the language model is less than or equal to the convergence threshold).

[0142] The embodiments of this application propose a pre-training task for a language model in the social conversation scenario, and train the language model in the social conversation scenario by invoking the language model to execute the pre-training task; the pre-training task may include at least one of the following: word recovery task, speaker prediction task, and speech order determination task; the word recovery task trains the language model by predicting words for the replacement identifiers used to replace the target words, so that the language model has the ability to learn word-level features in the social conversation scenario; the speaker prediction task trains the language model by predicting speakers for the replacement identifiers used to replace the target speaker identifiers, so that the language model has the ability to learn speaker features in the social conversation scenario; the speech order prediction task determines whether the speech order in the training sample is reversed, so that the language model has the ability to learn the speech order logical features in the social conversation scenario; in this way, the trained language model can better learn the language features in the social conversation scenario when encoding the conversation data in the social conversation.

[0143] The embodiments of this application propose a model training method, which mainly introduces the training process of the language model and the decoding model for specific language processing tasks, as well as the application of the language model and the decoding model in specific language processing tasks. This model training method can be executed by the aforementioned computer device. As Figure 4 shown, this model training method may include the following steps S401 to step S407:

[0144] S401, obtain the training data of the language model.

[0145] S402, obtain the pre-training task of the language model.

[0146] S403, perform transformation processing on the training data according to the task requirements of the pre-training task to obtain the training samples of the pre-training task.

[0147] S404, based on the training samples, invoke the language model to execute the pre-training task to obtain the trained language model.

[0148] In the embodiments of this application, the execution process of step S401 is the same as Figure 2 the execution process of step S201 in the embodiment shown, the execution process of step S402 is the same as Figure 2 the execution process of step S202 in the embodiment shown, the execution process of step S403 is the same as Figure 2 the execution process of step S203 in the embodiment shown, the execution process of step S404 is the same as Figure 2 the execution process of step S204 in the embodiment shown. The execution processes of each step from step S401 to step S404 can be referred to the above Figure 2Descriptions of corresponding steps in the illustrated embodiments are not elaborated herein.

[0149] S405. Obtain a decoding model for the language processing task.

[0150] The language processing task may include at least one of the following: session summary extraction task, session prediction task, and session retrieval task. Among them, the session summary extraction task is used to train the decoding model's ability to extract session summaries from the session data of social conversations; the session prediction task is used to train the decoding model to predict the content of the next round or multiple rounds of speech, that is, the ability to generate speech content; the session retrieval task is used to train the decoding model to retrieve speech content that matches the retrieval question from the session data of social conversations according to the retrieval question.

[0151] S406. Encode the session data in the social conversation using the trained language model to obtain a session encoding.

[0152] The process of encoding the session data in the social conversation using the trained language model can refer to Figure 2 the description of step S202 in the illustrated embodiments, which generally includes the processes of content splicing, embedding layer representation, and vector encoding, and will not be elaborated herein. For the convenience of introducing the content of the embodiments of the present application, the number of speech rounds included in the social conversation can be represented as N rounds, and the session encoding can include the encoding of the content of each round of speech in the N rounds of speech, where N is an integer greater than 1.

[0153] S407. Train the decoding model according to the session encoding in accordance with the task requirements of the language processing task.

[0154] The training processes and specific application scenarios of the decoding models for the session summary extraction task, session prediction task, and session retrieval task are different. The following will separately introduce the training processes and specific application scenarios of the decoding models for the session summary extraction task, session prediction task, and session retrieval task:

[0155] (1) When the language processing task is the session summary extraction task, the training data may further include the marked summary of the social conversation. The process of training the decoding model according to the session encoding in accordance with the requirements of the session summary extraction task may include: using the decoding model of the session summary extraction task to decode the session encoding to obtain the predicted summary of the social conversation; training the decoding model based on the difference between the marked summary and the predicted summary; and the trained language model can also be optimized or fine-tuned based on the difference between the marked summary and the predicted summary to strengthen the encoding ability of the language model in the session summary extraction task, that is, the language feature learning ability.

[0156] The decoding model trained according to the session summary extraction task can be applied to the social application scenario of session summary extraction. For example, the trained language model and the decoding model trained according to the session summary extraction task can be deployed in an online medical consultation platform or application, so that users of the online medical consultation platform or application can quickly obtain the consultation summary of the medical consultation and quickly understand the main content of the medical consultation; as Figure 5 shown in the service interface of the medical application, the speaker identification and speech content of each round of speech in the medical consultation are displayed. The service interface can also provide an access point for obtaining the consultation summary, and the consultation summary can be quickly obtained through this access point. Another example is that the trained language model and the decoding model trained according to the session summary extraction task can be deployed in an online meeting platform or application, so that users of the online meeting platform or application can quickly obtain the meeting summary and quickly understand the main content of the meeting.

[0157] (2) When the language processing task is a session prediction task, according to the requirements of the session prediction task, the process of training the decoding model according to the session encoding may include: using the decoding model of the session prediction task to decode the encoding of the speech content of the first M rounds of speech in N rounds of speech to obtain the predicted content of the subsequent N-M rounds of speech, where M is a positive integer less than N; training the decoding model according to the difference between the predicted content of the subsequent N-M rounds of speech and the speech content of the subsequent N-M rounds of speech; and according to the difference between the predicted content of the subsequent N-M rounds of speech and the speech content of the subsequent N-M rounds of speech, the trained language model can also be optimized or fine-tuned to strengthen the encoding ability of the language model under the session prediction task, that is, the language feature learning ability.

[0158] The decoding model trained according to the session prediction task can be applied to the social application scenario of session prediction. For example, the trained language model and the decoding model trained according to the session prediction task can be deployed in a social application program, so that the social application program can generate new speech content (for example, belonging to the same session subject) related to the original speech content in the group chat session based on the speech content of each speaker in the group chat session, and publish the new speech content to the group chat session in the identity of an intelligent group assistant or an intelligent group steward, enhancing the entertainment and interactivity.

[0159] (3) When the language processing task is a conversation retrieval task, the training data may further include retrieval questions for social conversations and marked answers to the retrieval questions; it is also necessary to encode the retrieval questions using the trained language model to obtain the encoded retrieval questions; among them, the conversation encoding includes the encoding of the speech content of each turn in the social conversation. According to the requirements of the conversation prediction task, the process of training the decoding model based on the conversation encoding may include: using the decoding model of the conversation retrieval task to calculate the similarity between the encoded retrieval questions and the encoded speech content of each turn in the social conversation, and performing decoding processing on the encoded speech content corresponding to the maximum similarity among the calculated similarities to obtain the predicted answer to the retrieval question; training the decoding model based on the difference between the marked answer and the predicted answer; it is also possible to optimize or fine-tune the trained language model based on the difference between the marked answer and the predicted answer to strengthen the encoding ability of the language model in the conversation retrieval task, that is, the language feature learning ability.

[0160] The decoding model trained according to the conversation retrieval task can be applied to the social application scenario of conversation retrieval. For example, the trained language model and the decoding model trained based on the conversation retrieval task can be deployed in a social application program, so that the social application program can find the speech content matching the retrieval question in the chat record based on the retrieval question.

[0161] It should be noted that the conversation data used for training the decoding model in the conversation summary extraction task may refer to the fourth conversation data in the training data, the conversation data used for training the decoding model in the conversation prediction task may refer to the fifth conversation data in the training data, and the conversation data used for training the decoding model in the conversation retrieval task may refer to the sixth conversation data in the training data; the fourth conversation data, the fifth conversation data, and the sixth conversation data may be the same conversation data as the aforementioned first conversation data, second conversation data, and third conversation data, or they may be different from each other. The embodiments of the present application do not limit this.

[0162] It should also be noted that the model training scheme described in the embodiments of the present application only introduces the process of training the decoding model once. In the actual training scenario, the decoding model needs to be trained iteratively multiple times until the loss information of the decoding model meets the convergence condition (for example, the loss value indicated by the loss information of the decoding model is less than or equal to the convergence threshold).

[0163] In the embodiments of the present application, after the language model is trained, a decoding model can be designed or trained for a specific language processing task, so that the language model and the decoding model can cooperate to implement the specific language processing task. And during the process of training the decoding model, the language model can be further optimized or adjusted, so that the language model can better adapt to the specific language processing task and strengthen the language feature learning ability of the language model under the specific language processing task.

[0164] The method of the embodiments of the present application is elaborated in detail above. To facilitate better implementation of the above solutions of the embodiments of the present application, correspondingly, the device of the embodiments of the present application is provided below.

[0165] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a model training device provided by the embodiments of the present application. The model training device can be set in the computer device provided by the embodiments of the present application. The computer device can be the intelligent terminal or server mentioned in the above method embodiments; in some embodiments, the model training device can be a computer program (including program code) running on the computer device, and the model training device can be used to execute Figure 2 or Figure 4 the corresponding steps in the method embodiments shown. Please refer to Figure 6 , the model training device may include the following units:

[0166] An obtaining unit 601, configured to obtain training data of the language model. The training data includes session data, and the session data includes speech contents generated by multiple rounds of speeches in a social session. Each round of speech is initiated by a speaker participating in the social session; and obtain a pre-training task of the language model. The pre-training task includes at least one of the following: a word recovery task, a speaker prediction task, and a speech order determination task;

[0167] A processing unit 602, configured to perform transformation processing on the training data according to the task requirements of the pre-training task to obtain training samples of the pre-training task; and based on the training samples, call the language model to execute the pre-training task to obtain a trained language model; the trained language model is used to encode the session data in the social session.

[0168] In one implementation manner, the pre-training task includes a word recovery task, and the speech content is composed of words; when the processing unit 602 performs transformation processing on the training data according to the task requirements of the pre-training task to obtain training samples of the pre-training task, it is specifically configured to perform the following steps:

[0169] According to the speech order of each round of speech in the social session, perform splicing processing on the speaker identifier and the speech content of each round of speech in the social session in sequence to obtain the reference content of the social session;

[0170] Replace the target word in the reference content with a replacement identifier to obtain a training sample for the word recovery task.

[0171] In one implementation, the processing unit 602 is configured to, based on the training sample, call a language model to perform a pre-training task. When obtaining a trained language model, it is specifically configured to perform the following steps:

[0172] Obtain a representation vector of the training sample for the word recovery task;

[0173] Use the language model to encode the representation vector of the training sample for the word recovery task to obtain an encoding of the replacement identifier in the training sample for the word recovery task;

[0174] Perform word prediction based on the encoding of the replacement identifier to obtain the probability that the replacement identifier is correctly predicted as the target word;

[0175] Determine the loss information for the word recovery task according to the probability that the replacement identifier is correctly predicted as the target word;

[0176] Update the model parameters of the language model according to the loss information for the word recovery task to obtain a trained language model.

[0177] In one implementation, the representation vector of the training sample for the word recovery task includes: the representation vector of the speaker identifier in the training sample for the word recovery task, the representation vector of the word, and the representation vector of the replacement identifier; any word in the training sample for the word recovery task is represented as a reference word; when the processing unit 602 is configured to obtain the representation vector of the reference word, it is specifically configured to perform the following steps:

[0178] Obtain the word vector of the reference word;

[0179] Determine the position vector of the reference word according to the arrangement position of the reference word in the speech content to which the reference word belongs;

[0180] Determine the speech content marker vector of the reference word according to the speech content to which the reference word belongs;

[0181] Determine the speaker representation vector of the reference word according to the speaker to which the reference word belongs;

[0182] Determine the representation vector of the reference word according to the word vector of the reference word, the position vector of the reference word, the speech content marker vector of the reference word, and the speaker representation vector of the reference word.

[0183] In one implementation, the pre-training task includes a speaker prediction task; when the processing unit 602 is used to perform transformation processing on the training data according to the task requirements of the pre-training task to obtain the training samples of the pre-training task, it is specifically used to perform the following steps:

[0184] According to the speaking order of each turn of speech in the social conversation, sequentially splice the speaker identifier and the speech content of each turn of speech in the social conversation to obtain the reference content of the social conversation;

[0185] Use the replacement identifier to replace the target speaker identifier in the reference content to obtain the training sample of the speaker prediction task.

[0186] In one implementation, when the processing unit 602 is used to call the language model to perform the pre-training task based on the training samples to obtain the trained language model, it is specifically used to perform the following steps:

[0187] Obtain the representation vector of the training sample of the speaker prediction task;

[0188] Use the language model to encode the representation vector of the training sample of the speaker prediction task to obtain the encoding of the replacement identifier in the training sample of the speaker prediction task;

[0189] Based on the encoding of the replacement identifier, perform speaker prediction to obtain the probability that the replacement identifier is correctly predicted as the target speaker identifier;

[0190] According to the probability that the replacement identifier is correctly predicted as the target speaker identifier, determine the loss information of the speaker prediction task;

[0191] According to the loss information of the speaker prediction task, update the model parameters of the language model to obtain the trained language model.

[0192] In one implementation, the pre-training task includes a speaking order determination task; when the processing unit 602 is used to perform transformation processing on the training data according to the task requirements of the pre-training task to obtain the training samples of the pre-training task, it is specifically used to perform the following steps:

[0193] Splice the classification symbol with the speaker identifier and the speech content of each turn of speech in the social conversation to obtain the spliced content of each turn of speech in the social conversation;

[0194] Perform multiple splicing processes in random order on the spliced content of each turn of speech in the social conversation to obtain the training samples of the speaking order determination task;

[0195] Among them, the arrangement order of the spliced content of each turn of speech in each random order splicing process is different, and each random order splicing process obtains a training sample of the speaking order determination task.

[0196] In one implementation, the processing unit 602, when used to call a language model to perform a pre-training task based on training samples and obtain a trained language model, is specifically used to perform the following steps:

[0197] Obtain the representation vectors of the training samples for the speech order determination task;

[0198] Use the language model to perform encoding processing on the representation vectors of the training samples for the speech order determination task to obtain the encodings of each classification symbol in the training samples for the speech order determination task;

[0199] Based on the encodings of each classification symbol, perform speech order prediction to obtain the prediction probability that the predicted order of each turn of speech in the training samples for the speech order determination task is consistent with the actual order;

[0200] According to the prediction probability of the training samples for the speech order determination task, determine the loss information for the speech order determination task;

[0201] According to the loss information for the speech order determination task, update the model parameters of the language model to obtain the trained language model.

[0202] In one implementation, the pre-training task includes a word recovery task, a speaker prediction task, and a speech order determination task; the processing unit 602, when used to call a language model to perform a pre-training task based on training samples and obtain a trained language model, is specifically used to perform the following steps:

[0203] Based on the training samples for the word recovery task, call the language model to perform the word recovery task to obtain the loss information for the word recovery task;

[0204] Based on the training samples for the speaker prediction task, call the language model to perform the speaker prediction task to obtain the loss information for the speaker prediction task;

[0205] Based on the training samples for the speech order determination task, call the language model to perform the speech order determination task to obtain the loss information for the speech order determination task;

[0206] According to the loss information for the word recovery task, the loss information for the speaker prediction task, and the loss information for the speech order determination task, update the model parameters of the language model to obtain the trained language model.

[0207] In one implementation, the acquisition unit 601 is further used to perform the following steps: acquire the decoding model for the language processing task;

[0208] The processing unit 602 is further configured to perform the following steps: encoding the conversation data in the social conversation by using the trained language model to obtain a conversation encoding;

[0209] Training the decoding model according to the conversation encoding according to the task requirements of the language processing task.

[0210] In one implementation, the language processing task is a conversation summary extraction task; the training data further includes the marked summary of the social conversation; when the processing unit 602 is used to train the decoding model according to the task requirements of the language processing task, it is specifically configured to perform the following steps:

[0211] Decoding the conversation encoding by using the decoding model of the conversation summary extraction task to obtain a predicted summary of the social conversation;

[0212] Training the decoding model based on the difference between the marked summary and the predicted summary.

[0213] In one implementation, the language processing task is a conversation prediction task; the social conversation includes the speech content generated by N rounds of speeches, where N is an integer greater than 1; the conversation encoding includes the encoding of the speech content of each round in the N rounds of speeches;

[0214] When the processing unit 602 is used to train the decoding model according to the task requirements of the language processing task, it is specifically configured to perform the following steps:

[0215] Decoding the encoding of the speech content of the first M rounds of speeches in the N rounds of speeches by using the decoding model of the conversation prediction task to obtain the predicted content of the last N - M rounds of speeches, where M is a positive integer less than N;

[0216] Training the decoding model according to the difference between the predicted content of the last N - M rounds of speeches and the speech content of the last N - M rounds of speeches.

[0217] In one implementation, the language processing task is a conversation retrieval task; the training data further includes the retrieval question for the social conversation and the marked answer to the retrieval question;

[0218] The processing unit 602 is further configured to perform the following steps: encoding the retrieval question by using the trained language model to obtain an encoding of the retrieval question; wherein, the conversation encoding includes the encoding of the speech content of each round in the social conversation;

[0219] When the processing unit 602 is used to train the decoding model according to the task requirements of the language processing task, it is specifically configured to perform the following steps:

[0220] The decoding model for the session retrieval task calculates the similarity between the encoding of the retrieval question and the encoding of the speech content of each turn in the social session, and decodes the encoding of the speech content corresponding to the maximum similarity among the calculated similarities to obtain the predicted answer to the retrieval question;

[0221] Based on the difference between the marked answer and the predicted answer, the decoding model is trained.

[0222] According to an embodiment of the present application, Figure 2 or Figure 4 Each method step involved in the method shown may be performed by Figure 6 each unit in the model training device shown. For example, Figure 2 the steps S201 to S202 shown may be performed by Figure 6 the obtaining unit 601 shown, Figure 2 the steps S203 to S204 shown may be performed by Figure 6 the processing unit 602 shown. Another example, Figure 4 the steps S401 to S402 and step S405 shown may be performed by Figure 6 the obtaining unit 601 shown, Figure 4 the steps S404 to S404 and steps S406 to S407 shown may be performed by Figure 6 the processing unit 602 shown.

[0223] According to another embodiment of the present application, Figure 6 each unit in the model training device shown may be separately or all combined into one or several other units to form, or some of the units may be further split into multiple smaller units in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In practical applications, the function of one unit may also be realized by multiple units, or the functions of multiple units may be realized by one unit. In other embodiments of the present application, the model training device may also include other units. In practical applications, these functions may also be assisted by other units and may be realized by the cooperation of multiple units.

[0224] According to another embodiment of the present application, it can be constructed, for example, by running a computer program (including program code) capable of executing the respective steps involved in the corresponding method shown in Figure 2 or Figure 4 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), Figure 6The model training device shown in the figure, and the model training method of the embodiments of the present application is implemented. The computer program can be recorded on, for example, a computer-readable storage medium, loaded into the above-mentioned computing device through the computer-readable storage medium, and run therein.

[0225] The embodiments of the present application propose a pre-training task for a language model for a social conversation scenario, and train the language model in the social conversation scenario by calling the language model to execute the pre-training task; the pre-training task can include at least one of the following: a word recovery task, a speaker prediction task, and a speaking order determination task; the word recovery task can be used to train the ability of the language model to learn word-level features in the social conversation scenario, the speaker prediction task can be used to train the ability of the language model to learn speaker features in the social conversation scenario, and the speaking order prediction task can be used to train the ability of the language model to learn the logical features of the speaking order in the social conversation scenario. The trained language model can be used to encode the conversation data in the social conversation, so that the trained language model can better learn the language features in the social conversation scenario.

[0226] Based on the above method and device embodiments, the embodiments of the present application provide a computer device, which can be the aforementioned intelligent terminal or server. Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of a computer device provided by the embodiments of the present application. Figure 7 The computer device shown at least includes a processor 701, an input interface 702, an output interface 703, and a computer-readable storage medium 704. Among them, the processor 701, the input interface 702, the output interface 703, and the computer-readable storage medium 704 can be connected through a bus or other means.

[0227] The input interface 702 can be used to obtain the training data of the language model, obtain the pre-training task of the language model, obtain the decoding model of the language processing task, etc.; the output interface 703 can be used for the encoding result of the language model and the decoding result of the decoding model, etc.

[0228] The computer-readable storage medium 704 can be stored in the memory of the computer device. The computer-readable storage medium 704 is used to store a computer program, and the computer program includes computer instructions. The processor 701 is used to execute the program instructions stored in the computer-readable storage medium 704. The processor 701 (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the computer device, and is suitable for implementing one or more computer instructions, and is specifically suitable for loading and executing one or more computer instructions to implement the corresponding method flow or corresponding function.

[0229] An embodiment of the present application further provides a computer-readable storage medium (Memory). A computer-readable storage medium is a memory device in a computer device, used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device, and of course can also include the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the computer device is stored in this storage space. And, one or more computer instructions suitable for being loaded and executed by the processor are also stored in this storage space. These computer instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory (Non-Volatile Memory), such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.

[0230] In one implementation, one or more computer instructions stored in the computer-readable storage medium 704 can be loaded and executed by the processor 701 to implement the corresponding steps of the model training method as described above Figure 2 or Figure 4 shown. Specifically, the computer instructions in the computer-readable storage medium 704 are loaded and executed by the processor 701 to perform the following steps:

[0231] Obtain the training data of the language model. The training data includes conversation data, and the conversation data includes the speech content generated by multiple rounds of speeches in a social conversation. Each round of speech is initiated by a speaker participating in the social conversation;

[0232] Obtain the pre-training tasks of the language model. The pre-training tasks include at least one of the following: word recovery task, speaker prediction task, and speech order determination task;

[0233] Perform transformation processing on the training data according to the task requirements of the pre-training task to obtain the training samples of the pre-training task;

[0234] Based on the training samples, call the language model to perform the pre-training task to obtain a trained language model; the trained language model is used to encode the conversation data in a social conversation.

[0235] In one implementation, the pre-training task includes a word recovery task, and the speech content is composed of words; when the computer instructions in the computer-readable storage medium 704 are loaded and executed by the processor 701 to perform transformation processing on the training data according to the task requirements of the pre-training task to obtain the training samples of the pre-training task, it is specifically used to perform the following steps:

[0236] According to the speaking order of each turn of speech in the social conversation, splice the speaker identification and the speech content of each turn of speech in the social conversation in sequence to obtain the reference content of the social conversation;

[0237] Replace the target word in the reference content with a replacement identifier to obtain the training sample of the word recovery task.

[0238] In one implementation, when the computer instructions in the computer-readable storage medium 704 are loaded and executed by the processor 701, and based on the training sample, call the language model to perform the pre-training task to obtain the trained language model, it is specifically used to perform the following steps:

[0239] Obtain the representation vector of the training sample of the word recovery task;

[0240] Use the language model to encode the representation vector of the training sample of the word recovery task to obtain the encoding of the replacement identifier in the training sample of the word recovery task;

[0241] Perform word prediction based on the encoding of the replacement identifier to obtain the probability that the replacement identifier is correctly predicted as the target word;

[0242] Determine the loss information of the word recovery task according to the probability that the replacement identifier is correctly predicted as the target word;

[0243] Update the model parameters of the language model according to the loss information of the word recovery task to obtain the trained language model.

[0244] In one implementation, the representation vector of the training sample of the word recovery task includes: the representation vector of the speaker identification in the training sample of the word recovery task, the representation vector of the word, and the representation vector of the replacement identifier; any word in the training sample of the word recovery task is represented as a reference word; when the computer instructions in the computer-readable storage medium 704 are loaded and executed by the processor 701 to obtain the representation vector of the reference word, it is specifically used to perform the following steps:

[0245] Obtain the word vector of the reference word;

[0246] Determine the position vector of the reference word according to the arrangement position of the reference word in the speech content to which the reference word belongs;

[0247] Determine the speech content marker vector of the reference word according to the speech content to which the reference word belongs;

[0248] Determine the speaker representation vector of the reference word according to the speaker to which the reference word belongs;

[0249] Determine the representation vector of the reference word based on the word vector of the reference word, the position vector of the reference word, the speech content marker vector of the reference word, and the speaker representation vector of the reference word.

[0250] In one implementation, the pre-training task includes a speaker prediction task; when the computer instructions in the computer-readable storage medium 704 are loaded and executed by the processor 701 to perform transformation processing on the training data according to the task requirements of the pre-training task to obtain the training samples of the pre-training task, it is specifically used to perform the following steps:

[0251] Sequentially splice the speaker identifiers and speech contents of each turn of speech in the social conversation according to the speech order of each turn of speech in the social conversation to obtain the reference content of the social conversation;

[0252] Use the replacement identifier to replace the target speaker identifier in the reference content to obtain the training samples of the speaker prediction task.

[0253] In one implementation, when the computer instructions in the computer-readable storage medium 704 are loaded and executed by the processor 701 to call the language model to perform the pre-training task based on the training samples and obtain the trained language model, it is specifically used to perform the following steps:

[0254] Obtain the representation vector of the training samples of the speaker prediction task;

[0255] Use the language model to encode the representation vector of the training samples of the speaker prediction task to obtain the encoding of the replacement identifier in the training samples of the speaker prediction task;

[0256] Perform speaker prediction based on the encoding of the replacement identifier to obtain the probability that the replacement identifier is correctly predicted as the target speaker identifier;

[0257] Determine the loss information of the speaker prediction task according to the probability that the replacement identifier is correctly predicted as the target speaker identifier;

[0258] Update the model parameters of the language model according to the loss information of the speaker prediction task to obtain the trained language model.

[0259] In one implementation, the pre-training task includes a speech order determination task; when the computer instructions in the computer-readable storage medium 704 are loaded and executed by the processor 701 to perform transformation processing on the training data according to the task requirements of the pre-training task to obtain the training samples of the pre-training task, it is specifically used to perform the following steps:

[0260] Splice the classification symbol with the speaker identifier and speech content of each turn of speech in the social conversation to obtain the spliced content of each turn of speech in the social conversation;

[0261] The concatenated content of each turn of speech in a social conversation is concatenated multiple times in a random order to obtain training samples for the speech order determination task;

[0262] Among them, the arrangement order of the concatenated content of each turn of speech in each random order concatenation process is different, and each random order concatenation process obtains a training sample for the speech order determination task.

[0263] In one implementation, when the computer instructions in the computer-readable storage medium 704 are loaded and executed by the processor 701 to call the language model to perform the pre-training task and obtain the trained language model based on the training samples, it is specifically used to perform the following steps:

[0264] Obtain the representation vector of the training samples for the speech order determination task;

[0265] Use the language model to encode the representation vector of the training samples for the speech order determination task to obtain the encodings of each classification symbol in the training samples for the speech order determination task;

[0266] Predict the speech order based on the encodings of each classification symbol to obtain the prediction probability that the predicted order of each turn of speech in the training samples for the speech order determination task is consistent with the actual order;

[0267] Determine the loss information for the speech order determination task according to the prediction probability of the training samples for the speech order determination task;

[0268] Update the model parameters of the language model according to the loss information for the speech order determination task to obtain the trained language model.

[0269] In one implementation, the pre-training task includes a word recovery task, a speaker prediction task, and a speech order determination task; when the computer instructions in the computer-readable storage medium 704 are loaded and executed by the processor 701 to call the language model to perform the pre-training task and obtain the trained language model based on the training samples, it is specifically used to perform the following steps:

[0270] Based on the training samples for the word recovery task, call the language model to perform the word recovery task to obtain the loss information for the word recovery task;

[0271] Based on the training samples for the speaker prediction task, call the language model to perform the speaker prediction task to obtain the loss information for the speaker prediction task;

[0272] Based on the training samples for the speech order determination task, call the language model to perform the speech order determination task to obtain the loss information for the speech order determination task;

[0273] Update the model parameters of the language model according to the loss information of the word recovery task, the loss information of the speaker prediction task, and the loss information of the speaking order determination task, to obtain a trained language model.

[0274] In one implementation, the computer instructions in the computer-readable storage medium 704 are loaded and further executed by the processor 701 to perform the following steps:

[0275] Obtain the decoding model of the language processing task;

[0276] Use the trained language model to encode the conversation data in the social conversation to obtain a conversation encoding;

[0277] According to the task requirements of the language processing task, train the decoding model based on the conversation encoding.

[0278] In one implementation, the language processing task is a conversation summary extraction task; the training data further includes the marked summary of the social conversation; when the computer instructions in the computer-readable storage medium 704 are loaded and executed by the processor 701 to train the decoding model according to the task requirements of the language processing task, it is specifically used to perform the following steps:

[0279] Use the decoding model of the conversation summary extraction task to decode the conversation encoding to obtain the predicted summary of the social conversation;

[0280] Train the decoding model based on the difference between the marked summary and the predicted summary.

[0281] In one implementation, the language processing task is a conversation prediction task; the social conversation includes the speech content generated by N rounds of speeches, where N is an integer greater than 1; the conversation encoding includes the encoding of the speech content of each round in the N rounds of speeches;

[0282] When the computer instructions in the computer-readable storage medium 704 are loaded and executed by the processor 701 to train the decoding model according to the task requirements of the language processing task, it is specifically used to perform the following steps:

[0283] Use the decoding model of the conversation prediction task to decode the encoding of the speech content of the first M rounds of speeches in the N rounds of speeches to obtain the predicted content of the last N - M rounds of speeches, where M is a positive integer less than N;

[0284] Train the decoding model based on the difference between the predicted content of the last N - M rounds of speeches and the speech content of the last N - M rounds of speeches.

[0285] In one implementation, the language processing task is a conversation retrieval task; the training data further includes the retrieval questions for the social conversation and the marked answers to the retrieval questions;

[0286] The computer instructions in the computer-readable storage medium 704 are loaded by the processor 701 and further execute the following steps: encoding the retrieval problem by using the trained language model to obtain the encoding of the retrieval problem; wherein, the session encoding includes the encoding of the speech content of each turn of speech in the social session.

[0287] When the computer instructions in the computer-readable storage medium 704 are loaded by the processor 701 and executed to train the decoding model according to the task requirements of the language processing task based on the session encoding, it is specifically used to execute the following steps:

[0288] Using the decoding model of the session retrieval task to calculate the similarity between the encoding of the retrieval problem and the encoding of the speech content of each turn of speech in the social session, and performing decoding processing on the encoding of the speech content corresponding to the maximum similarity among the calculated similarities to obtain the predicted answer to the retrieval problem.

[0289] Training the decoding model based on the difference between the marked answer and the predicted answer.

[0290] The embodiment of the present application proposes a pre-training task of a language model for the social session scenario, and trains the language model in the social session scenario by calling the language model to execute the pre-training task; the pre-training task may include at least one of the following: word recovery task, speaker prediction task, and speech order determination task; the word recovery task can be used to train the ability of the language model to learn word-level features in the social session scenario, the speaker prediction task can be used to train the ability of the language model to learn speaker features in the social session scenario, and the speech order prediction task can be used to train the ability of the language model to learn the speech order logic features in the social session scenario. The trained language model can be used to encode the session data in the social session, so that the trained language model can better learn the language features in the social session scenario.

[0291] According to one aspect of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the model training method provided in the above various optional manners.

[0292] As described above, it is only the specific implementation manner of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claimed rights.

Claims

1. A model training method, characterized in that, The method includes: Obtaining training data for a language model, where the training data includes conversation data, and the conversation data includes utterance content generated from multiple turns of speech in a social conversation, and each turn of speech is initiated by a speaker participating in the social conversation; Obtaining a pre-training task for the language model, where the pre-training task includes a word recovery task, a speaker prediction task, and a speech order determination task; Performing transformation processing on the training data according to the task requirements of the pre-training task to obtain training samples for the pre-training task; Based on the training samples for the word recovery task, invoking the language model to perform the word recovery task to obtain loss information for the word recovery task; Based on the training samples for the speaker prediction task, invoking the language model to perform the speaker prediction task to obtain loss information for the speaker prediction task; Based on the training samples for the speech order determination task, invoking the language model to perform the speech order determination task to obtain loss information for the speech order determination task; Updating the model parameters of the language model according to the loss information for the word recovery task, the loss information for the speaker prediction task, and the loss information for the speech order determination task to obtain a trained language model; the trained language model is used to encode the conversation data in the social conversation.

2. The method according to claim 1, wherein The pre-training task includes the word recovery task, and the utterance content is composed of words; the performing transformation processing on the training data according to the task requirements of the pre-training task to obtain training samples for the pre-training task includes: Sequentially splicing the speaker identifiers and utterance content of each turn of speech in the social conversation according to the speech order of each turn of speech in the social conversation to obtain reference content for the social conversation; Replacing the target word in the reference content with a replacement identifier to obtain training samples for the word recovery task.

3. The method according to claim 2, wherein The invoking the language model to perform the pre-training task based on the training samples to obtain a trained language model includes: Obtaining a representation vector of the training samples for the word recovery task; Encoding the representation vector of the training samples for the word recovery task using the language model to obtain an encoding of the replacement identifier in the training samples for the word recovery task; Performing word prediction based on the encoding of the replacement identifier to obtain the probability that the replacement identifier is correctly predicted as the target word; Determining the loss information for the word recovery task according to the probability that the replacement identifier is correctly predicted as the target word; Updating the model parameters of the language model according to the loss information for the word recovery task to obtain the trained language model.

4. The method according to claim 3, characterized in that, The representation vector of the training samples for the word recovery task includes: the representation vector of the speaker identifier in the training samples for the word recovery task, the representation vector of the word, and the representation vector of the replacement identifier; any word in the training samples for the word recovery task is represented as a reference word; obtaining the representation vector of the reference word includes: Obtaining the word vector of the reference word; Determine the position vector of the reference word according to the arrangement position of the reference word in the speech content to which the reference word belongs; Determine the speech content marker vector of the reference word according to the speech content to which the reference word belongs; Determine the speaker representation vector of the reference word according to the speaker to which the reference word belongs; Determine the representation vector of the reference word according to the word vector of the reference word, the position vector of the reference word, the speech content marker vector of the reference word, and the speaker representation vector of the reference word.

5. The method according to claim 1, wherein The pre-training task includes the speaker prediction task; the transformation process of the training data according to the task requirements of the pre-training task to obtain the training samples of the pre-training task includes: Sequentially splice the speaker identifiers and speech contents of each turn of speech in the social conversation according to the speech order of each turn of speech in the social conversation to obtain the reference content of the social conversation; Replace the target speaker identifier in the reference content with a replacement identifier to obtain the training sample of the speaker prediction task.

6. The method according to claim 5, wherein Based on the training samples, calling the language model to execute the pre-training task to obtain a trained language model includes: Obtain the representation vector of the training sample of the speaker prediction task; Encode the representation vector of the training sample of the speaker prediction task using the language model to obtain the encoding of the replacement identifier in the training sample of the speaker prediction task; Perform speaker prediction based on the encoding of the replacement identifier to obtain the probability that the replacement identifier is correctly predicted as the target speaker identifier; Determine the loss information of the speaker prediction task according to the probability that the replacement identifier is correctly predicted as the target speaker identifier; Update the model parameters of the language model according to the loss information of the speaker prediction task to obtain the trained language model.

7. The method according to claim 1, wherein The pre-training task includes the speech order determination task; the transformation process of the training data according to the task requirements of the pre-training task to obtain the training samples of the pre-training task includes: Splice the classification symbol with the speaker identifier and speech content of each turn of speech in the social conversation to obtain the spliced content of each turn of speech in the social conversation; Perform splicing processing on the spliced content of each turn of speech in the social conversation in multiple random orders to obtain the training sample of the speech order determination task; Among them, the arrangement order of the spliced content of each turn of speech in each random order splicing process is different, and each random order splicing process obtains a training sample of the speech order determination task.

8. The method according to claim 7, characterized in that Based on the training samples, calling the language model to execute the pre-training task to obtain a trained language model includes: Obtain the representation vector of the training sample of the speech order determination task; Encode the representation vector of the training sample of the speech order determination task using the language model to obtain the encoding of each classification symbol in the training sample of the speech order determination task; Predict the speaking order based on the encoding of each classification symbol, and obtain the prediction probability that the predicted order of each round of speaking in the training sample of the speaking order determination task is consistent with the actual order; Determine the loss information of the speaking order determination task according to the prediction probability of the training sample of the speaking order determination task; Update the model parameters of the language model according to the loss information of the speaking order determination task to obtain the trained language model.

9. The method according to claim 1, characterized in that The method further includes: Obtain the decoding model of the language processing task; Encode the conversation data in the social conversation using the trained language model to obtain a conversation encoding; Train the decoding model according to the conversation encoding according to the task requirements of the language processing task.

10. The method according to claim 9, characterized in that, The language processing task is a conversation summary extraction task; the training data further includes the marked summary of the social conversation; the training the decoding model according to the conversation encoding according to the task requirements of the language processing task includes: Decode the conversation encoding using the decoding model of the conversation summary extraction task to obtain the predicted summary of the social conversation; Train the decoding model based on the difference between the marked summary and the predicted summary.

11. The method according to claim 9, wherein The language processing task is a conversation prediction task; the social conversation includes the speech content generated by N rounds of speaking, where N is an integer greater than 1; the conversation encoding includes the encoding of the speech content of each round of speaking in the N rounds of speaking; The training the decoding model according to the conversation encoding according to the task requirements of the language processing task includes: Decode the encoding of the speech content of the first M rounds of speaking in the N rounds of speaking using the decoding model of the conversation prediction task to obtain the predicted content of the last N - M rounds of speaking, where M is a positive integer less than N; Train the decoding model based on the difference between the predicted content of the last N - M rounds of speaking and the speech content of the last N - M rounds of speaking.

12. The method according to claim 9, wherein, The language processing task is a conversation retrieval task; the training data further includes the retrieval question for the social conversation and the marked answer to the retrieval question; The method further includes: encoding the retrieval question using the trained language model to obtain the encoding of the retrieval question; where the conversation encoding includes the encoding of the speech content of each round of speaking in the social conversation; The training the decoding model according to the conversation encoding according to the task requirements of the language processing task includes: Calculate the similarity between the encoding of the retrieval question and the encoding of the speech content of each round of speaking in the social conversation using the decoding model of the conversation retrieval task, and decode the encoding of the speech content corresponding to the maximum similarity in the calculated similarities to obtain the predicted answer to the retrieval question; Train the decoding model based on the difference between the marked answer and the predicted answer.

13. A model training device, characterized in that, The device includes: An acquisition unit, configured to acquire training data of a language model, where the training data includes conversation data, the conversation data includes speech content generated from multiple turns of speech in a social conversation, and each turn of speech is initiated by a speaker participating in the social conversation; and acquire a pre-training task of the language model, where the pre-training task includes a word recovery task, a speaker prediction task, and a speech order determination task; A processing unit, configured to perform transformation processing on the training data according to the task requirements of the pre-training task to obtain training samples of the pre-training task; and based on the training samples of the word recovery task, call the language model to perform the word recovery task to obtain loss information of the word recovery task; based on the training samples of the speaker prediction task, call the language model to perform the speaker prediction task to obtain loss information of the speaker prediction task; based on the training samples of the speech order determination task, call the language model to perform the speech order determination task to obtain loss information of the speech order determination task; update model parameters of the language model according to the loss information of the word recovery task, the loss information of the speaker prediction task, and the loss information of the speech order determination task to obtain a trained language model; the trained language model is used to encode conversation data in the social conversation.

14. A computer device, characterized in that, The computer device includes: A processor, adapted to implement a computer program; A computer-readable storage medium storing a computer program, the computer program being adapted to be loaded and executed by the processor to perform the model training method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program being adapted to be loaded and executed by a processor to perform the model training method according to any one of claims 1 to 12.

16. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a processor, implement the model training method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Method for improving dialogue text generation based on text abstract generation and bidirectional corpus

    CN113158665A

  • Cross-language conversation understanding-oriented model pre-training system

    CN113312453A