Explanation model training method, explanation text generation method, device and equipment
By training the commentary model, using pre-trained dialogue model and sample commentary text to generate game commentary text, solving the problems of high cost and low efficiency of real-person commentary in the existing technology, and achieving efficient and diverse game commentary text generation.
Patent Information
- Application Number
- CN202311630441.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, game commentary mainly relies on real-person commentary, resulting in high cost and low efficiency, making it difficult to meet the diversity and scale of game commentary needs.
By providing a training method for interpreting models, using pre-trained dialogue models and sample interpretation text, adjusting model parameters to generate accurate interpretation text, reducing training data needs, and improving the training efficiency and generation efficiency of interpretation models.
The explanation model can automatically generate accurate explanation text, which improves the generation efficiency and diversity of game explanation texts and reduces training costs.
Smart Images

Figure CN120067665A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a training method for an explanation model, a method for generating an explanation text, a device, and a device. Background Art
[0002] In the digital Internet era, games have become a very important part of people's lives. For games, game explanations can not only enable players to learn from each other, but also make the game atmosphere more active. Therefore, game explanations are very important.
[0003] In related technologies, game explanations mainly rely on live human explanations, that is, on-site explanations of games by commentators. On-site explanations have relatively high requirements for commentators, not only requiring commentators to have a deep understanding of the game, but also rich content explanations. Therefore, the cost of artificial explanations is very expensive. Summary of the Invention
[0004] Embodiments of this application provide a training method for an explanation model, a method for generating an explanation text, a device, and a device. This method enables the explanation model to automatically generate accurate explanation texts based on indication information, and can reduce the data required for training, thereby improving the training efficiency of the explanation model. Obtaining the explanation text through the explanation model not only improves the generation efficiency of the explanation text, but also, since the explanation model is a generative model, on the basis of improving the generation efficiency, it can also ensure the diversity of the generated explanation texts. The technical solutions are as follows:
[0005] On the one hand, a training method for an explanation model is provided, and the method includes:
[0006] Obtain a training data set, where the training data set includes multiple groups of first sample pairs, and each group of the first sample pairs includes first indication information and a sample explanation text. The first indication information includes a sample explanation task, scenario data corresponding to the sample explanation task, and an explanation requirement. The sample explanation task is a task of explaining a target event in a target scenario, and the sample explanation text is used to describe the scenario data in the first indication information and meet the explanation requirement in the first indication information;
[0007] For each group of the first sample pairs, input the first indication information in the first sample pair into a pre-trained dialogue model to obtain a first predicted explanation text of the first sample pair;
[0008] Based on the first predicted explanation text of the first sample pair and the sample explanation text in the first sample pair, adjust the model parameters of the pre-trained dialogue model to obtain an explanation model for the target scenario, and the explanation model is used to generate an explanation text based on the first indication information.
[0009] On the other hand, a method for generating commentary text is provided, and the method includes:
[0010] During the operation of the target scenario, in response to reaching the commentary timing of the target commentary task, obtaining target scenario data corresponding to the target commentary task;
[0011] Based on the target commentary task, the target scenario data, and the target commentary requirements, obtaining target indication information, where the target indication information is used to indicate that the target commentary task is explained based on the target scenario data and the target commentary requirements;
[0012] Inputting the target indication information into a commentary model, where the commentary model is trained by the method described in any one of the above;
[0013] Outputting, through the commentary model, target commentary text corresponding to the target commentary task.
[0014] On the other hand, a training device for a commentary model is provided, and the device includes:
[0015] An acquisition module, configured to acquire a training data set, where the training data set includes multiple groups of first sample pairs, and each group of the first sample pairs includes first indication information and sample commentary text. The first indication information includes a sample commentary task, scenario data corresponding to the sample commentary task, and commentary requirements. The sample commentary task is a task of explaining a target event in a target scenario, and the sample commentary text is used to describe the scenario data in the first indication information and meet the commentary requirements in the first indication information;
[0016] An input-output module, configured to, for each group of the first sample pairs, input the first indication information in the first sample pair into a pre-trained dialogue model to obtain a first predicted commentary text of the first sample pair;
[0017] An adjustment module, configured to adjust model parameters of the pre-trained dialogue model based on the first predicted commentary text of the first sample pair and the sample commentary text in the first sample pair to obtain a commentary model for the target scenario, where the commentary model is used to generate commentary text based on the first indication information.
[0018] In some embodiments, the acquisition module is configured to:
[0019] Acquire multiple sample commentary texts;
[0020] Based on the multiple sample commentary texts, multiple sample commentary tasks, the scenario data and commentary requirements corresponding to each of the multiple sample commentary tasks, construct multiple second indication messages. Each second indication message includes a sample commentary task, the corresponding scenario data, commentary requirements, and a first quantity of sample commentary texts. The second indication message is used to indicate the generation of commentary texts for the included sample commentary tasks based on the text features of the first quantity of sample commentary texts.
[0021] Input the multiple second indication messages into a large language model respectively. Through the large language model, obtain the sample commentary texts corresponding to the multiple second indication messages respectively.
[0022] Determine the multiple first sample pairs from multiple groups of second sample pairs. Each group of second sample pairs includes a sample commentary task, scenario data, commentary requirements, and the corresponding sample commentary text in a second indication message.
[0023] In some embodiments, the obtaining module is configured to:
[0024] Obtain the quality values of multiple groups of third sample pairs. The multiple groups of third sample pairs are partial sample pairs among the multiple groups of second sample pairs. The quality value is used to indicate the quality level of the sample commentary text in the third sample pair.
[0025] Determine multiple groups of fourth sample pairs from the multiple groups of third sample pairs. The quality values of the multiple groups of fourth sample pairs are lower than a quality threshold.
[0026] Based on the multiple groups of fourth sample pairs, through the large language model, obtain the quality values of multiple groups of remaining sample pairs among the multiple groups of second sample pairs.
[0027] Use the second sample pairs among the multiple groups of second sample pairs whose quality values are not lower than the quality threshold as the first sample pairs.
[0028] In some embodiments, the obtaining module is configured to:
[0029] Based on the multiple groups of fourth sample pairs, construct third indication messages corresponding to the multiple groups of remaining sample pairs respectively. Each third indication message includes a remaining sample pair, a second quantity of fourth sample pairs, and the corresponding quality value. The third indication message is used to indicate the determination of the quality value of the included remaining sample pair based on the text features and quality values of the commentary texts in the second quantity of fourth sample pairs.
[0030] Input multiple third indication messages into the large language model respectively. Through the large language model, obtain the quality values of the multiple groups of remaining sample pairs respectively.
[0031] In some embodiments, the multiple sample commentary texts correspond to multiple event types, and the obtaining module is configured to:
[0032] For each sample commentary task, based on the event type of the target event in the sample commentary task, determine a first number of sample commentary texts from the multiple sample commentary texts that are the same as the event type corresponding to the sample commentary task;
[0033] Based on the multiple sample commentary tasks, the scenario data, commentary requirements, and the respective first number of sample commentary texts of the multiple sample commentary tasks, construct the multiple second indication information.
[0034] In some embodiments, the adjustment module is configured to:
[0035] Based on the first predicted commentary texts and the sample commentary texts of the multiple groups of first samples, adjust the model parameters of the pre-trained dialogue model to obtain an initial commentary model;
[0036] For each group of the multiple groups of first samples, input the first indication information in the first sample pair into the initial commentary model to obtain a second predicted commentary text of the first indication information;
[0037] Input the first indication information and the second predicted commentary text of the first indication information into a reward model to obtain a reward value, where the reward value is used to indicate the degree of preference of an object for the second predicted commentary text, and the reward model is used to predict the degree of preference of the object for the predicted commentary text;
[0038] Based on the reward value, adjust the model parameters of the initial commentary model to obtain the commentary model.
[0039] In some embodiments, the apparatus further includes a training module, configured to:
[0040] For each group of the multiple groups of first samples, input the first indication information in the first sample pair into the initial commentary model, and through the initial commentary model, obtain multiple predicted commentary texts of the first indication information;
[0041] Obtain the ranking among the multiple predicted commentary texts, where the ranking is used to indicate the degree of preference of an object for the multiple predicted commentary texts;
[0042] Determine multiple groups of commentary sample pairs from the multiple predicted commentary texts, where each group of commentary sample pairs includes two predicted commentary texts among the multiple predicted commentary texts;
[0043] Input each group of explanatory sample pairs and the corresponding first indication information into the initial reward model to obtain the reward values corresponding to the two predicted explanatory texts included;
[0044] Based on the reward values and rankings corresponding to the two predicted explanatory texts respectively, determine a loss, where the loss is used to indicate the difference between the ranking of the two predicted explanatory texts and the ranking corresponding to the reward values of the two predicted explanatory texts;
[0045] Based on the losses corresponding to the multiple groups of explanatory sample pairs respectively, adjust the initial reward model to obtain the reward model.
[0046] In some embodiments, the adjustment module is configured to:
[0047] Based on the reward value, determine the policy update gradient of the initial explanatory model;
[0048] Based on the policy update gradient, update the model parameters of the initial explanatory model to obtain the explanatory model.
[0049] On the other hand, a device for generating an explanatory text is provided, and the device includes:
[0050] An acquisition module, configured to, during the running of the target scenario, in response to reaching the explanatory timing of the target explanatory task, acquire the target scenario data corresponding to the target explanatory task;
[0051] A determination module, configured to obtain target indication information based on the target explanatory task, the target scenario data, and the target explanatory requirements, where the target indication information is used to indicate the explanation of the target explanatory task based on the target scenario data and the target explanatory requirements;
[0052] An input / output module, configured to input the target indication information into an explanatory model, where the explanatory model is trained by the training method described in any one of the above;
[0053] The input / output module is further configured to output, through the explanatory model, the target explanatory text corresponding to the target explanatory task.
[0054] In some embodiments, the acquisition module is further configured to:
[0055] Display a plurality of candidate explanatory requirements, where the explanatory language styles corresponding to the plurality of candidate explanatory requirements are different;
[0056] In response to a selection operation on any one of the candidate explanatory requirements, use the selected candidate explanatory requirement as the target explanatory requirement.
[0057] In some embodiments, there are multiple virtual teams in the target scenario, and each virtual team includes multiple virtual objects; the acquisition module is configured to:
[0058] In response to reaching the commentary timing of the target commentary task, acquire the scene data corresponding to the multiple virtual teams for the target commentary task respectively;
[0059] Based on the scene data corresponding to the multiple virtual teams respectively, use the scene data corresponding to the target virtual team that meets the preset conditions among the multiple virtual teams as the target scene data, and the target scene data is used to comment on the behavior of the target virtual team in the target scene.
[0060] In some embodiments, the commentary timing includes the timing when no virtual team battles are taking place, and the target commentary task includes the task of commenting on at least one of the position information of the virtual objects in the target virtual team, the terrain information of the location, the direction information, the virtual resource holding information, and the virtual battle information.
[0061] In some embodiments, there are multiple virtual teams in the target scenario, and each virtual team includes multiple virtual objects;
[0062] The commentary timing includes the timing when the virtual team battle ends, the timing when the virtual team battle is in progress, and the timing of entering the final circle; the types of virtual team battles include the team wipe type and the non-team wipe type; the target commentary task includes the task of commenting on at least one of the highlight moments, the reasons for victory or defeat, and the virtual tactics in the virtual team battle.
[0063] On the other hand, a computer device is provided, which includes a processor and a memory. The memory is used to store at least one program, and the at least one program is loaded and executed by the processor to implement the training method of the commentary model or the generation method of the commentary text in the embodiments of the present application.
[0064] On the other hand, a computer-readable storage medium is provided, in which at least one program is stored, and the at least one program is loaded and executed by a processor to implement the training method of the commentary model or the generation method of the commentary text in the embodiments of the present application.
[0065] On the other hand, a computer program product is provided, which includes at least one program. The at least one program is stored in a computer-readable storage medium. The processor of the computer device reads the at least one program from the computer-readable storage medium, and the processor executes the at least one program, so that the computer device executes the training method of the commentary model or the generation method of the commentary text in any of the above implementation manners.
[0066] An embodiment of the present application provides a method for training an explanation model. This method trains a pre-trained dialogue model based on indication information including an explanation task, scenario data, and explanation requirements to obtain an explanation model for the target scenario. Since the pre-trained dialogue model itself has good dialogue capabilities, using the basic dialogue capabilities of the pre-trained dialogue model and multiple groups of first samples to obtain the explanation model enables the explanation model to automatically generate accurate explanation texts based on the indication information, and can reduce the data required for training, thereby improving the training efficiency of the explanation model. And obtaining the explanation text through the explanation model not only improves the generation efficiency of the explanation text, but also, since the explanation model is a generative model, on the basis of improving the generation efficiency, it can also ensure the diversity of the generated explanation texts. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0068] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0069] Figure 2 is a flowchart of a method for training an explanation model provided by an embodiment of the present application;
[0070] Figure 3 is a flowchart of a method for generating an explanation text provided by an embodiment of the present application;
[0071] Figure 4 is a flowchart of another method for training an explanation model provided by an embodiment of the present application;
[0072] Figure 5 is a schematic structural diagram of a large language model provided by an embodiment of the present application;
[0073] Figure 6 is a schematic example diagram of a first indication information provided by an embodiment of the present application;
[0074] Figure 7 is a flowchart of a training process of an explanation model provided by an embodiment of the present application;
[0075] Figure 8 is a flowchart of a method for generating an explanation text provided by an embodiment of the present application;
[0076] Figure 9It is a flowchart of another method for generating explanatory text provided by an embodiment of the present application;
[0077] Figure 10 It is a schematic diagram for generating explanatory text provided by an embodiment of the present application;
[0078] Figure 11 It is a schematic diagram of a target indication information provided by an embodiment of the present application;
[0079] Figure 12 It is a block diagram of a training device for an explanatory model provided by an embodiment of the present application;
[0080] Figure 13 It is a block diagram of a device for generating explanatory text provided by an embodiment of the present application;
[0081] Figure 14 It is a block diagram of a terminal provided by an embodiment of the present application;
[0082] Figure 15 It is a block diagram of a server provided by an embodiment of the present application. Detailed implementation manners
[0083] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0084] In the present application, terms such as "first" and "second" are used to distinguish between identical or similar items with basically the same functions and effects. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor are the quantity and execution order limited.
[0085] In the present application, the term "at least one" means one or more, and the meaning of "multiple" means two or more.
[0086] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions. For example, the training data sets involved in the present application are all obtained under full authorization.
[0087] The following introduces the professional terms involved in the present application:
[0088] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0089] Machine Learning (ML) is an interdisciplinary subject that involves multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration. The pre-trained model is the latest development result of deep learning, integrating the above technologies.
[0090] Natural Language Processing (NLP) is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistics research; at the same time, it involves computer science and mathematics. The pre-trained model, an important technology for model training in the field of artificial intelligence, is developed from the large language model in the NLP field. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, etc.
[0091] The pre-training model (PTM), also known as the foundation model or large model, refers to a deep neural network (DNN) with a large number of parameters. It is trained on a vast amount of unlabeled data, and by leveraging the function approximation ability of the large-parameter DNN, the PTM extracts common features from the data. Through techniques such as fine-tuning, parameter-efficient fine-tuning (PEFT), and prompt-tuning, it is applicable to downstream tasks. Therefore, the pre-training model can achieve ideal results in few-shot or zero-shot scenarios. PTMs can be classified into language models (ELMO, BERT, GPT), vision models (swin-transformer, ViT, V-MOE), speech models (VALL-E), multi-modal models (ViBERT, CLIP, Flamingo, Gato), etc. according to the data modalities they process. Among them, multi-modal models refer to models that establish feature representations of two or more data modalities. The pre-training model is an important tool for outputting artificial intelligence-generated content (AIGC) and can also serve as a general interface connecting multiple specific task models.
[0092] The generative pre-training series models adopt the Transformer architecture, which consists of two parts: an encoder and a decoder. The encoder is composed of multiple layers of Transformer encoders and is used to encode the input text. The decoder is composed of multiple layers of Transformer decoders and is used to decode and generate the encoded text. The encoder and decoder structures of some generative pre-training models are relatively simple, both consisting of 12 Transformer encoders and decoders, and each encoder and decoder contains a multi-head self-attention mechanism and a feed-forward neural network layer. The structures of some other generative pre-training models are more complex, with more layers and more parameters. It contains 96 Transformer encoders and decoders, and each encoder and decoder contains a multi-head self-attention mechanism, a multi-head cross-attention mechanism, and a feed-forward neural network layer. Generally speaking, the structures of the generative pre-training series models are relatively similar. They all adopt the Transformer architecture and contain a multi-head self-attention mechanism and a feed-forward neural network layer in each layer. However, as the model scale increases, the generative pre-training model has more optimizations in its structure, making it more capable of generating natural language text. For example, the generative pre-training model has a structure with 12 layers of Decoder.
[0093] First-Person Shooting (FPS) game: A shooting game with the first-person perspective of the game object (i.e., the player) as the main perspective. Some FPS games may include a free perspective. Usually, several strongholds are provided in the virtual world, and game objects belonging to different teams (i.e., factions) control virtual objects to fight in the virtual environment. The victory conditions are to capture a stronghold, destroy the strongholds of the opposing faction, or kill all or part of the characters of the opposing faction.
[0094] Game commentary: Provide real-time commentary on the game match. Generally, it is necessary to select key game scenes in the game match and comment on the selected game scenes. Among them, commenting on the selected game scenes includes at least one of generating commentary text for the game scene and generating commentary audio for the game scene.
[0095] Next, the implementation environment related to this application will be introduced:
[0096] The training method of the commentary model provided by the embodiments of this application can be executed by a computer device, which can be provided as a server or a terminal. The following introduces the schematic diagram of the implementation environment of the training method of the commentary model provided by the embodiments of this application.
[0097] See Figure 1 , Figure 1 FIG. is a schematic diagram of the implementation environment of a training method of a commentary model provided by the embodiments of this application. The implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here. In some embodiments, the server 102 is used to train the commentary model, and the trained commentary model is used to comment on the events in the scene to obtain commentary text. The target application is installed on the terminal 101, and the target application is used to run the target scene. In some embodiments, the commentary model is embedded on the terminal 101, and the terminal 101 is used to construct indication information based on the scene data and commentary requirements in the target scene; based on the indication information, through the commentary model, generate commentary text for the target scene. In other embodiments, after the terminal 101 constructs the indication information, it sends it to the server 102, and the commentary text is generated through the commentary model on the server 102. Or, the terminal 101 sends the scene data and commentary requirements to the server 102, and the server 102 constructs the indication information; the server 102 generates commentary text based on the indication information through the commentary model.
[0098] In the embodiments of the present application, the target scenario can be a real scenario or a virtual scenario. For example, the virtual scenario is a game scenario. The commentary model can comment on the target event in the game scenario. The target event can be set and changed as needed. For example, the target event is a virtual battle event, a clearance event, etc.
[0099] In some embodiments, the terminal 101 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a VR (Virtual Reality) device, an AR (Augmented Reality) device, etc., but is not limited thereto. In some embodiments, the server 102 can be an independent server or a server cluster or a distributed system composed of multiple servers, and can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In some embodiments, the server 102 mainly undertakes computing work, and the terminal 101 undertakes secondary computing work; or, the server 102 undertakes secondary computing services, and the terminal 101 undertakes primary computing work; or, the server 102 and the terminal 101 adopt a distributed computing architecture for collaborative computing.
[0100] See Figure 2 , Figure 2 is a flowchart of a method for training a commentary model provided by an embodiment of the present application. This method is executed by a computer device, and this method includes the following steps.
[0101] 201. The computer device obtains a training data set. The training data set includes multiple groups of first sample pairs. Each group of first sample pairs includes first indication information and a sample commentary text. The first indication information includes a sample commentary task, scenario data corresponding to the sample commentary task, and commentary requirements. The sample commentary task is a task of commenting on a target event in a target scenario, and the sample commentary text is used to describe the scenario data in the first indication information and meets the commentary requirements in the first indication information.
[0102] In the embodiments of the present application, the target scenario can be a real scenario or a virtual scenario. For example, the real scenario can be a real sports competition scenario, and the virtual scenario can be a game scenario.
[0103] In the embodiments of the present application, the commentary requirements are used to indicate at least one requirement among the language style, number of words, structure layout, term usage, commentary detail, commentary logic, etc. of the to-be-generated commentary text.
[0104] In an embodiment of the present application, the sample explanation text is a text generated by explaining the sample explanation task according to the explanation requirements based on the scene data. In an embodiment of the present application, the scene data can be the underlying data (GameCore) of the target scene. In the target scene, the underlying data of each frame is saved to the server in a formatted structure. The scene data can also be event description data obtained based on the underlying data. If the target scene is a game scene, and the underlying data is the coordinate data of the virtual object, the event description data can be the description data of the virtual object at a certain position, which is also the position indicated by the coordinate data.
[0105] 202. For each group of first sample pairs, the computer device inputs the first indication information in the first sample pairs into the pre-trained dialogue model to obtain a first predicted interpretation text for the first sample pairs.
[0106] In the embodiment of the present application, the pre-trained dialogue model is trained based on multiple groups of dialogue sample pairs. The pre-trained dialogue model is used to output a reply text based on a question text. The multiple groups of dialogue sample pairs are obtained from an open domain dialogue corpus, which includes dialogue corpus in multiple fields. The pre-trained dialogue model already has good human-computer dialogue capabilities.
[0107] In the embodiment of the present application, the pre-trained dialogue model is a large language model, such as the pre-trained dialogue model can be any one of a plurality of mixed large language models. In this embodiment, the pre-trained dialogue model is taken as an example of a Chinese-English language model with initial question-answering and dialogue functions.
[0108] In the embodiment of the present application, the pre-trained dialogue model is used to output a first predicted interpretation text based on the first indication information. The first predicted interpretation text is used to describe the scene data in the first indication information and has a certain language style, word count, structure layout and other formats.
[0109] 203. The computer device adjusts the model parameters of the pre-trained dialogue model based on the first predicted interpretation text of the first sample pair and the sample interpretation text in the first sample pair to obtain an interpretation model of the target scenario, wherein the interpretation model is used to generate the interpretation text based on the first indication information.
[0110] In an embodiment of the present application, the computer device iteratively trains the pre-trained dialogue model based on multiple groups of first samples until preset requirements are met to obtain an explanation model.
[0111] Wherein, meeting the preset requirement may mean that the gap between the first predicted commentary text and the sample commentary text is less than a preset gap. Further, meeting the preset requirement means that the loss value determined based on the first predicted commentary text and the sample commentary text reaches convergence, or the loss value reaches a preset threshold, or the number of iterations reaches a preset number, which is not specifically limited here.
[0112] An embodiment of the present application provides a method for training an explanation model. This method trains a pre-trained dialogue model based on indication information including an explanation task, scenario data, and explanation requirements to obtain an explanation model for the target scenario. Since the pre-trained dialogue model itself has good dialogue capabilities, using the basic dialogue capabilities of the pre-trained dialogue model and multiple groups of first samples to obtain the explanation model enables the explanation model to automatically generate accurate explanation texts based on the indication information and reduces the data required for training, thereby improving the training efficiency of the explanation model. And obtaining the explanation text through the explanation model not only improves the generation efficiency of the explanation text, but also, since the explanation model is a generative model, on the basis of improving the generation efficiency, it can also ensure the diversity of the generated explanation texts.
[0113] The above Figure 2 is the basic process of the method for training the explanation model. Below, based on Figure 3 the process of using the explanation model will be introduced. Refer to Figure 3 , Figure 3 which is a flowchart of a method for generating an explanation text provided by an embodiment of the present application. This method is executed by a computer device and includes the following steps.
[0114] 301. During the operation of the computer device in the target scenario, in response to reaching the explanation timing of the target explanation task, the computer device acquires the target scenario data corresponding to the target explanation task.
[0115] In an embodiment of the present application, the target explanation task is a task of explaining a target event in the target scenario. The explanation timing of the target explanation task can be set and changed as needed. And the target scenario may include multiple target explanation tasks, and the multiple target explanation tasks may correspond to different explanation timings.
[0116] 302. The computer device obtains target indication information based on the target explanation task, the target scenario data, and the target explanation requirements. The target indication information is used to indicate the explanation of the target explanation task based on the target scenario data and the target explanation requirements.
[0117] In an embodiment of the present application, the computer device constructs the target indication information based on the target explanation task, the target scenario data, and the target explanation requirements. The target explanation requirements are pre-configured before the operation of the target scenario. Thus, when reaching the explanation timing, the target indication information can be directly constructed based on the target explanation requirements.
[0118] 303. The computer device inputs the target indication information into the explanation model.
[0119] In an embodiment of the present application, the explanation model is used to output an explanation text based on the target indication information.
[0120] 304. The computer device outputs the target commentary text corresponding to the target commentary task through the commentary model.
[0121] In the embodiment of the present application, the target commentary text is used to describe the target scene data and meets the target commentary requirements. Among them, the target commentary text includes the text content obtained by event description and event summary based on the target scene data, that is, the target commentary text includes at least one of the event description content, event summary content, etc. Further, the target commentary text also includes the predicted event content obtained by event prediction based on the target scene data.
[0122] The embodiment of the present application provides a method for generating a commentary text. After reaching the commentary timing of the target commentary task, the method constructs indication information based on the target commentary task, the corresponding scene data, and the commentary requirements, and then obtains the corresponding commentary text through the commentary model based on the indication information. This method generates the commentary text based on the scene data and the commentary requirements, ensuring the accuracy of the generated commentary text, and generates the commentary text based on the commentary model, improving the generation efficiency of the commentary text; and since the commentary model is a generative model, on the basis of improving the generation efficiency, it can also ensure the diversity of the generated commentary text.
[0123] The above Figure 2 is the basic process of the training method of the commentary model. The following is a further introduction to the training method of the commentary model based on Figure 4 See Figure 4 , Figure 4 is the flowchart of a training method of a commentary model provided by the embodiment of the present application. This method is executed by a computer device, and this method includes the following steps.
[0124] 401. The computer device obtains a pre-trained dialogue model, and the pre-trained dialogue model is trained based on multiple groups of dialogue samples.
[0125] In the embodiment of the present application, the pre-trained dialogue model in step 401 is the same as the pre-trained dialogue model in step 202, and will not be elaborated here.
[0126] 402. The computer device obtains multiple sample commentary texts.
[0127] In the embodiments of the present application, the number of multiple sample commentary texts can be set and changed as needed. Optionally, the multiple sample commentary texts are commentary texts obtained by an announcer or a player through analysis and summary based on the actual situation in the target scenario. Further, the multiple sample commentary texts are commentary texts that meet the commentary requirements determined from multiple candidate commentary texts, and thus the obtained multiple sample commentary texts are high-quality commentary texts. The multiple candidate commentary texts are commentary texts obtained by an announcer or a player through analysis and summary based on the actual situation in the target scenario.
[0128] In the embodiments of the present application, the multiple sample commentary texts correspond to multiple event types, and each event type corresponds to at least one sample commentary text. For example, referring to Table 1, Table 1 is an example of the sample commentary texts corresponding to multiple event types provided in the embodiments of the present application. The multiple event types are respectively a position analysis event, a terrain analysis event, a virtual battle analysis event, and a direction analysis event.
[0129] Table 1
[0130]
[0131] 403. The computer device constructs multiple second indication messages based on the multiple sample commentary texts, the multiple sample commentary tasks, the scenario data of each of the multiple sample commentary tasks, and the commentary requirements. Each second indication message includes a sample commentary task, the corresponding scenario data, the commentary requirements, and the first quantity of sample commentary texts. The second indication message is used to indicate the generation of commentary texts for the included sample commentary tasks based on the text features of the first quantity of sample commentary texts.
[0132] In the embodiments of the present application, the first quantity of sample commentary texts is a part of the multiple sample commentary texts. The first quantity can be set and changed as needed, and no specific limitation is provided here.
[0133] In the embodiments of the present application, the text features correspond to the commentary requirements, and the text features include at least one of the language style, the number of words, the structural layout, the term usage, the commentary detail, the commentary logic, etc. of the sample commentary texts.
[0134] In some embodiments, multiple sample commentary texts correspond to multiple event types. The process by which the computer device constructs multiple second indication messages based on multiple sample commentary texts, multiple sample commentary tasks, the scenario data and commentary requirements of each of the multiple sample commentary tasks includes the following steps: For each sample commentary task, the computer device determines, based on the event type of the target event in the sample commentary task, a first number of sample commentary texts from the multiple sample commentary texts that are of the same event type as the event type corresponding to the sample commentary task; The computer device constructs multiple second indication messages based on the multiple sample commentary tasks, the scenario data of each of the multiple sample commentary tasks, the commentary requirements, and the corresponding first number of sample commentary texts.
[0135] In this embodiment, the second prompt message is constructed based on the sample commentary text that is of the same event type as the event type of the target event to be commented on. Furthermore, based on the second prompt message, a commentary text of the corresponding event type can be generated, improving the accuracy of the commentary text.
[0136] It should be noted that the computer device can construct the second prompt message not only based on the event type, but also based on other detailed information in the target scenario to construct the second indication message, so as to generate a more accurate and in-depth sample commentary text based on the second indication message. Furthermore, the commentary model trained based on such a sample commentary text can generate a more accurate and in-depth commentary text.
[0137] 404. The computer device inputs the multiple second indication messages into the large language model respectively, and through the large language model, obtains the sample commentary texts corresponding to the multiple second indication messages respectively.
[0138] In the embodiment of the present application, after each second indication message is input into the large language model, the large language model outputs the sample commentary text corresponding to the second indication message.
[0139] In the embodiment of the present application, the large language model can be any one of multiple Hunyuan large models, and no specific limitation is made here. For example, see Figure 5 , Figure 5 is a schematic structural diagram of a large language model provided by an embodiment of the present application.
[0140] In the embodiment of the present application, the sample commentary text corresponding to the second indication message obtained through the large language model matches the text features of the first number of sample commentary texts. If the text features include the language style, the language style of the generated sample commentary text matches the language style of the first number of sample commentary texts.
[0141] 405. The computer device determines multiple groups of first sample pairs from multiple groups of second sample pairs. Each group of second sample pairs includes a sample commentary task, scenario data, commentary requirements, and the corresponding sample commentary text in one of the second indication messages.
[0142] In an embodiment of the present application, the sample interpretation task, scenario data, interpretation requirements, etc. in a second indication information constitute a first indication information.
[0143] In some embodiments, the process of the computer device determining multiple groups of first sample pairs from multiple groups of second sample pairs includes the following steps: The computer device obtains the quality values of multiple groups of third sample pairs, where the multiple groups of third sample pairs are partial sample pairs among the multiple groups of second sample pairs, and the quality value is used to indicate the quality level of the sample interpretation text in the third sample pair; determine multiple groups of fourth sample pairs from the multiple groups of third sample pairs, and the quality values of the multiple groups of fourth sample pairs are lower than the quality threshold; based on the multiple groups of fourth sample pairs, use a large language model to obtain the quality values of the remaining sample pairs in the multiple groups of second sample pairs; use the second sample pairs with quality values not lower than the quality threshold in the multiple groups of second sample pairs as the first sample pairs.
[0144] Optionally, the quality value is obtained based on multiple - dimensional evaluations of the generalized sample interpretation text by humans or computer devices. Among them, first determine the quality sub - values corresponding to the sample interpretation text in multiple dimensions respectively, and then use the mean or sum of the quality sub - values corresponding to the multiple dimensions as the quality value of the sample interpretation text. The multiple dimensions include, but are not limited to, dimensions such as text fluency, event compliance, and analysis depth.
[0145] It should be noted that in the generalized dataset obtained by the large language model, there may be some sample interpretation texts with low quality. In this embodiment, first, humans perform multi - dimensional scoring on some of the generalized sample interpretation texts, then select the sample interpretation texts with low quality values and use them as low - quality samples. Finally, the large language model determines the quality values of the generalized sample interpretation texts based on the low - quality samples, thereby improving the efficiency of determining the quality values without the need for humans to determine the quality values of the generalized sample interpretation texts one by one, saving time and effort. Moreover, through this embodiment, low - quality samples can be effectively filtered out, a training dataset with high quality can be obtained, and the accuracy of the trained interpretation model is high.
[0146] In some embodiments, the process by which the above computer device obtains the quality values of multiple remaining sample pairs in multiple groups of second sample pairs through a large language model based on multiple groups of fourth sample pairs includes the following steps: The computer device constructs third indication information corresponding to multiple groups of remaining sample pairs based on multiple groups of fourth sample pairs. Each third indication information includes a remaining sample pair, a second quantity of fourth sample pairs, and the corresponding quality value. The third indication information is used to indicate that the quality value of the included remaining sample pair is determined based on the text features and quality value of the explanatory text in the second quantity of fourth sample pairs; The computer device inputs multiple third indication information into the large language model respectively, and through the large language model, obtains the quality values of multiple groups of remaining sample pairs respectively.
[0147] In this embodiment, a small number of sample explanatory texts obtained by generalization are scored first, and then indication information is constructed based on the quality values of these sample explanatory texts. Generalization is performed through a large language model to obtain the quality values of multiple groups of sample pairs respectively, improving the acquisition efficiency of quality values, reducing the manual participation rate, and saving time and effort.
[0148] Optionally, the computer device also obtains low-score features of the sample explanatory text in the fourth sample pair in multiple dimensions, and constructs third indication information based on the low-score features in multiple dimensions. The low-score feature in any dimension is used to indicate the reason why the quality value of the sample explanatory text is low in this dimension. For example, for fluency, the low-score feature of the sample explanatory text in this dimension can be that the text is not fluent; Further, the low-score feature in this dimension can also be sentence segmentation errors, incorrect use of conjunctions, etc.
[0149] In some other embodiments, for each remaining sample pair, the large language model directly outputs whether the remaining sample pair is a sample pair with a quality value lower than the quality threshold based on the corresponding third indication information, without determining the quality value of the sample explanatory text, thereby further improving the screening efficiency.
[0150] In some other embodiments, the computer device directly constructs multiple third indication information based on multiple groups of third sample pairs and the corresponding quality values. This can not only determine the sample explanatory texts with low quality, but also determine the sample explanatory texts with high quality. By improving the diversity of reference samples in this way, the screening accuracy can be further improved.
[0151] In the embodiments of the present application, the process of obtaining the training data set is achieved through the above steps 402-405. In this embodiment, a small number of sample explanatory texts are first obtained, and these sample explanatory texts are used as a high-quality seed sample set. The indication information is input into the large language model in a few-shot manner to generalize the samples, obtaining a large number of sample explanatory texts, and then a large number of sample pairs are obtained. This embodiment can generalize the training data set from dozens of items to tens of thousands of items, improving the acquisition efficiency of the training data set, reducing the manual participation rate, and saving time and effort.
[0152] For example, refer to Figure 6 , Figure 6 is an example schematic diagram of a first indication information provided by an embodiment of the present application. Among them, the target scenario is taken as an example of a game scenario for illustration. The first indication information includes a sample explanation task, game data, and explanation requirements. Figure 6 The example of
[0153] In some embodiments, the above process of generalizing the sample pairs and removing low-quality texts can be performed in multiple rounds to obtain a large number of high-quality training samples.
[0154] 406. The computer device adjusts the model parameters of the pre-trained dialogue model based on the first predicted explanatory texts and sample explanatory texts of multiple groups of first sample pairs to obtain an initial explanatory model.
[0155] In the embodiments of the present application, the pre-trained dialogue model is taken as an example of a 100-billion Chinese-English language model with initial question-and-answer and dialogue functions for illustration. It is trained and generated based on a large number of data sets, has high performance in generating Chinese texts, and at the same time has a small model size and an efficient reasoning process.
[0156] In the embodiments of the present application, the computer device iteratively trains the pre-trained dialogue model based on the first predicted explanatory texts and sample explanatory texts of multiple groups of first sample pairs to obtain an initial explanatory model.
[0157] It should be noted that multiple first sample pairs can correspond to multiple language styles, and the preset requirements in the first indication information include language styles. Optionally, the computer device trains the pre-trained dialogue model based on the first sample pairs of multiple language styles, and then obtains an initial explanatory model that can generate explanatory texts of multiple language styles.
[0158] Among them, since the commentary requirements not only indicate the language style, but also indicate various commentary sub-requirements such as the structural layout, the number of words, and the use of terms, correspondingly, for each commentary sub-requirement, the computer device trains the pre-trained dialogue model respectively based on multiple types of sample pairs corresponding to the commentary sub-requirement, so that the initial commentary model can adapt to various commentary sub-requirements, thereby improving the generalization ability of the model.
[0159] In some other embodiments, if multiple groups of first sample pairs correspond to multiple event types, the computer device trains the pre-trained dialogue model respectively based on the first sample pairs of each of the multiple event types, so that the initial commentary model can adapt to multiple event types, thereby improving the generalization ability of the model.
[0160] In the embodiments of the present application, the training data set generated by generalization is initially screened, but the result output by the initial commentary model (M1) may not be an ideal commentary text. Moreover, the sample commentary text is not only affected by the training data, but also should be controllable and useful, that is, it is desired that the commentary text output by the commentary model is consistent with human preferences. For example, the generated commentary text not only has fluency and correct word order, but also needs to be useful and true. Therefore, it is necessary to optimize the initial commentary model based on the preferences of the object, and this optimization process includes the following steps 407-409.
[0161] 407. For each group of first sample pairs in multiple groups of first sample pairs, the computer device inputs the first indication information in the first sample pair into the initial commentary model to obtain a second predicted commentary text of the first indication information.
[0162] In the embodiments of the present application, the second predicted commentary text is used to describe the scenario data in the first indication information, and has certain formats such as a language style, the number of words, and a structural layout.
[0163] In the embodiments of the present application, the multiple groups of first sample pairs do not obtain their respective second predicted commentary texts simultaneously, but the multiple groups of first sample pairs iteratively execute steps 407-409 to obtain the commentary model. That is, steps 407-409 are an example of the usage process of a group of first sample pairs.
[0164] 408. The computer device inputs the first indication information and the second predicted commentary text of the first indication information into the reward model to obtain a reward value, where the reward value is used to indicate the degree of preference of the object for the second predicted commentary text, and the reward model is used to predict the degree of preference of the object for the predicted commentary text.
[0165] In some embodiments, the training process of the reward model includes the following steps: For each group of first sample pairs among multiple groups of first sample pairs, a computer device inputs the first indication information in the first sample pair into an initial interpretation model, and through the initial interpretation model, obtains multiple predicted interpretation texts of the first indication information; obtains the ranking among the multiple predicted interpretation texts, where the ranking is used to indicate the degree of preference of the object for the multiple predicted interpretation texts; determines multiple groups of interpretation sample pairs from the multiple predicted interpretation texts, and each group of interpretation sample pairs includes two predicted interpretation texts among the multiple predicted interpretation texts; inputs each group of interpretation sample pairs and the corresponding first indication information into the initial reward model to obtain the reward values corresponding to the two predicted interpretation texts included respectively; determines a loss based on the reward values corresponding to the two predicted interpretation texts and the ranking, where the loss is used to indicate the difference between the ranking of the two predicted interpretation texts and the ranking corresponding to the reward values of the two predicted interpretation texts; and adjusts the initial reward model based on the losses corresponding to the multiple groups of interpretation sample pairs respectively to obtain the reward model.
[0166] In the embodiments of the present application, the ranking corresponding to the reward values of two predicted interpretation texts refers to the ranking where the predicted interpretation text with a larger reward value is in front of the predicted interpretation text with a smaller reward value.
[0167] In the embodiments of the present application, iterative training is performed on the initial reward model based on multiple first indication information in multiple groups of first sample pairs until a preset requirement is met to obtain the reward model.
[0168] In the implementation of the present application, for each first indication information, it is input into the initial interpretation model multiple times to obtain multiple predicted interpretation texts of the first indication information. Since the initial interpretation model is a generative model, different multiple predicted interpretation texts can be obtained.
[0169] In the embodiments of the present application, for each first indication information, the initial interpretation model generates multiple outputs. For example, the outputs are four results A, B, C, and D. Then, the annotator sorts the four results, and then uses the pairwise ranking relationship to train the reward model for reward value regression.
[0170] Optionally, the network structure of the reward model is the same as the network structure of the initial interpretation model after removing the last embedding layer. The input of the initial reward model is the first prompt information and the predicted interpretation text, and the output is the reward value. Among them, for each first prompt information, the initial reward model randomly generates K outputs, where K is an integer greater than 1. Then their paired output results, that is, a total of output results for each first prompt information. During training, the initial interpretation model uses the output results of each first indication information as a training batch to calculate the comprehensive damage.
[0171] Accordingly, the process of the above computer device adjusting the initial reward model based on the losses corresponding to multiple groups of commentary samples to obtain the reward model includes the following two implementation manners.
[0172] In one implementation manner, the computer device uses the mean value of the losses corresponding to multiple groups of commentary samples as the comprehensive loss, and adjusts the initial reward model based on the comprehensive loss to obtain the reward model. In another implementation manner, the computer device uses the sum of the losses corresponding to multiple groups of commentary samples as the comprehensive loss, and adjusts the initial reward model based on the comprehensive loss to obtain the reward model.
[0173] For example, the loss function of the reward model is shown in the following formula (1). This reward function is used to maximize the difference in reward values between the predicted commentary text ranked ahead and the predicted commentary text ranked behind.
[0174]
[0175] Among them, loss(θ) represents the comprehensive loss of the initial reward model, and θ represents the model parameters of the initial reward model; represents the expectation on the data set D, D represents the data set composed of the first indication information and its multiple predicted commentary texts; x represents the first indication information; y w represents the text ranked ahead among the two predicted commentary texts; y l represents the text ranked behind among the two predicted commentary texts; σ represents the sigmoid function; M 2 (x, y w ) represents the reward value corresponding to the first indication information x and the predicted commentary text y w ; M 2 (x, y l ) represents the reward value corresponding to the first indication information x and the predicted commentary text y l .
[0176] In the embodiments of the present application, the multiple predicted commentary texts are ranked based on the degree of preference of the object for the multiple predicted commentary texts, and then the reward model is trained based on this ranking and the ranking corresponding to the reward value output by the reward model. The reward value predicted by the trained reward model can match the preferences of the object. Thus, based on this reward model, the degree of preference of the object corresponding to the commentary text can be accurately predicted, which is convenient for optimizing the initial commentary model.
[0177] 409. The computer device adjusts the model parameters of the initial commentary model based on the reward value to obtain the commentary model.
[0178] In the embodiments of the present application, the computer device uses the PPO (Proximal Policy Optimization) algorithm to optimize the initial commentary model. During the policy update process of this algorithm, it is necessary to ensure that the new and old policies are close enough to guarantee the stability of reinforcement learning.
[0179] For example, the objective function (loss function) of the PPO algorithm is shown in the following formula (2). This objective function is used to measure the gap between the new policy and the old policy of the initial commentary model, and a penalty term is added to prevent the update amplitude from being too large. Moreover, the policy parameters are updated using gradient descent to maximize the objective function. This objective function is used to maximize the reward value output after the predicted commentary text obtained by inputting the first prompt information into the initial commentary model is brought into the reward model, that is, to maximize the reward value of the commentary text output by the initial commentary model.
[0180]
[0181] Among them, objective(φ) represents the objective function of the initial commentary model; represents the trajectory distribution obtained by reinforcement learning of the initial commentary model; is initialized to be consistent with and represents the reinforcement learning policy, which is also the initial commentary model to be updated; represents the policy of the initial commentary model obtained through supervised learning; β represents the weight coefficient of the KL divergence, which is used to control the intensity of the KL penalty; is a regularization term. As the initial commentary model is updated, the difference between the data generated by the reinforcement learning model and the data for training the reward model will become larger and larger. Therefore, a KL penalty term is added to the loss function to ensure that the output of the model after reinforcement learning and the output gap is not very large. represents calculating the reward value M at each time step under the trajectory distribution obtained by reinforcement learning 2 (x,y) and the KL penalty and taking the mean. The role of the KL penalty is to limit the difference between the reinforcement learning policy and the supervised learning policy. x represents the input, and y represents the output.
[0182] In some embodiments, the process by which the above computer device adjusts the model parameters of the initial commentary model based on the reward value to obtain the commentary model includes the following steps: The computer device determines the policy update gradient of the initial commentary model based on the reward value; updates the model parameters of the initial commentary model based on the policy update gradient to obtain the commentary model.
[0183] Optionally, based on the PPO algorithm, the above process includes the following steps: The computer device obtains the objective function objective(φ) based on the reward value through the above formula (2) to measure the difference between the new policy and the old policy, and adds a penalty term to prevent the update amplitude from being too large; then uses the gradient descent method to determine the policy update gradient to update the policy parameters (i.e., model parameters) of the initial commentary model to maximize the objective function.
[0184] It should be noted that during the optimization process, since the training steps cannot be too large, that is, a KL penalty term is added to the update mechanism, so just one round of model optimization may not meet our expectations, because the results output by the fine-tuned model may not conform to the expectations either, and the quality of the commentary text output can be gradually improved through multiple rounds of optimization. Therefore, the commentary model is obtained through multiple rounds of iterative optimization.
[0185] In some embodiments, the training of the reward model and the optimization process of the initial commentary model are carried out alternately. That is, after adjusting the model parameters of the initial reward model once through a batch of data, the initial commentary model is optimized once based on the adjusted initial reward model, that is, the initial reward model and the initial commentary model are trained simultaneously. Repeat the above steps until the policy converges or reaches the preset number of iterations to obtain the reward model and the commentary model.
[0186] In other embodiments, the model parameters of the initial reward model are adjusted multiple times based on multiple batches of data. After obtaining the reward model, the initial commentary model is optimized multiple times based on the reward model, that is, the reward model is trained first and then the initial commentary model is optimized.
[0187] In other embodiments, after fine-tuning the model parameters of the pre-trained dialogue model once through a batch of data, based on the pre-trained dialogue model obtained from this adjustment, the initial reward model is trained, and then the pre-trained dialogue model is optimized based on the output of the initial reward model. Repeat the steps of fine-tuning the pre-trained dialogue model, training the initial reward model, and optimizing the pre-trained dialogue model until the policy converges or reaches the preset number of iterations to obtain the reward model and the commentary model.
[0188] In the embodiments of the present application, through the reinforcement learning strategy, the initial commentary model is optimized based on the degree of preference of the object for the commentary sample. The obtained commentary model not only has objective characteristics such as fluency and correct word order, but also is consistent with human preferences, is real and useful, conforms to subjective preferences, and further improves the quality of the commentary text output by the commentary model.
[0189] For example, see Figure 7 , Figure 7It is a training flow chart of an interpretation model provided by an embodiment of the present application. Among them, first construct second indication information based on manually annotated corpus, then generalize through a large language model to obtain a large amount of corpus. Then filter the corpus through the large language model to obtain high-quality corpus. This process can be iterated multiple times, and thus the quality of the obtained training data set is high. Then fine-tune the pre-trained dialogue model based on the obtained training data set to obtain an initial interpretation model, and then use the method of RLHF (Reinforcement Learning from Human Feedback, optimizing the language model using human feedback signals) reinforcement learning to optimize the initial interpretation model. Among them, train the reward model based on the output of the initial reward model. Determine the policy update gradient of the initial interpretation model based on the output of the reward model and the PPO algorithm, and then optimize the initial interpretation model. Repeat the above process to obtain the reward model and the interpretation model. Among them, θ represents the model parameters of the initial interpretation model, represents the policy update gradient; r θ (y|x) represents the reward value corresponding to the first indication information x and the predicted interpretation text y; π PPO (y|x) represents the policy parameters based on the PPO algorithm, π base (y|x) represents the policy parameters of the initial interpretation model; λ KL represents the weight coefficient of the KL divergence; D KL represents the trajectory distribution corresponding to the KL divergence.
[0190] The training method provided by the embodiment of the present application distills the analysis ability of existing large models onto a smaller LLM (Large Language Model) model, and then conducts domain learning on the LLM model, thereby training a dedicated large model for a vertical domain to achieve efficient inference performance and a relatively high level of domain knowledge analysis ability. And the large language model has already shown a human level in various specialties and academics. The embodiment of the present application uses the large language model as a baseline to train a large model with a smaller scale but similar capabilities to the large language. Among them, generalize and expand the training data set through the large language model, manually perform high-quality annotation on different events, and design different indication information (prompt) to be input into the large language model in the form of few-shot, thereby generating a large amount of training data set. Then, fine-tune the various interpretation tasks of the target scenario on the pre-trained dialogue language model, and finally obtain a small-scale large model with high performance and high precision.
[0191] The above Figure 3 is the basic process of the generation method of the interpretation text. Next, based on Figure 8 further introduce the generation method of the interpretation text. See Figure 8 , Figure 8The figure is a flowchart of a method for generating commentary text provided by an embodiment of the present application. The commentary model used in this method is trained by any of the above embodiments. The execution subject of this method is a computer device, and the method includes the following steps.
[0192] 801. The computer device displays multiple candidate commentary requirements, and the corresponding commentary language styles of the multiple candidate commentary requirements are different.
[0193] In the embodiments of the present application, the language style refers to different language materials and ways adopted according to different communication occasions, purposes, tasks, and the dispositions and qualities of communicators. Multiple language styles include daily spoken language style, applied writing style, artistic writing style, and personal language style, etc. Further, the language style includes various styles such as concise, concise, humorous, witty, cheerful and joyful, solemn and formal, etc.
[0194] 802. In response to a selection operation on any candidate commentary requirement, the computer device uses the selected candidate commentary requirement as the target commentary requirement.
[0195] In the embodiments of the present application, any candidate commentary requirement not only indicates the language style, but also indicates other commentary sub-requirements, such as structural layout, number of words, term usage, etc. Optionally, the computer device also displays multiple candidate sub-requirements for each of the multiple commentary sub-requirements, thereby facilitating the operator to make a free choice to select the commentary requirements that meet their preferences.
[0196] In the embodiments of the present application, the process of obtaining the target commentary requirement is realized through the above steps 801-802. In this embodiment, by providing rich and customizable commentary requirements, different commentary texts can be generated according to the habits and preferences of different objects, realizing a personalized customized commentary service for each person. That is, this embodiment greatly improves the expandability and interest of AI commentary, and thus can improve the public's interest in watching the game.
[0197] It should be noted that there can be multiple target commentary tasks in a target scene. The multiple target commentary tasks respectively correspond to different scene data, event types, and commentary requirements, that is, the multiple target commentary tasks respectively correspond to different target indication information.
[0198] Among them, the triggering times of the multiple target commentary tasks can be different. Before running the target scene, the commentary times of each of the multiple target commentary tasks and the commentary requirements of the multiple target commentary tasks are set. Then, during the running process of the target scene, after reaching the commentary time of any target commentary task, based on the target scene data corresponding to the target commentary task and the pre-set commentary requirements, the target indication information is directly generated, and then the target commentary text is obtained through the commentary model.
[0199] 803. During the operation of the computer device in the target scenario, in response to reaching the commentary opportunity of the target commentary task, the computer device obtains the target scenario data corresponding to the target commentary task.
[0200] In some embodiments, the target scenario is a game scenario, and there are multiple virtual teams in the target scenario. Each virtual team includes multiple virtual objects. The process by which the computer device obtains the target scenario data corresponding to the target commentary task in response to reaching the commentary opportunity of the target commentary task includes the following steps: In response to reaching the commentary opportunity of the target commentary task, the computer device obtains the scenario data corresponding to the target commentary task for each of the multiple virtual teams; based on the scenario data corresponding to each of the multiple virtual teams, the computer device uses the scenario data corresponding to the target virtual team that meets the preset conditions among the multiple virtual teams as the target scenario data, and the target scenario data is used to comment on the behavior of the target virtual team in the target scenario.
[0201] In the embodiments of the present application, the preset conditions can be set and changed as needed. Optionally, the target virtual team that meets the preset conditions refers to the virtual team selected by the event selection method based on reinforcement learning. Optionally, the computer device obtains the scenario data corresponding to each of the multiple virtual teams and stores them in the event pool. The target virtual team is selected from the event pool.
[0202] For example, refer to Figure 9 , Figure 9 is a flowchart of a method for generating commentary text provided by an embodiment of the present application. Based on the event trigger mechanism, after the computer device reaches the commentary opportunity of the target commentary task, it triggers the commentary on the target event in the target commentary task. Among them, multiple virtual teams are added to the event pool, and the target virtual team is selected through reinforcement learning. The commentary requirements of the selected language style and the scenario data of the target virtual team are input into the commentary model to obtain a commentary sample.
[0203] In the embodiments of the present application, the target scenario is taken as an example of a game scenario for illustration. Further, there is a circle shrinking event in the game scenario. The circle shrinking event means that as the game progresses, the movable area of the virtual objects in the game scenario will continuously shrink in a circular form according to a preset mechanism. Each time the circle shrinking event occurs, the movable area will shrink once. Optionally, the event types of the target event in the target commentary task include the circle entry analysis event and the team summary event.
[0204] Among them, the circle entry analysis event includes the position analysis event, the terrain analysis event, the direction analysis event, the virtual resource holding situation analysis event, and the virtual battle analysis event. Correspondingly, the target commentary task includes the task of commenting on at least one of the position information, terrain information, direction information, virtual resource holding information, and virtual battle information of the virtual objects in the target virtual team.
[0205] In the embodiments of the present application, a virtual team battle refers to a group battle in which virtual objects belonging to different virtual teams occur within a certain area, and the area where the virtual team battle occurs can be referred to as a team battle area. For any virtual team battle, it can occur at any moment during the game operation and can also end at any moment during the game operation.
[0206] Among them, the commentary timing includes the timing when no virtual team battle is taking place. In some embodiments, the commentary timing is the timing when no virtual team battle occurs within a preset duration before a circle brush event, or the timing when no virtual team battle occurs within a preset duration after the circle brush event. Or both are commentary timings. The preset duration before the circle brush event refers to the time period with the occurrence time of the circle brush event as the end point, and the preset duration after the circle brush event refers to the time period with the occurrence time of the circle brush event as the starting point.
[0207] In other embodiments, determine a first time point that is a preset duration from the occurrence time of the circle brush event before the circle brush event. If no virtual team battle occurs at this first time point, then take the first time point as the commentary timing. Or, determine a second time point that is a preset duration from the occurrence time of the circle brush event after the circle brush event. If no virtual team battle occurs at this second time point, then take the second time point as the commentary timing. Or both are commentary timings.
[0208] Among them, the team summary event refers to summarizing the events that occur in the virtual team battle. The types of virtual team battles include the team wipeout type and the non-team wipeout type, and the non-team wipeout type is also the friction team type. The application scenarios of the team summary event include but are not limited to the following three application scenarios. One application scenario is to review and analyze the battle situation and outstanding performances of a certain virtual team in the virtual team battle, that is, the highlight review. Another application scenario is to analyze the reasons for the victory or defeat of a certain virtual team, that is, the attribution of the team result. Another application scenario is to analyze the tactics, gameplay, etc. of a certain virtual team for teaching. Correspondingly, the target commentary task includes the task of commenting on at least one of the highlight moments, victory and defeat reasons, and virtual tactics in the virtual team battle.
[0209] Among them, the commentary timing includes the timing when the virtual team battle ends, the timing when the virtual team battle is in progress, and the timing of entering the final circle. Among them, the timing when the virtual team battle is in progress refers to the timing of team tug-of-war. The timing of entering the final circle refers to the timing of tug-of-war in the final circle.
[0210] In the embodiments of the present application, tasks for explaining various events, various explanation timings, and various team battle types are improved. The application scenarios, team battle types, and trigger timings in the summarized events can be combined arbitrarily, and multiple different explanation tasks can be obtained. Each explanation task independently corresponds to a first indication information to control the output of the explanation text. For example, the explanation task can indicate that after the end of a virtual team battle of the group annihilation type, analyze the reasons why the winning virtual team annihilated the opponent, when there is tug-of-war within the friction map group, review the battle situation of a certain virtual team in the virtual team battle, or when there is tug-of-war in the final circle, analyze the outstanding performance of a certain virtual team in the previous virtual team battle, etc. In the embodiments of the present application, the event types and the content of the speech explanations are greatly expanded, bringing new possibilities and creativity to game explanations.
[0211] In some other embodiments, there are levels in the game scenario. The explanation timing includes the timing when the virtual object passes through the level. The explanation task includes the task of explaining at least one of the outstanding performance of the virtual object during the clearance process, the reason for successful clearance, etc.
[0212] In some other embodiments, there are NPC (non-player character) roles in the game scenario. Then the explanation timing includes the timing when the virtual object defeats the NPC role. The explanation task includes the task of explaining at least one of the outstanding performance of the virtual object during the battle with the NPC role, the reason for successful battle, etc.
[0213] 804. The computer device obtains target indication information based on the target explanation task, the target scenario data, and the target explanation requirements. The target indication information is used to indicate the explanation of the target explanation task based on the target scenario data and the target explanation requirements.
[0214] In the embodiments of the present application, the computer device constructs the target indication information based on the target explanation task, the target scenario data, and the target explanation requirements.
[0215] For example, refer to Figure 10 , Figure 10 which is a schematic diagram of a target indication information provided by the embodiments of the present application. Among them, taking the target scenario as the game scenario and the target explanation task as explaining the position information of the virtual object in a certain virtual team as an example. The target indication information includes the target explanation task, multiple explanation requirements, and game data. Among them, the data input by the user part is the scenario data, which is further the event description data obtained based on the underlying data. The system part is set according to different explanation tasks. During reasoning, this method can output the explanation text in a specific style end-to-end without the need for a large number of engineering strategies to implement, and has high migration.
[0216] In an embodiment of the present application, taking the target scenario as a game scenario as an example, a game content analysis method based on a large model and game underlying data Gamecore is proposed. This method is driven by the game underlying data Gamecore, and customizes a vertical commentary model in the game. This commentary model provides game analysis of various event types, such as circle entry analysis, team highlight summary, team victory attribution, etc., and provides rich and customizable events and commentary scripts for the AI commentary mechanism, greatly improving the scalability and interest of the AI commentary, and enhancing the public's interest in watching games.
[0217] In an embodiment of the present application, driven by a large model, by constructing different indication information (prompt), commentary scripts for game events and content analysis with very different styles are generated, improving the diversity. And in the process of generating the commentary scripts, the personalized needs of players are also considered. According to the game habits and preferences of different players, different commentary scripts are generated to achieve a customized commentary service for thousands of people.
[0218] 805. The computer device inputs the target indication information into the commentary model, and through the commentary model, outputs the target commentary text corresponding to the target commentary task.
[0219] For example, refer to Figure 11 , Figure 11 FIG. is a schematic diagram of the generation of a commentary text provided by an embodiment of the present application. Among them, taking the target scenario as a game scenario and the target commentary task as an example of commenting on the performance of a virtual team in a virtual team battle for illustration. The target indication information includes the target commentary task, multiple commentary requirements, and game data. After inputting the target indication information into the commentary model, the target commentary text of the target indication information is obtained.
[0220] In an embodiment of the present application, after the computer device obtains the target commentary text, it plays the target commentary text through an audio playback device so that the object can receive the target commentary text.
[0221] In an embodiment of the present application, the game content commentary solution based on a large model and GameCore, relying on the capabilities of the large model, comprehensively improves the quality of the AI commentary text, making the commentary content more rich, vivid, and customized. And the system theoretical framework and model algorithm provided by the embodiment of the present application have high migratability and can be easily applied to the commentary of various scenarios, such as applied to the commentary of various FPS games, and have broad application prospects and application values.
[0222] An embodiment of the present application provides a method for generating commentary text. After reaching the commentary timing of the target commentary task, based on the target commentary task, the corresponding scenario data, and the commentary requirements, indication information is constructed. Then, based on the indication information, through a commentary model, the corresponding commentary text is obtained. This method generates commentary text based on scenario data and commentary requirements, ensuring the accuracy of the generated commentary text. And it generates commentary text based on a commentary model, improving the generation efficiency of the commentary text. Moreover, since the commentary model is a generative model, on the basis of improving the generation efficiency, it can also ensure the diversity of the generated commentary text.
[0223] Figure 12 is a block diagram of a training device for a commentary model provided according to an embodiment of the present application. Refer to Figure 12 , the device includes:
[0224] An acquisition module 1201, configured to acquire a training data set. The training data set includes multiple groups of first sample pairs. Each group of first sample pairs includes first indication information and a sample commentary text. The first indication information includes a sample commentary task, the scenario data corresponding to the sample commentary task, and the commentary requirements. The sample commentary task is a task of commenting on a target event in a target scenario. The sample commentary text is used to describe the scenario data in the first indication information and conforms to the commentary requirements in the first indication information;
[0225] An input / output module 1202, configured to, for each group of first sample pairs, input the first indication information in the first sample pair into a pre-trained dialogue model to obtain a first predicted commentary text of the first sample pair;
[0226] An adjustment module 1203, configured to adjust the model parameters of the pre-trained dialogue model based on the first predicted commentary text of the first sample pair and the sample commentary text in the first sample pair to obtain a commentary model for the target scenario. The commentary model is used to generate commentary text based on the first indication information.
[0227] In some embodiments, the acquisition module 1201 is configured to:
[0228] Acquire multiple sample commentary texts;
[0229] Based on multiple sample commentary texts, multiple sample commentary tasks, the scenario data and commentary requirements of each of the multiple sample commentary tasks, construct multiple second indication information. Each second indication information includes a sample commentary task, the corresponding scenario data, the commentary requirements, and a first quantity of sample commentary texts. The second indication information is used to indicate the generation of commentary text for the included sample commentary task based on the text features of the first quantity of sample commentary texts;
[0230] Input multiple pieces of second indication information into the large language model respectively, and obtain the sample commentary texts corresponding to the multiple pieces of second indication information respectively through the large language model;
[0231] Determine multiple groups of first sample pairs from multiple groups of second sample pairs. Each group of second sample pairs includes a sample commentary task, scenario data, commentary requirements, and the corresponding sample commentary text in one piece of second indication information.
[0232] In some embodiments, the acquisition module 1201 is configured to:
[0233] Obtain the quality values of multiple groups of third sample pairs respectively. The multiple groups of third sample pairs are partial sample pairs in the multiple groups of second sample pairs, and the quality value is used to indicate the quality level of the sample commentary text in the third sample pair;
[0234] Determine multiple groups of fourth sample pairs from the multiple groups of third sample pairs. The quality values of the multiple groups of fourth sample pairs are lower than the quality threshold;
[0235] Based on the multiple groups of fourth sample pairs, obtain the quality values of the multiple groups of remaining sample pairs in the multiple groups of second sample pairs through the large language model;
[0236] Use the second sample pairs in the multiple groups of second sample pairs whose quality values are not lower than the quality threshold as the first sample pairs.
[0237] In some embodiments, the acquisition module 1201 is configured to:
[0238] Based on the multiple groups of fourth sample pairs, construct third indication information corresponding to the multiple groups of remaining sample pairs respectively. Each third indication information includes a remaining sample pair, a second quantity of fourth sample pairs, and the corresponding quality value. The third indication information is used to indicate the quality value of the included remaining sample pair determined based on the text features and quality value of the commentary text in the second quantity of fourth sample pairs;
[0239] Input the multiple pieces of third indication information into the large language model respectively, and obtain the quality values of the multiple groups of remaining sample pairs respectively through the large language model.
[0240] In some embodiments, the multiple sample commentary texts correspond to multiple event types. The acquisition module 1201 is configured to:
[0241] For each sample commentary task, based on the event type of the target event in the sample commentary task, determine a first quantity of sample commentary texts with the same event type as the sample commentary task from the multiple sample commentary texts;
[0242] Based on the multiple sample commentary tasks, the scenario data, commentary requirements of the multiple sample commentary tasks respectively, and the first quantity of sample commentary texts corresponding to each of them, construct multiple pieces of second indication information.
[0243] In some embodiments, the adjustment module 1203 is configured to:
[0244] Based on multiple groups of first samples, their respective first predicted explanatory texts and sample explanatory texts, adjust the model parameters of the pre-trained dialogue model to obtain an initial explanatory model;
[0245] For each group of first samples in the multiple groups of first samples, input the first indication information in the first sample into the initial explanatory model to obtain a second predicted explanatory text of the first indication information;
[0246] Input the first indication information and the second predicted explanatory text of the first indication information into the reward model to obtain a reward value, where the reward value is used to indicate the degree of preference of the object for the second predicted explanatory text, and the reward model is used to predict the degree of preference of the object for the predicted explanatory text;
[0247] Based on the reward value, adjust the model parameters of the initial explanatory model to obtain an explanatory model.
[0248] In some embodiments, the device further includes a training module, configured to:
[0249] For each group of first samples in the multiple groups of first samples, input the first indication information in the first sample into the initial explanatory model, and through the initial explanatory model, obtain multiple predicted explanatory texts of the first indication information;
[0250] Obtain the ranking among the multiple predicted explanatory texts, where the ranking is used to indicate the degree of preference of the object for the multiple predicted explanatory texts;
[0251] Determine multiple groups of explanatory sample pairs from the multiple predicted explanatory texts, where each group of explanatory sample pairs includes two predicted explanatory texts among the multiple predicted explanatory texts;
[0252] Input each group of explanatory sample pairs and the corresponding first indication information into the initial reward model to obtain the reward values corresponding to the two predicted explanatory texts included respectively;
[0253] Based on the reward values corresponding to the two predicted explanatory texts respectively and the ranking, determine a loss, where the loss is used to indicate the difference between the ranking of the two predicted explanatory texts and the ranking corresponding to the reward values of the two predicted explanatory texts;
[0254] Based on the losses corresponding to the multiple groups of explanatory sample pairs respectively, adjust the initial reward model to obtain a reward model.
[0255] In some embodiments, the adjustment module 1203 is configured to:
[0256] Based on the reward value, determine the policy update gradient of the initial explanatory model;
[0257] Update the model parameters of the initial commentary model based on the policy update gradient to obtain the commentary model.
[0258] An embodiment of the present application provides a training device for a commentary model. The device trains a pre-trained dialogue model based on indication information including a commentary task, scenario data, and commentary requirements to obtain a commentary model for the target scenario. Since the pre-trained dialogue model itself has good dialogue capabilities, using the basic dialogue capabilities of the pre-trained dialogue model and multiple sets of first sample pairs to obtain the commentary model enables the commentary model to automatically generate accurate commentary text based on the indication information and reduces the data required for training, thereby improving the training efficiency of the commentary model. And obtaining the commentary text through the commentary model not only improves the generation efficiency of the commentary text, but also, since the commentary model is a generative model, on the basis of improving the generation efficiency, it can also ensure the diversity of the generated commentary text.
[0259] Figure 13 It is a block diagram of a device for generating a commentary text according to an embodiment of the present application. Refer to Figure 13 The device includes:
[0260] An acquisition module 1301, configured to, during the operation of the target scenario, in response to reaching the commentary timing of the target commentary task, acquire target scenario data corresponding to the target commentary task;
[0261] A determination module 1302, configured to obtain target indication information based on the target commentary task, the target scenario data, and the target commentary requirements, where the target indication information is used to indicate the commentary on the target commentary task based on the target scenario data and the target commentary requirements;
[0262] An input / output module 1303, configured to input the target indication information into the commentary model, where the commentary model is trained by any one of the above training methods;
[0263] The input / output module 1303 is further configured to output, through the commentary model, a target commentary text corresponding to the target commentary task.
[0264] In some embodiments, the acquisition module 1301 is further configured to:
[0265] Display multiple candidate commentary requirements, where the commentary language styles corresponding to the multiple candidate commentary requirements are different;
[0266] In response to a selection operation on any one of the candidate commentary requirements, use the selected candidate commentary requirement as the target commentary requirement.
[0267] In some embodiments, there are multiple virtual teams in the target scenario, and each virtual team includes multiple virtual objects; the acquisition module 1301 is configured to:
[0268] In response to reaching the commentary timing of the target commentary task, obtain the scenario data of multiple virtual battle teams corresponding to the target commentary task respectively;
[0269] Based on the scenario data corresponding to multiple virtual battle teams respectively, use the scenario data corresponding to the target virtual battle team that meets the preset conditions among the multiple virtual battle teams as the target scenario data, and the target scenario data is used to comment on the behavior of the target virtual battle team in the target scenario.
[0270] In some embodiments, the commentary timing includes the timing when no virtual team battle is in progress, and the target commentary task includes the task of commenting on at least one of the position information of virtual objects in the target virtual battle team, the terrain information of the location where they are located, the direction information, the virtual resource holding information, and the virtual battle information.
[0271] In some embodiments, there are multiple virtual battle teams in the target scenario, and each virtual battle team includes multiple virtual objects;
[0272] The commentary timing includes the timing when the virtual team battle ends, the timing when the virtual team battle is in progress, and the timing of entering the final circle; the types of virtual team battles include the team wipe type and the non-team wipe type; the target commentary task includes the task of commenting on at least one of the highlight moments, the reasons for victory or defeat, and the virtual tactics in the virtual team battle.
[0273] The embodiment of the present application provides a device for generating commentary text. After reaching the commentary timing of the target commentary task, the device constructs indication information based on the target commentary task, the corresponding scenario data, and the commentary requirements, and then obtains the corresponding commentary text through the commentary model based on the indication information. The device generates the commentary text based on the scenario data and the commentary requirements, ensuring the accuracy of the generated commentary text, and generating the commentary text based on the commentary model, improving the generation efficiency of the commentary text; and since the commentary model is a generative model, on the basis of improving the generation efficiency, it can also ensure the diversity of the generated commentary text.
[0274] In the embodiment of the present application, the computer device can be a terminal or a server. When the computer device is a terminal, the terminal is used as the execution subject to implement the technical solution provided by the embodiment of the present application; when the computer device is a server, the server is used as the execution subject to implement the technical solution provided by the embodiment of the present application; or, the technical solution provided by the present application is implemented through the interaction between the terminal and the server, and the embodiment of the present application does not limit this.
[0275] Figure 14 The structural block diagram of the terminal 1400 provided by an exemplary embodiment of the present application is shown.
[0276] Generally, the terminal 1400 includes: a processor 1401 and a memory 1402.
[0277] The processor 1401 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1401 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1401 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0278] The memory 1402 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1402 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1402 is used to store at least one program code, and the at least one program code is used to be executed by the processor 1401 to implement the training method of the interpretation model or the generation method of the interpretation text provided in the method embodiments of the present application.
[0279] In some embodiments, the terminal 1400 may further optionally include: a peripheral device interface 1403 and at least one peripheral device. The processor 1401, the memory 1402, and the peripheral device interface 1403 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1403 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 1404, a display screen 1405, a camera assembly 1406, an audio circuit 1407, and a power supply 1408.
[0280] The peripheral device interface 1403 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1401 and the memory 1402. In some embodiments, the processor 1401, the memory 1402, and the peripheral device interface 1403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1401, the memory 1402, and the peripheral device interface 1403 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0281] The radio frequency circuit 1404 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1404 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1404 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1404 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1404 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1404 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0282] The display screen 1405 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1405 is a touch display screen, the display screen 1405 also has the ability to collect touch signals on or above the surface of the display screen 1405. The touch signals can be input to the processor 1401 as control signals for processing. At this time, the display screen 1405 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1405, which is provided on the front panel of the terminal 1400; in other embodiments, there may be at least two display screens 1405, which are respectively provided on different surfaces of the terminal 1400 or are in a foldable design; in other embodiments, the display screen 1405 may be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal 1400. Even, the display screen 1405 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1405 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0283] The camera module 1406 is used to collect images or videos. Optionally, the camera module 1406 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to realize the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera module 1406 may further include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. The dual-color-temperature flash refers to the combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.
[0284] The audio circuit 1407 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1401 for processing, or input to the radio frequency circuit 1404 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 1400. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1401 or the radio frequency circuit 1404 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1407 may also include a headphone jack.
[0285] The power supply 1408 is used to supply power to each component in the terminal 1400. The power supply 1408 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 1408 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0286] In some embodiments, the terminal 1400 further includes one or more sensors 1409. The one or more sensors 1409 include but are not limited to: an acceleration sensor 1410, a gyroscope sensor 1411, a pressure sensor 1412, an optical sensor 1413, and a proximity sensor 1414.
[0287] The acceleration sensor 1410 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the terminal 1400. For example, the acceleration sensor 1410 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1401 can control the display screen 1405 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1410. The acceleration sensor 1410 can also be used for collecting game or user's motion data.
[0288] The gyroscope sensor 1411 can detect the body direction and rotation angle of the terminal 1400. The gyroscope sensor 1411 can cooperate with the acceleration sensor 1410 to collect the 3D actions of the user on the terminal 1400. According to the data collected by the gyroscope sensor 1411, the processor 1401 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0289] The pressure sensor 1412 can be disposed on the side frame of the terminal 1400 and / or the lower layer of the display screen 1405. When the pressure sensor 1412 is disposed on the side frame of the terminal 1400, it can detect the holding signal of the user on the terminal 1400, and the processor 1401 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 1412. When the pressure sensor 1412 is disposed on the lower layer of the display screen 1405, the processor 1401 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1405. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0290] The optical sensor 1413 is used to collect the ambient light intensity. In one embodiment, the processor 1401 can control the display brightness of the display screen 1405 according to the ambient light intensity collected by the optical sensor 1413. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1405 is increased; when the ambient light intensity is low, the display brightness of the display screen 1405 is decreased. In another embodiment, the processor 1401 can also dynamically adjust the shooting parameters of the camera module 1406 according to the ambient light intensity collected by the optical sensor 1413.
[0291] The proximity sensor 1414, also known as a distance sensor, is usually disposed on the front panel of the terminal 1400. The proximity sensor 1414 is used to collect the distance between the user and the front of the terminal 1400. In one embodiment, when the proximity sensor 1414 detects that the distance between the user and the front of the terminal 1400 is gradually decreasing, the processor 1401 controls the display screen 1405 to switch from the lit state to the off state; when the proximity sensor 1414 detects that the distance between the user and the front of the terminal 1400 is gradually increasing, the processor 1401 controls the display screen 1405 to switch from the off state to the lit state.
[0292] Those skilled in the art can understand that Figure 14 the structure shown in does not limit the terminal 1400, and it may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.
[0293] Figure 15It is a schematic structural diagram of a server provided according to an embodiment of the present application. The server 1500 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 1501 and one or more memories 1502. Among them, the memory 1502 is used to store executable program codes, and the processor 1501 is configured to execute the above-mentioned executable program codes to implement the training method of the explanation model or the generation method of the explanation text provided by each of the above method embodiments. Of course, the server may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input / output. The server may also include other components for implementing device functions, which will not be elaborated here.
[0294] The embodiment of the present application also provides a computer-readable storage medium. At least one segment of program is stored in the computer-readable storage medium and is loaded and executed by a processor to implement the training method of the explanation model or the generation method of the explanation text in any of the above implementation manners.
[0295] The embodiment of the present application also provides a computer program product. The computer program product includes at least one segment of program. The at least one segment of program is stored in a computer-readable storage medium. The processor of the computer device reads the at least one segment of program from the computer-readable storage medium, and the processor executes the at least one segment of program, so that the computer device executes the training method of the explanation model or the generation method of the explanation text in any of the above implementation manners.
[0296] In some embodiments, the computer program product involved in the embodiment of the present application may be deployed to be executed on one computer device, or on multiple computer devices located at one place. Or, it may be executed on multiple computer devices distributed at multiple places and interconnected through a communication network. The multiple computer devices distributed at multiple places and interconnected through a communication network may form a blockchain system.
[0297] All the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present application, which will not be elaborated one by one here. The above are only optional embodiments of the present application and are not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A training method for an explanation model, characterized in that, the method includes: Obtain a training data set, which includes multiple groups of first sample pairs. Each group of the first sample pairs includes first indication information and a sample explanation text. The first indication information includes a sample explanation task, the scenario data corresponding to the sample explanation task, and an explanation requirement. The sample explanation task is a task of explaining a target event in a target scenario, and the sample explanation text is used to describe the scenario data in the first indication information and meet the explanation requirement in the first indication information; For each group of the first sample pairs, input the first indication information in the first sample pair into a pre-trained dialogue model to obtain a first predicted explanation text of the first sample pair; Based on the first predicted explanation text of the first sample pair and the sample explanation text in the first sample pair, adjust the model parameters of the pre-trained dialogue model to obtain an explanation model for the target scenario. The explanation model is used to generate an explanation text based on the first indication information.
2. The method according to claim 1, characterized in that, the obtaining of the training data set includes: Obtain multiple sample explanation texts; Based on the multiple sample explanation texts, multiple sample explanation tasks, the scenario data and explanation requirements corresponding to the multiple sample explanation tasks, construct multiple second indication information. Each second indication information includes a sample explanation task, the corresponding scenario data, an explanation requirement, and a first quantity of sample explanation texts. The second indication information is used to indicate the generation of an explanation text for the included sample explanation task based on the text features of the first quantity of sample explanation texts; Input the multiple second indication information into a large language model respectively, and through the large language model, obtain the sample explanation texts corresponding to the multiple second indication information respectively; Determine the multiple groups of first sample pairs from multiple groups of second sample pairs. Each group of second sample pairs includes a sample explanation task, scenario data, an explanation requirement, and the corresponding sample explanation text in one of the second indication information.
3. The method according to claim 2, characterized in that, the determining of the multiple groups of first sample pairs from the multiple groups of second sample pairs includes: Obtain the quality values of multiple groups of third sample pairs respectively. The multiple groups of third sample pairs are part of the multiple groups of second sample pairs, and the quality value is used to indicate the quality level of the sample explanation text in the third sample pair; Determine multiple groups of fourth sample pairs from the multiple groups of third sample pairs. The quality values of the multiple groups of fourth sample pairs are lower than a quality threshold; Based on the multiple groups of fourth sample pairs, through the large language model, obtain the quality values of multiple groups of remaining sample pairs in the multiple groups of second sample pairs; Use the second sample pairs in the multiple groups of second sample pairs whose quality values are not lower than the quality threshold as the first sample pairs.
4. The method according to claim 3, characterized in that, the obtaining of the quality values of multiple groups of remaining sample pairs in the multiple groups of second sample pairs through the large language model based on the multiple groups of fourth sample pairs includes: Based on the multiple groups of fourth sample pairs, construct third indication information corresponding to the multiple groups of remaining sample pairs respectively. Each third indication information includes a remaining sample pair, a second quantity of fourth sample pairs, and a corresponding quality value. The third indication information is used to indicate the quality value of the included remaining sample pair determined based on the text features and quality values of the explanatory texts in the second quantity of fourth sample pairs; Input the multiple third indication information into the large language model respectively, and through the large language model, obtain the quality values of the multiple groups of remaining sample pairs respectively.
5. The method according to claim 2, wherein, the multiple sample explanatory texts correspond to multiple event types. The construction of multiple second indication information based on the multiple sample explanatory texts, multiple sample explanatory tasks, the scenario data and explanatory requirements of the multiple sample explanatory tasks respectively includes: For each sample explanatory task, based on the event type of the target event in the sample explanatory task, determine a first quantity of sample explanatory texts from the multiple sample explanatory texts that have the same event type as the sample explanatory task; Based on the multiple sample explanatory tasks, the scenario data, explanatory requirements of the multiple sample explanatory tasks respectively, and the first quantity of sample explanatory texts corresponding to each of them, construct the multiple second indication information.
6. The method according to claim 1, wherein, the adjustment of the model parameters of the pre-trained dialogue model based on the first predicted explanatory text of the first sample pair and the sample explanatory text in the first sample pair to obtain the explanatory model for the target scenario includes: Based on the first predicted explanatory texts and sample explanatory texts of the multiple groups of first sample pairs respectively, adjust the model parameters of the pre-trained dialogue model to obtain an initial explanatory model; For each group of the multiple groups of first sample pairs, input the first indication information in the first sample pair into the initial explanatory model to obtain a second predicted explanatory text of the first indication information; Input the first indication information and the second predicted explanatory text of the first indication information into a reward model to obtain a reward value, where the reward value is used to indicate the degree of preference of an object for the second predicted explanatory text, and the reward model is used to predict the degree of preference of an object for the predicted explanatory text; Based on the reward value, adjust the model parameters of the initial explanatory model to obtain the explanatory model.
7. The method according to claim 6, wherein, the training process of the reward model includes: For each group of the multiple groups of first sample pairs, input the first indication information in the first sample pair into the initial explanatory model, and through the initial explanatory model, obtain multiple predicted explanatory texts of the first indication information; Obtain the ranking among the multiple predicted explanatory texts, where the ranking is used to indicate the degree of preference of an object for the multiple predicted explanatory texts; Determine multiple groups of explanatory sample pairs from the multiple predicted explanatory texts, and each group of explanatory sample pairs includes two predicted explanatory texts among the multiple predicted explanatory texts; Input each group of explanatory sample pairs and the corresponding first indication information into the initial reward model to obtain the reward values corresponding to the two predicted explanatory texts included; Based on the reward values and rankings corresponding to the two predicted explanatory texts respectively, determine a loss, where the loss is used to indicate the difference between the rankings of the two predicted explanatory texts and the rankings corresponding to the reward values of the two predicted explanatory texts; Based on the losses corresponding to the multiple groups of explanatory sample pairs respectively, adjust the initial reward model to obtain the reward model.
8. The method according to claim 6, wherein, the adjusting the model parameters of the initial explanatory model based on the reward value to obtain the explanatory model includes: determining a policy update gradient of the initial explanatory model based on the reward value; updating the model parameters of the initial explanatory model based on the policy update gradient to obtain the explanatory model.
9. A method for generating an explanatory text, wherein, the method includes: During the operation of the target scenario, in response to reaching the explanatory timing of the target explanatory task, obtain the target scenario data corresponding to the target explanatory task; Based on the target explanatory task, the target scenario data, and the target explanatory requirements, obtain target indication information, where the target indication information is used to indicate explaining the target explanatory task based on the target scenario data and the target explanatory requirements; Input the target indication information into an explanatory model, where the explanatory model is trained by the method according to any one of claims 1-8; Output the target explanatory text corresponding to the target explanatory task through the explanatory model.
10. The generation method according to claim 9, wherein, the obtaining process of the target explanatory requirements includes: displaying a plurality of candidate explanatory requirements, where the explanatory language styles corresponding to the plurality of candidate explanatory requirements are different; In response to a selection operation on any one of the candidate explanatory requirements, use the selected candidate explanatory requirement as the target explanatory requirement.
11. The generation method according to claim 9, wherein, there are multiple virtual teams in the target scenario, and each virtual team includes multiple virtual objects; the obtaining the target scenario data corresponding to the target explanatory task in response to reaching the explanatory timing of the target explanatory task includes: In response to reaching the explanatory timing of the target explanatory task, obtain the scenario data corresponding to the target explanatory task for each of the multiple virtual teams; Based on the scenario data corresponding to the multiple virtual teams respectively, use the scenario data corresponding to the target virtual team that meets the preset conditions among the multiple virtual teams as the target scenario data, where the target scenario data is used to explain the behavior of the target virtual team in the target scenario.
12. The generation method according to claim 11, wherein, the explanatory timing includes a timing when no virtual team battle is in progress, and the target explanatory task includes a task of explaining at least one of the position information of virtual objects in the target virtual team, the terrain information of the location where they are located, the direction information, the virtual resource holding information, and the virtual battle information.
13. The generation method according to claim 9, wherein, there are multiple virtual teams in the target scenario, and each virtual team includes multiple virtual objects; the commentary timing includes one of the timing of the end of a virtual team battle, the timing of a virtual team battle, and the timing of entering the final circle; the type of the virtual team battle includes one of the team wipe type and the non-team wipe type; the target commentary task includes the task of commenting on at least one of the highlight moments, the reasons for victory or defeat, and the virtual tactics in the virtual team battle.
14. A training device for a commentary model, wherein, the device includes: an acquisition module, configured to acquire a training data set, where the training data set includes multiple groups of first sample pairs, and each group of the first sample pairs includes first indication information and a sample commentary text. The first indication information includes a sample commentary task, scenario data corresponding to the sample commentary task, and a commentary requirement. The sample commentary task is a task of commenting on a target event in a target scenario, and the sample commentary text is used to describe the scenario data in the first indication information and meets the commentary requirement in the first indication information; an input-output module, configured to, for each group of the first sample pairs, input the first indication information in the first sample pair into a pre-trained dialogue model to obtain a first predicted commentary text of the first sample pair; an adjustment module, configured to adjust the model parameters of the pre-trained dialogue model based on the first predicted commentary text of the first sample pair and the sample commentary text in the first sample pair, so as to obtain a commentary model for the target scenario, and the commentary model is used to generate a commentary text based on the first indication information.
15. A generation device for a commentary text, wherein, the device includes: an acquisition module, configured to, during the running of a target scenario, in response to reaching the commentary timing of a target commentary task, acquire target scenario data corresponding to the target commentary task; a determination module, configured to obtain target indication information based on the target commentary task, the target scenario data, and a target commentary requirement, where the target indication information is used to indicate commenting on the target commentary task based on the target scenario data and the target commentary requirement; an input-output module, configured to input the target indication information into a commentary model, and the commentary model is trained according to the method described in any one of claims 1-8; the input-output module is further configured to output a target commentary text corresponding to the target commentary task through the commentary model.
16. A computer device, wherein, the computer device includes a processor and a memory. The memory is used to store at least one segment of program, and the at least one segment of program is loaded and executed by the processor to perform the training method of the commentary model described in any one of claims 1-8 or the generation method of the commentary text described in any one of claims 9-13.
17. A computer-readable storage medium, wherein, The computer-readable storage medium is used to store at least one program, and the at least one program is used to execute the training method of the interpretation model according to any one of claims 1-8 or the generation method of the interpretation text according to any one of claims 9-13.
18. A computer program product, characterized in that the computer program product includes at least one program, the at least one program is stored in a computer-readable storage medium, a processor of a computer device reads the at least one program from the computer-readable storage medium, and the processor executes the at least one program, so that the computer device executes the training method of the interpretation model according to any one of claims 1-8 or the generation method of the interpretation text according to any one of claims 9-13.