Explanation generation model training method and device, equipment and storage medium
By generalizing and training large language models to generate a two-person commentary framework, the problem of monotonous AI single-person commentary style is solved, the interactivity and information transmission efficiency of e-sports commentary are improved, and the cost and delay of model call is reduced.
Patent Information
- Application Number
- CN202410012436.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-03
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, AI single-player commentary has a monotonous style in e-sports competitions, lacks interactivity, and has poor audio-visual experience for the audience.
The seed instruction set is generalized through the first large language model, context text containing alternating primary and secondary speech speeches is generated, and the second large language model is trained to output secondary speech speeches to form a two-person interpretation framework.
It realizes the interactiveness and information transmission of two-person commentary in e-sports competitions, reduces the cost and delay of model call, and enriches the audience's audio-visual experience.
Smart Images

Figure CN120258165A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and particularly to a training method, device, equipment and storage medium for an explanation generation model. Background Art
[0002] In various esports competitions, it is necessary to provide real-time commentary on esports competitions.
[0003] In the related art, a method for realizing automatic single-person commentary through AI (Artificial Intelligence) is provided. However, the speech style of AI single-person commentary is relatively monotonous, lacking interactivity, and the audio-visual experience of the audience is not good. Summary of the Invention
[0004] This application provides a training method, device, equipment and storage medium for an explanation generation model, provides a generation method for deputy commentary, and combines it with the main commentary, thereby providing a two-person commentary framework.
[0005] According to one aspect of this application, a training method for an explanation generation model is provided. The method includes the following steps.
[0006] Generalize the seed instruction set through a first large language model to obtain a generation instruction set. The number of generation instructions included in the generation instruction set is greater than the number of seed instructions included in the seed instruction set. Both the generation instructions and the seed instructions include context texts with alternating main and deputy commentary.
[0007] Based on the generation instruction set, train a second large language model to obtain an explanation generation model. The scale of the second large language model is smaller than that of the first large language model. The task of the explanation generation model is to take the main commentary as input and output the deputy commentary.
[0008] According to another aspect of this application, a training device for an explanation generation model is provided. The device includes the following modules.
[0009] Generalization module, which is used to generalize the seed instruction set through a first large language model to obtain a generation instruction set. The number of generation instructions included in the generation instruction set is greater than the number of seed instructions included in the seed instruction set. Both the generation instructions and the seed instructions include context texts with alternating main and deputy commentary.
[0010] Training module, which is used to train a second large language model based on the generation instruction set to obtain an explanation generation model. The scale of the second large language model is smaller than that of the first large language model. The task of the explanation generation model is to take the main commentary as input and output the deputy commentary.
[0011] According to one aspect of the present application, a computer device is provided. The computer device includes: a processor and a memory. The memory stores a computer program, and the computer program is loaded and executed by the processor to implement the training method of the above-described commentary generation model.
[0012] According to another aspect of the present application, a computer-readable storage medium is provided. The storage medium stores a computer program, and the computer program is loaded and executed by the processor to implement the training method of the above-described commentary generation model.
[0013] According to another aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the training method of the above-described commentary generation model.
[0014] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include the following.
[0015] The solution of the present application will obtain a seed instruction set, generalize the seed instruction set using a first large language model to obtain a generation instruction set. Both the generation instructions in the generation instruction set and the seed instructions in the seed instruction set include context texts with alternating main and sub-commentary utterances. Then, based on the generation instruction set, a second large language model is trained to obtain a commentary generation model. The task of the commentary generation model is to take the main commentary utterance as input and output the sub-commentary utterance. The scale of the second large language model is smaller than that of the first large language model.
[0016] In the present application, the sub-commentary utterance output by the commentary generation model can be combined with the input main commentary utterance to form a dual-person commentary utterance. Therefore, the present application provides a dual-person commentary framework, which disassembles the game commentary task into a dual-person commentary task, is applicable to various commentary scenarios, and has good portability.
[0017] Different from the AI single-person commentary provided by the related technology, the dual-person commentary can convey more effective information while fully mobilizing the emotions of the audience. Moreover, in the above method, the training data set (generation instruction set) is generalized by a large language model with a relatively large scale, and the commentary generation model is obtained by training a large language model with a smaller scale. In the real-time commentary scenario, generating commentary utterances through a small-scale model reduces the model call cost, and ensures the model call rate in the real-time commentary scenario with low latency.
[0018] Moreover, in some scenarios, the main commentary speech input is the real-time commentary content of a human. The secondary commentary speech generated by this application can relieve the commentary pressure of a single person. The machine and the human cooperate with each other for commentary, which can enrich the audio-visual experience of the audience. In other scenarios, the main commentary speech input is generated by artificial intelligence. At this time, both commentaries of the two people are generated by AI, which completely releases human resources and reduces the labor cost of commentary. Brief Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0020] Figure 1 It is a schematic diagram of the training and use principles of the commentary generation model provided by an embodiment of the present application.
[0021] Figure 2 It is a flowchart of the training method of the commentary generation model provided by an embodiment of the present application.
[0022] Figure 3 It is a schematic diagram of a seed instruction provided by an embodiment of the present application.
[0023] Figure 4 It is a flowchart of the generation method of the generation instruction set provided by an embodiment of the present application.
[0024] Figure 5 It is a schematic diagram of the generation method of the generation instruction set provided by an embodiment of the present application.
[0025] Figure 6 It is a schematic diagram of the one-step generation method provided by an embodiment of the present application.
[0026] Figure 7 It is a schematic diagram of the two-step generation method provided by an embodiment of the present application.
[0027] Figure 8 It is a flowchart of the generation method of the seed instruction set provided by an embodiment of the present application.
[0028] Figure 9 It is a schematic diagram of the generation method of the seed instruction set provided by an embodiment of the present application.
[0029] Figure 10 It is a flowchart of the training method of the commentary generation model provided by an embodiment of the present application.
[0030] Figure 11It is a flowchart of a post - processing method for an explanation generation model provided by an embodiment of the present application.
[0031] Figure 12 It is a structural block diagram of a training device for an explanation generation model provided by an embodiment of the present application.
[0032] Figure 13 It is a structural block diagram of a computer device provided by an embodiment of the present application.
[0033] Figure 14 It is a structural block diagram of a computer device provided by another embodiment of the present application. Detailed implementation manners
[0034] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0035] First, a brief introduction to the terms involved in the embodiments of the present application is given.
[0036] Artificial Intelligence (AI): It is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision - making.
[0037] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware - level technologies and software - level technologies. Artificial intelligence basic technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre - trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre - trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine - tuning. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0038] Large language model: Usually refers to a model with a large number of parameters and deep network layers. A large model is a machine learning model with a large number of parameters and computing resources. These models require a large amount of data and computing power during the training process and have millions to billions of parameters. The design purpose of large models is to improve the model's representation ability and performance, and be able to better capture patterns and regularities in data when dealing with complex tasks.
[0039] E-sports event: It is an event organized by the event organizer for game competitions. In this application, the e-sports event can be an event organized for any game. The game can be any one of First-Person Shooting (FPS) games, Third-Personal Shooting (TPS) games, MOBA (Multiplayer Online Battle Arena) games, tactical competition games, and simulation games (SLG).
[0040] In the game, the participating players can control the virtual characters located in the virtual environment of the game session for activities. The activities of the virtual characters include but are not limited to at least one of adjusting the body posture, crawling, walking, running, cycling, flying, jumping, driving, picking up, shooting, attacking, and throwing. Schematically, the virtual character is a virtual human character, such as a simulated human character or an anime character.
[0041] Live broadcast scenario of the e-sports event: An independent signal acquisition device is installed at the e-sports event site to collect the audio and video data of the e-sports event in real time and import it into the director's terminal (director device or platform). Then, the director's terminal publishes the real-time collected audio and video of the e-sports event to the Internet for users to watch through the network. For example, it is published to the client for users to watch. Optionally, the game session data provided by the game server of the e-sports event is also imported into the director's terminal, and the director's terminal publishes the game session data to the Internet for users to watch in a way of visualizing the game session data.
[0042] Figure 1 A schematic diagram showing the training and usage principles of the commentary generation model provided by an exemplary embodiment of this application is shown. Among them, the computer system includes a training device 10 for the commentary generation model and a usage device 11 for the commentary generation model. The training device 10 and the usage device 11 are transmitted in a wired or wireless manner, and the training device 10 sends the trained commentary generation model to the usage device 11.
[0043] The training device 10 is used to execute the training process 100 of the commentary generation model, and the usage device 11 is used to execute the inference process 110 of the commentary generation model. In the training process 100, the following steps are executed:
[0044] 1. Get the seed instruction set 101. The seed instruction set 101 is a small-scale but high-quality instruction set. The seed instruction set 101 includes multiple seed instructions, and each seed instruction includes a context text in which the main and secondary commentators’ speeches alternate. The speech is the content of the commentary broadcast, and one event corresponds to at least one speech. An event is the smallest selected broadcast unit, such as "early stage-reduction in personnel", "few total team throws", "poor team equipment-protection", "story line-final circle" and other commentary events. The main commentator undertakes more broadcasting tasks in the two-person commentary and is responsible for the main topic guidance. The secondary commentator replies, analyzes and echoes the content of the main commentator's broadcast to fill the silence. Schematically, a seed instruction is as follows:
[0045]
[0046] In an illustrative manner, the process of obtaining the seed instruction set 101 is as follows: extract a single-player commentary segment from the single-player commentary of an existing complete game, input the single-player commentary segment as the main commentary segment into the third large language model, prompt the third large language model to output a secondary commentary segment, manually evaluate the secondary commentary segment output by the third large language model, and if qualified, combine the secondary commentary segment with the input main commentary segment to obtain a seed instruction. If unqualified, obtain evaluation content, such as "insufficient cohesion between commentaries", etc., merge the evaluation content with the original prompt instruction, and prompt the third large language model to output a secondary commentary segment again, until the secondary commentary segment generated by the manual evaluation is qualified, and the iteration ends.
[0047] Optionally, determine whether the current event belongs to an event in the event whitelist. If so, determine to generate a secondary commentary speech for the current event; if not, continue to determine whether the next event belongs to the event whitelist.
[0048] Second, the seed instruction set 101 is generalized to obtain the generation instruction set 102. Each generation instruction in the generation instruction set 102 includes a context text in which the main and secondary interpretation speech words alternate, and the above example of the "seed instruction" can be referred to. The number of generation instructions included in the generation instruction set 102 is greater than the number of seed instructions in the seed instruction set 101. The generalization operation is to obtain a larger corpus.
[0049] In an optional embodiment, the seed instructions in the seed instruction set 101 are imitated by the first large language model to obtain generated instructions. Optionally, data cleaning is also performed on the instructions obtained by the imitation.
[0050] III. Train the second large language model 103 through the generated instruction set 102 to obtain the commentary generation model 104. The scale of the second large language model is smaller than that of the first large language model for data generalization described above, and the scale of the second large language model is smaller than that of the third large language model for generating the seed instruction set 101. The task of the commentary generation model 104 is to take the main commentary speech as the input and output the secondary commentary speech.
[0051] It can be understood that the secondary commentary speech output by the commentary generation model 104 can be combined with the input main commentary speech to form a dual - commentary speech. Therefore, the present application provides a dual - commentary framework. Different from the AI single - commentary provided by the related technology, the dual - commentary can convey more effective information while fully mobilizing the emotions of the audience. Moreover, the above - mentioned commentary generation model is obtained by training a large language model with a smaller scale. The small - scale model reduces the model calling cost and ensures the model calling rate in the real - time commentary scenario with a lower latency.
[0052] In the inference process 110, some post - processing operations will also be performed.
[0053] In the inference process 110, obtain the target main commentary speech 111, input the target main commentary speech 111 into the commentary generation model 104, and the commentary generation model 104 infers to obtain multiple candidate secondary commentary speeches 112; score the multiple candidate secondary commentary speeches 112, and select the candidate secondary commentary speech with the highest score as the target secondary commentary speech 113. The target secondary commentary speech 113 is the final inference result. Optionally, the scoring dimensions include at least one of the word type in the speech, the speech length, and the repetition degree with the target main commentary speech 111.
[0054] It can be understood that the trained commentary generation model has learned how to generate secondary commentary speeches, but the quality of the secondary commentary speeches is still not perfect. Therefore, the above content designs a scoring algorithm to select the candidate secondary commentary speech with the highest score as the final inference result.
[0055] In the above text, the training device 10 of the commentary generation model and the using device 11 of the commentary generation model can be computer devices with machine - learning capabilities. The computer devices can be terminals or servers.
[0056] The above-mentioned training device 10 and usage device 11 may be the same computer device. Alternatively, the training device 10 and usage device 11 may also be different computer devices. Moreover, when the training device 10 and usage device 11 are different devices, the training device 10 and usage device 11 may be of the same type of device. For example, both the training device 10 and usage device 11 may be servers. Or, the training device 10 and usage device 11 may also be of different types of devices. The above-mentioned server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The above-mentioned terminal may be a mobile phone, computer, intelligent voice interaction device, intelligent household appliance, vehicle-mounted terminal, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions in this regard.
[0057] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the main interpretation speech and large language models involved in this application are obtained under full authorization.
[0058] Moreover, regarding the relevant information, the relevant information processors will follow the principles of legality, legitimacy, and necessity, clarify the purpose, method, and scope of relevant information processing, obtain the consent of the relevant information subjects, and take necessary technical and organizational measures to ensure the security of relevant information.
[0059] Figure 2 The flowchart of the training method of the interpretation generation model provided by an exemplary embodiment of this application is shown. Taking the method as being executed by Figure 1 the training device 10 shown for example, the method includes:
[0060] Step 220, generalize the seed instruction set through a first large language model to obtain a generated instruction set. The number of generated instructions included in the generated instruction set is greater than the number of seed instructions included in the seed instruction set. The generated instructions include context texts with alternating main and secondary interpretation speeches;
[0061] The seed instruction set is a small-scale but high-quality data set. The seed instruction includes the context text of the alternating main and secondary commentary speech. The speech is the content of the commentary. Figure 3 , Figure 3 A seed instruction is shown. The seed instruction includes alternating broadcast contents of the main commentary and the secondary commentary.
[0062] Optionally, the seed command includes the main and secondary commentary speech of multiple consecutive events. An event can correspond to the main commentary speech and the secondary commentary speech, or it can only correspond to the main commentary speech. Events are the smallest selected broadcast units, such as "early stage - reduction in personnel", "few total throws of the team", "team equipment - poor armor", "story line - finals", etc.
[0063] The main commentator is responsible for more reporting tasks and guiding the main topics in a two-person commentary. The assistant commentator responds, analyzes and echoes the main commentator's content and fills in the silence.
[0064] It is worth noting that the seed instruction is the context text. In the following text, the context text will be used as training data to train the explanation generation model. The secondary explanation speech generated by the explanation generation model will have contextual connection, coherence and logic.
[0065] In one embodiment, the two-person interpretation speech of a real person is used as a seed instruction. In one embodiment, the single-person interpretation speech of a real person is obtained, and the single-person interpretation speech is used as the main interpretation speech. A large language model is used to generate a secondary interpretation speech, and the main interpretation speech and the secondary interpretation speech are combined to obtain a seed instruction.
[0066] In one embodiment, a real person's single-person commentary speech is obtained and used as the main commentary speech. Based on the keywords in the main commentary speech, relevant professional knowledge is searched in the game knowledge base. Based on the retrieved professional knowledge and the main commentary speech, a large language model is used to generate a secondary commentary speech. The main commentary speech and the secondary commentary speech are combined to obtain a seed instruction.
[0067] The embodiment of the present application provides a two-person commentary framework based on a large language model, which can be applied to various commentary scenarios, such as competitive games of FPS games, competitive games of MOBA games, football games, basketball games, etc. For the two-person commentary framework provided by the present application, any commentary scenario can be trained to obtain a commentary generation model through the method of the present application.
[0068] The generalization operation is used to obtain a larger corpus. In one embodiment, through a first large language model, the seed instructions in the seed instruction set are imitated to obtain generated instructions, and then a generated instruction set is obtained. The generated instructions are also the context texts with alternating main and secondary commentary utterances.
[0069] Optionally, the generated instructions include the main and secondary commentary utterances of a continuous series of events. An event can correspond to both a main commentary utterance and a secondary commentary utterance, or only a main commentary utterance. An event is the smallest broadcast unit selected, such as events like "Early stage - Staff reduction", "Team's total throwables are few", "Team equipment - Poor armor", "Storyline - Final circle", and so on.
[0070] The first large language model can be an extra-large model such as GPT-4.
[0071] Step 240: Based on the generated instruction set, train a second large language model to obtain a commentary generation model. The scale of the second large language model is smaller than that of the first large language model. The task of the commentary generation model is to take the main commentary utterance as input and output the secondary commentary utterance.
[0072] The generated instruction set includes multiple generated instructions, and each generated instruction includes the context text with alternating main and secondary commentary utterances. In this application, each generated instruction is used as training data to train the second large language model to obtain a commentary generation model. The scale of the second large language model is smaller than that of the first large language model.
[0073] Optionally, the second large language model is an open-source small and medium-sized language model with 6B - 13B parameters.
[0074] In summary, in the above embodiments, the seed instruction set is obtained, the seed instruction set is generalized using the first large language model to obtain a generated instruction set. Both the generated instructions in the generated instruction set and the seed instructions in the seed instruction set include the context text with alternating main and secondary commentary utterances; then, based on the generated instruction set, the second large language model is trained to obtain a commentary generation model. The task of the commentary generation model is to take the main commentary utterance as input and output the secondary commentary utterance, and the scale of the second large language model is smaller than that of the first large language model.
[0075] In the above embodiments, the secondary commentary utterance output by the commentary generation model can be combined with the input main commentary utterance to form a two-person commentary utterance. Therefore, this application provides a two-person commentary framework, which decomposes the game commentary task into a two-person commentary task, is applicable to various commentary scenarios, and has good transferability.
[0076] Different from the AI single-person commentary provided by related technologies, the dual-person commentary can convey more effective information while fully mobilizing the emotions of the audience. Moreover, in the above method, the training dataset (generation instruction set) is generalized by a large-scale large language model, and the commentary generation model is obtained by training a small-scale large language model. In the real-time commentary scenario, generating commentary utterances through a small-scale model reduces the model invocation cost, and moreover, ensures the model invocation rate in the real-time commentary scenario with low latency.
[0077] Moreover, in some scenarios, the input main commentary utterance is the real-time commentary content of a human, and the generated secondary commentary utterance of the present application can relieve the commentary pressure of a single person. The cooperation between the machine and the human for commentary can enrich the audio-visual experience of the audience. In other scenarios, the input main commentary utterance is generated by artificial intelligence. At this time, both the dual-person commentaries are generated by AI, which fully releases human resources and reduces the labor cost of commentary.
[0078] Based on Figure 2 In the optional embodiment shown, Figure 2 Step 240 in Figure 4 includes the steps 410 to 450 shown in Figure 1 Taking the method executed by the computer device 10 shown in
[0079] as an example for illustration.
[0080] In the present application, the seed instruction set will be generalized through multiple instruction generation cycles to obtain the generation instruction set. In each instruction generation cycle, one generation instruction will be generated.
[0081] With reference to Figure 5 Figure 5 shows that in one instruction generation cycle, three seed instructions 503 are extracted from the seed instruction set 501, and two generation instructions 504 are extracted from the first generation instruction set 502. Optionally, when the number of instructions in the first generation instruction set 502 is less than two, all the instructions in the first generation instruction set 502 are extracted. When there is no generation instruction in the first generation instruction set 502, no instruction is extracted.
[0082] Step 420, combine at least one seed instruction and at least one generation instruction to obtain a reference example;
[0083] With reference to Figure 5 Figure 5 Take the three extracted seed instructions 503 and the two generated instructions 504 together as a reference example 505 and input it into the first large language model 506. Optionally, when the number of instructions in the first generated instruction set 502 is less than two, combine all the extracted instructions with at least one seed instruction. When there is no generated instruction in the first generated instruction set 502, directly use at least one seed instruction as the reference example.
[0084] Step 430, prompt the first large language model to imitate the reference example to obtain new generated instructions; add the new generated instructions to the first generated instruction set;
[0085] Combined reference Figure 5 , Figure 5 Figure 10 shows inputting the reference example 505 into the first large language model 506, and prompting the first large language model 506 to imitate the reference example 505. The first large language model 506 outputs a new generated instruction 507.
[0086] In one embodiment, combined reference Figure 5 , the new generated instruction 507 also needs to perform a data cleaning step to determine whether the new generated instruction 507 is of high quality. If so, add the new generated instruction 507 to the first generated instruction 502, if not, discard the new generated instruction 507. Schematically, determine whether the new generated instruction 507 is of high quality through the following steps.
[0087] ① Determine that the new generated instruction does not contain preset failure words. Schematically, if the new generated instruction contains words such as "sorry" or "apology" that clearly indicate a request failure, then the new generated instruction is irrelevant noise data and should be discarded.
[0088] ② Determine that the number of words in the new generated instruction is greater than the first preset value. Schematically, if the number of words in the new generated instruction is less than 50 words, it is considered that the new generated instruction has insufficient information and is low-quality data, and it will be discarded.
[0089] ③ Determine that the number of words contained in the new generated instruction is greater than the second preset value. Schematically, perform word segmentation statistics on the new generated instruction. If the number of words contained in the new generated instruction is too small, discard it.
[0090] ④ Determine that the new generated instruction does not contain words with an appearance frequency greater than the third preset value. Schematically, perform word segmentation statistics on the new generated instruction. If the appearance frequency of a certain word is too high, it means that the new generated instruction is just a repetition of some words and should be discarded.
[0091] ⑤ Determine that the similarity between the new generated instruction and any one of the at least one seed instruction is less than a fourth preset value, and determine that the similarity between the new generated instruction and any one of the at least one generated instruction is less than a fourth preset value.
[0092] Calculate the similarity between the new generated instruction and the input reference example. If it is too high, discard it.
[0093] Step 440, execute multiple instruction generation cycles so that the number of generated instructions included in the first generated instruction set is greater than the number of seed instructions included in the seed instruction set;
[0094] Execute multiple instruction generation cycles, and the scale of the first generated instruction set will be greater than the scale of the seed instruction set.
[0095] Step 450, determine the first generated instruction set as the generated instruction set.
[0096] Determine the first generated instruction set as the generated instruction set after generalization is completed.
[0097] In summary, the above embodiments provide a method for generalizing to obtain a generated instruction set. By extracting seed instructions from the seed instruction set, extracting generated instructions from the first generated instruction set, and combining the seed instructions and the generated instructions as a common reference example, the second large language model is prompted to imitate and generate new generated instructions.
[0098] Extracting seed instructions from the seed instruction set can prevent the newly generated instructions from being too divergent and deviating from the normal data distribution; extracting generated instructions from the generated instruction set can generate different generated instructions and cover as much explanatory content as possible.
[0099] Based on Figure 4 In the optional embodiment shown, step 430 prompts the first large language model to imitate the reference example to generate new generated instructions, including two methods: one-step generation and two-step generation.
[0100] One-step generation method: directly prompt the first large language model to imitate the reference example to generate a two-person commentary segment to obtain a new generated instruction.
[0101] Combined with reference Figure 6 , Figure 6 shows that according to at least one seed instruction in the seed instruction set 601 and at least one generated instruction in the first generated instruction set 602, a reference example 603 is combined. The reference example 603 and the prompt 605 are input into the first large language model 604. The prompt 605 is used to prompt the first large language model 604 to imitate the reference example 603 and output a two-person commentary segment. The first large language model 604 outputs a new generated instruction 606, and the new generated instruction 606 is added to the first generated instruction set 602.
[0102] In one embodiment, both the seed instructions and the generation instructions in the reference example include the main and secondary commentary utterance pairs of a continuous plurality of events. An event is the smallest unit of a broadcast. Combining the reference Figure 6 , the prompt 605 is used to prompt the first large language model 604 to imitate the reference example 603 and output in the format of the main commentary utterance of the previous event, the main commentary utterance of the current event, and the secondary commentary utterance of the current event according to the type of the current event, so as to generate a two-person commentary segment and obtain a new generation instruction.
[0103] It can be understood that prompting the first large language model 604 to output in the format of the main commentary utterance of the previous event, the main commentary utterance of the current event, and the secondary commentary utterance of the current event is beneficial to guiding the model to generate the main and secondary commentary utterances of the current event according to the main commentary utterance of the context, and the generated main and secondary commentary utterances are more coherent and logical and more in line with the human expression method.
[0104] It can be understood that prompting the first large language model 604 to output the type of the current event is beneficial to guiding the model to generate the main and secondary commentary utterances closely related to the event, and then the generated main and secondary commentary utterances can be calibrated to avoid the generated main and secondary commentary utterances deviating from the event.
[0105] In summary, in the above one-step generation method, the first large language model is directly prompted to generate a two-person commentary segment, avoiding error accumulation, and the generated secondary commentary utterance will not deviate from the central idea.
[0106] Two-step generation method: Prompt the first large language model to imitate the reference example to generate a main commentary segment to obtain a first main commentary segment, and the first main commentary segment includes a plurality of first main commentary utterances; based on the first main commentary segment, prompt the first large language model to generate a secondary commentary segment to obtain a first secondary commentary segment, and the first secondary commentary segment includes a plurality of first secondary commentary utterances; alternately combine the plurality of first main commentary utterances and the plurality of first secondary commentary utterances to obtain a new generation instruction.
[0107] Combining the reference Figure 7 , Figure 7 shows a reference example 703 obtained by combining at least one seed instruction in the seed instruction set 701 and at least one generation instruction in the first generation instruction set 702. Inputting the reference example 703 and the prompt 705 into the first large language model 704, a first main commentary segment 706 is obtained. The prompt 705 is used to prompt the first large language model 704 to imitate the reference example 703 and output the main commentary segment.
[0108] The first main commentary segment 706 and the prompt 707 are input into the first large language model 704 to obtain a first secondary commentary segment 708. The prompt 707 is used to prompt the first large language model 704 to output a secondary commentary segment based on the first main commentary segment 706. After the first main commentary segment 706 and the first secondary commentary segment 708 are combined, they are added to the first generation instruction set 702 as a new generation instruction.
[0109] In one embodiment, the seed instruction and the generation instruction in the reference example both include a main and a secondary commentary speech pair of a plurality of consecutive events. An event is the smallest unit of a broadcast. Figure 7 The prompt 705 is used to prompt the first large language model 604 to imitate the reference example 603, output according to the type of the current event, the format of the main commentary speech of the previous event, and the main commentary speech of the current event, and generate a main commentary segment.
[0110] At this time, the first main commentary segment includes the first main commentary speech of multiple consecutive events. Schematically, the first main commentary segment includes {Event 1: main commentary speech A; Event 2: main commentary speech B; Event 3: main commentary speech C}. The first secondary commentary segment includes the first secondary commentary speech of multiple consecutive events. Schematically, the first secondary commentary segment includes {Event 1: secondary commentary speech a; Event 2: main commentary speech b; Event 3: main commentary speech c}. Multiple first main commentary speeches and multiple first secondary commentary speeches are combined event by event to obtain new generation instructions. The new generation instructions are shown in the following table:
[0111]
[0112] It can be understood that prompting the first large language model 604 to output in the format of the main interpretation speech of the previous event and the main interpretation speech of the current event is conducive to guiding the model to generate a coherent main interpretation speech. The generated main interpretation speech is more coherent and logical, and more in line with human expression.
[0113] It can be understood that prompting the first large language model 604 to output the current event type is conducive to guiding the model to generate a main interpretation utterance that is closely related to the event, and then calibrating the generated main interpretation utterance to avoid the generated main interpretation utterance deviating from the event.
[0114] In summary, the above two-step generation method first prompts the first large language model to generate a main commentary segment, and then prompts the first large language model to generate a secondary commentary segment, so that the generated secondary commentary speech is more logical, and the secondary commentary speech is a supplement and echo of the main commentary speech.
[0115] based on Figure 2 In the optional embodiment shown, a step of obtaining a seed instruction set is also performed. Specifically, it includes Figure 8Steps 810 to 840 shown. For example, the method is illustrated by being executed by the computer device 10 shown in Figure 1 shown.
[0116] Step 810, obtain a second main commentary segment, the second main commentary segment including a plurality of second main commentary utterances;
[0117] The second main commentary segment, as seed data, has relatively high quality. The second main commentary segment can be a single-person commentary segment of a real person, or a main commentary segment generated by other artificial intelligence models. For example, by other artificial intelligence models to identify game scenes, on-site audio, in-depth game data (number of defeats in the game, number of times being defeated, equipment status), etc., to generate a main commentary segment.
[0118] In one embodiment, the second main commentary segment includes main commentary utterances of a plurality of consecutive events. A generation process of a second main commentary segment is as follows: extract a single-person commentary segment from the single-person commentary text of a complete game, the single-person commentary segment including single-person commentary utterances of a plurality of consecutive events; use the single-person commentary segment as the second main commentary segment.
[0119] Step 820, based on the second main commentary segment, through a prompting instruction, prompt a third large language model to generate a secondary commentary segment, obtaining a second secondary commentary segment, the second secondary commentary segment including a plurality of second secondary commentary utterances; wherein, the scale of the third large language model is larger than the scale of the second large language model;
[0120] Optionally, the third large language model is an extra-large language model such as GPT-4.
[0121] In one embodiment, the quality of the secondary commentary segment generated according to a single prompting instruction is usually not very high. Therefore, the present application adopts an iterative process of evaluation and generation until the quality of the generated secondary commentary segment is relatively high and the iteration stops.
[0122] Specifically, in the i-th iteration process, based on the second main commentary segment, through the i-th prompting instruction, prompt the third large language model to generate a secondary commentary segment, obtaining the i-th secondary commentary segment; when i is equal to 1, the i-th prompting instruction is the initial prompting instruction.
[0123] In the case where the quality of the i-th secondary commentary segment is qualified, determine the i-th secondary commentary segment as the second secondary commentary segment; in the case where the quality of the i-th secondary commentary segment is unqualified, combine the evaluation content of the i-th secondary commentary segment and the i-th prompting instruction to obtain the (i + 1)-th prompting instruction, the (i + 1)-th prompting instruction being used to execute the (i + 1)-th iteration process.
[0124] Optionally, the determination of qualified quality is performed manually. Optionally, the evaluation content is generated manually.
[0125] Schematically, the evaluation content can be "insufficient connection between commentary rooms", "the commentary is too general and does not conform to conventional commentary terms", "forced causal explanation", etc.
[0126] In one embodiment, a single-person commentary segment is extracted from the single-person commentary text of a complete game. The single-person commentary segment includes the single-person commentary words of multiple consecutive events; the single-person commentary segment is used as the second main commentary segment. Generate secondary commentary words event by event to obtain a second secondary commentary segment; when the current event in the second main commentary segment does not belong to the events in the event white list, continue to search for the next event in the second main commentary segment; when the current event belongs to the events in the event white list, based on the second main commentary segment, through a prompt instruction, prompt the third large language model to generate the secondary commentary words of the current event.
[0127] It can be understood that not all events in the game are suitable for dual commentary. For example, for a team battle event when the battle is intense, if the secondary commentator interrupts the main commentator, it will reduce the excitement of the commentary. The event white list includes pre-defined events that support dual commentary.
[0128] Combined reference Figure 9 , Figure 9 shows the generation process of the secondary commentary segment in a seed instruction. A single-person commentary segment 902 is extracted from the single-person commentary text 901 of a complete game. The single-person commentary segment 902 is judged event by event. If the current event is suitable for dual commentary, based on the single-person commentary segment 902, the third large language model 903 is used to generate the secondary commentary words of the current event. If the current event is not suitable for dual commentary, no secondary commentary words are generated for the current event, and the next event is continued to be searched to judge whether the next event is suitable for dual commentary. Figure 9 In, by the white list events suitable for commentary input, it is judged whether the current event is suitable for dual commentary. Figure 9 Finally, the second secondary commentary segment 904 is generated.
[0129] Step 830, alternately combine multiple second main commentary words and multiple second secondary commentary words to obtain a seed instruction;
[0130] In one embodiment, the second main commentary segment includes the main commentary words of multiple consecutive events, and the second secondary commentary segment includes the secondary commentary words of multiple consecutive events. Alternately combine multiple second main commentary words and multiple second secondary commentary words event by event to obtain a seed instruction. A seed instruction includes the main and secondary commentary words of multiple consecutive events.
[0131] Step 840, repeat the above steps to obtain a seed instruction set;
[0132] Steps 810 to 830 provide a method for generating seed instructions. By repeatedly executing steps 810 - 830, a seed instruction set can be obtained.
[0133] In summary, the above embodiments provide a way to generate a seed instruction set. By using a large - scale large - language model to generate seed instructions, the quality of the generated seed instructions is relatively high, which is conducive to improving the quality of the generated commentary model.
[0134] Moreover, the above embodiments provide a cross - iterative process of evaluation and generation. When the quality of the generated secondary commentary utterances is relatively high, the iteration is stopped. The cross - iterative process is conducive to improving the quality of the generated seed instructions.
[0135] Furthermore, when generating seed instructions, events not in the event whitelist are deleted. The event whitelist includes preset events that support dual - person commentary, ensuring that the generated seed instructions are reasonable dual - person commentary segments, which is conducive to improving the quality of the secondary commentary utterances generated by the commentary generation model.
[0136] Based on Figure 2 In the optional embodiment shown, step 260 "Based on the generated instruction set, train a second large - language model to obtain a commentary generation model" includes Figure 10 Steps 1010 to 1030 shown.
[0137] Step 1010, based on the main commentary utterance sequence included in the generated instruction and the secondary commentary utterance sequences of the first t - 1 events, predict the secondary commentary utterance of the t - th event through a second large - language model; output the probability that the prediction result is the secondary commentary utterance of the t - th event included in the generated instruction;
[0138] In the previous steps, a generated instruction set is generalized. The generated instruction set includes multiple generated instructions, and each generated instruction includes main - secondary commentary utterance pairs of multiple consecutive events.
[0139] In this application, based on the real main commentary utterance sequence (the main commentary utterance sequence included in the generated instruction) and the real secondary commentary utterance sequences of the first t - 1 events (the secondary commentary utterance sequences of the first t - 1 events included in the generated instruction), the secondary commentary utterance of the t - th event is predicted.
[0140] Step 1020, construct a loss function by accumulating the multiple probabilities corresponding to the multiple events included in the generated instruction;
[0141] In this application, a loss function will be constructed according to the probability that the predicted secondary commentary utterance is the real secondary commentary utterance (the secondary commentary utterance of the t - th event included in the generated instruction). Schematically, the main commentary utterance sequence in the generated instruction is denoted as The secondary explanation utterance sequence in the generation instruction is represented as y = where M represents the number of events, and x i represents the primary explanation utterance of event i, and y i represents the secondary explanation utterance of event i. The loss function is represented as:
[0142]
[0143] where x represents the true primary explanation utterance sequence, and y <t represents the secondary explanation utterance sequence of the first t - 1 true events, and y t represents the secondary explanation utterance of the true t-th event, and log G(y t |x,y <t ) represents the probability that the prediction result of the model is the secondary explanation utterance of the true t-th event.
[0144] By accumulating the probabilities of M events and taking the negation of the accumulated result, the loss function is constructed.
[0145] In one embodiment, the event types of multiple events corresponding to the primary explanation utterance sequence are also input. At this time, the primary explanation utterance sequence and the event types of each event in the sequence are combined to obtain the input
[0146] Step 1030, based on the generation instruction set, train the second large language model through the loss function to obtain the explanation generation model.
[0147] According to multiple generation instructions in the generation instruction set, each generation instruction includes the primary and secondary explanation utterances of consecutive multiple events. Through the loss function constructed above, train the second large language model to obtain the explanation generation model. The second large language model can be a language model with a relatively small scale, such as 13B, etc.
[0148] In summary, the above embodiments provide a training method for the explanation generation model, which predicts the secondary explanation utterance of the t-th event through the true primary explanation utterance sequence and the secondary explanation utterance sequence of the first t - 1 true events, and the obtained prediction result will maintain the coherence and logic of the context.
[0149] Based on Figure 2 In the optional embodiment shown, after step 260, it further includes Figure 10 the steps 1110, 1120, and 1130 shown. Taking the example of this method being executed by Figure 1 the device 110 shown, this method includes:
[0150] Step 1110, input the target primary explanation utterance into the explanation generation model, and infer multiple candidate secondary explanation utterances;
[0151] The target main commentary speech, that is, the main commentary speech input during model inference. The target main commentary speech can be the commentary speech of a real person or the commentary speech generated by artificial intelligence.
[0152] Step 1120: Score multiple candidate secondary commentary speeches based on scoring dimensions, where the scoring dimensions include at least one of the word type in the speech, the length of the speech, and the degree of repetition with the target main commentary speech.
[0153] The above-mentioned trained commentary generation model has learned how to generate secondary commentary speeches, but the quality of the secondary commentary speeches is still not perfect. Therefore, the embodiments of the present application also design a scoring algorithm. It is hoped that good secondary commentary speeches meet the following characteristics:
[0154] (1) It does not contain meaningless words, words that do not conform to the commentary vocabulary, or words with obvious tendencies.
[0155] For example, "wait and see", "our side", etc. Therefore, the present application provides a method for scoring candidate secondary commentary speeches based on the word type in the speech.
[0156] (2) The length of the secondary commentary speech is not too long.
[0157] If the speech is too long, it is easy to have a situation where the broadcast content has passed for a long time. Therefore, the present application provides a method for scoring candidate secondary commentary speeches based on the length of the speech.
[0158] (3) The secondary commentary speech cannot be too repetitive with the input target main commentary speech.
[0159] Therefore, the present application provides a method for scoring candidate secondary commentary speeches based on the degree of repetition with the target main commentary speech.
[0160] Step 1130: Determine the candidate secondary commentary speech with the highest score as the inference result.
[0161] In summary, the above embodiments give the inference process of applying the scoring algorithm, and selecting the candidate secondary commentary speech with the highest score as the final inference result is beneficial to ensuring the quality of the finally output inference result.
[0162] Based on Figure 11 in Step 1120, for characteristic (1), the following scoring method is designed.
[0163] For any candidate secondary solution utterance among multiple candidate secondary solution utterances, determine the number of words in the multiple words of the candidate secondary solution utterance that intersect with the blacklist word list to obtain a first value; use the number of words included in the blacklist word list as a second value; based on the ratio of the first value to the second value, obtain the score of the candidate secondary solution utterance.
[0164] Schematically, it is expressed by the formula:
[0165]
[0166] Among them, the blacklist word list P represents the number of words in the blacklist word list, and multiple candidate secondary solution utterances are denoted as W represents the number of multiple candidate secondary solution utterances, and S v represents the score of the candidate secondary solution utterance.
[0167] Based on Figure 11 in step 1120, for feature (2), the following scoring method is designed.
[0168] Based on the lengths of multiple candidate secondary solution utterances, score the multiple candidate secondary solution utterances, so that the utterances with lengths less than the preset utterance length among the multiple candidate secondary solution utterances obtain the same score, and, so that the longer the utterances among the multiple candidate secondary solution utterances, the lower the score.
[0169] Specifically, for any candidate secondary solution utterance among the multiple candidate secondary solution utterances, calculate the difference between the length of the candidate secondary solution utterance and the preset utterance length; input the difference into the sign function to obtain a third value, and the sign function satisfies that when the input is greater than zero, the output is a fixed positive number, and when the input is less than or equal to zero, the output is zero; multiply the difference by the third value and then take the negative to obtain a fourth value; based on the calculation result with e as the base and the fourth value as the exponent, obtain the score of the candidate secondary solution utterance.
[0170] Schematically, it is expressed by the formula:
[0171]
[0172] Among them, sign() is the sign function, and when the input is greater than 0, the output is 1, and when the input is less than or equal to 0, the output is 0. L i is the length of the i-th candidate secondary solution utterance, L is the preset utterance length, that is, the ideal utterance length, and S l is the score of the candidate secondary solution utterance calculated at this time. The above formula satisfies that only when the candidate secondary solution utterance exceeds the ideal utterance length, the score changes, and the more the length exceeds, the lower the score value.
[0173] Based on Figure 11In step 1120, for feature (3), the following scoring method is designed.
[0174] For any one of the multiple candidate secondary solution utterances, calculate the similarity between the candidate secondary solution utterance and the target primary solution utterance; based on the similarity, obtain the score of the candidate secondary solution utterance.
[0175] Schematically, it is expressed by the formula:
[0176] S s =-ROUGE(X,Y i )
[0177] where the ROUGE (Recall-Oriented Understudy for Gisting Evaluation) function is used to measure the similarity between the target primary solution utterance X and the i-th candidate secondary solution utterance Y i and S s is the score of the candidate secondary solution utterance calculated at this time.
[0178] Based on the above features (1), (2), and (3), at least two scores can be weighted and summed to obtain the final output score of the candidate secondary solution utterance. Taking the weighted sum of three scores as an example, the formula is expressed as:
[0179] S i =α*S v +β*S l +γ*S s ;
[0180] where S v represents the sub-score calculated based on the utterance word type, S l represents the sub-score calculated based on the utterance length, S s represents the sub-score calculated based on the similarity with the target primary solution utterance, and α, β, and γ are the respective weights of the three.
[0181] Finally, the target secondary solution utterance determined from multiple candidate secondary solution utterances is:
[0182]
[0183] That is, the target secondary solution utterance Y * is the Y i for which S i obtains the maximum value.
[0184] This application provides a solution for automatically generating dual-commentary for games, which fine-tunes an open-source large model by extracting generalized instructions. The open-source large model can be used to generate interactions for the secondary commentator.
[0185] The specific framework is as follows:
[0186] 1. Seed data generation: By extracting the context of the single-player commentary as the main commentator, the GPT-4 model and instructions of relevant commands are used to generate the secondary commentator's lines. Through the feedback of manual evaluation, the instructions are continuously optimized to construct a smaller but higher-quality instruction set as the seed data set.
[0187] 2. Instruction generalization: Generalize the seed data set. The GPT-4 model can imitate and generate more data based on the input examples. Extract some examples from the seed data set and the generated data respectively, and input them into the large model to generate more main and secondary commentator data of different event types. And combine some evaluation criteria to clean the generated data to construct a larger corpus data set.
[0188] 3. Fine-tune the large model: Use the larger corpus data set as the training data set, with the main commentator's lines as the input and the secondary commentator's lines as the output, and fine-tune the parameters of the open-source generative language model to enable it to have the ability to generate commentary.
[0189] 4. Result post-processing: Let the fine-tuned large model generate multiple outputs for each input, and use the designed scoring function to evaluate each output, and the secondary commentator that meets the requirements can be selected as the final output.
[0190] In an optional embodiment, for seed data generation: The construction of the seed instruction set first requires extracting the context from the single-player commentary of the existing complete game as the main commentator input. The specific extraction method is to judge according to the event corresponding to each commentary whether it is in the defined event whitelist. If so, it is determined that the secondary commentator's lines can be generated for this event. Because not all events are suitable for inserting dual-commentary. For example, in a FPS game, if it is an in-group event with intense combat, the secondary commentator interrupting the main commentator's words will reduce the excitement of the commentary.
[0191] Then, using the GPT-4 model and corresponding human prompting instructions, relatively standardized and reasonable alternative commentary scripts are generated. To enable the GPT-4 model to generate commentaries that meet actual requirements, a method of cross-iteration of evaluation and generation is adopted. Initially, a human prompting instruction is used to generate a corresponding set of alternative commentaries in a small evaluation set. Then, these results are manually evaluated to identify and summarize deficiencies, such as "insufficient connection between commentaries", "the commentary is too general and does not conform to conventional commentary terms", "forced causal explanations", etc. These evaluation contents are converted into new prompting instructions and merged with the original human prompting instructions until the manual evaluation is relatively reasonable, and the iteration ends.
[0192] Using the finally obtained instructions and the extracted input, an initial set of seed instructions corresponding to the main commentary script and the alternative commentary script can be generated. These seed instructions in this part are obtained from real games and have been manually evaluated, belonging to a high-quality supervised data set.
[0193] In an optional embodiment, for instruction generalization: The seed instruction set generated above needs to be collected from real games, and the sample size is small, while training a generative model requires a relatively large amount of data as supervision. Therefore, the instruction set needs to be generalized, specifically, to generate more similar data. This part of the work can also be completed using large language models. Define the Seed dataset as the seed instruction set and the Generated dataset as the generated instruction set. Each time a new instruction is generated, samples are taken from the Seed dataset and the Generated dataset respectively as reference examples for the language model. Selecting data from the Seed dataset can prevent the generated instructions from being too divergent and deviating from the normal data distribution, while selecting examples from the data generated by the Generated dataset can generate different examples and cover as much commentary content as possible.
[0194] This application designs two methods for generalizing instructions, namely one-step and two-step methods. The one-step method is to directly imitate the reference examples to generate a pair of commentary scripts for the main commentary script and the alternative commentary script. This method is relatively direct and reduces the accumulation of errors. The two-step method is to first construct the main commentary script as the input, and then, according to the human prompting instructions, let the language model generate the alternative commentator's script, which can make the alternative commentator's script more in line with the required logic.
[0195] In addition, for the generated instructions, this application also designs a judgment criterion to determine whether the generated instructions are high-quality data and avoid introducing too much data noise. The evaluation criteria are as follows:
[0196] (1) When the words generated by the large language model contain obvious words indicating request failure such as "sorry" and "apology", the generated examples are irrelevant noise data.
[0197] (2) When the total number of words is less than 50, the information is insufficient and it is regarded as low-quality data.
[0198] (3) When performing word segmentation and statistics on the instructions, if the number of words appearing is too small or the frequency of a certain word is too high, it indicates that the generated content is just a repetition of some words and should be discarded.
[0199] (4) Calculate the similarity with the input examples, and discard those with too high similarity.
[0200] In an optional embodiment, for fine-tuning the large model: the generation of the previous seed instructions and data generalization are both completed by calling super-large models such as GPT-4. These models are pre-trained in the general field, so they have a large number of parameters and high calling costs. Now, many open-source small and medium-sized language models with 6B - 13B parameters can also achieve good results after a certain amount of data instruction fine-tuning. Fine-tuning refers to adaptively training the data of the model that has been pre-trained and whose parameters are not randomly initialized on specific tasks. Moreover, since the number of parameters of small and medium-sized language models is relatively much smaller, the inference speed of the model is faster and the calling cost is smaller. The open-source models include Llama, Baichuan, Glm, etc.
[0201] Define the main solution speaking sequence, and the merged input of event type and generation instruction, etc. is Corresponding to the sub-solution speaking sequence to be generated Denote the fine-tuned model as G, and the objective function of fine-tuning is:
[0202]
[0203] Among them, t represents the t-th event, M represents a total of M events, x represents the true main solution speaking sequence, y <t represents the sub-solution speaking sequence of the true first t - 1 events, y t represents the sub-solution speaking of the true t-th event, logG(y t |x,y <t ) represents the probability that the prediction result of the model is the sub-solution speaking of the true t-th event.
[0204] In an optional embodiment, for post-processing of results: After fine-tuning, the model can learn how to generate alternative explanatory utterances. However, it is very difficult to construct the data perfectly. Therefore, there are still certain differences between the actually generated explanatory utterances and the required explanatory utterances. This application designs a scoring algorithm that uses the fine-tuned model to generate multiple possible results, scores and ranks them, and selects the best result for output.
[0205] A good alternative explanatory utterance satisfies the following characteristics:
[0206] a. Do not contain some meaningless words, words that do not conform to the explanatory vocabulary, or words with obvious tendencies, such as "wait and see", "our side", etc.
[0207] b. The length of the alternative explanatory utterance should not be too long, otherwise it is easy to have a situation where the broadcast content has passed for a long time.
[0208] c. The generated alternative explanatory utterance should not be too repetitive with the input main explanatory utterance.
[0209] Therefore, this application designs a scoring function. Let the model be G, and an input X corresponds to the output of W alternative explanatory utterances, denoted as L i is the length of Y i In addition, a blacklist word list that is not expected to appear in the explanatory utterance is manually constructed The ideal length of the explanatory utterance is denoted as L. In order to minimize the number of words in the blacklist in the explanatory utterance, the validity score S v is defined to measure, and the calculation formula is as follows:
[0210]
[0211] The more words in B appear in the alternative explanatory utterance, the smaller the score.
[0212] Then there is the measurement of the explanatory length, which is evaluated and calculated using the score S l :
[0213]
[0214] where sign is the sign function, which outputs 1 when the input is greater than 0 and 0 when the input is less than or equal to 0. In this way, the score only changes when the length exceeds the ideal utterance length, and the more it exceeds, the lower the score. As for the repetition degree between the main and alternative explanations, it is measured using the ROUGE function in natural language understanding, and the score is calculated:
[0215] S s =-ROUGE(X,Y i );
[0216] The final total score is obtained by weighting these three, and the weights are denoted as α, β, and γ respectively. The calculation of the total score S and the final output Y * are as follows:
[0217] S i = α * S v + β * S l + γ * S s ;
[0218]
[0219] Figure 12 The block diagram of the training device of the commentary generation model provided by an exemplary embodiment of the present application is shown. The device includes:
[0220] The generalization module 1210 is used to generalize the seed instruction set through the first large language model to obtain a generation instruction set. The number of generation instructions included in the generation instruction set is greater than the number of seed instructions included in the seed instruction set. Both the generation instructions and the seed instructions include the context text with alternating main and sub-commentary speeches;
[0221] The training module 1220 is used to train the second large language model based on the generation instruction set to obtain a commentary generation model. The scale of the second large language model is smaller than that of the first large language model. The task of the commentary generation model is to take the main commentary speech as the input and output the sub-commentary speech.
[0222] In an optional embodiment, the generalization module 1210 is used to extract at least one seed instruction from the seed instruction set in one instruction generation cycle of multiple instruction generation cycles; and, extract at least one generation instruction from the first generation instruction set;
[0223] Combine at least one seed instruction and at least one generation instruction to obtain a reference example;
[0224] Prompt the first large language model to imitate the reference example to obtain new generation instructions; add the new generation instructions to the first generation instruction set;
[0225] Execute multiple instruction generation cycles so that the number of generation instructions included in the first generation instruction set is greater than the number of seed instructions included in the seed instruction set;
[0226] Determine the first generation instruction set as the generation instruction set.
[0227] In an optional embodiment, the generalization module 1210 is used to prompt the first large language model to imitate the reference example to generate a two-person commentary segment to obtain new generation instructions.
[0228] In an optional embodiment, the generalization module 1210 is configured to prompt the first large language model to imitate a reference example to generate a main commentary segment, obtaining a first main commentary segment, where the first main commentary segment includes multiple first main commentary utterances; based on the first main commentary segment, prompt the first large language model to generate a secondary commentary segment, obtaining a first secondary commentary segment, where the first secondary commentary segment includes multiple first secondary commentary utterances; alternately combine the multiple first main commentary utterances and the multiple first secondary commentary utterances to obtain a new generation instruction.
[0229] In an optional embodiment, both the seed instruction and the generation instruction in the reference example include main-secondary commentary utterance pairs of multiple consecutive events; the generalization module 1210 is configured to prompt the first large language model to imitate the reference example and output in the format of the main commentary utterance of the previous event, the main commentary utterance of the current event, and the secondary commentary utterance of the current event, generating a dual-person commentary segment to obtain a new generation instruction.
[0230] In an optional embodiment, both the seed instruction and the generation instruction in the reference example include main-secondary commentary utterance pairs of multiple consecutive events; the generalization module 1210 is configured to prompt the first large language model to imitate the reference example and output in the format of the main commentary utterance of the previous event and the main commentary utterance of the current event, generating a main commentary segment to obtain a first main commentary segment, where the first main commentary segment includes the first main commentary utterances of multiple consecutive events; the first secondary commentary segment includes the first secondary commentary utterances of multiple consecutive events; combine the multiple first main commentary utterances and the multiple first secondary commentary utterances event by event to obtain a new generation instruction.
[0231] In an optional embodiment, the generalization module 1210 is configured to perform at least one of the following steps:
[0232] Determine that the new generation instruction does not contain a preset failure word;
[0233] Determine that the number of words in the new generation instruction is greater than a first preset value;
[0234] Determine that the number of words contained in the new generation instruction is greater than a second preset value;
[0235] Determine that the new generation instruction does not contain a word with an occurrence frequency greater than a third preset value;
[0236] Determine that the similarity between the new generation instruction and any one of at least one seed instruction is less than a fourth preset value, and determine that the similarity between the new generation instruction and any one of at least one generation instruction is less than a fourth preset value.
[0237] In an alternative embodiment, the apparatus further includes an acquisition module 1230. The acquisition module 1230 is configured to acquire a second main commentary segment, where the second main commentary segment includes a plurality of second main commentary utterances; based on the second main commentary segment, through a prompting instruction, prompt a third large language model to generate a secondary commentary segment, obtaining a second secondary commentary segment, where the second secondary commentary segment includes a plurality of second secondary commentary utterances; alternately combine the plurality of second main commentary utterances and the plurality of second secondary commentary utterances to obtain a seed instruction; repeat the above steps to obtain a seed instruction set; wherein the scale of the third large language model is larger than the scale of the second large language model.
[0238] In an alternative embodiment, the acquisition module 1230 is configured to extract a single-player commentary segment from the single-player commentary text of a complete game, where the single-player commentary segment includes the single-player commentary utterances of a continuous plurality of events; use the single-player commentary segment as the second main commentary segment; generate secondary commentary utterances event by event to obtain a second secondary commentary segment; when the current event in the second main commentary segment does not belong to the events in the event white list, continue to search for the next event in the second main commentary segment; when the current event belongs to the events in the event white list, based on the second main commentary segment, through a prompting instruction, prompt the third large language model to generate the secondary commentary utterance of the current event.
[0239] In an alternative embodiment, the acquisition module 1230 is configured to, in the i-th iteration process, based on the second main commentary segment, through the i-th prompting instruction, prompt the third large language model to generate a secondary commentary segment, obtaining an i-th secondary commentary segment; when the quality of the i-th secondary commentary segment is qualified, determine the i-th secondary commentary segment as the second secondary commentary segment; when the quality of the i-th secondary commentary segment is unqualified, combine the evaluation content of the i-th secondary commentary segment and the i-th prompting instruction to obtain an (i + 1)-th prompting instruction, where the (i + 1)-th prompting instruction is used to execute the (i + 1)-th iteration process; wherein when i is equal to 1, the i-th prompting instruction is an initial prompting instruction.
[0240] In an alternative embodiment, each generation instruction in the generation instruction set includes a main-secondary commentary utterance pair of a continuous plurality of events. The training module 1220 is configured to, based on the main commentary utterance sequence included in the generation instruction and the secondary commentary utterance sequences of the first t - 1 events, predict the secondary commentary utterance of the t-th event through the second large language model; output the probability that the prediction result is the secondary commentary utterance of the t-th event included in the generation instruction; construct a loss function by accumulating the plurality of probabilities corresponding to the plurality of events included in the generation instruction; based on the generation instruction set, through the loss function, train the second large language model to obtain a commentary generation model.
[0241] In an optional embodiment, the device further includes a post-processing module 1240. The post-processing module 1240 is configured to input the target main interpretation utterance into the interpretation generation model, and infer a plurality of candidate secondary interpretation utterances; score the plurality of candidate secondary interpretation utterances based on scoring dimensions, where the scoring dimensions include at least one of the word type in the utterance, the utterance length, and the repetition degree with the target main interpretation utterance; and determine the candidate secondary interpretation utterance with the highest score as the inference result.
[0242] In an optional embodiment, the post-processing module 1240 is configured to, for any one of the plurality of candidate secondary interpretation utterances, determine the number of words in the candidate secondary interpretation utterance that intersect with the blacklist word list to obtain a first value; use the number of words included in the blacklist word list as a second value; and obtain the score of the candidate secondary interpretation utterance based on the ratio of the first value to the second value.
[0243] In an optional embodiment, the post-processing module 1240 is configured to score the plurality of candidate secondary interpretation utterances based on the lengths of the plurality of candidate secondary interpretation utterances, such that the utterances with lengths less than the preset utterance length among the plurality of candidate secondary interpretation utterances obtain the same score, and the longer the utterance length among the plurality of candidate secondary interpretation utterances, the lower the score.
[0244] In an optional embodiment, the post-processing module 1240 is configured to, for any one of the plurality of candidate secondary interpretation utterances, calculate the difference between the length of the candidate secondary interpretation utterance and the preset utterance length; input the difference into a sign function, and obtain a third value, where the sign function satisfies that when the input is greater than zero, the output is a fixed positive number, and when the input is less than or equal to zero, the output is zero; multiply the difference by the third value and then take the inverse to obtain a fourth value; and obtain the score of the candidate secondary interpretation utterance based on the calculation result with e as the base and the fourth value as the exponent.
[0245] In an optional embodiment, the post-processing module 1240 is configured to, for any one of the plurality of candidate secondary interpretation utterances, calculate the similarity between the candidate secondary interpretation utterance and the target main interpretation utterance; and obtain the score of the candidate secondary interpretation utterance based on the similarity.
[0246] In summary, the solution of the present application will obtain a seed instruction set, generalize the seed instruction set using a first large language model to obtain a generation instruction set, where the generation instructions in the generation instruction set and the seed instructions in the seed instruction set both include context texts with alternating main and secondary interpretation utterances; then train a second large language model based on the generation instruction set to obtain an interpretation generation model, and the task of the interpretation generation model is to take the main interpretation utterance as the input and output the secondary interpretation utterance, and the scale of the second large language model is smaller than that of the first large language model.
[0247] In this application, the secondary commentary utterances output by the commentary generation model can be combined with the input primary commentary utterances to form dual-person commentary utterances. Therefore, this application provides a dual-person commentary framework that decomposes the game commentary task into a dual-person commentary task, is applicable to various commentary scenarios, and has good transferability.
[0248] Different from the AI single-person commentary provided by the related art, the dual-person commentary can convey more effective information while fully mobilizing the emotions of the audience. Moreover, in the above method, the training data set (generation instruction set) is generalized by a large-scale large language model, and the commentary generation model is obtained by training a small-scale large language model. In the real-time commentary scenario, generating commentary utterances through a small-scale model reduces the model call cost, and moreover, ensures the model call rate in the real-time commentary scenario with low latency.
[0249] Moreover, in some scenarios, the input primary commentary utterances are the real-time commentary content of a human. The secondary commentary utterances generated by this application can relieve the commentary pressure of a single person, and the machine and the human cooperate with each other to provide commentary, which can enrich the audio-visual experience of the audience. In other scenarios, the input primary commentary utterances are generated by artificial intelligence. At this time, both the dual-person commentaries are generated by AI, which fully releases human resources and reduces the labor cost of commentary.
[0250] Figure 13 It is a schematic structural diagram of a computer device shown according to an exemplary embodiment. The computer device 1300 includes a central processing unit (CPU) 1301, a system memory 1304 including a random access memory (RAM) 1302 and a read-only memory (ROM) 1303, and a system bus 1305 connecting the system memory 1304 and the central processing unit 1301. The computer device 1300 further includes a basic input / output system (Input / Output, I / O system) 1306 that helps transmit information between various components within the computer device, and a mass storage device 1307 for storing an operating system 1313, application programs 1314, and other program modules 1315.
[0251] The basic input / output system 1306 includes a display 1308 for displaying information and input devices 1309 such as a mouse, keyboard, etc. for user input of information. Both the display 1308 and the input devices 1309 are connected to the central processing unit 1301 through an input / output controller 1310 connected to the system bus 1305. The basic input / output system 1306 may further include an input / output controller 1310 for receiving and processing inputs from a plurality of other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1310 also provides outputs to a display screen, printer, or other types of output devices.
[0252] The mass storage device 1307 is connected to the central processing unit 1301 through a mass storage controller (not shown) connected to the system bus 1305. The mass storage device 1307 and its associated computer-readable media provide non-volatile storage for the computer device 1300. That is, the mass storage device 1307 may include computer-readable media (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.
[0253] Without loss of generality, the computer-readable media may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, digital video disc (DVD), or other optical storage, magnetic tape cartridges, tapes, magnetic disk storage, or other magnetic storage devices. Of course, those skilled in the art will know that the computer storage media is not limited to the above several types. The above-mentioned system memory 1304 and the mass storage device 1307 may be collectively referred to as memory.
[0254] According to various embodiments of the present disclosure, the computer device 1300 may also run by connecting to a remote computer device on the network through a network such as the Internet. That is, the computer device 1300 may be connected to the network 1311 through the network interface unit 1312 connected to the system bus 1305, or in other words, the network interface unit 1312 may also be used to connect to other types of networks or remote computer device systems (not shown).
[0255] The memory further includes one or more programs. The one or more programs are stored in the memory, and the central processing unit 1301 implements all or part of the steps of the above-described training method of the interpretation generation model by executing the one or more programs.
[0256] Figure 14 FIG. shows a structural block diagram of a computer device 1400 provided by an exemplary embodiment of the present application. The computer device 1400 may be a portable mobile terminal, such as: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer, or a desktop computer. The computer device 1400 may also be referred to by other names such as a user device, a portable terminal, a laptop terminal, a desktop terminal, etc.
[0257] Generally, the computer device 1400 includes: a processor 1401 and a memory 1402.
[0258] The processor 1401 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1401 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1401 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0259] The memory 1402 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1402 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1401 to implement the training method of the explanation generation model provided in the method embodiments of the present application.
[0260] In some embodiments, the computer device 1400 may also optionally include: a peripheral device interface 1403 and at least one peripheral device. The processor 1401, the memory 1402, and the peripheral device interface 1403 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1403 through a bus, signal lines, or a circuit board. Exemplarily, the peripheral device may include at least one of a radio frequency circuit 1404, a display screen 1405, a camera component 1406, an audio circuit 1407, and a power supply 1408.
[0261] The peripheral device interface 1403 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1401 and the memory 1402. In some embodiments, the processor 1401, the memory 1402, and the peripheral device interface 1403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1401, the memory 1402, and the peripheral device interface 1403 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0262] The radio frequency circuit 1404 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1404 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1404 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1404 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1404 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1404 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0263] The display screen 1405 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1405 is a touch display screen, the display screen 1405 also has the ability to collect touch signals on or above the surface of the display screen 1405. The touch signals can be input to the processor 1401 as control signals for processing. At this time, the display screen 1405 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1405, which is provided on the front panel of the computer device 1400; in other embodiments, there may be at least two display screens 1405, which are respectively provided on different surfaces of the computer device 1400 or are in a folding design; in other embodiments, the display screen 1405 may be a flexible display screen, which is provided on the curved surface or the folding surface of the computer device 1400. Even, the display screen 1405 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1405 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0264] The camera module 1406 is used to collect images or videos. Optionally, the camera module 1406 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are respectively any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to realize the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera module 1406 may further include a flash. The flash can be a single-color-temperature flash or a two-color-temperature flash. The two-color-temperature flash refers to the combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.
[0265] The audio circuit 1407 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1401 for processing, or input to the radio frequency circuit 1404 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the computer device 1400. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1401 or the radio frequency circuit 1404 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1407 may further include a headphone jack.
[0266] The power supply 1408 is used to supply power to each component in the computer device 1400. The power supply 1408 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 1408 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0267] In some embodiments, the computer device 1400 further includes one or more sensors 1409. The one or more sensors 1409 include but are not limited to: an acceleration sensor 1410, a gyroscope sensor 1411, a pressure sensor 1412, an optical sensor 1413, and a proximity sensor 1414.
[0268] The acceleration sensor 1410 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the computer device 1400. For example, the acceleration sensor 1410 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1401 can control the display screen 1405 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1410. The acceleration sensor 1410 can also be used for the collection of game or user movement data.
[0269] The gyroscope sensor 1411 can detect the body direction and rotation angle of the computer device 1400. The gyroscope sensor 1411 can cooperate with the acceleration sensor 1410 to collect the 3D actions of the user on the computer device 1400. According to the data collected by the gyroscope sensor 1411, the processor 1401 can achieve the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0270] The pressure sensor 1412 can be disposed on the side frame of the computer device 1400 and / or the lower layer of the display screen 1405. When the pressure sensor 1412 is disposed on the side frame of the computer device 1400, a grip signal of the user on the computer device 1400 can be detected, and the processor 1401 can perform left / right hand recognition or a quick operation according to the grip signal collected by the pressure sensor 1412. When the pressure sensor 1412 is disposed on the lower layer of the display screen 1405, the processor 1401 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1405. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0271] The optical sensor 1413 is used to collect the ambient light intensity. In one embodiment, the processor 1401 can control the display brightness of the display screen 1405 according to the ambient light intensity collected by the optical sensor 1413. Exemplarily, when the ambient light intensity is high, the display brightness of the display screen 1405 is increased; when the ambient light intensity is low, the display brightness of the display screen 1405 is decreased. In another embodiment, the processor 1401 can also dynamically adjust the shooting parameters of the camera assembly 1406 according to the ambient light intensity collected by the optical sensor 1413.
[0272] The proximity sensor 1414, also known as a distance sensor, is usually disposed on the front panel of the computer device 1400. The proximity sensor 1414 is used to collect the distance between the user and the front of the computer device 1400. In one embodiment, when the proximity sensor 1414 detects that the distance between the user and the front of the computer device 1400 is gradually decreasing, the processor 1401 controls the display screen 1405 to switch from the lit state to the off state; when the proximity sensor 1414 detects that the distance between the user and the front of the computer device 1400 is gradually increasing, the processor 1401 controls the display screen 1405 to switch from the off state to the lit state.
[0273] Those skilled in the art can understand that Figure 14 the structure shown in does not limit the computer device 1400, and may include more or fewer components than shown in the figure, or combine some components, or adopt a different component layout.
[0274] This application also provides a computer-readable storage medium, in which at least one instruction, at least one segment of program, a code set or an instruction set is stored, and the at least one instruction, the at least one segment of program, the code set or the instruction set is loaded and executed by a processor to implement the training method of the interpretation generation model provided by the above method embodiment.
[0275] The present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the training method of the explanation generation model provided in the above method embodiment.
[0276] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0277] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.
[0278] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A training method for an explanation generation model, characterized in that, The method includes: Generalizing a seed instruction set through a first large language model to obtain a generated instruction set, where the number of generated instructions included in the generated instruction set is greater than the number of seed instructions included in the seed instruction set, and both the generated instructions and the seed instructions include context texts with alternating main and sub-commentary utterances; Training a second large language model based on the generated instruction set to obtain the commentary generation model, where the scale of the second large language model is smaller than that of the first large language model, and the task of the commentary generation model is to take the main commentary utterance as input and output the sub-commentary utterance.
2. The method according to claim 1, wherein The generalizing the seed instruction set through the first large language model to obtain a generated instruction set includes: In one instruction generation cycle among multiple instruction generation cycles, extracting at least one seed instruction from the seed instruction set; and extracting at least one generated instruction from a first generated instruction set; Combining the at least one seed instruction and the at least one generated instruction to obtain a reference example; Prompting the first large language model to imitate the reference example to obtain new generated instructions; adding the new generated instructions to the first generated instruction set; Executing the multiple instruction generation cycles such that the number of generated instructions included in the first generated instruction set is greater than the number of seed instructions included in the seed instruction set; Determining the first generated instruction set as the generated instruction set.
3. The method according to claim 2, wherein The prompting the first large language model to imitate the reference example to obtain new generated instructions includes: Prompting the first large language model to imitate the reference example to generate a two-person commentary segment to obtain the new generated instructions; Or, Prompting the first large language model to imitate the reference example to generate a main commentary segment to obtain a first main commentary segment, where the first main commentary segment includes multiple first main commentary utterances; based on the first main commentary segment, prompting the first large language model to generate a sub-commentary segment to obtain a first sub-commentary segment, where the first sub-commentary segment includes multiple first sub-commentary utterances; alternately combining the multiple first main commentary utterances and the multiple first sub-commentary utterances to obtain the new generated instructions.
4. The method according to claim 3, characterized in that, Both the seed instructions and the generated instructions in the reference example include pairs of main and sub-commentary utterances for a continuous series of events; The prompting the first large language model to imitate the reference example to generate a two-person commentary segment to obtain the new generated instructions includes: Prompting the first large language model to imitate the reference example and output in the format of the main commentary utterance of the previous event, the main commentary utterance of the current event, and the sub-commentary utterance of the current event to generate a two-person commentary segment to obtain the new generated instructions.
5. The method according to claim 3, wherein Both the seed instructions and the generated instructions in the reference example include pairs of main and sub-commentary utterances for a continuous series of events; The prompting the first large language model to imitate the reference example to generate a main commentary segment to obtain a first main commentary segment, where the first main commentary segment includes multiple first main commentary utterances, includes: Prompt the first large language model to imitate the reference example and output in the format of the main explanation speech of the previous event and the main explanation speech of the current event to generate a main explanation segment, obtaining the first main explanation segment, where the first main explanation segment includes the first main explanation speeches of a continuous plurality of events; The first sub-explanation segment includes the first sub-explanation speeches of a continuous plurality of events; The alternating combination of the plurality of first main explanation speeches and the plurality of first sub-explanation speeches to obtain the new generation instruction includes: Combining the plurality of first main explanation speeches and the plurality of first sub-explanation speeches event by event to obtain the new generation instruction.
6. The method according to claim 2, characterized in that Before adding the new generation instruction to the first generation instruction set, at least one of the following steps is further included: Determine that the new generation instruction does not contain a preset failure word; Determine that the number of words in the new generation instruction is greater than a first preset value; Determine that the number of words included in the new generation instruction is greater than a second preset value; Determine that the new generation instruction does not contain words with an occurrence frequency greater than a third preset value; Determine that the similarity between the new generation instruction and any one of the at least one seed instruction is less than a fourth preset value, and determine that the similarity between the new generation instruction and any one of the at least one generation instruction is less than the fourth preset value.
7. According to the method described in any one of claims 1 to 6, characterized in that, The method further includes: Obtain a second main explanation segment, where the second main explanation segment includes a plurality of second main explanation speeches; Based on the second main explanation segment, through a prompt instruction, prompt a third large language model to generate a sub-explanation segment, obtaining a second sub-explanation segment, where the second sub-explanation segment includes a plurality of second sub-explanation speeches; Alternately combine the plurality of second main explanation speeches and the plurality of second sub-explanation speeches to obtain one of the seed instructions; Repeat the above steps to obtain the seed instruction set; Wherein, the scale of the third large language model is larger than the scale of the second large language model.
8. The method according to claim 7, wherein The method further includes: Extract a single-player explanation segment from the single-player explanation text of a complete game, where the single-player explanation segment includes the single-player explanation speeches of a continuous plurality of events; Use the single-player explanation segment as the second main explanation segment; The step of, based on the second main explanation segment, through a prompt instruction, prompting a third large language model to generate a sub-explanation segment, obtaining a second sub-explanation segment, includes: Generate sub-explanation speeches event by event to obtain the second sub-explanation segment; When the current event in the second main explanation segment does not belong to the events in the event whitelist, continue to search for the next event in the second main explanation segment; When the current event belongs to the events in the event whitelist, based on the second main explanation segment, through a prompt instruction, prompt a third large language model to generate the sub-explanation speech of the current event.
9. The method according to claim 7, characterized in that, The step of, based on the second main explanation segment, through a prompt instruction, prompting a third large language model to generate a sub-explanation segment, obtaining a second sub-explanation segment, includes: In the i-th iteration process, based on the second main commentary segment, the third large language model is prompted by the i-th prompt instruction to generate a sub-commentary segment, obtaining the i-th sub-commentary segment; When the quality of the i-th sub-commentary segment is qualified, the i-th sub-commentary segment is determined as the second sub-commentary segment; When the quality of the i-th sub-commentary segment is unqualified, combining the evaluation content of the i-th sub-commentary segment and the i-th prompt instruction to obtain the (i + 1)-th prompt instruction, where the (i + 1)-th prompt instruction is used to perform the (i + 1)-th iteration process; Among them, when i is equal to 1, the i-th prompt instruction is the initial prompt instruction.
10. The method according to any one of claims 1 to 6, characterized in that, Each generation instruction in the generation instruction set includes the main and sub-commentary utterance pairs of a continuous plurality of events; training the second large language model based on the generation instruction set to obtain the commentary generation model includes: Based on the main commentary utterance sequence included in the generation instruction and the sub-commentary utterance sequences of the first t - 1 events, predicting the sub-commentary utterance of the t-th event through the second large language model; outputting the probability that the prediction result is the sub-commentary utterance of the t-th event included in the generation instruction; Constructing a loss function by accumulating the probabilities corresponding to a plurality of events included in the generation instruction; Training the second large language model based on the generation instruction set through the loss function to obtain the commentary generation model.
11. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Inputting the target main commentary utterance into the commentary generation model to infer multiple candidate sub-commentary utterances; Scoring the multiple candidate sub-commentary utterances based on scoring dimensions, where the scoring dimensions include at least one of the word type in the utterance, the utterance length, and the repetition degree with the target main commentary utterance; Determining the candidate sub-commentary utterance with the highest score as the inference result.
12. The method according to claim 11, wherein The scoring the multiple candidate sub-commentary utterances based on the scoring dimensions includes: For any one of the multiple candidate sub-commentary utterances, determining the number of words in the candidate sub-commentary utterance that intersect with the words in the blacklist word list to obtain a first value; Taking the number of words included in the blacklist word list as a second value; Obtaining the score of the candidate sub-commentary utterance based on the ratio of the first value to the second value.
13. The method according to claim 11, wherein The scoring the multiple candidate sub-commentary utterances based on the scoring dimensions includes: Scoring the multiple candidate sub-commentary utterances based on the lengths of the multiple candidate sub-commentary utterances, such that the utterances with lengths less than the preset utterance length among the multiple candidate sub-commentary utterances obtain the same score, and the longer the utterance length among the multiple candidate sub-commentary utterances, the lower the score.
14. The method according to claim 13, wherein The scoring the multiple candidate sub-commentary utterances based on the lengths of the multiple candidate sub-commentary utterances includes: For any one of the multiple candidate sub-commentary utterances, calculating the difference between the length of the candidate sub-commentary utterance and the preset utterance length; Input the difference into a sign function to obtain a third value. The sign function satisfies that when the input is greater than zero, the output is a fixed positive number, and when the input is less than or equal to zero, the output is zero; Multiply the difference by the third value and then take the negation to obtain a fourth value; Based on the calculation result with e as the base and the fourth value as the exponent, obtain the score of the candidate secondary solution utterance.
15. The method according to claim 11, wherein The scoring of the multiple candidate secondary solution utterances based on the scoring dimension includes: For any one of the multiple candidate secondary solution utterances, calculate the similarity between the candidate secondary solution utterance and the target primary solution utterance; Based on the similarity, obtain the score of the candidate secondary solution utterance.
16. A training device for an explanation generation model, characterized in that, The device includes: A generalization module for generalizing a seed instruction set through a first large language model to obtain a generated instruction set. The number of generated instructions included in the generated instruction set is greater than the number of seed instructions included in the seed instruction set. Both the generated instructions and the seed instructions include context texts with alternating primary and secondary solution utterances; A training module for training a second large language model based on the generated instruction set to obtain the explanation generation model. The scale of the second large language model is smaller than the scale of the first large language model. The task of the explanation generation model is to take the primary solution utterance as the input and output the secondary solution utterance.
17. A computer device, characterized in that, The computer device includes: a processor and a memory. The memory stores a computer program, and the computer program is loaded and executed by the processor to implement the training method of the explanation generation model according to any one of claims 1 to 15.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the training method of the explanation generation model according to any one of claims 1 to 15.
19. A computer program product, characterized in that, The computer program product includes computer instructions. The computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute the training method of the explanation generation model according to any one of claims 1 to 15.