Game dialogue generation and game dialogue model training method and device

By filling the large language model with reasoning data corresponding to game rules, generating logically strong thought chain data and training it, the problem of insufficient complex reasoning of large language models in dialogue game tasks is solved, and the accuracy of dialogue responses and user experience are improved.

CN117217317BActive Publication Date: 2026-04-28IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2023-10-23
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing large language models lack the ability to perform complex reasoning in dialogue game tasks, resulting in insufficient accuracy in responses.

Method used

By acquiring historical dialogues of the target game, a game dialogue model is trained based on thought chain data, inference data corresponding to the game rules is populated, logically strong thought chain data is generated, and supervised fine-tuning is performed to improve the model's logical reasoning ability.

Benefits of technology

It improves the logical thinking and accuracy of dialogue responses in the game's dialogue model, thereby enhancing the user's game dialogue experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117217317B_ABST
    Figure CN117217317B_ABST
Patent Text Reader

Abstract

The application provides a game dialogue generation method and device and a game dialogue model training method and device. The method comprises the following steps: obtaining historical dialogues of a target game; and generating dialogue replies of the historical dialogues based on a game dialogue model of the target game. The game dialogue model of the target game is obtained based on thinking chain data of the target game. The thinking chain data is obtained by filling reasoning data corresponding to game rules of the target game in original dialogues of the target game. The method and device provided by the application fill the reasoning data corresponding to the game rules of the target game in the original dialogues to obtain the thinking chain data of the target game, and generate the dialogue replies based on the game dialogue model obtained by training based on the thinking chain data, thereby improving the logical thinking and accuracy of the dialogue replies, and greatly improving the use experience of the user in the game dialogue based on the game dialogue model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and apparatus for generating game dialogues and training game dialogue models. Background Technology

[0002] The emergence of Large Language Models (LLMs) has provided a novel and convenient interaction method—a unified and user-friendly natural language interface. LLMs have achieved excellent results in various complex tasks such as text generation and brainstorming, demonstrating powerful emergent thinking capabilities in these tasks. However, their performance in certain dialogue game tasks is only average. Dialogue game tasks include important ones such as rock-paper-scissors, number guessing, dice rolling, and word chain games. Currently, the main approach is to learn prompts from the large model (constructing prompts), enabling it to perform many tasks, such as text generation and comprehension, and multi-turn dialogues.

[0003] However, the reasoning ability of traditional prompts is currently insufficient, resulting in inaccurate answers when users interact with large models for tasks such as rock-paper-scissors and number guessing. Summary of the Invention

[0004] This invention provides a method and apparatus for generating game dialogues and training game dialogue models, in order to solve the shortcomings of existing models in terms of insufficient complex reasoning ability, which leads to insufficient accuracy in responses.

[0005] This invention provides a method for generating game dialogue, comprising:

[0006] Retrieve the target game's historical dialogue;

[0007] Based on the game dialogue model of the target game, dialogue responses to the historical dialogues are generated. The game dialogue model of the target game is trained based on the thought chain data of the target game. The thought chain data is obtained by filling in the original dialogue of the target game with reasoning data corresponding to the game rules of the target game.

[0008] According to a game dialogue generation method provided by the present invention, the step of acquiring the thought chain data includes:

[0009] Extract the step dialogues for each step in the target game from the original dialogue, and fill the step dialogues with reasoning data corresponding to the game rules of the target game to obtain the step reasoning dialogues;

[0010] Based on the step-by-step reasoning dialogue in each step of the target game, the mind chain data of the target game is generated.

[0011] According to a game dialogue generation method provided by the present invention, the step of extracting step dialogues of each step in the target game from the original dialogue and filling the step dialogues with reasoning data corresponding to the game rules of the target game to obtain step reasoning dialogues includes:

[0012] From the step-by-step dialogues of each step in the target game, select the step-by-step game process dialogues that are in the middle section of the target game as intermediate step dialogues;

[0013] The intermediate step dialogue is filled with reasoning data corresponding to the game rules of the target game to obtain the step reasoning dialogue.

[0014] According to a game dialogue generation method provided by the present invention, the step of generating thought chain data of the target game based on reasoning dialogue at each step in the target game includes:

[0015] Based on the execution order of each step in the target game, the thought process flow of the target game is combined;

[0016] Randomly select reasoning dialogues from each step and fill them into the corresponding steps in the thought chain process to generate the thought chain data of the target game.

[0017] According to a game dialogue generation method provided by the present invention, the step of combining the thought process flow of the target game based on the execution order of each step in the target game includes:

[0018] The intermediate steps in the target game are looped to obtain the intermediate loop process.

[0019] By combining the steps at the beginning and end of the target game, as well as the intermediate steps of the loop, the thought process flow of the target game is obtained.

[0020] According to a game dialogue generation method provided by the present invention, the training steps of the game dialogue model include:

[0021] Based on the thought chain data and the original dialogue, a game dialogue model for the target game is trained.

[0022] The present invention also provides a game dialogue generation device, comprising:

[0023] The acquisition unit retrieves the historical dialogue of the target game.

[0024] The generation unit generates dialogue responses to the historical dialogues based on the game dialogue model of the target game. The game dialogue model of the target game is trained based on the thought chain data of the target game. The thought chain data is obtained by filling in the original dialogue of the target game with reasoning data corresponding to the game rules of the target game.

[0025] The present invention also provides a training device for a game dialogue model, comprising:

[0026] The data acquisition unit acquires the original dialogue of the target game;

[0027] The data construction unit fills the original dialogue with reasoning data corresponding to the game rules of the target game to obtain the thought chain data of the target game;

[0028] The training unit trains the game dialogue model of the target game based on the thought chain data.

[0029] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the game dialogue generation methods described above, or a game dialogue model training method.

[0030] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the game dialogue generation method as described above, or the game dialogue model training method.

[0031] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the game dialogue generation methods described above, or a game dialogue model training method.

[0032] The present invention provides a method and apparatus for generating game dialogues and training game dialogue models. Based on filling the original dialogue with reasoning data corresponding to the game rules of the target game, the system obtains the thought chain data of the target game. The system then trains a game dialogue model based on the thought chain data to conduct game dialogues and generate dialogue responses. This improves the logical thinking and accuracy of the dialogue responses, thereby greatly enhancing the user experience of conducting game dialogues based on the game dialogue model. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0034] Figure 1 This is a flowchart illustrating the game dialogue generation method provided by the present invention;

[0035] Figure 2 This is one of the flowcharts illustrating the training method for the game dialogue model provided by this invention;

[0036] Figure 3 This is a schematic diagram of the process for generating thought chain data provided by the present invention;

[0037] Figure 4 This is the second flowchart illustrating the training method for the game dialogue model provided by this invention;

[0038] Figure 5 This is a schematic diagram of the game dialogue generation device provided by the present invention;

[0039] Figure 6 This is a schematic diagram of the structure of the training device for the game dialogue model provided by the present invention;

[0040] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0041] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0042] The essence of human-computer interaction is the relationship between humans and data. Traditional human-computer interaction methods require different applications to handle different types of data, making the interaction complex and cumbersome. The emergence of large language models (LLMs) provides a completely new and convenient interaction method: a unified and easy-to-use natural language interface. Large LLMs, trained on massive unsupervised data, can perform many tasks after supervised fine-tuning (SFT) and alignment, such as text generation and understanding, and multi-turn dialogue. Instruction tuning can be considered a special form of SFT; it's a process of further training the LLM on a dataset consisting of (instruction, output) pairs to enhance its capabilities and controllability. High-quality SFT data is constructed to enable the model to understand and follow human instructions. LLMs have achieved good results in various complex tasks such as text generation and brainstorming, demonstrating powerful emergent thinking capabilities in these tasks, but their performance in certain dialogue game tasks is only average.

[0043] As language models grow larger, the cost of fine-tuning becomes increasingly prohibitive. With such a large number of parameters, traditional fine-tuning alone is insufficient for effective model transfer, and the sheer volume of parameters also drastically increases the cost of backpropagation of gradients. Therefore, prompting learning has emerged. Prompting learning modifies downstream tasks and incorporates expert knowledge to make the input and output of the target task more closely match the data used during the original language model training. Large models possess strong in-context few-shot capabilities; however, traditional prompting methods lag behind in mathematical calculations, common-sense reasoning, and logical reasoning, resulting in lower accuracy in dialogue data.

[0044] To address the aforementioned problems, this invention provides a game dialogue generation method that achieves game dialogue generation with strong reasoning capabilities and accurate model dialogue data. Figure 1 This is a flowchart illustrating the game dialogue generation method provided by the present invention, as shown below. Figure 1 As shown, the method includes:

[0045] Step 110: Obtain the target game's historical dialogue;

[0046] Here, target games are games that allow users to complete the game through dialogue. This type of game differs from traditional video or physical games; it doesn't require players to perform complex operations or possess special equipment. Instead, players only need to use language to participate in the game. For example, we can consider a game of rock-paper-scissors. In this game, players can decide who goes first or who wins through dialogue. They can say "rock," "scissors," or "paper," and then judge the result according to the rules. This approach makes target games more interactive and interesting. Furthermore, historical dialogue can be user-input dialogue data or dialogue data actively output by the game generation model, such as "Let's play rock-paper-scissors, I'll choose scissors."

[0047] Step 120: Based on the game dialogue model of the target game, generate dialogue responses for the historical dialogues. The game dialogue model of the target game is trained based on the thought chain data of the target game. The thought chain data is obtained by filling in the original dialogue of the target game with reasoning data corresponding to the game rules of the target game.

[0048] Specifically, the historical dialogue of the target game is input into the corresponding game dialogue model. The game dialogue model outputs dialogue responses from the historical dialogue, and this dialogue format is repeated until the game dialogue of the target game ends. It can be understood that the dialogue responses here can be the various steps of the target game output by the game dialogue model. Furthermore, the game dialogue model here is obtained based on the training method of the game dialogue model described in any of the above embodiments.

[0049] It should be noted that during the training phase of the target game's dialogue model, supervised fine-tuning can be performed on the pre-trained model using thought chain data. The thought chain data is used as input to the initial model, and the reasoning data within the thought chain data is used as labels for supervised learning. The parameters of the initial model are adjusted based on the difference between the model's actual output and the labels, resulting in a game dialogue model with logical reasoning capabilities. Furthermore, the game dialogue model trained based on thought chain data can learn the thought logic specific to the target game, enabling it to output dialogue data with strong and accurate logical reasoning abilities.

[0050] Here, the inference data corresponding to the target game rules refers to the data obtained by logically reasoning from the original dialogue based on the game rules. For example, the original dialogue could be, "User: Let's play rock-paper-scissors, I choose rock; Game dialogue model: I choose scissors, you win!" According to the rules of the target game, "rock beats scissors," we can deduce that "rock beats scissors." Therefore, we can logically reason from the original dialogue to obtain the inference data that "rock beats scissors." It should be noted that if the inference data corresponding to the target game rules is not filled into the original dialogue, the game dialogue model may obtain dialogue data that does not conform to the target game rules based on the user's dialogue. For example, if the user's dialogue data is "I choose rock," the game dialogue model's dialogue data might be "I choose paper, rock can break paper, congratulations, you win!" Obviously, according to common sense, rock can break paper, but the reasoning logic of this game dialogue model does not conform to the target game rules, resulting in lower accuracy of the obtained dialogue data and potentially leading to irrelevant answers. Furthermore, for target games with complex logic, the logical reasoning ability of the game dialogue model trained on the original data is even more insufficient, further compromising the accuracy of the output dialogue data. Therefore, by filling the original dialogue with reasoning data corresponding to the game rules of the target game, the resulting thought chain data of the target game can have a stronger logical reasoning ability that conforms to the game rules.

[0051] The method provided in this invention obtains the target game's thought chain data by filling in reasoning data corresponding to the game rules of the target game into the original dialogue, and then uses the game dialogue model trained based on the thought chain data to conduct game dialogue and generate dialogue responses. This improves the logical thinking and accuracy of the dialogue responses, thereby greatly enhancing the user experience of conducting game dialogues based on the game dialogue model.

[0052] Based on any of the above embodiments, the steps for acquiring the thought chain data include:

[0053] Extract the step dialogues for each step in the target game from the original dialogue, and fill the step dialogues with reasoning data corresponding to the game rules of the target game to obtain the step reasoning dialogues;

[0054] Based on the step-by-step reasoning dialogue in each step of the target game, the mind chain data of the target game is generated.

[0055] Specifically, the step-by-step dialogues for each step in the target game can be extracted from the collected original dialogues, resulting in separate sets of step-by-step dialogues for each step. This means extracting the step-by-step dialogues that belong to the same game step from the original dialogue. For example, the set of step-by-step dialogues for the starting step could be {"Let's play rock-paper-scissors," ..., "Let's play rock-paper-scissors together"}. It should be noted that these step-by-step dialogues can be user-input dialogues or dialogues output by the game dialogue model.

[0056] Furthermore, the step-by-step dialogue is populated with reasoning data corresponding to the game rules of the target game. This can be achieved by directly concatenating the reasoning data of the step-by-step dialogue at the beginning or end, resulting in a step-by-step reasoning dialogue. It is understandable that this step-by-step reasoning dialogue is more interpretable and logically consistent with the game rules of the target game compared to the regular step-by-step dialogue. It is worth noting that extracting the step-by-step dialogues from the original dialogue and populating them with reasoning data corresponding to the game rules of the target game ensures the accuracy and logical consistency of the populated reasoning data.

[0057] Finally, the step-by-step reasoning dialogues of each step in the target game can be combined. For example, the step-by-step reasoning dialogues can be combined according to the execution order of each step in the target game, or they can be combined according to the actual execution scenario of the target game to obtain the mind chain data of the target game.

[0058] The method provided in this invention is based on extracting step dialogues of each step in the target game from the original dialogue, filling the step dialogues with reasoning data corresponding to the game rules of the target game to obtain step reasoning dialogues, and generating mind chain data of the target game based on the step reasoning dialogues. This ensures the accuracy of the position of reasoning data in the mind chain data and enhances the logic of the mind chain data.

[0059] Based on any of the above embodiments, the step of extracting step dialogues for each step in the target game from the original dialogue, and filling the step dialogues with reasoning data corresponding to the game rules of the target game to obtain step reasoning dialogues, includes:

[0060] From the step-by-step dialogues of each step in the target game, select the step-by-step game process dialogues that are in the middle section of the target game as intermediate step dialogues;

[0061] The intermediate step dialogue is filled with reasoning data corresponding to the game rules of the target game to obtain the step reasoning dialogue.

[0062] Specifically, dialogues in the middle stages of the target game can be selected from the dialogues at each step of the game. These are considered intermediate dialogues. It should be noted that only dialogues in the middle stages typically involve data related to the game rules, requiring logical reasoning. Therefore, intermediate dialogues generally contain stronger logical relationships than dialogues in other stages.

[0063] Furthermore, the intermediate game process dialogues can be filtered out and used as intermediate step dialogues to populate the reasoning data. It should be noted that in each intermediate step dialogue containing the game result, reasoning data corresponding to the game rules of the target game can be populated to obtain step-by-step reasoning dialogues. These step-by-step reasoning dialogues can include the game result and the logic for obtaining that result.

[0064] Understandably, given the large volume of the original data, the amount of dialogue data extracted for each step is also substantial. Furthermore, in typical target games, the opening and closing steps generally do not involve logical reasoning; they are mostly opening remarks and closing statements. Therefore, compared to filling in reasoning data for all step dialogues, filling in reasoning data for the intermediate step dialogues not only improves the efficiency and accuracy of obtaining step-by-step reasoning dialogues but also makes the logic of the step-by-step reasoning dialogues stronger.

[0065] Based on any of the above embodiments, the step-by-step reasoning dialogue based on each step in the target game, generating the thought chain data of the target game, includes:

[0066] Based on the execution order of each step in the target game, the thought process flow of the target game is combined;

[0067] Randomly select reasoning dialogues from each step and fill them into the corresponding steps in the thought chain process to generate the thought chain data of the target game.

[0068] Here, the execution order of each step in the target game reflects its execution logic. Therefore, the start step, intermediate steps, and end step in the execution order can be combined to obtain the target game's thought process flow. For example, the start and end steps can be retained, while the intermediate steps are iterated to obtain the target game's thought process flow. This thought process flow can be considered the entire process of the game dialogue model interacting with the user, or it can be an automated script that can be executed automatically. For example, the thought process flow can include a start process, multiple rounds of intermediate processes, and an end process. It is understandable that combining the execution order of each step in the target game to obtain the target game's thought process flow makes the resulting thought process data more standardized and logically sound.

[0069] Additionally, a random step-by-step reasoning dialogue can be selected from each step's reasoning dialogue and inserted into the corresponding step in the thought chain process to generate the target game's thought chain data. It should be noted that each step's reasoning dialogue contains various types of dialogue data, such as dialogue data adapted to different interaction scenarios. Taking the start step as an example, the reasoning dialogue can include dialogue data from the game dialogue model starting the game, as well as dialogue data from the user starting the game. Similarly, the end step can also end with dialogue data using different phrases, such as "Okay, if you want to play another round, just let me know. Have fun," or "I'm done playing, I'll play again next time." That is, the game can end with dialogue data output by the game dialogue model, or it can end with dialogue data output by the user. It is understandable that randomly selecting reasoning dialogues from each step and inserting them into the corresponding step in the thought chain process to generate thought chain data improves the diversity and comprehensiveness of the thought chain data, covering thought chain data that closely resembles the interaction scenarios between the user and the game dialogue model. Furthermore, thought chain data can be generated in batches based on the thought chain process, improving the efficiency of thought chain data generation.

[0070] The method provided in this invention combines the thought process flow of the target game based on the execution order of each step, and randomly selects reasoning dialogues from each step to fill in the corresponding steps in the thought process flow, generating thought process data for the target game, thus ensuring the logical flow of the thought process data. At the same time, it also improves the efficiency of generating thought process data, further enhancing the diversity and comprehensiveness of the thought process data, thereby making the logical reasoning ability of the game dialogue model trained based on the thought process data stronger and the output dialogue data more accurate.

[0071] Based on any of the above embodiments, the step of combining the thought process flow of the target game based on the execution order of each step in the target game includes:

[0072] The intermediate steps in the target game are looped to obtain the intermediate loop process.

[0073] By combining the steps at the beginning and end of the target game, as well as the intermediate steps of the loop, the thought process flow of the target game is obtained.

[0074] Specifically, game dialogues may involve multiple rounds of conversation, meaning there may be multiple rounds of executing steps from the target game within the game dialogue. For example, the game dialogue might involve the user and the game dialogue model playing rock-paper-scissors for three rounds. Therefore, the steps in the middle section of the target game can be iterated to obtain a cyclical intermediate flow, which can contain different rounds of game flow. It is understandable that the steps in the middle section of the target game may contain up to multiple rounds of game dialogue.

[0075] Furthermore, by combining the steps at the beginning and end of the target game, as well as the intermediate steps in the loop, a more complete thought chain can be derived. For example, the thought chain could be: "Game Start: input = random(s... input1 )||random(S input1 )+random(S input2 ); In the middle round of the game, for iinrange(iteration): if the model goes out first and the user goes out later, input = input + <s>+random(M1)+ <end>+random(M1), determine the result, input = input + <s>+judge(M input1 ), ask whether to continue the game, input = input + random(M2) + <end>+random(M3); If the user outputs the model first and then outputs the model later, input = input + <end>+random(M1)+ <s>+random(M1), based on the result, determine whether to continue, input = input + judge(M1). model1 )+random(M2)+ <end>+random(M3); Game over: target ends with a polite phrase, input = input + <end>+random(F input1 ), target = random(F model1 ); target ends with the game answer, input = input + <end>+random(M1), target=random(M1)+judge(M model1 ")". Among them, S input1 S input2 M1, M2, and M3 represent the initial step-by-step reasoning dialogues; F model1 This indicates a step-by-step reasoning dialogue in the middle section that includes the game outcome; F model1 The final segment represents the step-by-step reasoning dialogue; random() represents the dialogue data output by the game model, S input1 This represents the first step of the reasoning dialogue, where random represents random selection, and random(S) input1 () indicates that a step-by-step reasoning dialogue is randomly selected from the first step-by-step reasoning dialogue; iteration indicates the number of iterations; <s>、 <end>"+" is a multi-round identifier and "+" is a concatenation symbol.

[0076] Based on the method provided in the embodiments of the present invention, the steps in the middle section of the target game are looped to obtain the loop intermediate process. The steps in the beginning and end sections of the target game, as well as the loop intermediate process, are combined to obtain the thought chain process of the target game, thus ensuring the flow logic of the thought chain data.

[0077] Based on any of the above embodiments, the training steps of the game dialogue model include:

[0078] Based on the thought chain data and the original dialogue, a game dialogue model for the target game is trained.

[0079] Specifically, first, an initial model is obtained. The original dialogue can be used as training data for the initial model. This initial model is pre-trained using a Decoder-Only architecture, and validation can be performed based on pre-trained models from 13B and 65B. Further, supervised fine-tuning of the pre-trained model can be performed using thought chain data. The thought chain data is used as input to the initial model, and the reasoning data within the thought chain data is used as labels for supervised learning. The parameters of the initial model are adjusted based on the difference between the model's actual output and the labels, resulting in a game dialogue model with logical reasoning capabilities. It should be noted that during training, the learning rate can be set to 0.00005 and decayed to 0.1 to ensure the model can more fully learn the logical reasoning capabilities from the thought chain data. Finally, model reasoning can be performed on the game dialogue model based on high-quality thought chain data.

[0080] The method provided in this invention trains a game dialogue model for the target game based on thought chain data and original dialogues. The model is coarsely tuned using a large amount of original dialogue data, and then finely tuned using instructions based on thought chain data to obtain the final game dialogue model, thereby improving the training effect and efficiency of the game dialogue model.

[0081] Based on any of the above embodiments Figure 2 This is one of the flowcharts illustrating the training method for the game dialogue model provided by this invention, such as... Figure 2 As shown, the method includes:

[0082] Step 210: Collect and organize the original dialogue of the target game;

[0083] Step 220: Clean the original dialogue data;

[0084] Step 230: Generate the mind chain data of the target game based on the original dialogue;

[0085] Specifically, Figure 3 This is a schematic diagram of the process for generating thought chain data provided by the present invention, such as... Figure 3 As shown, the method includes: step 310, obtaining the thought process flow of the target game based on the execution order of each step in the target game; step 320, generating thought process data of the target game in batches based on the thought process flow of the target game. For example, the original dialogue could be "input: Let's play rock-paper-scissors, I choose scissors." <s>Okay, I play paper, you win, shall we continue? <end>I'll play the stone. <s>Okay, I'll play paper. I'm sorry, you lost. Shall we play another round? <end>Okay, I'll provide the cloth. <s>I threw stone. Congratulations, you won. Shall we continue playing? <end>That's all. Target: Okay, no problem. If you'd like to play again, feel free to ask. Have a nice day. The thought chain data generated based on the original dialogue could be: "Input: Let's play rock-paper-scissors, I choose scissors." <s>Okay, I'll play paper. Scissors beat paper, you win. Shall we continue? <end>I'll play the stone. <s>Okay, I play cloth, Booker Stone. Sorry, you lost. Shall we play another round? <end>Okay, I'll provide the cloth. <s>I play rock, and paper can cover the rock. Congratulations, you win! Want to continue playing? <end>"I'm not continuing, I'll play again tomorrow." target: "Okay, I'll continue tomorrow. Have a nice day."

[0086] Step 240: Based on the thought chain data, train the game dialogue model of the target game.

[0087] Based on any of the above embodiments Figure 4 This is the second flowchart illustrating the training method for the game dialogue model provided by this invention, as shown below. Figure 4 As shown, the method includes:

[0088] Step 410: Obtain the original dialogue of the target game;

[0089] Specifically, this can be achieved by collecting dialogue game data from the target game, as well as using the existing data from the target game's dialogue model during previous training, as the original dialogue for the current target game's dialogue model. This allows for the construction of game dialogue data with logical reasoning capabilities through the original dialogue. It's worth noting that using the target game's logical reasoning dialogue data as training data enables the game dialogue model to learn the logical reasoning abilities from the training data, resulting in a game dialogue model with strong reasoning capabilities and accurate responses.

[0090] It should be noted that the original dialogue obtained here only includes the game steps and results of the target game, excluding the reasoning process of deriving the game result from the game process. Furthermore, the original data obtained using the above method contains some redundant sample data, so fuzzy deduplication can be prioritized for the original dialogue data. For example, deduplication can first be performed using text similarity metrics such as edit distance and cosine similarity, followed by clustering algorithms. Finally, the original dialogue is manually sampled and checked to ensure logical correctness. Further, the deduplicated original dialogue can undergo further data cleaning to filter out dirty data, data with formatting issues, and low-quality data with unqualified content. For example, low-quality data can be eliminated using preset filtering rules, such as language-based filtering rules, metric-based filtering rules, keyword-based filtering rules, and other dimension-based filtering rules. Additionally, data with incomplete sentence semantics can be filtered out by generating truncation functions; data cleaning scripts can also be used to filter out dirty data such as punctuation marks at the beginning or illegal characters. It is understandable that data cleaning of the raw data improves the quality of the original dialogue data, thereby improving the accuracy of the training data obtained based on the original dialogue, and thus improving the training effect of the game dialogue model.

[0091] Step 420: Fill the original dialogue with reasoning data corresponding to the game rules of the target game to obtain the thought chain data of the target game;

[0092] Here, the inference data corresponding to the target game rules refers to the data obtained by logically reasoning from the original dialogue based on the game rules. For example, the original dialogue could be, "User: Let's play rock-paper-scissors, I choose rock; Game dialogue model: I choose scissors, you win!" According to the rules of the target game, "rock beats scissors," we can deduce that "rock beats scissors." Therefore, we can logically reason from the original dialogue to obtain the inference data that "rock beats scissors." It should be noted that if the inference data corresponding to the target game rules is not filled into the original dialogue, the game dialogue model may obtain dialogue data that does not conform to the target game rules based on the user's dialogue. For example, if the user's dialogue data is "I choose rock," the game dialogue model's dialogue data might be "I choose paper, rock can break paper, congratulations, you win!" Obviously, according to common sense, rock can break paper, but the reasoning logic of this game dialogue model does not conform to the target game rules, resulting in lower accuracy of the obtained dialogue data and potentially leading to irrelevant answers. Furthermore, for target games with complex logic, the logical reasoning ability of the game dialogue model trained on the original data is even more insufficient, further compromising the accuracy of the output dialogue data. Therefore, by filling the original dialogue with reasoning data corresponding to the game rules of the target game, the resulting thought chain data of the target game can have a stronger logical reasoning ability that conforms to the game rules.

[0093] Furthermore, the thought chain data of the target game is obtained based on the reasoning data. Here, thought chain data refers to chain-like prompt text containing step-by-step hints. Prompts with strong logical reasoning abilities can be derived from the intermediate reasoning steps in the thought chain data. For example, the game steps of the target game can be determined from the original dialogue. Reasoning data corresponding to the game rules of the target game can be filled into the dialogue data of each game step to obtain the filled dialogue data, and thus, the thought chain data of the target game.

[0094] Step 430: Based on the thought chain data, train the game dialogue model of the target game.

[0095] Specifically, supervised fine-tuning of a pre-trained model can be performed using thought chain data. The thought chain data is used as input to the initial model, and the reasoning data within it is used as labels for supervised learning. The parameters of the initial model are adjusted based on the difference between the model's actual output and the labels, resulting in a game dialogue model with logical reasoning capabilities. It should be noted that the game dialogue model trained based on thought chain data can learn the thought logic specific to the target game, enabling it to output dialogue data with strong and accurate logical reasoning abilities.

[0096] The method provided in this invention obtains the target game's thought chain data by filling the original dialogue with reasoning data corresponding to the game rules of the target game, and trains the game dialogue model of the target game based on the thought chain data, thereby improving the logical reasoning ability of the game dialogue model. Furthermore, the logical thinking of the game dialogue model conforms to the game rules of the target game, thus improving the accuracy of the dialogue data generated based on the game dialogue model.

[0097] Based on any of the above embodiments Figure 5 This is a schematic diagram of the game dialogue generation device provided by the present invention, as shown below. Figure 5 As shown, the device includes:

[0098] Get Unit 510, retrieve the target game's historical dialogue;

[0099] The generation unit 520 generates dialogue responses to the historical dialogues based on the game dialogue model of the target game. The game dialogue model of the target game is trained based on the thought chain data of the target game. The thought chain data is obtained by filling in the original dialogue of the target game with reasoning data corresponding to the game rules of the target game.

[0100] The device provided in this invention obtains the target game's thought chain data by filling in reasoning data corresponding to the game rules of the target game into the original dialogue, and then uses the game dialogue model trained based on the thought chain data to conduct game dialogue and generate dialogue responses. This improves the logical thinking and accuracy of the dialogue responses, thereby greatly enhancing the user experience of conducting game dialogues based on the game dialogue model.

[0101] Based on any of the above embodiments, the generating unit is specifically used for:

[0102] Extract the step dialogues for each step in the target game from the original dialogue, and fill the step dialogues with reasoning data corresponding to the game rules of the target game to obtain the step reasoning dialogues;

[0103] Based on the step-by-step reasoning dialogue in each step of the target game, the mind chain data of the target game is generated.

[0104] Based on any of the above embodiments, the generating unit is further specifically used for:

[0105] From the step-by-step dialogues of each step in the target game, select the step-by-step game process dialogues that are in the middle section of the target game as intermediate step dialogues;

[0106] The intermediate step dialogue is filled with reasoning data corresponding to the game rules of the target game to obtain the step reasoning dialogue.

[0107] Based on any of the above embodiments, the generating unit is further specifically used for:

[0108] Based on the execution order of each step in the target game, the thought process flow of the target game is combined;

[0109] Randomly select reasoning dialogues from each step and fill them into the corresponding steps in the thought chain process to generate the thought chain data of the target game.

[0110] Based on any of the above embodiments, the generating unit is further specifically used for:

[0111] The intermediate steps in the target game are looped to obtain the intermediate loop process.

[0112] By combining the steps at the beginning and end of the target game, as well as the intermediate steps of the loop, the thought process flow of the target game is obtained.

[0113] Based on any of the above embodiments, the generation unit further includes a training unit, which is specifically used for:

[0114] Based on the thought chain data and the original dialogue, a game dialogue model for the target game is trained.

[0115] Based on any of the above embodiments Figure 6 This is a schematic diagram of the structure of the training device for the game dialogue model provided by the present invention, as shown below. Figure 6 As shown, the device includes:

[0116] Data acquisition unit 610 acquires the original dialogue of the target game;

[0117] The data construction unit 620 fills the original dialogue with reasoning data corresponding to the game rules of the target game to obtain the thought chain data of the target game;

[0118] Training unit 630 trains the game dialogue model of the target game based on the thought chain data.

[0119] The apparatus provided in this invention obtains the thought chain data of the target game by filling in reasoning data corresponding to the game rules of the target game in the original dialogue, and trains the game dialogue model of the target game based on the thought chain data, thereby improving the logical reasoning ability of the game dialogue model. Moreover, the logical thinking of the game dialogue model conforms to the game rules of the target game, thereby improving the accuracy of the dialogue data generated based on the game dialogue model.

[0120] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740, wherein the processor 710, communications interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute a training method for a game dialogue model. This method includes: acquiring the original dialogue of the target game; filling the original dialogue with reasoning data corresponding to the game rules of the target game to obtain the thought chain data of the target game; and training the game dialogue model of the target game based on the thought chain data.

[0121] Alternatively, a game dialogue generation method can be executed, which includes: acquiring historical dialogues of a target game; and generating dialogue responses to the historical dialogues based on a game dialogue model of the target game, wherein the game dialogue model of the target game is obtained based on the training method of the game dialogue model as described in any of the above embodiments.

[0122] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0123] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the training method for the game dialogue model provided by the above methods. The method includes: acquiring the original dialogue of the target game; filling the original dialogue with reasoning data corresponding to the game rules of the target game to obtain the thought chain data of the target game; and training the game dialogue model of the target game based on the thought chain data.

[0124] The computer can also execute a game dialogue generation method, which includes: acquiring historical dialogues of a target game; and generating dialogue responses to the historical dialogues based on a game dialogue model of the target game, wherein the game dialogue model of the target game is obtained based on a training method for the game dialogue model as described in any of the above embodiments.

[0125] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a training method for a game dialogue model provided by the methods described above. The method includes: acquiring the original dialogue of a target game; filling the original dialogue with reasoning data corresponding to the game rules of the target game to obtain the thought chain data of the target game; and training the game dialogue model of the target game based on the thought chain data.

[0126] Alternatively, a game dialogue generation method can be executed, which includes: acquiring historical dialogues of a target game; and generating dialogue responses to the historical dialogues based on a game dialogue model of the target game, wherein the game dialogue model of the target game is obtained based on the training method of the game dialogue model as described in any of the above embodiments.

[0127] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0128] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.< / end> < / s> < / end> < / s> < / end> < / s> < / end> < / s> < / end> < / s> < / end> < / s> < / end> < / s> < / end> < / end> < / end> < / s> < / end> < / end> < / s> < / end> < / s>

Claims

1. A method for generating game dialogue, characterized in that, include: Retrieve the target game's historical dialogue; Based on the game dialogue model of the target game, the dialogue response of the historical dialogue is generated. The game dialogue model of the target game is trained based on the thought chain data of the target game. The training steps of the game dialogue model include: using the thought chain data as input to the pre-trained initial model, using the reasoning data in the thought chain data as labels, and adjusting the parameters of the initial model based on the difference between the actual output of the initial model and the labels, so as to train the game dialogue model. The thought chain data is obtained by filling in reasoning data corresponding to the game rules of the target game into the original dialogue of the target game.

2. The game dialogue generation method according to claim 1, characterized in that, The steps for acquiring the thought chain data include: Extract the step dialogues for each step in the target game from the original dialogue, and fill the step dialogues with reasoning data corresponding to the game rules of the target game to obtain the step reasoning dialogues; Based on the step-by-step reasoning dialogue in each step of the target game, the mind chain data of the target game is generated.

3. The game dialogue generation method according to claim 2, characterized in that, The step-by-step dialogue is extracted from the original dialogue, and reasoning data corresponding to the game rules of the target game is filled into the step-by-step dialogue to obtain the step-by-step reasoning dialogue, including: From the step-by-step dialogues of each step in the target game, select the step-by-step game process dialogues that are in the middle section of the target game as intermediate step dialogues; The intermediate step dialogue is filled with reasoning data corresponding to the game rules of the target game to obtain the step reasoning dialogue.

4. The game dialogue generation method according to claim 2, characterized in that, The step-by-step reasoning dialogue based on each step in the target game generates the target game's thought chain data, including: Based on the execution order of each step in the target game, the thought process flow of the target game is combined; Randomly select reasoning dialogues from each step and fill them into the corresponding steps in the thought chain process to generate the thought chain data of the target game.

5. The game dialogue generation method according to claim 4, characterized in that, The process of combining the thought chain of the target game based on the execution order of each step in the target game includes: The intermediate steps in the target game are looped to obtain the intermediate loop process. By combining the steps at the beginning and end of the target game, as well as the intermediate steps of the loop, the thought process flow of the target game is obtained.

6. The game dialogue generation method according to any one of claims 1 to 5, characterized in that, The training steps for the game dialogue model include: Based on the thought chain data and the original dialogue, a game dialogue model for the target game is trained.

7. A method for training a game dialogue model, characterized in that, include: Obtain the original dialogue of the target game; Fill the original dialogue with reasoning data corresponding to the game rules of the target game to obtain the thought chain data of the target game; Based on the aforementioned thought chain data, train the game dialogue model for the target game; The step of training the game dialogue model for the target game based on the thought chain data includes: The thought chain data is used as input to the pre-trained initial model, and the reasoning data in the thought chain data is used as labels. Based on the difference between the actual output of the initial model and the labels, the parameters of the initial model are adjusted to train the game dialogue model.

8. A game dialogue generation device, characterized in that, include: The acquisition unit retrieves the historical dialogue of the target game. The generation unit generates dialogue responses for the historical dialogues based on the game dialogue model of the target game. The game dialogue model of the target game is trained based on the thought chain data of the target game. The training steps of the game dialogue model include: using the thought chain data as input to the pre-trained initial model, using the inference data in the thought chain data as labels, and adjusting the parameters of the initial model based on the difference between the actual output of the initial model and the labels to train the game dialogue model. The thought chain data is obtained by filling in reasoning data corresponding to the game rules of the target game into the original dialogue of the target game.

9. A training device for a game dialogue model, characterized in that, include: The data acquisition unit acquires the original dialogue of the target game; The data construction unit fills the original dialogue with reasoning data corresponding to the game rules of the target game to obtain the thought chain data of the target game; The training unit trains the game dialogue model of the target game based on the thought chain data. The training unit is specifically used for: The thought chain data is used as input to the pre-trained initial model, and the reasoning data in the thought chain data is used as labels. Based on the difference between the actual output of the initial model and the labels, the parameters of the initial model are adjusted to train the game dialogue model.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the game dialogue generation method as described in any one of claims 1 to 6, or the game dialogue model training method as described in claim 7.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the game dialogue generation method as described in any one of claims 1 to 6, or the game dialogue model training method as described in claim 7.

Citation Information

Patent Citations

  • Multi-view relation network chart question and answer method and system based on attention mechanism

    CN115329079A

  • Information extraction method, and method and device for training question and answer processing model

    CN116662496A