Decision-making method and device for card game and electronic equipment
By generating a mixed training sample dataset and training the initial decision model, the problem of poor generalization ability of card game decision model among multiple games is solved, achieving higher decision-making ability and wider applicability.
Patent Information
- Application Number
- CN202510342615.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art has poor generalization ability between multiple card games, and prompt-based methods can only utilize the inherent knowledge of the language model, resulting in the limitation of decision-making ability by the model.
By filtering multiple card games, obtaining game trajectory data of each game, generating a mixed training sample data set, training a pre-built initial decision model, and evaluating and adjusting the universality through preset benchmarks, an actual decision model that meets the preset universality conditions is obtained.
It improves the generalization of decision-making models in various card games, improves the upper limit of decision-making ability, and solves the problems of poor generalization ability and limited decision-making ability.
Smart Images

Figure CN120114843A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a decision-making method, device, and electronic device for card games. Background Art
[0002] In related technologies, the decision-making capabilities of language models in board games and card games, such as Texas Hold'em, Blackjack, and Guandan, have been explored to demonstrate the performance of artificial intelligence algorithms in the game field.
[0003] However, there are still some limitations in related technologies.
[0004] First, the generality of related technologies is insufficient. Most of the work in related technologies focuses on designing elaborate prompts for a single game to improve performance, and its generalization ability among multiple games remains to be verified.
[0005] Second, related technologies usually adopt a prompt-based method, which can only utilize the inherent knowledge of the language model, thus limiting its performance ceiling. Like humans, mastering complex game strategy knowledge requires repeatedly playing the game and summarizing successful experiences. However, most related technologies ignore this point, resulting in poor decision-making capabilities for complex games.
[0006] In summary, in related technologies, the generalization ability among multiple games is poor, and the prompt-based method can only utilize the inherent knowledge of the language model, making the decision-making ability limited by the model, which urgently needs to be improved. Summary of the Invention
[0007] This application provides a decision-making method, device, and electronic device for card games to solve the problems in related technologies that the generalization ability among multiple games is poor, and the prompt-based method can only utilize the inherent knowledge of the language model, resulting in the decision-making ability being limited by the model.
[0008] The first aspect of the embodiments of this application provides a decision-making method for card games, which is applied to the model construction stage. The method includes the following steps: obtaining multiple card games that meet preset screening conditions; obtaining the game trajectory data of each card game to generate a mixed training sample data set using the game trajectory data; training a pre-constructed initial decision model using the mixed training sample data set, and performing a generality evaluation on the trained initial decision model using a preset benchmark to obtain an evaluation result, and adjusting the trained initial decision model based on the evaluation result and the mixed general data set corresponding to the preset benchmark to obtain an actual decision model that meets the preset generality conditions.
[0009] Optionally, in an embodiment of the present application, the obtaining of the game trajectory data of each card game to generate a mixed training sample data set by using the game trajectory data includes: matching a corresponding teacher model, opponent model, and number of games for each card game based on the game characteristics of each card game; generating the game trajectory data by using the multiple game data of the teacher model and the opponent model of each card game under the number of games; screening the game trajectory data to obtain the decision data of the winning party; constructing a training sample data set corresponding to each card game based on the decision data of the winning party, and generating the mixed training sample data set based on the training sample data set corresponding to each card game.
[0010] Optionally, in an embodiment of the present application, after obtaining the decision data of the winning party, it further includes: obtaining the game description of each card game; obtaining each data instance of the observation-action pair in the decision data of the winning party of each game; determining the compliance of each data instance of the observation-action pair based on the game description to obtain a determination result; screening the decision data of the winning party by using the determination result to obtain compliant decision data, and constructing a training sample data set corresponding to each card game by using the compliant decision data, and generating the mixed training sample data set based on the training sample data set corresponding to each card game.
[0011] Optionally, in an embodiment of the present application, before training an initial decision model that meets preset general conditions by using the training sample data set, it further includes: defining the instructions and outputs of a language model by using the observation-action pairs of each card game to obtain the initial decision model, where the instructions include the game description, state data, and output format description of each card game.
[0012] Optionally, in an embodiment of the present application, the training of a pre-constructed initial decision model by using the mixed training sample data set includes: calculating the cross-entropy loss for the outputs in the mixed training sample data set; using the cross-entropy loss as a loss function to train the initial decision model by using the loss function to obtain a trained initial decision model; where the expression of the loss function is:
[0013]
[0014] where i represents the sample index, o i represents the instruction, a i represents the output, t represents the current predicted step, p represents the probability of each character in the output predicted by the initial decision model, represents the loss function.
[0015] Optionally, in an embodiment of the present application, it further includes: generating a simulation instruction by using any one of the multiple card games to simulate a game of the any one of the card games; inputting the simulation instruction into the actual decision-making model to output a corresponding simulated decision-making action; updating the simulation instruction based on the simulated decision-making action until the game of the any one of the card games is completed and the game result of the game is obtained; and optimizing the actual decision-making model by using the game result.
[0016] An embodiment of the second aspect of the present application provides a decision-making method for a card game, which is applied to the model usage stage. The method includes the following steps: obtaining the current state information and game description of the card game to be decided; generating a corresponding decision instruction based on the current state information and the game description; and inputting the decision instruction into a pre-constructed decision-making model to output a corresponding decision-making action, where the decision-making model is trained by instruction data and game trajectories of multiple card games.
[0017] An embodiment of the third aspect of the present application provides a decision-making device for a card game, which is applied to the model construction stage. The device includes: an obtaining module, configured to obtain multiple card games that meet a preset screening condition; a generating module, configured to obtain game trajectory data of each card game to generate a mixed training sample data set by using the game trajectory data; and a training module, configured to train a pre-constructed initial decision-making model by using the mixed training sample data set, perform a generality evaluation on the trained initial decision-making model by using a preset benchmark to obtain an evaluation result, and adjust the trained initial decision-making model based on the evaluation result and a mixed general data set corresponding to the preset benchmark to obtain an actual decision-making model that meets a preset generality condition.
[0018] Optionally, in an embodiment of the present application, the generating module includes: a matching unit, configured to match a corresponding teacher model, opponent model, and number of games for each card game based on the game characteristics of each card game; a generating unit, configured to generate the game trajectory data by using the multiple game data of the teacher model and the opponent model of each card game under the number of games; a first screening unit, configured to screen the game trajectory data to obtain decision data of the winning party; and a first constructing unit, configured to construct a training sample data set corresponding to each card game based on the decision data of the winning party, and generate the mixed training sample data set based on the training sample data set corresponding to each card game.
[0019] Optionally, in an embodiment of the present application, the generation module further includes: a first acquisition unit configured to acquire the game instructions of each card game; a second acquisition unit configured to acquire each data instance of the observation-action pair in the decision data of the winning party in each game round; a determination unit configured to determine the compliance of each data instance of the observation-action pair based on the game instructions to obtain a determination result; a second construction unit configured to filter the decision data of the winning party by using the determination result to obtain compliant decision data, and use the compliant decision data to construct a training sample data set corresponding to each card game, so as to generate the mixed training sample data set based on the training sample data set corresponding to each card game.
[0020] Optionally, in an embodiment of the present application, it further includes: a definition module configured to define the instructions and outputs of the language model by using the observation-action pairs of each card game to obtain the initial decision model, where the instructions include the game instructions, state data, and output format instructions of each card game.
[0021] Optionally, in an embodiment of the present application, the training module includes: a calculation unit configured to calculate the cross-entropy loss for the output in the mixed training sample data set; a training unit configured to use the cross-entropy loss as a loss function to train the initial decision model by using the loss function to obtain a trained initial decision model; where the expression of the loss function is:
[0022]
[0023] where i represents the sample index, o i represents the instruction, a i represents the output, t represents the current predicted step, p represents the probability of each character in the output predicted by the initial decision model, represents the loss function.
[0024] Optionally, in an embodiment of the present application, it further includes: a simulation module configured to generate simulation instructions by using any one of the multiple card games to simulate the game round of the any one card game; a calculation module configured to input the simulation instructions into the actual decision model to output corresponding simulated decision actions; an update module configured to update the simulation instructions based on the simulated decision actions until the game round of the any one card game is completed and obtain the game result of the game round; an optimization module configured to optimize the actual decision model by using the game result.
[0025] A fourth - aspect embodiment of the present application provides a decision - making device for a card game, which is applied to the model - using stage. Wherein, the device includes: an acquisition module, configured to acquire the current state information and game instructions of the card game to be decided; a generation module, configured to generate corresponding decision instructions based on the current state information and the game instructions; a decision - making module, configured to input the decision instructions into a pre - constructed decision model to output corresponding decision actions, where the decision model is trained by instruction data of multiple card games and corresponding game trajectories.
[0026] A fifth - aspect embodiment of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the program to implement the decision - making method for the card game as described in the above - mentioned embodiment.
[0027] A sixth - aspect embodiment of the present application provides a computer - readable storage medium, where the computer - readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the decision - making method for the card game as described in the above - mentioned embodiment.
[0028] A seventh - aspect embodiment of the present application provides a computer program product, including a computer program, which is used to implement the decision - making method for the card game as above when executed.
[0029] The embodiments of the present application can screen multiple card games, acquire the game trajectory data of each card game, generate a mixed training sample data set by using the game trajectory data to explore the mutual influence between different card games and the influence on the general ability of the model, train a pre - constructed initial decision model by using the mixed training sample data set to obtain a decision model that can be generally applied to multiple card games, perform a generality evaluation on the trained initial decision model by using a preset benchmark to obtain an evaluation result, and adjust the trained initial decision model based on the evaluation result and the mixed general data set corresponding to the preset benchmark to restore the general ability of the trained initial decision model, so that the actual decision model is not limited to the decision of card games, thereby improving the upper limit of the decision - making ability. Thus, it solves the problem in the related art that the generalization ability between multiple games is poor, and the method based on prompts can only utilize the inherent knowledge of the language model, resulting in the decision - making ability being limited by the model.
[0030] The additional aspects and advantages of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The above - mentioned and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0032] Figure 1 Flowchart of a decision-making method for a card game provided according to an embodiment of the present application;
[0033] Figure 2 Flowchart of a decision-making method for a card game provided according to an embodiment of the present application;
[0034] Figure 3 Schematic diagram of a landlord instruction template provided according to an embodiment of the present application;
[0035] Figure 4 Schematic diagram of the evaluation result of the general capabilities of a model provided according to an embodiment of the present application;
[0036] Figure 5 Schematic diagram of the structure of a decision-making device for a card game provided according to an embodiment of the present application;
[0037] Figure 6 Flowchart of another decision-making method for a card game provided according to an embodiment of the present application;
[0038] Figure 7 Schematic diagram of the structure of another decision-making device for a card game provided according to an embodiment of the present application;
[0039] Figure 8 Schematic diagram of the structure of an electronic device provided according to an embodiment of the present application. Detailed implementation manners
[0040] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.
[0041] The decision-making method, device, and electronic device for a card game according to an embodiment of the present application will be described below with reference to the accompanying drawings. In view of the problem in the related art mentioned in the above background art that the generalization ability between multiple games is poor, and the method based on prompts can only utilize the inherent knowledge of the language model, resulting in the decision-making ability being limited by the model, the present application provides a decision-making method for a card game. In this method, multiple card games can be screened, and the game trajectory data of each card game can be obtained to generate a mixed training sample data set by using the game trajectory data, so as to explore the mutual influence between different card games and the influence on the general ability of the model. The pre-constructed initial decision model is trained by using the mixed training sample data set to obtain a decision model that can be generally applied to multiple card games, and the trained initial decision model is evaluated for generality by using a preset benchmark to obtain an evaluation result, and the trained initial decision model is adjusted based on the evaluation result and the mixed general data set corresponding to the preset benchmark to restore the general ability of the trained initial decision model, so that the actual decision model can be not limited to the decision-making of card games, thereby improving the upper limit of the decision-making ability. Thus, the problem in the related art that the generalization ability between multiple games is poor, and the method based on prompts can only utilize the inherent knowledge of the language model, resulting in the decision-making ability being limited by the model is solved.
[0042] Specifically, Figure 1 is a schematic flowchart of a decision-making method for a card game provided by an embodiment of the present application.
[0043] As Figure 1 shown, this decision-making method for a card game is applied to the model construction stage, where the method includes the following steps:
[0044] In step S101, multiple card games that meet the preset screening conditions are obtained.
[0045] In the actual execution process, since the purpose of the embodiment of the present application is to fine-tune and train the language model through the high-quality battle data of multiple card games to improve the decision-making ability of the language model in multiple card games, therefore, the embodiment of the present application needs to preliminarily screen the card games. The specific screening conditions can be obtained according to the final use purpose of the model. For example, the embodiment of the present application can consider the popularity, complexity, and availability of high-quality models or data of the card games. Based on the above screening conditions, the embodiment of the present application can obtain multiple card games that meet the screening conditions, such as Dou Di Zhu, Guan Dan, Japanese Mahjong (Riichi Mahjong), Uno, Gin Rummy, Leduc Poker, Limit Hold'em, and No-Limit Hold'em.
[0046] In step S102, the game trajectory data of each card game is obtained to generate a mixed training sample data set by using the game trajectory data.
[0047] Further, for model training, it is necessary to generate instruction-output pairs for game decisions. To this end, the embodiments of the present application can collect data based on the above-selected card games to obtain game trajectory data for each card game, and then establish a mixed training sample data set. The mixed training sample data set contains data of all the selected card games to improve the general ability of the model through the mixed sample data.
[0048] Optionally, in an embodiment of the present application, obtaining the game trajectory data for each card game to generate a mixed training sample data set by using the game trajectory data includes: matching a corresponding teacher model, opponent model, and number of game rounds for each card game based on the game characteristics of each card game; generating game trajectory data by using the multiple game round data of the teacher model and opponent model of each card game under the number of game rounds; filtering the game trajectory data to obtain the decision data of the winning party; constructing a training sample data set corresponding to each card game based on the decision data of the winning party, and generating a mixed training sample data set based on the training sample data set corresponding to each card game.
[0049] Among them, the embodiments of the present application can obtain game instruction data through three steps: game trajectory data generation, trajectory data filtering, and data format conversion, and then generate a mixed training sample data set.
[0050] Specifically, in the game trajectory data generation, the embodiments of the present application can generate game trajectory data by having the teacher model and the opponent fight multiple times in the game simulator. Among them, for different card games, the embodiments of the present application can match corresponding teacher models and opponent models to ensure the validity of the obtained data.
[0051] In the trajectory data filtering, only the decision data of the winning player is filtered and retained.
[0052] In the data format conversion, an instruction template is formulated for each card game, and the game observation data and action data are filled into the template to obtain the game instruction data.
[0053] Optionally, in an embodiment of the present application, after obtaining the decision data of the winning party, it further includes: obtaining the game description of each card game; obtaining each observation-action pair data instance in the decision data of the winning party of each game round; determining the compliance of each observation-action pair data instance based on the game description to obtain a determination result; using the determination result to filter the decision data of the winning party to obtain compliant decision data, and using the compliant decision data to construct a training sample data set corresponding to each card game, and generating a mixed training sample data set based on the training sample data set corresponding to each card game.
[0054] When performing data screening, the embodiments of the present application can, after retaining the decision data of the winning party, determine the legal actions for each action from the observation and action data in the decision data.
[0055] The embodiments of the present application can determine the card game corresponding to the data instance of each observation-action pair, obtain the legal action options of the card game, and then determine the compliance of the actions according to the legal action options, and retain the data samples with the legal action option data greater than a certain threshold to construct the training sample data set for each card game. Then, a mixed training sample data set is obtained by combining the training sample data sets of multiple card games.
[0056] It should be noted that the number of training sample data for each card game in the mixed training sample data set of the embodiments of the present application is not the same. For example, the embodiments of the present application can train a language model using the training sample data of each card game respectively to determine the amount of sample data required when the performance of each card game converges. Then, data sampling is performed on the corresponding card game based on the required amount of sample data (i.e., using the teacher model and the opponent model to simulate the game). Finally, the mixed training sample data is obtained based on the sum of the training sample data of each card data.
[0057] In step S103, the pre-constructed initial decision model is trained using the mixed training sample data set, and the trained initial decision model is evaluated for generality using a preset benchmark to obtain an evaluation result. Then, the trained initial decision model is adjusted based on the evaluation result and the mixed general data set corresponding to the preset benchmark to obtain an actual decision model that meets the preset generality conditions.
[0058] In the actual execution process, the embodiments of the present application can train the initial decision model using the mixed training sample data set, thereby improving the versatility of the trained initial decision model in multiple card games. When facing different card games, there is no need to switch models.
[0059] Based on the trained initial decision model, the embodiments of the present application can also use preset benchmarks, such as three common evaluation benchmarks, MMLU-Pro, MATH-500, and HumanEval, to evaluate the capabilities of the trained initial decision model in three aspects: knowledge answering, mathematics, and programming, in order to determine the changes in the general capabilities of the initial decision model, that is, the evaluation result.
[0060] To restore the capabilities of the initial decision model after training in these three aspects, the embodiments of the present application can also collect open-source instruction datasets for three tasks: knowledge Q&A, mathematics, and programming, that is, the mixed general dataset corresponding to the preset benchmark, and restore the general capabilities of the initial decision model after training through fine-tuning on these three datasets to obtain the actual decision model, so that the actual decision model not only ensures universality in various card games but also has general capabilities such as knowledge Q&A, mathematics problem-solving, and code generation. That is to say, based on the above dual training, the embodiments of the present application can obtain an actual decision model with better universality.
[0061] Among them, the general conditions can be adjusted according to the actual requirements of the model. For example, the general conditions can be that the actual decision model can be applicable to various card games without the need to adjust parameters for different card games. The general conditions can also be that while outputting decision actions in card games, responses in aspects such as knowledge, mathematics, and programming can also be made.
[0062] Optionally, in an embodiment of the present application, before training the initial decision model that meets the preset general conditions using the training sample dataset, it further includes: defining the instructions and outputs of the language model using the observation-action pairs of each card game to obtain the initial decision model, where the instructions include the game description, state data, and output format description of each card game.
[0063] As a possible implementation, the embodiments of the present application can first convert the observation-action pairs of the card game into the input (instructions) and output of the language model. Among them, the instructions mainly consist of three parts: game description, state data, and output format description. The game description includes the game rules and the player's goals; the state data includes information such as the player's hand cards, community cards, historical action sequences, and legal actions; the output format specifies that the model should output actions in the target format, and the specific format can be set accordingly according to actual requirements.
[0064] Optionally, in an embodiment of the present application, training the pre-constructed initial decision model using the mixed training sample dataset includes: calculating the cross-entropy loss for the outputs in the mixed training sample dataset; using the cross-entropy loss as the loss function to train the initial decision model using the loss function to obtain the initial decision model after training; where the expression of the loss function is:
[0065]
[0066] Furthermore, in the model training of the embodiments of the present application, for each instruction-output pair, only the cross-entropy loss of the output is calculated for model training.
[0067] Among them, the expression of the loss function is:
[0068]
[0069] Among them, i is the sample index, o i represents an instruction, a i represents an output, t represents the current prediction step, that is, the current prediction is the t-th character in the output, and p represents the probability of each character in the output predicted by the initial decision model. represents the loss function.
[0070] Optionally, in an embodiment of the present application, it further includes: generating a simulated instruction using any one of multiple card games to simulate the game of any one of the card games; inputting the simulated instruction into the actual decision model to output corresponding simulated decision actions; updating the simulated instruction based on the simulated decision actions until the game of any one of the card games is completed and the game result of the game is obtained; optimizing the actual decision model using the game result.
[0071] After the model training is completed, the embodiment of the present application can evaluate the actual decision model in the game simulator. First, deploy and start the language model server for the actual decision model. Then start the game simulator and the language model client. The evaluation of the actual decision model is achieved by calling the server through the language model client. In addition, the embodiment of the present application can also evaluate the general ability of the actual decision model on three benchmarks: MMLU-Pro, MATH-500, and HumanEval.
[0072] Furthermore, according to the evaluation results, the embodiment of the present application can optimize the actual decision model.
[0073] Combined with Figures 2 to 4 , the working principle of the decision method for the card game in the embodiment of the present application is described with an embodiment.
[0074] As Figure 2 shown, the embodiment of the present application may include the following steps:
[0075] Step S201: Game selection.
[0076] Embodiments of the present application can improve the decision-making ability of language in multiple card games by fine-tuning a language model with high-quality battle data from multiple card games. In terms of card game selection, embodiments of the present application can consider the popularity, complexity, and availability of high-quality models or data of the game. Based on these three aspects, embodiments of the present application have selected eight card games: Dou Di Zhu, Guan Dan, Japanese Mahjong (Riichi Mahjong), Uno, Gin Rummy, Leduc Poker, Limit Hold'em, and No-Limit Hold'em. Among them, the complexity of the eight card games is shown in Table 1, where Table 1 is a table of card game complexity. It can be seen that most games have a high complexity.
[0077] Table 1
[0078] Card game Number of information sets Average information set size Action space DouDizhu 10^53~10^83 10^23 10^4 GuanDan 10^118 10^36 10^6 Riichi Mahjong 10^121 10^48 10^2 Uno 10^163 10^10 10^1 Gin Rummy 10^52 10^9 10^1 Leduc Hold’em 10^2 10^2 10^0 Limit Texas Hold’em 10^14 10^3 10^0 No-limit Texas Hold’em 10^162 10^3 10^4
[0079] Among the above eight card games, Dou Di Zhu and Guan Dan have publicly accessible strong game AIs as teacher models; Japanese Mahjong has publicly available human expert player data.
[0080] Step S202: Game instruction data synthesis.
[0081] The game instruction data synthesis aims to synthesize instruction data available for the language model, mainly including three steps: game trajectory data generation, trajectory data filtering, and data format conversion. The following is an introduction to each step:
[0082] Game trajectory data generation: Embodiments of the present application generate game interaction data by having the teacher model compete with the opponent. The teacher model and opponent information for each game can be shown in Table 2, where Table 2 is a table of data synthesis information.
[0083] Table 2
[0084]
[0085] For Dou Di Zhu, the embodiments of this application can use DouZero as the teacher model and a rule-based model as the opponent model. For Guan Dan, the embodiments of this application can use DanZero as the teacher model and a rule-based model as the opponent model. For Japanese Mahjong, the embodiments of this application downloaded the game data of human professional players in 2020 from the Tenhou platform. The opponent model is Mortal, a powerful Mahjong AI, which is only used during evaluation. For Uno and Gin Rummy, the embodiments of this application use the rule models in RLCard as the teacher models and a random model as the opponent. For Leduc Poker, Limit Hold'em, and No-Limit Hold'em, the embodiments of this application use the DQN model trained by RLCard as the teacher model. According to the complexity of different games, the embodiments of this application conduct a different number of game rounds for each game, as shown in Table 2.
[0086] Trajectory data filtering: The embodiments of this application can filter the generated data according to two criteria to obtain high-quality data. First, only the observation and action data of winning players are retained. In addition, each observation-action pair of a player is regarded as a separate data instance. Second, for all eight card games, the environment provides legal action options for each action. The embodiments of this application only retain the data samples with the number of legal action options greater than 1. The amount of data obtained after filtering is shown in Table 2. It can also be seen from the table that the number of legal candidate actions for each sample of Dou Di Zhu, Guan Dan, and Mahjong exceeds that of the other five card games, making these three card games relatively more complex.
[0087] Data format conversion: In order to perform instruction fine-tuning on the model, the embodiments of this application design prompts for each game to convert the observation-action pairs into instructions and corresponding outputs. The instructions mainly consist of three parts: game description, state data, and output format description. The game description includes the game rules and the player's goal. The state data includes information such as the player's hand cards, community cards, historical action sequence, and legal actions. The output format specifies that the model should output actions in JSON format. Among them, the instructions for Dou Di Zhu Figure 3 are shown.
[0088] The embodiments of this application can use the data obtained above to train language models respectively to determine the amount of data required when the performance of each game converges. According to this preliminary experiment, the embodiments of this application sampled 700,000, 950,000, 650,000, 200,000, 50,000, 250,000, 200,000, and 100,000 for 8 card games respectively to obtain a mixed game dataset of 3.1 million, that is, a mixed training sample dataset. The mixed training sample dataset is used to improve the decision-making ability of the language model in card games to obtain an actual decision-making model.
[0089] Step S203: General benchmark and general data preparation.
[0090] To evaluate the changes in the general capabilities of the language model, three commonly used evaluation benchmarks, MMLU-Pro, MATH-500, and HumanEval, are selected in the embodiments of this application to evaluate the changes in the capabilities of the language model in three aspects: knowledge Q&A, mathematics, and programming. At the same time, in order to restore the capabilities of the model in these three aspects, the embodiments of this application collect open-source instruction datasets for the three tasks of knowledge Q&A, mathematics, and programming. The general capabilities of the language model are restored through fine-tuning on these three datasets to obtain an actual decision-making model after training. Specifically, in the embodiments of this application, 20,000, 20,000, 20,000, and 8,000 samples are respectively sampled for knowledge, mathematics, programming, and games to obtain a mixed general dataset of 68,000. Fine-tuning is performed on this mixed general dataset to restore the general capabilities of the model.
[0091] Step S204: Model training.
[0092] The initial decision-making model after training that simultaneously masters multiple card games is obtained by fine-tuning the language model through the mixed training sample dataset synthesized in step S202; the capabilities of the model in the three aspects of knowledge Q&A, mathematics, and programming are restored by further fine-tuning this model using the mixed general dataset in step S203 to obtain the actual decision-making model.
[0093] During model training, for each instruction-output pair, only the cross-entropy loss is calculated for the output for model training. In the embodiments of this application, multiple open-source language models can be used as the initial model. Specifically, in the embodiments of this application, three models, Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, and Glm-4-9B-Chat, are tested. These basic models all adopt a structure similar to the Transformer decoder, mainly including multiple stacked self-attention layers, and each self-attention layer mainly includes a multi-head attention layer and a feed-forward layer. Given the input o i containing game observation data i and the corresponding action output a
[0094]
[0095] Considering that the number of parameters of the language model is large and the cost of full-parameter training is high. In the embodiments of this application, the LoRA technique is used to fine-tune the language model. LoRA introduces a small number of additional new parameters. During fine-tuning, the main model is fixed and only the additional parameters are updated.
[0096] Step S205: Model evaluation.
[0097] After the model training is completed, the language model can be evaluated in the game simulator. First, deploy the trained model and start the language model server. Then start the game simulator and the language model client. The game simulator is responsible for the entire game running logic and notifies each player to make actions in a timely manner. The language model client, as a player, calls the language model server to obtain the actions corresponding to the observation data. Finally, the game win / loss or game score is used as the evaluation metric for the language model. Additionally, the general capabilities of the language model are evaluated on three benchmarks: MMLU-Pro, MATH-500, and HumanEval.
[0098] Based on the above steps, relevant experiments were conducted in the embodiments of this application. Among them, the experimental results can be shown in Table 3, which is the evaluation result table of the card game.
[0099] Table 3
[0100]
[0101] The embodiments of this application evaluated the proposed method with some API models, and the results are shown in Table 3. Among them, the top 5 rows are the results of the API models, and the bottom 3 rows are the results of the proposed method. It can be seen from the table that through fine-tuning on the mixed game data, the performance of the language model has been greatly improved on multiple games, especially in Dou Di Zhu, Guan Dan, and Japanese Mahjong.
[0102] Figure 4 For the evaluation results of the model on the general benchmarks. Among them, in Figure 4 , MMLU-Pro, MATH-500, and HumanEval are respectively used to evaluate the capabilities of the model in three aspects: knowledge Q&A, mathematics, and programming. In the figure, Llama-3.1-8B-Instruct and Glm-4-9B-Chat are two open-source language models, and mix represents the model trained with the game mixed data, that is, the initial decision-making model after training; mix-general represents the model trained first with the game mixed data and then with the general mixed data, that is, the actual decision-making model.
[0103] From Figure 4 it can be seen that after training with the mixed training sample data, the general capabilities of the initial decision-making model have decreased significantly. However, after fine-tuning on the mixed general data, the general capabilities of the actual decision-making model have been greatly restored.
[0104] According to the decision-making method of the card game proposed in the embodiments of the present application, multiple card games can be screened, and the game trajectory data of each card game can be obtained to generate a mixed training sample data set by using the game trajectory data, so as to explore the mutual influence between different card games and the influence on the general ability of the model. The pre-constructed initial decision-making model is trained by using the mixed training sample data set to obtain a decision-making model that can be widely used in multiple card games, and the trained initial decision-making model is evaluated for generality by using a preset benchmark to obtain an evaluation result, and the trained initial decision-making model is adjusted based on the evaluation result and the mixed general data set corresponding to the preset benchmark to restore the general ability of the trained initial decision-making model, so that the actual decision-making model can be not limited to the decision-making of card games, thereby improving the upper limit of the decision-making ability. Thus, the problem in the related art that the generalization ability between multiple games is poor and the method based on prompts can only utilize the inherent knowledge of the language model, resulting in the decision-making ability being limited by the model, is solved.
[0105] Next, a decision-making device for a card game proposed in the embodiments of the present application will be described with reference to the accompanying drawings.
[0106] Figure 5 It is a block diagram of a decision-making device for a card game according to an embodiment of the present application.
[0107] As Figure 5 shown, the decision-making device 10 for the card game is applied to the model construction stage. Among them, the device 10 includes: an acquisition module 101, a generation module 102, and a training module 103.
[0108] Specifically, the acquisition module 101 is used to acquire multiple card games that meet the preset screening conditions.
[0109] The generation module 102 is used to obtain the game trajectory data of each card game to generate a mixed training sample data set by using the game trajectory data.
[0110] The training module 103 is used to train the pre-constructed initial decision-making model by using the mixed training sample data set, and evaluate the generality of the trained initial decision-making model by using a preset benchmark to obtain an evaluation result, and adjust the trained initial decision-making model based on the evaluation result and the mixed general data set corresponding to the preset benchmark to obtain an actual decision-making model that meets the preset generality conditions.
[0111] Optionally, in an embodiment of the present application, the generation module 102 includes: a matching unit, a generation unit, a first screening unit, and a first construction unit.
[0112] Among them, the matching unit is used to match a corresponding teacher model, opponent model, and number of games for each card game based on the game characteristics of each card game.
[0113] A generation unit configured to generate game trajectory data by using multiple game data of each card game in the number of game rounds by a teacher model and an opponent model.
[0114] A first screening unit configured to screen the game trajectory data to obtain decision data of the winning party.
[0115] A first construction unit configured to construct a training sample data set corresponding to each card game based on the decision data of the winning party, and generate a mixed training sample data set based on the training sample data set corresponding to each card game.
[0116] Optionally, in an embodiment of the present application, the generation module 102 further includes: a first acquisition unit, a second acquisition unit, a determination unit, and a second construction unit.
[0117] The first acquisition unit is configured to acquire game descriptions of each card game.
[0118] The second acquisition unit is configured to acquire data instances of each observation-action pair in the decision data of the winning party in each game round.
[0119] The determination unit is configured to determine the compliance of the data instance of each observation-action pair based on the game description to obtain a determination result.
[0120] The second construction unit is configured to screen the decision data of the winning party by using the determination result to obtain compliant decision data, and construct a training sample data set corresponding to each card game by using the compliant decision data, and generate a mixed training sample data set based on the training sample data set corresponding to each card game.
[0121] Optionally, in an embodiment of the present application, the decision device 10 of the card game further includes: a definition module.
[0122] The definition module is configured to define instructions and outputs of a language model by using the observation-action pairs of each card game to obtain an initial decision model, where the instructions include game descriptions, state data, and output format descriptions of each card game.
[0123] Optionally, in an embodiment of the present application, the training module 103 includes: a calculation unit and a training unit.
[0124] The calculation unit is configured to calculate cross-entropy loss for the outputs in the mixed training sample data set.
[0125] The training unit is configured to use the cross-entropy loss as a loss function to train the initial decision model by using the loss function to obtain a trained initial decision model.
[0126] Wherein, the expression of the loss function is:
[0127]
[0128] Among them, i represents the sample index, and o i represents the instruction, and a i represents the output, t represents the current predicted step number, p represents the probability of each character in the output predicted by the initial decision model, represents the loss function.
[0129] Optionally, in an embodiment of the present application, the decision-making device 10 of the card game further includes: a simulation module, a calculation module, an update module, and an optimization module.
[0130] Among them, the simulation module is used to generate a simulation instruction by using any one of multiple card games to simulate the game of any one of the card games.
[0131] The calculation module is used to input the simulation instruction into the actual decision model to output a corresponding simulated decision action.
[0132] The update module is used to update the simulation instruction based on the simulated decision action until the game of any one of the card games is completed, and the game result of the game is obtained.
[0133] The optimization module is used to optimize the actual decision model by using the game result.
[0134] It should be noted that the foregoing explanation of the embodiment of the decision-making method for the card game also applies to the decision-making device of the card game in this embodiment, and will not be repeated here.
[0135] According to the decision-making device of the card game proposed in the embodiment of the present application, multiple card games can be screened, and the game trajectory data of each card game can be obtained, so as to generate a mixed training sample data set by using the game trajectory data, to explore the mutual influence between different card games and the influence on the general ability of the model, train the pre-constructed initial decision model by using the mixed training sample data set to obtain a decision model that can be generalized to multiple card games, and use a preset benchmark to evaluate the generality of the trained initial decision model to obtain an evaluation result, and adjust the trained initial decision model based on the evaluation result and the mixed general data set corresponding to the preset benchmark to restore the general ability of the trained initial decision model, so that the actual decision model can be not limited to the decision of the card game, thereby improving the upper limit of the decision-making ability. Thus, the problem in the related art that the generalization ability between multiple games is poor, and the method based on prompts can only utilize the inherent knowledge of the language model, resulting in the decision-making ability being limited by the model is solved.
[0136] The above is the description of the embodiment of the present application in the model construction stage. The following elaborates on the embodiment of the present application in the model usage stage.
[0137] Specifically, Figure 6 It is a schematic flowchart of a decision-making method for a card game provided by an embodiment of the present application.
[0138] As Figure 6 shown, the decision-making method for the card game is applied to the model usage stage. Among them, the method includes the following steps:
[0139] In step S601, obtain the current state information and game description of the card game to be decided.
[0140] In step S602, generate corresponding decision instructions based on the current state information and game description.
[0141] In step S603, input the decision instructions into a pre-constructed decision model to output corresponding decision actions, where the decision model is trained by instruction data and game trajectories of multiple card games.
[0142] According to the decision-making method for the card game proposed by the embodiment of the present application, multiple card games can be screened, and the game trajectory data of each card game can be obtained to generate a mixed training sample data set using the game trajectory data to explore the mutual influence between different card games and the influence on the general ability of the model. The pre-constructed initial decision model is trained using the mixed training sample data set to obtain a decision model that can be generalized to multiple card games, and the trained initial decision model is evaluated for generality using a preset benchmark to obtain an evaluation result, and the trained initial decision model is adjusted based on the evaluation result and the mixed general data set corresponding to the preset benchmark to restore the general ability of the trained initial decision model, so that the actual decision model can be not limited to the decision of card games, thereby improving the upper limit of the decision-making ability. Thus, it solves the problem in the related art that the generalization ability between multiple games is poor, and the method based on prompts can only utilize the inherent knowledge of the language model, resulting in the decision-making ability being limited by the model.
[0143] Secondly, refer to the drawings to describe the decision-making device for the card game proposed by the embodiment of the present application.
[0144] Figure 7 It is a block diagram of the decision-making device for the card game of the embodiment of the present application.
[0145] As Figure 7 shown, the decision-making device 20 for the card game is applied to the model usage stage. Among them, the device 20 includes: an acquisition module 201, a generation module 202, and a decision module 203.
[0146] Specifically, an acquisition module 201 is configured to acquire the current state information and game instructions of a card game to be decided.
[0147] A generation module 202 is configured to generate corresponding decision instructions based on the current state information and the game instructions.
[0148] A decision module 203 is configured to input the decision instructions into a pre-constructed decision model to output corresponding decision actions, where the decision model is trained by instruction data of multiple card games and corresponding game trajectories.
[0149] It should be noted that the foregoing explanation of the embodiments of the decision method for card games also applies to the decision device for card games in this embodiment, and will not be elaborated here.
[0150] According to the decision device for card games provided by the embodiments of the present application, multiple card games can be screened, and the game trajectory data of each card game can be acquired to generate a mixed training sample data set by using the game trajectory data, so as to explore the mutual influence between different card games and the influence on the general ability of the model. The pre-constructed initial decision model is trained by using the mixed training sample data set to obtain a decision model that can be generalized to multiple card games, and the trained initial decision model is evaluated for generality by using a preset benchmark to obtain an evaluation result, and the trained initial decision model is adjusted based on the evaluation result and the mixed general data set corresponding to the preset benchmark to restore the general ability of the trained initial decision model, so that the actual decision model can be not limited to the decision of card games, thereby improving the upper limit of the decision-making ability. Thus, the problem in the related art that the generalization ability between multiple games is poor, and the method based on prompts can only utilize the inherent knowledge of the language model, resulting in the decision-making ability being limited by the model is solved.
[0151] Figure 8 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device may include:
[0152] A memory 801, a processor 802, and a computer program stored on the memory 801 and executable on the processor 802.
[0153] When the processor 802 executes the program, it implements the decision method for card games provided in the foregoing embodiment.
[0154] Further, the electronic device further includes:
[0155] A communication interface 803 for communication between the memory 801 and the processor 802.
[0156] The memory 801 is used to store a computer program executable on the processor 802.
[0157] The memory 801 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.
[0158] If the memory 801, the processor 802, and the communication interface 803 are implemented independently, the communication interface 803, the memory 801, and the processor 802 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity in representation, Figure 8 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0159] Optionally, in a specific implementation, if the memory 801, the processor 802, and the communication interface 803 are integrated on a single chip, the memory 801, the processor 802, and the communication interface 803 can communicate with each other through an internal interface.
[0160] The processor 802 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0161] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the decision-making method of the card game as described above is implemented.
[0162] This application embodiment also provides a computer program product, including a computer program. When the computer program is executed by a processor, the decision-making method of the card game provided by the embodiments of this application is implemented.
[0163] In the description of this specification, the descriptions with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0164] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0165] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or N executable instructions for realizing a customized logic function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application belong.
[0166] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or N wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0167] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0168] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0169] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0170] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A decision-making method for a card game, characterized in that: Applied to the model building stage, wherein the method comprises the following steps: Get a variety of card games that meet the preset filtering criteria; Acquire game track data of each card game to generate a mixed training sample data set using the game track data; The pre-constructed initial decision model is trained using the mixed training sample data set, and the universality of the trained initial decision model is evaluated using a preset benchmark to obtain an evaluation result, and the trained initial decision model is adjusted based on the evaluation result and the mixed universal data set corresponding to the preset benchmark to obtain an actual decision model that meets the preset universality conditions.
2. The method according to claim 1, characterized in that: The obtaining of game track data of each card game to generate a mixed training sample data set using the game track data includes: Based on the game characteristics of each card game, matching a corresponding teacher model, opponent model and number of games for each card game; Generate the game trajectory data using multiple game data of the teacher model and the opponent model of each card game at the game number; Filtering the game trajectory data to obtain decision data of the winner; A training sample data set corresponding to each card game is constructed based on the decision data of the winning party, so as to generate the mixed training sample data set based on the training sample data set corresponding to each card game.
3. The method according to claim 2, characterized in that After obtaining the decision data of the winner, it also includes: Obtaining game instructions for each of the card games; Get the data instance of each observation-action pair in the decision data of the winner of each game; Determining compliance of the data instance of each observation-action pair based on the game instructions to obtain a determination result; The decision data of the winning party is screened using the determination result to obtain compliant decision data, and the compliant decision data is used to construct a training sample data set corresponding to each card game, so as to generate the mixed training sample data set based on the training sample data set corresponding to each card game.
4. The method according to claim 1, characterized in that: Before using the training sample data set to train an initial decision model that meets the preset general conditions, the method further includes: The observation-action pairs of each card game are used to define the instructions and outputs of the language model to obtain the initial decision model, wherein the instructions include game instructions, state data and output format instructions of each card game.
5. The method according to claim 4, characterized in that The using the mixed training sample data set to train the pre-built initial decision model includes: Calculating cross entropy loss for the outputs in the mixed training sample data set; Using the cross entropy loss as a loss function, and using the loss function to train the initial decision model to obtain the trained initial decision model; Among them, the expression of the loss function is: Where i represents the sample index, o i Indicates the instruction, a i represents the output, t represents the number of steps currently predicted, p represents the probability of each character in the output predicted by the initial decision model, Represents the loss function.
6. The method according to claim 4, characterized in that Also includes: Using any card game among the plurality of card games to generate simulation instructions to simulate a game of the any card game; Inputting the simulation instruction into the actual decision model to output a corresponding simulation decision action; updating the simulation instruction based on the simulation decision action until the game of any card game is completed and the game result of the game is obtained; The actual decision model is optimized using the game results.
7. A decision-making method for a card game, characterized in that: The decision-making method of the card game as claimed in any one of claims 1 to 6 is applied to the model use stage, wherein the method comprises the following steps: Get the current status information and game description of the card game to be decided; Generate corresponding decision instructions based on the current state information and the game instructions; The decision instruction is input into a pre-built decision model to output a corresponding decision action, wherein the decision model is trained by instruction data and game trajectories of a variety of card games.
8. A decision-making device for a card game, characterized in that: Applied to the model building stage, wherein the device comprises: An acquisition module, used to acquire a variety of card games that meet preset screening conditions; A generation module, used to obtain game track data of each card game, so as to generate a mixed training sample data set using the game track data; A training module is used to train a pre-constructed initial decision model using the mixed training sample data set, and to evaluate the universality of the trained initial decision model using a preset benchmark to obtain an evaluation result, and to adjust the trained initial decision model based on the evaluation result and the mixed universal data set corresponding to the preset benchmark to obtain an actual decision model that meets the preset universality conditions.
9. A decision-making device for a card game, characterized in that: Applied to the model use stage, wherein the device comprises: The acquisition module is used to obtain the current status information and game description of the card game to be decided; A generation module, used for generating corresponding decision instructions based on the current state information and the game instructions; The decision module is used to input the decision instruction into a pre-built decision model to output a corresponding decision action, wherein the decision model is trained by instruction data of multiple card games and corresponding game trajectories.
10. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the card game decision method according to any one of claims 1 to 6 or 7.