A method, device, storage medium and electronic device for training a game decision-making model
By using the data and decision information in historical game videos, the game decision model is trained, and the problem of insufficient speed and accuracy of game decision generation in the existing technology is solved, and more efficient game decisions are achieved.
Patent Information
- Application Number
- CN202411717226.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-11-27
AI Technical Summary
The prior art is difficult to effectively train game decision models, resulting in insufficient speed and accuracy of game decision generation.
By obtaining the historical game video of the sample players, extracting game data and decision information for the specified time period, inputting a general large language model as a training sample, annotating and training, and generating a model for game decision making.
The generation speed and accuracy of game decision models are improved, so that the trained models can make decisions more efficiently in the game environment.
Smart Images

Figure CN119202730B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and particularly to a method, device, storage medium, and electronic device for training a game decision-making model. Background Art
[0002] With the continuous development of technology, machine learning is applied more and more widely, especially in decision-making games, such as a real-time strategy game and a card game.
[0003] Currently, when playing a decision-making game, a player can directly select a machine learning model to play the game, that is, a pre-trained machine learning model directly replaces the player. The machine learning model can determine the game decision for the next moment according to the current game environment, and directly control the game character corresponding to the player according to the game decision. Generally, the performance of the machine learning model will affect the generation speed and accuracy of the game decision, thus affecting the development and result of the game. Therefore, how to train a game decision-making model to improve the generation speed and accuracy of the game decision is a very important issue.
[0004] Based on this, this specification provides a method for training a game decision-making model. Summary of the Invention
[0005] This specification provides a method, device, storage medium, and electronic device for training a game decision-making model to partially solve the above problems existing in the prior art.
[0006] This specification adopts the following technical solutions:
[0007] This specification provides a method for training a game decision-making model, including:
[0008] Obtain the historical game videos of sample players;
[0009] Extract data from the historical game videos, determine the game data of the sample players within a specified time period as training samples, and determine the first decision information corresponding to the decisions made by the sample players in the game states corresponding to the training samples as the first annotation of the training samples;
[0010] Determine the first prompt text corresponding to the training samples, and input the first prompt text and the training samples into a general large language model to determine the first information output by the general large language model;
[0011] Use the first annotation and the first information as the second annotation of the training samples;
[0012] Train the game decision-making model to be trained according to the training samples and the second annotation; the trained game decision-making model is used to determine game decisions according to the game data of the player to be decided; wherein, the game decision-making model is a large language model, and the number of parameters of the game decision-making model is smaller than that of the general large language model.
[0013] Optionally, extract data from the historical game video to determine the game data of the sample player within a specified time period as training samples, and determine the first decision information corresponding to the decision made by the sample player in the game state corresponding to the training samples as the first annotation of the training samples, specifically including:
[0014] Use a pre-set data extraction tool to extract data from the historical game video according to a preset time length to determine the game data of the sample player within several time periods;
[0015] From each game data, determine the game data of the sample player within a specified time period as training samples, and determine the first decision information corresponding to the decision made by the sample player in the game state corresponding to the training samples as the first annotation of the training samples.
[0016] Optionally, from each game data, determine the game data of the sample player within a specified time period as training samples, and determine the first decision information corresponding to the decision made by the sample player in the game state corresponding to the training samples, specifically including:
[0017] Determine the information content corresponding to each game data;
[0018] Filter each game data according to a preset filtering rule according to each information content;
[0019] From the filtered game data, determine the game data of the sample player within a specified time period as training samples, and determine the first decision information corresponding to the decision made by the sample player in the game state corresponding to the training samples as the first annotation of the training samples.
[0020] Optionally, input the first prompt text and the training samples into the general large language model to determine the first information output by the general large language model, specifically including:
[0021] Convert the training samples according to a preset template to determine the first text;
[0022] Concatenate the first prompt text and the first text to obtain the second text;
[0023] Input the second text into a general large language model to determine the first information output by the general large language model.
[0024] Optionally, using the first annotation and the first information as the second annotation of the training sample specifically includes:
[0025] Determine the decision information included in the first information and use it as the second decision information;
[0026] Determine the difference between the first annotation and the second decision information;
[0027] When the difference is greater than a preset threshold, use the information in the first information other than the second decision information and the first annotation as the second annotation of the training sample;
[0028] When the difference is not greater than the preset threshold, use the first information and the first annotation as the second annotation of the training sample.
[0029] Optionally, the first prompt text includes at least a game development status prompt and a game decision prompt;
[0030] Determine the first prompt text corresponding to the training sample, specifically including:
[0031] Display the training sample to the operator;
[0032] In response to the input operation of the operator, determine the first prompt text corresponding to the training sample.
[0033] Optionally, the method further includes:
[0034] Obtain the game data of the player to be decided and use it as the first data;
[0035] Determine the second prompt text corresponding to the first data;
[0036] Input the first data and the second prompt text into the trained game decision model to determine the prediction information output by the game decision model;
[0037] Determine the game decision instruction according to the prediction information;
[0038] Control the game character corresponding to the player to be decided according to the game decision instruction.
[0039] This specification provides a game decision model training device, including:
[0040] An acquisition module for acquiring the historical game videos of sample players;
[0041] An extraction module for extracting data from the historical game video, determining the game data of the sample player within a specified time period as training samples, and determining the first decision information corresponding to the decisions made by the sample player in the game state corresponding to the training samples as the first annotation of the training samples;
[0042] A first determination module for determining the first prompt text corresponding to the training samples, inputting the first prompt text and the training samples into a general large language model, and determining the first information output by the general large language model;
[0043] A second determination module for using the first annotation and the first game information as the second annotation of the training samples;
[0044] A training module for training a game decision model to be trained according to the training samples and the second annotation; the trained game decision model is used to determine game decisions based on the game data of the player to be decided; wherein, the game decision model is a large language model.
[0045] This specification provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above-mentioned game decision model training method.
[0046] This specification provides an electronic device including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor implements the above-mentioned game decision model training method when executing the program.
[0047] At least one of the technical solutions adopted in this specification can achieve the following beneficial effects:
[0048] The game decision model training method provided in this specification obtains the historical game video of the sample player, extracts data from the historical game video, determines the game data of the sample player within a specified time period as training samples, and determines the first decision information corresponding to the decisions made by the sample player in the game state corresponding to the training samples as the first annotation of the training samples. Determine the first prompt text corresponding to the training samples, input the first prompt text and the training samples into a general large language model, and determine the first information output by the general large language model. Use the first annotation and the first information as the second annotation of the training samples, and train the game decision model to be trained according to the training samples and the second annotation.
[0049] As can be seen from the above method, when training the game decision-making model in this application, the historical game videos of sample players are obtained, data extraction is performed on the historical game videos, the game data of the sample players within a specified time period is determined and used as training samples, and the first decision-making information corresponding to the decisions made by the sample players in the game states corresponding to the training samples is determined and used as the first annotation of the training samples. The first prompt text corresponding to the training samples is determined, and the first prompt text and the training samples are input into the general large language model to determine the first information output by the general large language model. The first annotation and the first information are used as the second annotation of the training samples, and the game decision-making model to be trained is trained according to the training samples and the second annotation, so that the trained game decision-making model can be used to determine game decisions based on the game data of the players to be decided, improving the generation speed and accuracy of game decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The drawings described herein are used to provide a further understanding of this specification, form a part of this specification, and the schematic embodiments and descriptions thereof are used to explain this specification and do not constitute an improper limitation to this specification. In the drawings:
[0051] Figure 1 is a schematic flowchart of a method for training a game decision-making model provided in this specification;
[0052] Figure 2 is a schematic diagram of training a game decision-making model to be trained provided in this specification;
[0053] Figure 3 is a schematic diagram of the training process of a game decision-making model provided in this specification;
[0054] Figure 4 is a schematic diagram of the application of a game decision-making model provided in this specification;
[0055] Figure 5 is a schematic diagram of a game decision-making model training device provided in this specification;
[0056] Figure 6 corresponding to Figure 1 is a schematic diagram of the structure of an electronic device. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of them. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0058] The following will, with reference to the drawings, elaborate on the technical solutions provided by each embodiment of this specification.
[0059] Figure 1 The following is a schematic flowchart of a method for training a game decision-making model provided in this specification, including the following steps:
[0060] S100: Obtain the historical game videos of sample players.
[0061] In this specification, the device for training the game decision-making model can obtain the historical game videos of sample players. Among them, the device for training the game decision-making model can be a server or an electronic device such as a desktop computer or a laptop computer. For the sake of convenience of description, the following will only use the server as the execution entity to illustrate the method for training the game decision-making model provided in this specification.
[0062] The above sample players are players of decision-making type games. The decision-making type games can be real-time strategy type games, or card type games. Of course, they can also be other decision-making type games, which are not specifically limited in this specification. The above sample players can be real players, such as ordinary players or professional players, or intelligent agents for executing decision-making type games, that is, any existing machine learning model for executing decision-making type games, which are not specifically limited in this specification. The above historical game videos are historical replay data of sample players when playing decision-making type games, that is, historical replay files. The historical replay data can be replay data of any game in the history of sample players, and the historical replay data can be generated by the client of the decision-making type game played by the sample players. Of course, the above historical game videos can also be video data of sample players when playing decision-making type games, such as screen recording data, which are not specifically limited in this specification.
[0063] The above historical game videos can be pre-collected by the server. Specifically, the server can pre-collect the historical game videos of players from data sources such as the official website, client, and forum of decision-making type games. Of course, the server can also pre-collect the historical game videos of players from other data sources including historical game videos, and this specification does not make specific limitations. The data volume of the above historical game videos can be 10,000. Of course, it can also be other volumes, and this specification does not make specific limitations.
[0064] S102: Extract data from the historical game videos, determine the game data of the sample player within a specified time period as a training sample, and determine the first decision information corresponding to the decision made by the sample player in the game state corresponding to the training sample as the first annotation of the training sample.
[0065] In this specification, the server can extract data from historical game videos, determine the game data of the sample player within a specified time period as a training sample, and determine the first decision information corresponding to the decision made by the sample player in the game state corresponding to the training sample as the first annotation of the training sample. Among them, the specified time period is pre-set. The game data is all the data related to the game of the sample player within the specified time period. The game data can include the game time and the game resource information of the sample player, and can also include historical decision information, and can also include other information of the other player who plays against the sample player. This specification does not make specific limitations. The above game resource information can be the information of the resources owned by the sample player. For example, taking a real-time strategy type game as an example, the game resource information can be the information such as minerals, gas, worker supply, worker units, buildings, population, and technology owned by the sample player. The above historical decision information can be the decision information corresponding to the decisions that the sample player has made within the specified time period. The decision is made by the sample player in the game state corresponding to the game data of the previous time period within the specified time period. The decision information can include resource management decisions. For example, taking a real-time strategy type game as an example, the decision information can include the information corresponding to decisions such as resource collection, economic expansion, building construction, technology research and development, unit production, and base defense. The game state is the game environment, that is, the game environment corresponding to the game data. The above other information of the other player who plays against the sample player can be the information obtained by the sample player as the game develops. The other information can include the game resource information and historical decision information of the other player.
[0066] Specifically, the server can use a pre-set data extraction tool to extract data from historical game videos according to a pre-set time length to determine the game data of sample players within a number of time periods. Then, from each game data, determine the game data of sample players within a specified time period and use it as a training sample, and determine the first decision information corresponding to the decisions made by the sample players in the game state corresponding to the training sample and use it as the first annotation of the training sample. Among them, the above data extraction tool can be a pre-set algorithm or program. Of course, it can also be any existing data extraction tool, and this specification does not make specific limitations. The above time length is pre-set, and this time length can be 1 second or 5 seconds, and this specification does not make specific limitations.
[0067] The above specified time period can be any one of the above-mentioned several time periods. Therefore, when determining the game data of sample players within the specified time period from each game data and using it as a training sample, the server can randomly select a time period from each game data and use the game data of the sample players within the selected time period as the training sample. Subsequently, the server can use the first decision information corresponding to the decisions made by the sample players in the game state corresponding to the game data of the sample players within the selected time period as the first annotation of the training sample. This first decision information is generally obtained from the game data of the next time period after the selected time period, that is, the game data of the next time period after the selected time period includes the decisions made by the sample players in the game state corresponding to the game data within the selected time period.
[0068] In addition, the above specified time period can also be any continuous preset number of time periods among the above-mentioned several time periods. The preset number is pre-set, and the preset number can be 5. Therefore, when determining the game data of sample players within the specified time period from each game data and using it as a training sample, the server can first divide each game data according to the preset number to obtain several data groups, and then randomly select one data group from each data group. Use the game data included in the selected data group as the training sample. Subsequently, the server can use the first decision information corresponding to the decisions made by the sample players in the game state corresponding to the game data included in the selected data group as the first annotation of the training sample. This first decision information is generally obtained from the game data of the next data group after the selected data group, that is, the game data of the next data group after the selected data group includes the decisions made by the sample players in the game state corresponding to the game data included in the selected data group. Of course, the server can also directly randomly determine the game data within a continuous preset number of time periods from each game data and use it as the training sample. Subsequently, the server can use the decision information included in the game data of the next time period after the determined continuous preset number of time periods, that is, the first decision information, as the first annotation of the training sample.
[0069] In this specification, in order to reduce the use of repeated training samples to train the game decision model to be trained, save training time, improve training efficiency, so that the game decision model to be trained can learn more knowledge, and further improve the accuracy of the trained game decision model. Therefore, the server can determine the information content corresponding to the game data in each time period. The information content refers to the difference between the game data in any time period and the game data in the previous time period of that time period. That is, the difference between the game data in any two consecutive time periods can be used as the information content corresponding to the game data in the later time period of any two consecutive time periods. When the information content of the game data in a certain time period is low, it means that the difference between the game data in this time period and the game data in the previous time period is small. On the contrary, when the information content of the game data in a certain time period is high, it means that the difference between the game data in this time period and the game data in the previous time period is large.
[0070] Based on this, when determining the game data of the sample player in the specified time period from each game data as the training sample, and determining the first decision information corresponding to the decision made by the sample player in the game state corresponding to the training sample, the server can determine the information content corresponding to each game data, and filter each game data according to the preset filtering rules according to each information content. From the filtered game data, determine the game data of the sample player in the specified time period as the training sample, and determine the first decision information corresponding to the decision made by the sample player in the game state corresponding to the training sample as the first annotation of the training sample.
[0071] Among them, the above information content represents the difference between the game data in any two consecutive time periods. Therefore, when determining the information content corresponding to each game data, the server can sequentially determine the game data in the previous time period of that time period for the game data in each time period as the first game data, and determine the difference between the game data in that time period and the first game data as the information content corresponding to the game data in that time period. The information content can be determined according to the similarity between the game data in that time period and the first game data. That is, the information content can be the difference between 1 and the above similarity.
[0072] The above filtering rule can be used to filter game data with an information content less than a preset value, where the preset value is a pre-set value. Therefore, when filtering each game data according to the preset filtering rule based on each information content, the server can, for each game data, filter the game data when the information content of the game data is less than the preset value, and use the game data as the filtered game data when the information content of the game data is not less than the preset value.
[0073] S104: Determine the first prompt text corresponding to the training sample, and input the first prompt text and the training sample into the general large language model to determine the first information output by the general large language model.
[0074] In this specification, the server can determine the first prompt text corresponding to the training sample, and input the first prompt text and the training sample into the general large language model to determine the first information output by the general large language model. Among them, the first prompt text can be pre-set, and the first information is the analysis result obtained by the general large language model analyzing the training sample under the prompt of the first prompt text. The content included in the first information is related to the prompt content included in the first prompt text. The general large language model can be a model pre-trained by the server, or any existing large language model. The general large language model can be a GPT (Generative Pre-trained Transformer) model. Of course, it can also be a large language model with other architectures, which is not specifically limited in this specification. The general large language model can be a model pre-trained on data in multiple fields.
[0075] The above first prompt text at least includes a game development status prompt and a game decision prompt. The game development status prompt is used to prompt the general large language model to analyze the training sample to determine the current game development status. Therefore, the first information can include the game development status, which is obtained by the general large language model analyzing the training sample. The game development status can include a game overview, game stage, game status quo, and important resources, etc. The game overview can be a brief overview of the game status corresponding to the training sample. The game stage can be one of the early stage, middle stage, and late stage of the game. The game status quo can include resource situations. For example, in the case of a real-time strategy game, the game status quo can include the situations of resources such as units, buildings, economy, and technology. The important resources can be the important resources among the resources owned by the sample player.
[0076] The above game decision prompts are used to prompt the general large language model to analyze training samples to determine the decisions that the sample player can execute in the game state corresponding to the training samples. Therefore, the above first information may also include game decision information, which may include the potential decisions that the sample player can execute, and may also include the potential decisions of the other player who plays the game against the sample player. The game decision information can be obtained by the general large model analyzing the training samples.
[0077] In addition, the above first prompt text may also include game background prompts and game task prompts. The game background prompts may include the background, gameplay rules, etc. of the game of the decision type, and the game task prompts may include game task information such as task names and task contents. Of course, the above first prompt text may also include the game character information corresponding to the sample player, and the game character information may include information such as character type, name, and gameplay.
[0078] In this specification, the languages of the above first prompt text and first information may be Chinese, English, or other types of languages, and this specification does not make specific limitations. For example, assuming the first prompt text is in Chinese, the first prompt text may be "You are a trained AI that can analyze and summarize Game A. You understand the nuances and strategies of the Protoss."
[0079] Based on the summary of multiple rounds in the game, we hope that you can structurally analyze the game process. Your analysis should include the following aspects:
[0080] 1. Game overview: Provide a brief overview of the current situation based on all rounds.
[0081] 2. Current game stage: Determine the game stage based on the information of all rounds. Is it the early stage, mid-stage, or late stage of the game?
[0082] 3. Our current situation: Describe our current situation in the following ways:
[0083] 3.1 Units and buildings: Analyze the status of our units and buildings. Ensure that sufficient assimilator construction is used for gas collection, especially in the expansion bases and main bases with unused vespene geysers.
[0084] 3.2 Economy: Evaluate our economic situation, including the collection and use of resources. As Protoss units, prioritize gas collection, and advanced technologies highly rely on gas.
[0085] 3.3 Technology: Describe the current status of our technology research. Which technologies have we unlocked so far? Analyze our technology tree and point out the available and potential upgrades or units. Ensure that gas-intensive technologies are supported by sufficient assimilator structures.
[0086] 4. Our strategy: Infer our potential strategies based on our current situation and information from all parties.
[0087] 5. The enemy's strategy: Infer the enemy's potential strategies based on the available information.
[0088] 6. Key information: Highlight the most important aspects that have a significant impact on the game in all rounds. "Of course, the above first prompt text can also be in English, so it will not be elaborated here.
[0089] In this specification, when determining the first prompt text corresponding to the training sample, the server can display the training sample to the operator, and in response to the operator's input operation, determine the first prompt text corresponding to the training sample. Of course, the server can also directly determine the first prompt text corresponding to the training sample in response to the operator's input operation, and this first prompt text is directly input by the operator and sent to the server.
[0090] In this specification, when inputting the first prompt text and the training sample into the general large language model to determine the first information output by the general large language model, the server can directly splice the first prompt text and the training sample, and then input the spliced result into the general large language model to determine the first information output by the general large language model.
[0091] S106: Use the first annotation and the first information as the second annotation of the training sample.
[0092] S108: Train the game decision model to be trained according to the training sample and the second annotation; the trained game decision model is used to determine game decisions based on the game data of the player to be decided; wherein, the game decision model is a large language model, and the number of parameters of the game decision model is smaller than the number of parameters of the general large language model.
[0093] In this specification, the server may use the first annotation and the first information as the second annotation of the training sample. Then, based on the training sample and the second annotation, the game decision-making model to be trained is trained. Among them, the above-mentioned second annotation may include all the contents of the first annotation and the first information. When the first information only includes the game development state, the second annotation includes the first annotation and the first information. However, when the first information only includes game decision-making information, or the first information includes both the game development state and game decision-making information, the server may determine the decision-making information included in the first information (i.e., the above-mentioned game decision-making information) and use it as the second decision-making information. Then, the difference between the first annotation and the second decision-making information is determined. When the difference is greater than the preset threshold, the information in the first information except the second decision-making information and the first annotation are used as the second annotation of the training sample. When the difference is not greater than the preset threshold, the first information and the first annotation are used as the second annotation of the training sample. The preset threshold is set in advance. By using the first decision-making information in the first annotation as the main basis and the content included in the first information as the supplement, the game decision-making model to be trained is trained, so that the game decision-making model can better learn how to generate game decisions and improve the accuracy of the output results of the game decision-making model. Moreover, for professional fields with strong professionalism, the proportion of data on specific scenarios in the training data of general large language models is relatively low, resulting in low decision-making effectiveness of these general large language models. Compared with general large language models, the above-mentioned game decision-making model is trained based on training samples determined from historical game videos of games based on decision types. The trained game decision-making model can learn knowledge in specific fields, and thus the game decisions output by the game decision-making model are more accurate.
[0094] In addition, the game decision-making model to be trained mentioned above may be a pre-trained large language model or an untrained large language model, and this specification does not make specific limitations. If the game decision-making model to be trained is a pre-trained large language model, then the process of the server training the game decision-making model to be trained based on the training sample and the second annotation is actually a process of fine-tuning the game decision-making model to be trained based on the training sample and the second annotation.
[0095] When training the game decision-making model to be trained based on the training sample and the second annotation, as Figure 2 shown, Figure 2This is a schematic diagram of training a game decision model to be trained. The server can input training samples and the first prompt text into the game decision model to be trained, and determine the first result output by the game decision model to be trained. Then, according to the difference between the first result and the second annotation, the game decision model to be trained is trained. Among them, the first result may include the game development state and may also include game decision information, which is not specifically limited in this specification. The content included in the first result is related to the prompt content included in the first prompt text. The process of inputting the training samples and the first prompt text into the game decision model to be trained and determining the first result output by the game decision model to be trained is similar to the process of inputting the first prompt text and training samples into the general large language model in step S104 above to determine the first information output by the general large language model, and will not be elaborated here.
[0096] In this specification, the game decision model to be trained or the trained game decision model is a large language model, and the number of parameters of the game decision model to be trained or the trained game decision model is smaller than the number of parameters of the general large language model. And the number of parameters of the game decision model to be trained or the trained game decision model can be 7B, and the game decision model to be trained or the trained game decision model is the Qwen2 model. Of course, it can also be a large language model with other architectures, which is not specifically limited in this specification. In addition, since the number of parameters of the trained game decision model is smaller than the number of parameters of the general large language model, this makes the game decision model easier to deploy and apply than the general large language model, and at the same time, it also improves the speed of the game decision model to generate game decisions.
[0097] As can be seen from the above method, when training the game decision-making model in this application, the server can obtain the historical game videos of sample players, extract data from the historical game videos, determine the game data of the sample players within a specified time period, and use it as training samples. In addition, the server can determine the first decision information corresponding to the decisions made by the sample players in the game states corresponding to the training samples, and use it as the first annotation of the training samples. The server can determine the first prompt text corresponding to the training samples, and input the first prompt text and the training samples into a general large language model to determine the first information output by the general large language model. The server can use the first annotation and the first information as the second annotation of the training samples, and train the game decision-making model to be trained according to the training samples and the second annotation, so that the trained game decision-making model can be used to determine game decisions based on the game data of the players to be decided. By training the game decision-making model to be trained with the training samples determined from the historical game videos of games based on decision types, the trained game decision-making model can learn knowledge in specific fields, and thus the game decisions output by the game decision-making model are more accurate. At the same time, the game decision-making model is a large language model, and the number of parameters of the game decision-making model is smaller than that of the general large language model, making the game decision-making model easy to deploy and improving the speed of generating game decisions by the game decision-making model.
[0098] In addition, the above-trained game decision-making model can also be used as an opponent for game players. By playing games against the game players, the game level of the game players can be improved. That is, the game decision-making model can be used to play games against game players, and the game players can be ordinary players or professional players. Of course, the above-trained game decision-making model can also be used as a substitute for game players to play games against other game players. That is, real game players can choose the game decision-making model to replace themselves and play games against other game players, so as to save the time of real game players and improve the game level of real game players. In addition, the above-trained game decision-making model can also be used to play games against other machine learning models to assist in the training or fine-tuning of other machine learning models, which helps to improve the accuracy of the output results of other machine learning models.
[0099] Furthermore, the number of parameters of the game decision-making model is smaller than that of the general large language model, making the game decision-making model easy to be deployed on the game client, and it can quickly generate game decisions to play games against game players, thereby improving the game level of game players, and it is also convenient to replace real game players to play games against other game players. At the same time, the game decision-making model with a small number of parameters is easy to deploy, has a fast speed and high accuracy in generating game decisions, making the training or fine-tuning of other machine learning models based on the game decision-making model fast, and the accuracy of the output results of other machine learning models after training or fine-tuning is high.
[0100] In this specification, in order to better splice the first prompt text and the training sample so that the game decision-making model to be trained can better learn the features of the training sample and the first prompt text, and thus better train the game decision-making model to be trained. Therefore, when the first prompt text and the training sample are input into the general large language model in step S104 above to determine the first information output by the general large language model, the server can also convert the training sample according to a preset template to determine the first text. The first prompt text and the first text are spliced to obtain a second text. The second text is input into the general large language model to determine the first information output by the general large language model. Among them, the above preset template can be set in advance. This preset template is used to convert the training sample into text. The languages of the above training sample, preset template, first text, and second text can be Chinese, English, or other types of languages, which are not specifically limited in this specification. The above preset template can include fixed template content and several slots to be filled. Specifically, the preset template can be "At game time 【】, our current game situation is as follows: Resources: 【】 - Game time: 【】 - Worker supply: 【】 - Minerals: 【】 - Gas: 【】 - Remaining supply: 【】 - Supply limit: 【】 - Used supply: 【】." Among them, the text part in the above preset template is the fixed template content, and the above "【】" are the slots to be filled, and the content in these slots to be filled needs to be determined by the server according to the training sample. The server can convert the training sample according to the preset template to determine the first text. This first text can be generated by the server based on the content corresponding to the slots to be filled and the preset template, that is, the text obtained after filling the content corresponding to the slots to be filled into the preset template. The content corresponding to the slots to be filled is determined by the server based on the training sample. Therefore, the first text can be "At game time 【00:00】, our current game situation is as follows: Resources:
n
[12] - Minerals:
[50] - Gas: 【0】 - Remaining supply: 【3】 - Supply limit:
[15] - Used supply:
[12] ." It should be noted that the above preset model and the first text are only examples, and this specification does not limit the specific content included in the template and the first text.
[0101] In this specification, as Figure 3 shown, Figure 3 is a schematic diagram of the training process of a game decision-making model provided in this specification. The server can first obtain the historical game videos of sample players, Figure 3Taking the historical game video as an example of the historical replay file. Extract data from the historical game video, determine the game data of the sample players within the specified time period, and use it as the training sample. Also, determine the first decision information corresponding to the decisions made by the sample players in the game state corresponding to the training sample, and use it as the first annotation of the training sample. Determine the first prompt text corresponding to the training sample, and input the first prompt text and the training sample into the general large language model to determine the first information output by the general large language model. Use the first annotation and the first information as the second annotation of the training sample. Train the game decision model to be trained based on the training sample, the first prompt text, and the second annotation.
[0102] In this specification, through the above Figure 1 The game decision model trained by the training method shown is a large language model with macroscopic decision-making capabilities. That is, this game decision model can act as the commander role in the game and generate macroscopic decisions in real time based on the current game data, rather than microscopic instructions. For example, macroscopic decisions can be upgrading buildings, developing the population, and attacking, while microscopic instructions can be attacking building A, advancing to point B, etc. The above game can be a simulation environment.
[0103] In this specification, after training the game decision model, the server can obtain the game data of the player to be decided and use it as the first data. Determine the second prompt text corresponding to the first data. Input the first data and the second prompt text into the trained game decision model to determine the predicted information output by the game decision model. Determine the game decision instruction based on the predicted information. Control the game character corresponding to the player to be decided according to the game decision instruction. Among them, the above player to be decided can be a real game player, or a game decision model replacing the real game player, or other machine learning models, which are not specifically limited in this specification. The above second prompt text can be pre-set, and the content included in the second prompt text is the same as that of the above first prompt text. And the process of determining the second prompt text can be similar to the process of determining the first prompt text, which will not be elaborated here. The above predicted information can include the game development status and game decision information. The above game decision instruction is an instruction generated based on the game decision information in the predicted information. After obtaining the game decision instruction, the server can directly control the game character corresponding to the player to be decided.
[0104] The process of inputting the first data and the second prompt text into the trained game decision model to determine the predicted information output by the game decision model can be as Figure 4 shown. Figure 4It is a schematic diagram of the application of a game decision-making model provided in this specification. Of course, the process of inputting the first data and the second prompt text into the trained game decision-making model to determine the predicted information output by the game decision-making model is similar to the process of inputting the training sample and the first prompt text into the game decision-making model to be trained to determine the first result output by the game decision-making model to be trained, which will not be elaborated here.
[0105] In this specification, after obtaining the trained game decision-making model, the performance of the game decision-making model can also be tested. The performance can include the winning rate of the game decision-making model, that is, the winning rate when the game decision-making model conducts a game confrontation with other machine learning models. Specifically, the server can, under different game difficulties, use the built-in AI (Artificial Intelligence) model and the general large language model in the game to conduct game confrontations with the game decision-making model respectively to evaluate the game decision-making model and obtain an evaluation result. The evaluation result is the winning rate of the game decision-making model and can also include the confrontation duration, that is, the game duration. The above-mentioned difficulty refers to the game difficulty, which can include very easy, easy, medium, and difficult, etc. It should be noted that the above-mentioned built-in AI model, that is, the AI model built in the game, is not limited to any type of machine learning model, and the above-mentioned general large language model can be a GPT model.
[0106] The above evaluation results can be presented in the form of a table, as shown in Table 1 for example. It should be noted that Table 1 is only an example. Table 1 shows the game time (i.e., duration) when the built-in AI model conducts a game confrontation with the game decision-making model under difficulties such as very easy, easy, medium, and difficult, as well as the winning rate of the game decision-making model. Moreover, Table 1 also shows the game time (i.e., duration) when the general large language model conducts a game confrontation with the game decision-making model without restricting the game difficulty, as well as the winning rate of the game decision-making model.
[0107] Table 1
[0108]
[0109] As shown in Table 1 above, when the built-in AI model plays against the game decision-making model at a very simple difficulty level, the game time is 15 minutes, and the winning rate of the game decision-making model is 100%. When the built-in AI model plays against the game decision-making model at a simple difficulty level, the game time is 15 minutes, and the winning rate of the game decision-making model is 100%. When the built-in AI model plays against the game decision-making model at a medium difficulty level, the game time is 20 minutes, and the winning rate of the game decision-making model is 69%. When the built-in AI model plays against the game decision-making model at a difficult difficulty level, the game time is 20 minutes, and the winning rate of the game decision-making model is 67%. When the general large language model plays against the game decision-making model without restricting the game difficulty level, the game time is 20 minutes, and the winning rate of the game decision-making model is 65%. It should be noted that the above general large language model is exemplified by GPT-4o.
[0110] After obtaining the evaluation results, the server can continue to train the game decision-making model according to the evaluation results. Of course, the server can also directly deploy the game decision-making model to the target device. The target device can be the client or server where the game is located. In addition, game data during the application of the game decision-making model can be continuously obtained later, and the game decision-making model can be continuously trained based on the game data.
[0111] The above is the method of one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding game decision-making model training device, as Figure 5 shown.
[0112] Figure 5 is a schematic diagram of a game decision-making model training device provided by this specification, including:
[0113] An acquisition module 200, configured to acquire historical game videos of sample players;
[0114] An extraction module 202, configured to extract data from the historical game videos, determine the game data of the sample player within a specified time period as a training sample, and determine the first decision information corresponding to the decision made by the sample player in the game state corresponding to the training sample as the first annotation of the training sample;
[0115] A first determination module 204, configured to determine the first prompt text corresponding to the training sample, input the first prompt text and the training sample into the general large language model, and determine the first information output by the general large language model;
[0116] A second determination module 206, configured to use the first annotation and the first game information as the second annotation of the training sample;
[0117] A training module 208 for training a game decision-making model to be trained according to the training samples and the second annotation; the trained game decision-making model is used to determine game decisions based on the game data of the player to be decided; wherein, the game decision-making model is a large language model.
[0118] Optionally, the extraction module 202 is specifically configured to use a pre-set data extraction tool to extract data from the historical game video according to a preset time length, and determine the game data of the sample player within a plurality of time periods; from each piece of game data, determine the game data of the sample player within a specified time period as a training sample, and determine the first decision information corresponding to the decision made by the sample player in the game state corresponding to the training sample as the first annotation of the training sample.
[0119] Optionally, the extraction module 202 is specifically configured to determine the information content corresponding to each piece of game data; filter each piece of game data according to a preset filtering rule according to the information content; from the filtered game data, determine the game data of the sample player within a specified time period as a training sample, and determine the first decision information corresponding to the decision made by the sample player in the game state corresponding to the training sample as the first annotation of the training sample.
[0120] Optionally, the first determination module 204 is specifically configured to convert the training sample according to a preset template to determine a first text; splice the first prompt text and the first text to obtain a second text; input the second text into a general large language model to determine the first information output by the general large language model.
[0121] Optionally, the second determination module 206 is specifically configured to determine the decision information included in the first information as the second decision information; determine the difference between the first annotation and the second decision information; when the difference is greater than a preset threshold, use the information in the first information other than the second decision information and the first annotation as the second annotation of the training sample; when the difference is not greater than the preset threshold, use the first information and the first annotation as the second annotation of the training sample.
[0122] Optionally, the first prompt text includes at least a game development state prompt and a game decision prompt;
[0123] Optionally, the first determination module 204 is specifically configured to display the training sample to an operator; in response to the input operation of the operator, determine the first prompt text corresponding to the training sample.
[0124] Optionally, the device further includes:
[0125] An application module 210 is configured to obtain game data of a player to be decided and use it as first data; determine a second prompt text corresponding to the first data; input the first data and the second prompt text into a trained game decision model to determine prediction information output by the game decision model; determine a game decision instruction according to the prediction information; and control a game character corresponding to the player to be decided according to the game decision instruction.
[0126] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the Figure 1 game decision model training method provided above.
[0127] This specification also provides Figure 6 a schematic structural diagram of an electronic device corresponding to Figure 1 as shown. As Figure 6 shown, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the Figure 1 game decision model training method described above.
[0128] Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logical devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or a logical device.
[0129] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logic function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, today, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not just one type of HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0130] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same functions. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0131] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0132] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0133] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0134] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 means for implementing the functions specified in one or more of the blocks or multiple blocks.
[0135] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 means for implementing the functions specified in one or more of the blocks or multiple blocks.
[0136] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 means for implementing the functions specified in one or more of the blocks or multiple blocks.
[0137] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0138] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0139] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0140] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0141] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0142] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0143] The various embodiments in this specification are described in a progressive manner. For the parts that are the same or similar among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the corresponding descriptions in the method embodiments.
[0144] The above are only the embodiments of this specification and are not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.
Claims
1. A method for training a game decision model, characterized in that: The method comprises: Get historical game videos of sample players; Extracting data from the historical game video, determining the game data of the sample player within a specified time period and using the data as a training sample, and determining first decision information corresponding to a decision executed by the sample player in a game state corresponding to the training sample and using the first label of the training sample; Determine a first prompt text corresponding to the training sample, input the first prompt text and the training sample into a universal large language model, and determine first information output by the universal large language model, wherein the first prompt text at least includes a game development status prompt and a game decision prompt; Using the first annotation and the first information as a second annotation of the training sample; The game decision model to be trained is trained according to the training samples and the second annotations; the trained game decision model is used to determine the game decision according to the game data of the player to be decided; wherein the game decision model is a large language model, and the parameter amount of the game decision model is smaller than the parameter amount of the general large language model.
2. The method according to claim 1, characterized in that Extracting data from the historical game video, determining the game data of the sample player within a specified time period and using the data as a training sample, and determining first decision information corresponding to a decision executed by the sample player in a game state corresponding to the training sample and using the first annotation of the training sample, specifically includes: Using a pre-set data extraction tool, extract data from the historical game video according to a preset time length, and determine the game data of the sample players in several time periods; From each game data, the game data of the sample player within a specified time period is determined and used as a training sample, and the first decision information corresponding to the decision executed by the sample player in the game state corresponding to the training sample is determined and used as a first annotation of the training sample.
3. The method according to claim 2, characterized in that Determining the game data of the sample player within a specified time period from each game data as a training sample, and determining first decision information corresponding to the decision executed by the sample player in the game state corresponding to the training sample, specifically includes: Determine the information content corresponding to each game data; According to the content of each information and the preset filtering rules, the game data are filtered; From the filtered game data, the game data of the sample player within a specified time period is determined and used as a training sample, and the first decision information corresponding to the decision executed by the sample player in the game state corresponding to the training sample is determined and used as a first annotation of the training sample.
4. The method according to claim 1, characterized in that Inputting the first prompt text and the training sample into a universal large language model, and determining first information output by the universal large language model specifically includes: According to a preset template, the training sample is converted to determine a first text; Concatenate the first prompt text and the first text to obtain a second text; The second text is input into a general large language model, and first information output by the general large language model is determined.
5. The method according to claim 1, characterized in that Using the first annotation and the first information as the second annotation of the training sample specifically includes: Determining decision information included in the first information and using the decision information as second decision information; Determining a difference between the first annotation and the second decision information; When the difference is greater than a preset threshold, using the information in the first information except the second decision information and the first annotation as the second annotation of the training sample; When the difference is not greater than a preset threshold, the first information and the first annotation are used as a second annotation of the training sample.
6. The method according to claim 1, characterized in that The first prompt text at least includes a game development status prompt and a game decision prompt; Determining a first prompt text corresponding to the training sample specifically includes: Displaying the training sample to an operator; In response to the input operation of the operator, a first prompt text corresponding to the training sample is determined.
7. The method according to claim 1, characterized in that The method further comprises: Obtaining game data of the player to be decided and using it as the first data; Determine a second prompt text corresponding to the first data; Inputting the first data and the second prompt text into a trained game decision model to determine prediction information output by the game decision model; Determining game decision instructions based on the prediction information; According to the game decision instruction, the game character corresponding to the player to be decided is controlled.
8. A game decision model training device, characterized in that: include: An acquisition module is used to obtain historical game videos of sample players; An extraction module, configured to extract data from the historical game video, determine the game data of the sample player within a specified time period, and use the data as a training sample, and determine first decision information corresponding to a decision executed by the sample player in a game state corresponding to the training sample, and use the first annotation of the training sample; A first determination module is used to determine a first prompt text corresponding to the training sample, and input the first prompt text and the training sample into a universal large language model to determine first information output by the universal large language model, wherein the first prompt text at least includes a game development status prompt and a game decision prompt; A second determining module, configured to use the first annotation and the first information as a second annotation of the training sample; A training module, used for training the game decision model to be trained according to the training samples and the second annotations; The trained game decision model is used to determine the game decision according to the game data of the player to be decided; wherein the game decision model is a large language model.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Speech generation method and device based on AI model, equipment and storage medium
CN117194626A