A method for processing landlord fighting data and related devices
By using the landlord model and card game model to process the hand card and game status data in the Doudizhu game, the data processing problem under non-perfect information game is solved, and the processing effect of Doudizhu data and the accuracy of card game actions are improved.
Patent Information
- Application Number
- CN202111487148.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-07
AI Technical Summary
In Doudizhu game, due to the lack of perfect information among players, technologies such as Monte Carlo tree search cannot be effectively applied, which affects the rationality and accuracy of card-playing action data in the data and reduces the processing effect of Doudizhu data.
The initial hand data and game status data are processed using the landlord model and the card model. Through the structures such as the self-attention network layer, the standardization layer along the layer, the 1x1 convolutional network layer, the feature expansion layer and the output layer, execution instructions and card play data are generated to realize automatic card playing operations.
It improves the processing effect of Doudizhu data, enhances the rationality and accuracy of card-playing action data, and solves the data processing problem under non-perfect information game.
Smart Images

Figure CN114146422B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly relates to a method for processing Dou Di Zhu data, a device for processing Dou Di Zhu data, a server, and a computer-readable storage medium. Background Art
[0002] With the continuous development of information technology, in the field of game technology, due to the problem of insufficient participants in the game, it is necessary to automatically process the data in the game in order to realize the virtual game process, improve the sense of participation of actual game participants, and reduce the waiting time of actual participants due to the lack of game players.
[0003] In the related art, generally, Monte Carlo tree search is used in card games to make decision operations for the card-playing logic. However, in the Dou Di Zhu game, any two players do not know the card values of each other, and only know each other's roles among the three players. Such a game belongs to an imperfect information game and cannot meet the requirements of Monte Carlo tree search. This affects the rationality and accuracy of the card-playing action data in the data, and reduces the processing effect of Dou Di Zhu data.
[0004] Therefore, how to improve the processing effect of Dou Di Zhu data is a key issue that those skilled in the art are concerned about. Summary of the Invention
[0005] The purpose of this application is to provide a method for processing Dou Di Zhu data, a device for processing Dou Di Zhu data, a server, and a computer-readable storage medium, so as to solve the problem of poor processing effect of Dou Di Zhu data.
[0006] To solve the above technical problems, this application provides a method for processing Dou Di Zhu data, including obtaining initial hand card data from a Dou Di Zhu game server;
[0007] Processing the initial hand card data with a landlord calling model to obtain an execution instruction, and sending the execution instruction so that the Dou Di Zhu game server sends game state data according to the execution instruction; wherein, the landlord calling model is a model trained according to landlord calling training data;
[0008] Processing the game state data with a card-playing model to obtain card-playing data, and sending the card-playing data; wherein, the card-playing model is a model trained according to card-playing training data; wherein, both the landlord calling model and the card-playing model include an input layer, a self-attention network layer, a layer normalization layer, a 1x1 convolutional network layer, a feature expansion layer, and an output layer.
[0009] Optionally, processing the initial hand card data with a landlord calling model to obtain an execution instruction includes:
[0010] Execute the above instructions in sequence according to the order of the input layer, self-attention network layer, layer normalization layer, 1x1 convolutional network layer, feature expansion layer, and landlord classification layer for the initial hand card data.
[0011] Optionally, process the game state data using a card-playing model to obtain card-playing data, including:
[0012] Execute the above instructions in sequence according to the order of the input layer, self-attention network layer, layer normalization layer, 1x1 convolutional network layer, and card-playing layer for the game state data to obtain the card-playing data.
[0013] Optionally, the training process of the landlord model includes:
[0014] Obtain landlord training data;
[0015] Perform feature extraction processing on the landlord training data to obtain landlord training features;
[0016] Train the self-attention network based on the landlord training features to obtain the landlord model.
[0017] Optionally, the training process of the card-playing model includes:
[0018] Obtain card-playing training data;
[0019] Perform feature extraction processing on the card-playing training data to obtain card-playing training features;
[0020] Train the self-attention network based on the card-playing training features to obtain the card-playing model.
[0021] Optionally, performing feature extraction processing on the card-playing training data to obtain card-playing training features includes:
[0022] Extract data from the card-playing training data according to the landlord game rule model to obtain a card-playing data matrix;
[0023] Use the card-playing data matrix as the card-playing training features.
[0024] This application also provides a landlord game data processing device, including:
[0025] A hand card data acquisition module, used to obtain initial hand card data from the landlord game server;
[0026] The landlord calling processing module is used to process the initial hand card data by using the landlord calling model, obtain an execution instruction, and send the execution instruction, so that the landlord game server sends game status data according to the execution instruction; wherein, the landlord calling model is a model trained according to the landlord calling training data;
[0027] The card playing processing module is used to process the game status data by using the card playing model, obtain card playing data, and send the card playing data; wherein, the card playing model is a model trained according to the card playing training data; wherein, both the landlord calling model and the card playing model include an input layer, a self-attention network layer, a layer normalization layer, a 1x1 convolutional network layer, a feature expansion layer, and an output layer.
[0028] Optionally, the landlord calling processing module includes:
[0029] The training data acquisition unit is used to acquire the landlord calling training data;
[0030] The training feature extraction unit is used to perform feature extraction processing on the landlord calling training data to obtain the landlord calling training features;
[0031] The network training unit is used to train the self-attention network according to the landlord calling training features to obtain the landlord calling model.
[0032] This application also provides a server, including:
[0033] A memory for storing a computer program;
[0034] A processor for implementing the steps of the landlord game data processing method as described above when executing the computer program.
[0035] This application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the landlord game data processing method as described above are implemented.
[0036] A method for processing Dou Di Zhu data provided by the present application includes obtaining initial hand card data from a Dou Di Zhu game server; processing the initial hand card data using a calling landlord model to obtain an execution instruction, and sending the execution instruction so that the Dou Di Zhu game server sends game status data according to the execution instruction; wherein, the calling landlord model is a model trained according to calling landlord training data; processing the game status data using a playing card model to obtain playing card data, and sending the playing card data; wherein, the playing card model is a model trained according to playing card training data; wherein, both the calling landlord model and the playing card model include an input layer, a self-attention network layer, a layer normalization layer, a 1x1 convolutional network layer, a feature expansion layer, and an output layer.
[0037] By first obtaining the initial hand card data, then processing it through the calling landlord model to obtain an execution instruction to determine whether to execute the calling landlord operation, and finally processing the game status data through the playing card model to obtain playing card data, automated card-playing operations are finally realized, improving the processing effect of Dou Di Zhu data.
[0038] The present application also provides a Dou Di Zhu data processing device, a server, and a computer-readable storage medium, which have the above effects and will not be elaborated here. Description of the Drawings
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0040] Figure 1 It is a flowchart of a method for processing Dou Di Zhu data provided by an embodiment of the present application;
[0041] Figure 2 It is a schematic structural diagram of a Dou Di Zhu data processing device provided by an embodiment of the present application;
[0042] Figure 3 It is a schematic diagram of the architecture of a calling landlord model provided by an embodiment of the present application;
[0043] Figure 4 It is a schematic diagram of the architecture of a playing card model provided by an embodiment of the present application. Detailed Embodiments
[0044] The core of the present application is to provide a method for processing Dou Di Zhu data, a Dou Di Zhu data processing device, a server, and a computer-readable storage medium to solve the problem of poor processing effect of Dou Di Zhu data.
[0045] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.
[0046] In related technologies, Monte Carlo tree search is generally used in card games to make decision operations for card-playing logic. However, in the Dou Di Zhu game, any two players do not know the card values of each other, and only the three players know their respective roles. Such a game belongs to an imperfect information game and cannot meet the requirements of Monte Carlo tree search. This affects the rationality and accuracy of the card-playing action data in the data, and reduces the processing effect of Dou Di Zhu data.
[0047] Therefore, this application provides a method for processing Dou Di Zhu data. By first obtaining the initial hand card data, then processing it through the landowner-calling model to obtain an execution instruction to determine whether to execute the landowner-calling operation, and finally processing the game state data through the card-playing model to obtain the card-playing data, automated card-playing operations are finally realized, improving the processing effect of Dou Di Zhu data.
[0048] The following illustrates a method for processing Dou Di Zhu data provided by this application through an embodiment.
[0049] Please refer to Figure 1 , Figure 1 which is a flowchart of a method for processing Dou Di Zhu data provided by the embodiments of this application.
[0050] In this embodiment, the method may include:
[0051] S101, obtain the initial hand card data from the Dou Di Zhu game server;
[0052] This step aims to obtain the initial hand card data from the Dou Di Zhu game server. Since the Dou Di Zhu game is a game in which all card types are dealt and no additional card type data is added. Therefore, the Dou Di Zhu game server directly obtains all the initial hand card data and then makes decisions based on the game state data during the game process.
[0053] S102, process the initial hand card data using the landowner-calling model to obtain an execution instruction, and send the execution instruction so that the Dou Di Zhu game server sends the game state data according to the execution instruction; among them, the landowner-calling model is a model trained based on the landowner-calling training data;
[0054] Based on S101, this step aims to process the initial hand card data using the "calling the landlord" model to obtain execution instructions. It can be seen that this step mainly simulates the processing process of the data in the landlord game, which includes two stages: the "calling the landlord" operation and the "playing cards" operation. The operation methods, operation logics, input data, and output data of the "calling the landlord" operation and the "playing cards" operation are all different. Therefore, different models are used for operation in this embodiment.
[0055] Furthermore, the input data of the "calling the landlord" model is the initial hand card data, and the output is the execution instruction. The execution instruction can be to execute the "calling the landlord" operation, or to execute the "grabbing the landlord" operation, or not to execute any operation.
[0056] Furthermore, this step may include:
[0057] The initial hand card data is sequentially executed according to the order of the input layer, self-attention network layer, layer normalization layer, 1x1 convolutional network layer, feature expansion layer, and "calling the landlord" classification layer to obtain the execution instruction.
[0058] It can be seen that this alternative solution mainly explains how to obtain the execution instruction. In this alternative solution, the initial hand card data is sequentially executed according to the order of the input layer, self-attention network layer, layer normalization layer, 1x1 convolutional network layer, feature expansion layer, and "calling the landlord" classification layer to obtain the execution instruction.
[0059] Furthermore, the training process of the "calling the landlord" model may include:
[0060] Step 1, obtain the "calling the landlord" training data;
[0061] Step 2, perform feature extraction processing on the "calling the landlord" training data to obtain the "calling the landlord" training features;
[0062] Step 3, train the self-attention network according to the "calling the landlord" training features to obtain the "calling the landlord" model.
[0063] It can be seen that this alternative solution mainly explains how to train and obtain the landlord model. In this alternative solution, the "calling the landlord" training data is obtained; then, feature extraction processing is performed on the "calling the landlord" training data to obtain the "calling the landlord" training features; finally, the self-attention network is trained according to the "calling the landlord" training features to obtain the "calling the landlord" model. Among them, the training method of the network can adopt any one of the training means provided by the prior art, and no specific limitation is made here.
[0064] Among them, the "calling the landlord" training data is the game data obtained by the server during the actual game process in the "calling the landlord" session.
[0065] S103. Process the game state data using the card-playing model to obtain card-playing data and send the card-playing data. The card-playing model is a model trained based on card-playing training data. Both the landlord-calling model and the card-playing model include an input layer, a self-attention network layer, a layer normalization layer, a 1x1 convolutional network layer, a feature expansion layer, and an output layer.
[0066] Based on S102, this step aims to process the game state data using the card-playing model to obtain card-playing data. It can be seen that this step aims to obtain the card-playing data through the card-playing model to simulate the card-playing operations of the robot and improve the accuracy and precision of the card-playing data.
[0067] Furthermore, this step may include:
[0068] Execute the game state data in sequence according to the order of the input layer, the self-attention network layer, the layer normalization layer, the 1x1 convolutional network layer, and the card-playing layer to obtain card-playing data.
[0069] It can be seen that this alternative solution mainly explains how to obtain the card-playing data. In this alternative solution, the game state data is executed in sequence according to the order of the input layer, the self-attention network layer, the layer normalization layer, the 1x1 convolutional network layer, and the card-playing layer to obtain card-playing data.
[0070] Furthermore, the training process of the card-playing model may include:
[0071] Step 1. Obtain card-playing training data;
[0072] Step 2. Perform feature extraction processing on the card-playing training data to obtain card-playing training features;
[0073] Step 3. Train the initial card-playing model based on the card-playing training features to obtain the card-playing model.
[0074] It can be seen that this alternative solution mainly explains how to train the obtained card-playing model. In this alternative solution, first, obtain the card-playing training data; then, perform feature extraction processing on the card-playing training data to obtain card-playing training features; finally, train the initial card-playing model based on the card-playing training features to obtain the card-playing model.
[0075] Among them, the card-playing training data is the game data obtained by the server during the actual game process in the card-playing session.
[0076] Furthermore, step 2 in the previous alternative solution may include:
[0077] Step 1. Extract data from the card-playing training data according to the landlord-calling rules model to obtain a card-playing data matrix;
[0078] Step 2: Use the played card data matrix as the training feature for playing cards.
[0079] It can be seen that this alternative solution mainly explains how to extract the training features for playing cards. In this alternative solution, the played card training data is extracted according to the Dou Di Zhu rule model to obtain the played card data matrix; finally, the played card data matrix is used as the training feature for playing cards.
[0080] Among them, the Dou Di Zhu rule model is a model representing the game rules of Dou Di Zhu. Specifically, according to the game rules, the quantities of 15 types of cards from 2, 3, 4, 5, 6, 7, 8, 9, 10, J, Q, K, A, the little king, and the big king are represented by the number of elements in each of the 15 positions from 1 to 15 in 15 dimensions. And the total number of quantities of different types of playing cards has 23 fine classifications, including: the quantity of the player's current cards; the quantities of the cards played by each player in the first four rounds in the order of playing cards, a total of 16; the cards that the 4 players have each played, a total of 4; the cards that all players have played in total; the remaining cards of the other three players excluding the current player. Finally, a matrix of 15x23 composed of different card quantities and an array of 15x1 composed of different card position information are obtained. In this way, the model that extracts data according to the Dou Di Zhu game rules is the Dou Di Zhu rule model.
[0081] In summary, in this embodiment, the initial hand card data is first obtained, then processed by the landlord calling model to obtain an execution instruction to determine whether to perform the landlord calling operation, and finally the game state data is processed by the playing card model to obtain the played card data, and finally the automated card-playing operation is realized, improving the processing effect of Dou Di Zhu data.
[0082] The following further illustrates a Dou Di Zhu data processing method provided by this application through a specific embodiment.
[0083] This embodiment mainly solves the intelligent decision-making problems in the landlord calling and playing card links of the Dou Di Zhu game. The core modules of the present invention are two deep learning models based on the self-attention mechanism, namely the landlord calling model and the playing card model. The landlord calling model is responsible for outputting an execution instruction at the start of the game, and this execution instruction indicates whether to call the landlord. When the landlord calling model outputs 1, it represents calling the landlord; when the landlord calling model outputs 0, it represents not calling the landlord. The playing card model is responsible for outputting the card type data that the player should play each time it is the player's turn to play cards. Among them, the card type data can be a specific card type or a pass operation.
[0084] Among them, the landlord calling model includes an input layer, a self-attention network layer, a layer normalization layer, a 1x1 convolutional network layer, a feature expansion layer, and a landlord calling classification layer. When the game starts, the game environment will call the landlord calling model, and the model will execute different network layers in the following Figure 1 order.
[0085] Among them, the input of the input layer is a 15x1 array. Each position from the 1st to the 15th position in the array represents the quantity of 15 types of cards, namely 2, 3, 4, 5, 6, 7, 8, 9, 10, J, Q, K, A, the little king, and the big king.
[0086] Among them, the self-attention network layer uses the attention mechanism to perform card splitting, card combination, and card face size estimation on the input, providing basic features for further judging whether to call the landlord. The input of the self-attention network layer is a 15x1 array, and the output is a 15x64 feature matrix. The self-attention mechanism simulates the thinking mode of players during the process of calling the landlord, improving the accuracy of extracting features related to calling the landlord.
[0087] Please refer to Figure 3 , Figure 3 , which is a schematic diagram of a landlord-calling model architecture provided by an embodiment of the present application.
[0088] Please refer to Figure 4 , Figure 4 , which is a schematic diagram of a card-playing model architecture provided by an embodiment of the present application.
[0089] It can be seen that through Figure 3 and Figure 4 the model architectures, the corresponding model steps can be realized, further improving the effect of playing cards in the landlord game.
[0090] Among them, the layer normalization performs layer-wise normalization processing on the input of the previous layer, making the input data more standardized and improving the convergence speed and generalization ability of the model.
[0091] Among them, the 1x1 convolutional network layer performs weighted scaling on the output of the previous layer along the channel direction, enhancing the abstraction degree of the output. The input of this layer is a 15x64 feature matrix, and the output is a 15x64 feature matrix.
[0092] Among them, the feature expansion layer is responsible for expanding the input 15x64 features into a 960x1 array.
[0093] Among them, the landlord-calling classification layer compresses the input 960x1 array into 1 real number through a fully-connected neural network, and then converts it into a probability of calling the landlord between 0 and 1 through the sigmoid function.
[0094] Among them, the architecture of the card-playing model can include: an input layer, a self-attention network layer, a layer normalization layer, a 1x1 convolutional network layer, a feature expansion layer, and a card-playing layer. When it is the player's turn to play cards each time, the game environment will call the card-playing model, and the model will execute each module according to the process and finally calculate the card type that the player should play.
[0095] Among them, the input layer is a 15x23 matrix. In the dimension with 15 elements, each position from the 1st to the 15th position represents the quantity of 15 types of cards, namely 2, 3, 4, 5, 6, 7, 8, 9, 10, J, Q, K, A, the little king, and the big king. The total number of different types of poker cards has 23 sub - classifications: the quantity of the player's current cards; the quantity of cards played by each player in the first four rounds according to the order of playing cards, a total of 16; the cards that each of the 4 players has played, a total of 4; the total cards played by all players; the remaining cards of the other three players excluding the current player. The different card quantities form a 15x23 matrix, and the different card position information forms a 15x1 array.
[0096] Among them, the self - attention mechanism module performs hierarchical abstraction processing on the input features. As the depth of the deep learning network increases, the network layer's understanding of the current game situation gradually deepens, and the relevance between the output features and the card - playing decision gradually increases. The self - attention mechanism can better focus on the combination card types that are more beneficial to the current situation from the current hand of cards. At the same time, it can better concentrate the data focus on the features with better relevance to the current situation in the first 4 rounds and the player's historical card - playing.
[0097] Among them, the card - playing layer is composed of a fully - connected layer and a softmax layer in series. The fully - connected layer further abstracts the game situation, and the softmax layer normalizes the probabilities of different card types and selects the action with the highest probability as the game action adopted for the current situation.
[0098] It can be seen that in this embodiment, by first obtaining the initial hand - card data, then processing it through the landlord - calling model to obtain an execution instruction to determine whether to execute the landlord - calling operation, and finally processing the game state data through the card - playing model to obtain card - playing data, the automated card - playing operation is finally realized, improving the processing effect of the landlord - fighting data.
[0099] Next, the landlord - fighting data processing device provided in the embodiments of the present application will be introduced. The landlord - fighting data processing device described below can be mutually corresponding and referred to with the landlord - fighting data processing method described above.
[0100] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a landlord - fighting data processing device provided in the embodiments of the present application.
[0101] In this embodiment, the device may include:
[0102] A hand - card data acquisition module 100, configured to obtain initial hand - card data from the landlord - fighting game server;
[0103] The landlord calling processing module 200 is used to process the initial hand card data by using the landlord calling model, obtain an execution instruction, and send the execution instruction, so that the landlord fighting game server sends game status data according to the execution instruction; wherein, the landlord calling model is a model trained according to the landlord calling training data;
[0104] The card playing processing module 300 is used to process the game status data by using the card playing model, obtain card playing data, and send the card playing data; wherein, the card playing model is a model trained according to the card playing training data; wherein, both the landlord calling model and the card playing model include an input layer, a self-attention network layer, a layer normalization layer, a 1x1 convolutional network layer, a feature expansion layer, and an output layer.
[0105] Optionally, the landlord calling processing module 200 may include:
[0106] The training data acquisition unit is used to acquire the landlord calling training data;
[0107] The training feature extraction unit is used to perform feature extraction processing on the landlord calling training data to obtain the landlord calling training features;
[0108] The network training unit is used to train the self-attention network according to the landlord calling training features to obtain the landlord calling model.
[0109] An embodiment of the present application further provides a server, including:
[0110] A memory for storing a computer program;
[0111] A processor for implementing the steps of the landlord fighting data processing method as described in the above embodiments when executing the computer program.
[0112] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the landlord fighting data processing method as described in the above embodiments are implemented.
[0113] The various embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0114] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0115] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0116] The above has introduced in detail a method for processing Dou Di Zhu data, a device for processing Dou Di Zhu data, a server, and a computer-readable storage medium provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for processing Dou Di Zhu data, characterized in that, it includes obtaining initial hand card data from a Dou Di Zhu game server; processing the initial hand card data using a calling landlord model to obtain an execution instruction, and sending the execution instruction so that the Dou Di Zhu game server sends game state data according to the execution instruction; wherein, the calling landlord model is a model trained according to calling landlord training data; processing the game state data using a playing card model to obtain playing card data, and sending the playing card data; wherein, the playing card model is a model trained according to playing card training data; wherein, both the calling landlord model and the playing card model include an input layer, a self-attention network layer, a layer normalization layer, a 1x1 convolutional network layer, a feature expansion layer, and an output layer; wherein, the playing card model is a model trained according to playing card training data, and includes: extracting data from the playing card training data according to a Dou Di Zhu rule model to obtain a playing card data matrix, and using the playing card data matrix as a playing card training feature; the Dou Di Zhu rule model is a model representing the game rules of Dou Di Zhu, and includes arrays and matrices.
2. The Dou Di Zhu data processing method according to claim 1, characterized in that, processing the initial hand card data using a calling landlord model to obtain an execution instruction, including: sequentially executing the initial hand card data in the order of an input layer, a self-attention network layer, a layer normalization layer, a 1x1 convolutional network layer, a feature expansion layer, and a calling landlord classification layer to obtain the execution instruction.
3. The Dou Di Zhu data processing method according to claim 1, characterized in that, processing the game state data using a playing card model to obtain playing card data, including: sequentially executing the game state data in the order of an input layer, a self-attention network layer, a layer normalization layer, a 1x1 convolutional network layer, and a playing card layer to obtain the playing card data.
4. The Dou Di Zhu data processing method according to claim 1, characterized in that, the process of training the calling landlord model includes: obtaining calling landlord training data; performing feature extraction processing on the calling landlord training data to obtain calling landlord training features; training the self-attention network according to the calling landlord training features to obtain the calling landlord model.
5. The Dou Di Zhu data processing method according to claim 1, characterized in that, the process of training the playing card model includes: obtaining playing card training data; performing feature extraction processing on the playing card training data to obtain playing card training features; training the self-attention network according to the playing card training features to obtain the playing card model.
6. The Dou Di Zhu data processing method according to claim 5, characterized in that, performing feature extraction processing on the playing card training data to obtain playing card training features, including: extracting data from the playing card training data according to a Dou Di Zhu rule model to obtain a playing card data matrix; using the playing card data matrix as the playing card training feature.
7. A Dou Di Zhu data processing device, characterized in that, it includes: A hand card data acquisition module for acquiring initial hand card data from a Dou Di Zhu game server; A landlord calling processing module for processing the initial hand card data using a landlord calling model to obtain an execution instruction and sending the execution instruction so that the Dou Di Zhu game server sends game state data according to the execution instruction; wherein, the landlord calling model is a model trained according to landlord calling training data; A card playing processing module for processing the game state data using a card playing model to obtain card playing data and sending the card playing data; wherein, the card playing model is a model trained according to card playing training data; wherein, both the landlord calling model and the card playing model include an input layer, a self-attention network layer, a layer normalization layer, a 1x1 convolutional network layer, a feature expansion layer, and an output layer; wherein, the card playing model is a model trained according to card playing training data, including: extracting data from the card playing training data according to a Dou Di Zhu rule model to obtain a card playing data matrix, and using the card playing data matrix as card playing training features; the Dou Di Zhu rule model is a model representing the game rules of Dou Di Zhu, including arrays and matrices.
8. The Dou Di Zhu data processing device according to claim 7, wherein, the landlord calling processing module includes: a training data acquisition unit for acquiring landlord calling training data; a training feature extraction unit for performing feature extraction processing on the landlord calling training data to obtain landlord calling training features; a network training unit for training a self-attention network according to the landlord calling training features to obtain the landlord calling model.
9. A server, wherein, it includes: a memory for storing a computer program; a processor for implementing the steps of the Dou Di Zhu data processing method according to any one of claims 1 to 6 when executing the computer program.
10. A computer-readable storage medium, wherein, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the Dou Di Zhu data processing method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Game match method and device based on artificial intelligence, equipment and storage medium
CN111437608A
Universal multi-agent game algorithm
CN112755538A