Method, device, electronic device and storage medium for learning to capture pieces and review the game
By using the policy network model and valuation network model to filter out the bad hand points of the user and recommend the target point, and generate review demonstration data, it solves the problem that users find it difficult to independently improve the wrong point points in the review learning process, and improves learning efficiency and independent learning ability.
Patent Information
- Application Number
- CN202111268992.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-28
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-10-28
AI Technical Summary
In the review and learning of eating, it is difficult for users to independently identify and improve the wrong chess pieces, resulting in inefficient learning and relying on the teacher's guidance.
By obtaining the game data after human-computer interactive gameplay, using the pre-trained strategy network model and valuation network model, the bad hand points of the user's landing point are selected, and the corresponding target landing point is recommended, and the review demonstration data is generated for users to learn.
It realizes user-independent identification and improvement of errors, reduces dependence on teacher guidance, and improves users' initiative in independent learning and Go skills.
Smart Images

Figure CN113975783B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of game review and learning, and in particular to a game review and learning method, device, electronic device and storage medium. Background Art
[0002] At present, there are many Go capture products on the market. Their main purpose is to play chess alternately between users and artificial intelligence (AI) on a 9-way or 13-way chessboard. The first party to capture several pieces of the opponent wins. There are many similar games and game variants. Users need to improve their Go skills by replaying and learning to capture pieces. However, in the process of replaying and learning to capture pieces, users often have questions such as "Which pieces were played incorrectly?", "Where should the wrong pieces be placed?", "How should the subsequent chess game proceed?", etc. Effective replay learning requires the guidance of a teacher, which will affect the autonomy of learning. Summary of the invention
[0003] In view of this, the purpose of this application is to propose a method, device, electronic device and storage medium for learning to capture pieces and review the game to solve the above-mentioned technical problems.
[0004] The exemplary embodiments of the present disclosure provide a method for learning to capture a piece by reviewing a game, including:
[0005] Obtaining game data of the user after the human-computer interactive game, and determining multiple user placement points of the user in the game according to the game data;
[0006] For each of the user's placement points, determine at least one recommended placement point corresponding to the user's placement point; and determine a recommended value for the user's placement point based on the user's placement point and the recommended placement point corresponding to the user's placement point;
[0007] Determine the user's placement point that meets the predetermined condition among all the recommended values as a bad move point;
[0008] For each of the bad move points, determine the winning rate value of each of the recommended placement points corresponding to the bad move point, determine the highest of the winning rate values as the highest winning rate value, and determine the recommended placement point corresponding to the highest winning rate value as the target placement point; based on the target placement point, generate review demonstration data, and display the review demonstration data to the user.
[0009] In some exemplary embodiments, determining at least one recommended placement point corresponding to the user's placement point specifically includes:
[0010] Determine the chessboard layout data before the user places the chess piece;
[0011] Inputting the chessboard layout data into a pre-trained strategy network model to obtain at least one recommended move point output by the strategy network model;
[0012] Among them, at least one of the recommended placement points is the same as the user's placement point.
[0013] In some exemplary embodiments, the output of the policy network model also includes a recommendation degree corresponding to each of the recommended placement points; the recommendation degree of the user placement point is equal to the recommendation degree of the recommended placement point that is the same as the user placement point;
[0014] The calculation process of the recommended value is:
[0015] Recommendation value = recommendation degree of the user's placement point / total recommendation degrees of recommended placement points with higher recommendation degrees than the user's placement point.
[0016] In some exemplary embodiments, for each bad hand point, determining the winning rate value of each recommended placement point corresponding to the bad hand point, and determining the recommended placement point corresponding to the highest winning rate value as the target placement point includes:
[0017] Using the pre-trained valuation network model and the fast move network model, searching for the winning rate value of each of the recommended move points in a fast search manner;
[0018] Each of the recommended placement points is sorted according to all of the winning rate values, the highest of the winning rate values is determined as the highest winning rate value, and the recommended placement point corresponding to the highest winning rate value is determined as the target placement point.
[0019] In some exemplary embodiments, generating replay demonstration data according to the target drop point, and displaying the replay demonstration data to the user includes:
[0020] Determine the chessboard layout data of the target drop point;
[0021] Inputting the chessboard layout data of the target chessboard placement point into a pre-trained strategy network model to obtain at least one next recommended chessboard placement point output by the strategy network model;
[0022] Using the pre-trained valuation network model and the fast move network model, searching for the winning rate value of each of the recommended next move points in a fast search manner;
[0023] Sort each of the recommended move points for the next step according to all the win rate values, determine the highest of the win rate values as the highest win rate value, and determine the recommended move point corresponding to the highest win rate value as the target move point for the next step;
[0024] Based on the next target drop point, re-determine the chessboard layout data;
[0025] In response to determining that the move ending condition is met, replay demonstration data is generated and displayed to the user.
[0026] In some exemplary embodiments, the fast search method is MCTS search.
[0027] In some exemplary embodiments, the end condition of the MCTS search includes:
[0028] The search time exceeds 5s;
[0029] or,
[0030] The highest win rate value is 50% higher than the second highest win rate value;
[0031] Among them, the second highest winning rate value is the winning rate value closest to the highest winning rate value.
[0032] Based on the same inventive concept, the exemplary embodiment of the present disclosure also provides a device for learning to capture a piece and review the game, including:
[0033] An acquisition module is used to acquire game data of a user after a human-computer interactive game, and determine multiple user placement points of the user in the game according to the game data;
[0034] A calculation module, for each of the user's placement points, determines at least one recommended placement point corresponding to the user's placement point; and determines a recommended value of the user's placement point based on the user's placement point and the recommended placement point corresponding to the user's placement point;
[0035] A bad-hand point module determines the user's placement point that meets a predetermined condition among all the recommended values as a bad-hand point;
[0036] The display module determines, for each bad move point, the winning rate value of each recommended move point corresponding to the bad move point, and determines the recommended move point corresponding to the highest winning rate value as the target move point; generates review demonstration data based on the target move point, and displays the review demonstration data to the user.
[0037] Based on the same inventive concept, an exemplary embodiment of the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above methods when executing the program.
[0038] Based on the same inventive concept, an exemplary embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to enable a computer to execute any of the above methods.
[0039] From the above, it can be seen that the present application provides a method, device, electronic device and storage medium for learning to capture pieces through reviewing the game, which first screens out the bad moves among the moves made by users in human-computer games, and recommends target moves with high winning rates corresponding to the bad moves, and generates review demonstration data based on the target moves for users to learn; it gets rid of the dependence on teacher guidance in the review learning process and improves the user's initiative in independent learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the present application or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0041] Figure 1 A schematic diagram of an application scenario of an exemplary embodiment of the present disclosure;
[0042] Figure 2 It is a flowchart of a method for learning to capture a piece and review a game according to an exemplary embodiment of the present disclosure;
[0043] Figure 3 A schematic diagram of a process flow for selecting a recommended drop point according to an exemplary embodiment of the present disclosure;
[0044] Figure 4 Another flowchart of the process of selecting a recommended drop point according to an exemplary embodiment of the present disclosure is shown;
[0045] Figure 5 A schematic diagram of a process flow for selecting a target drop point according to an exemplary embodiment of the present disclosure;
[0046] Figure 6 A flowchart of a process for generating replay demonstration data according to an exemplary embodiment of the present disclosure;
[0047] Figure 7 It is a structural schematic diagram of a device for learning to capture pieces and review a game according to an exemplary embodiment of the present disclosure;
[0048] Figure 8 It is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0049] The principles and spirit of the present application will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present disclosure, and are not intended to limit the scope of the present application in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0050] According to an embodiment of the present disclosure, a text processing model training method, a text processing method and related equipment are proposed.
[0051] It should be understood herein that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction rather than having any limiting meaning.
[0052] For ease of understanding, the terms involved in the embodiments of the present disclosure are explained below:
[0053] Neural Networks (ANNs): Based on the principles of biological neural networks and the needs of practical applications, practical artificial neural network models are built, corresponding learning algorithms are designed, certain intelligent activities of the human brain are simulated, and then technically implemented to solve practical problems.
[0054] Policy Network Model (PolicyNet): After training (learning) using a neural network, it can screen and determine the next move based on the current chessboard.
[0055] ValueNet: A neural network model that is trained (learned) using a large number of samples and can predict the winning rate of various moves in the next step based on the current state of the board. It has the characteristics of a neural network and has a certain self-learning ability.
[0056] FastNet (Fast Moving Network): A neural network model obtained through training (learning) with a neural network that can perform subsequent rapid alternating moves based on the current state of the chessboard.
[0057] Monte Carlo Tree Search (MCTS): also known as MCTS search, is a heuristic search algorithm based on tree data structure that is still relatively effective when the search space is huge.
[0058] Review: Review, a term in Go, is also called "replaying the game". It refers to replaying the game after the game is over to check the pros and cons of the moves and the key gains and losses in the game. It is generally used for self-study or asking experts to give guidance and analysis. If you rehearse according to the chess record, it is called "playing the score" or "reading the chess record".
[0059] Bad move point: The wrong point of placement on the chessboard is a bad move point.
[0060] User placement point: refers to the position where the user places the piece during the human-computer game.
[0061] AI move point: the position where the AI moves during the human-computer game.
[0062] The principle and spirit of the present application are explained in detail below with reference to several representative embodiments of the present disclosure. SUMMARY OF THE INVENTION
[0064] Go-related replay learning means that after each game, both players repeat the previous game again. This can effectively deepen the impression of the game and find loopholes in the offense and defense of both sides to improve their own chess skills. However, when replaying, users often do not know "Which pieces were played incorrectly?", "Where should the wrong pieces be played?", "How should the subsequent games proceed?", that is, they do not know "Which pieces are bad moves?", "Where is the target drop point with a high winning rate corresponding to the bad move point?", "How will the subsequent game proceed based on the target drop point?", and the existing AI does not have the function of pointing out the bad move points in the replay game and demonstrating the subsequent several steps of the target drop point corresponding to the bad move point, so teachers need to provide relevant guidance, that is, related replay learning requires the assistance of teachers, especially for junior students, who are highly dependent on teachers during replay learning, and their learning autonomy will naturally be affected.
[0065] In response to the problems existing in the above-mentioned prior art, the present disclosure provides a method, device, electronic device and storage medium for learning to capture pieces and review the game. In this method, the multiple user placement points in the game data after the human-computer interactive game are first screened for bad move points; then, the target placement points corresponding to the bad move points are obtained through calculation and analysis, and based on the target placement points, review demonstration data that can display the subsequent chess game trends are generated for user learning and research. That is, the device corresponding to this method can have related functions such as pointing out bad move points for user placement points in human-computer games, target placement points with high winning rates corresponding to the bad move points, and displaying the subsequent chess game trends based on the target placement points, thereby getting rid of the dependence on teacher guidance and improving the initiative of users in independent learning.
[0066] After introducing the basic principles of the present disclosure, various non-limiting embodiments of the present disclosure are described in detail below.
[0067] Application Scenario Overview
[0068] refer to Figure 1 , which is a schematic diagram of an application scenario of the method for learning to eat pieces and review the game provided in an embodiment of the present disclosure. The application scenario includes a terminal device 101, a server 102, and a data storage system 103. Among them, the terminal device 101, the server 102 and the data storage system 103 can be connected through a wired or wireless communication network. The terminal device 101 includes but is not limited to a desktop computer, a mobile phone, a mobile computer, a tablet computer, a media player, a smart wearable device, a personal digital assistant (PDA) or other electronic devices that can achieve the above functions. The server 102 and the data storage system 103 can both be independent physical servers, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0069] The server 102 is used to provide a game review and learning service to the user of the terminal device 101 , and a client for communicating with the server 102 is installed in the terminal device 101 .
[0070] First, the server 102 sends the game data after the human-computer interactive game to the client of the terminal device 101 through the communication network, and displays it on the display interface corresponding to the client. At the same time, the server 102 determines multiple user placement points of the user in the game based on the game data, and determines at least one recommended placement point corresponding to each user placement point; obtains the recommended value of the user placement point by calculation, defines the user placement point corresponding to the recommended value that meets the preset conditions as a bad move point, and sends the bad move point to the client of the terminal device 101, and displays it on the display interface corresponding to the client.
[0071] Then, the server 102 uses the recommended placement point with the highest winning rate value corresponding to a bad move point as the target placement point, and generates review demonstration data based on the target placement point, stores it in the data storage system 103, and sends the review demonstration data to the client of the terminal device 101 through the communication network, and displays it through the display interface.
[0072] Repeat the above process to realize the process of reviewing and demonstrating the target placement points corresponding to all bad moves, thereby completing the user's capture review and learning task.
[0073] Combine the following Figure 1The present invention describes the method, device, electronic device and storage medium for learning to capture a piece according to an exemplary embodiment of the present invention by using the application scenario. It should be noted that the above application scenario is only shown to facilitate understanding of the spirit and principle of the present invention, and the embodiments of the present invention are not limited in this respect. On the contrary, the embodiments of the present invention can be applied to any applicable scenario.
[0074] Exemplary Methods
[0075] Some embodiments of the present application provide a method for learning to capture a piece and then replay the game, such as Figure 2 As shown, including:
[0076] S201, obtaining game data of a user after a human-computer interactive game, and determining multiple user placement points of the user in the game according to the game data;
[0077] S202: for each of the user's placement points, determine at least one recommended placement point corresponding to the user's placement point; and determine a recommended value for the user's placement point based on the user's placement point and the recommended placement point corresponding to the user's placement point;
[0078] S203, determining the user's placement points that meet a predetermined condition among all the recommended values as bad move points;
[0079] S204. For each bad move point, determine the winning rate value of each recommended move point corresponding to the bad move point, determine the highest of the winning rate values as the highest winning rate value, and determine the recommended move point corresponding to the highest winning rate value as the target move point; generate replay demonstration data based on the target move point, and display the replay demonstration data to the user.
[0080] Repeat step S204 to generate replay demonstration data for target drop points corresponding to all bad move points for user learning.
[0081] Among them, the game data is the complete chess game data after the human-computer interactive game, and the game data includes the AI placement points and the user placement points.
[0082] Among them, steps S201 to S203 are the process of screening the user's move points in the game data and finding the bad move points, which can be regarded as a shallow analysis of the game data; step S204 is the process of finding the target move points corresponding to the bad move points and generating the replay demonstration data based on the target move points, which can be regarded as a deep analysis of the game data. Both the shallow analysis and the deep analysis are performed with the help of a pre-trained neural network model.
[0083] Among them, the replay demonstration data can be used for users to study and learn to improve their chess skills. The replay demonstration data can be full-game demonstration data of the entire chess game based on the target drop point, or it can be replay demonstration data of the subsequent steps based on the target drop point. Whether it is full-game demonstration data or replay demonstration data of the subsequent steps, as long as it can reflect the subsequent trend of the chess game, it is not limited here for user learning and research. When the replay demonstration data is replay demonstration data of the subsequent steps based on the target drop point, in order to reflect the subsequent trend of the chess game, it is generally replay demonstration data of at least 5 subsequent steps including the target drop point, that is, AI uses the neural network model to play at least 5 steps by itself to reflect the subsequent trend of the chess game.
[0084] In the specific implementation, the game data after the human-computer interactive game is first obtained in the step, and the obtained game data can be displayed on the corresponding display interface of the client, and the determined multiple user placement points can be highlighted. Afterwards, according to the sequence of the game process, the AI will input the chessboard data before each user's move into the pre-trained neural network model for calculation; the neural network model will output at least one recommended placement point for the chessboard data, and the recommendation degree of the user's placement point is compared with the recommendation degree of the recommended placement point to obtain a recommended value. The one with a lower recommended value can be defined as a bad move point. The determination of the bad move point can be adopted in one of the following two ways: (1) All user placement points below the preset recommended value are defined as bad move points; (2) The recommended values can be sorted from large to small / from small to large, and the last few / first few in the sorting are defined as bad move points. For example, the recommended values are sorted from large to small, and the last five in the sorting are bad move points; for another example, the recommended values are sorted from small to large, and the first five in the sorting are bad move points.
[0085] Afterwards, the neural network model will estimate the winning rate of each recommended placement point corresponding to each bad move point, and determine the recommended placement point corresponding to the highest winning rate value as the target placement point; the neural network model will generate review demonstration data based on the target placement point and display it to the user for learning and research.
[0086] Of course, in the specific implementation process, the neural network model can generate replay demonstration data that can reflect the subsequent chess game trends based on the target drop point corresponding to each bad move point; it can also generate replay demonstration data for the target drop points corresponding to one or more bad move points selected by the user. When the neural network model is to generate replay demonstration data based on the target drop points corresponding to multiple bad move points, in the absence of user agreement, the neural network model will display the corresponding replay demonstration data one by one in the order of the multiple bad move points in the game process; in the case of user agreement, the corresponding replay demonstration data will be displayed in the order agreed by the user. No limitation is made here.
[0087] The capture review learning method of this embodiment first screens out the bad move points among the user's moves in human-computer games, and recommends target moves with high winning rates corresponding to the bad move points, and generates review demonstration data based on the target moves for user study and research, thus getting rid of the reliance on teacher guidance in the review learning process and improving the students' initiative in independent learning.
[0088] In some exemplary embodiments, Figure 3 As shown, for each of the user's placement points, determining at least one recommended placement point corresponding to the user's placement point includes:
[0089] S301, determining the chessboard layout data before the user places a chess piece;
[0090] S302: Input the chessboard layout data into a pre-trained strategy network model to obtain at least one recommended move point output by the strategy network model.
[0091] Among them, the output of the strategy network model also includes the recommendation degree of the user's placement point and the recommendation degree corresponding to each of the recommended placement points; so as to form a recommendation degree comparison between the user's placement point and the corresponding recommended placement point, and determine whether the user's placement point is a bad move.
[0092] Among them, the game data is the complete chess game data after the human-computer interactive game, and the user will use the complete chess game data as the basis for the re-learning of capturing pieces; the chessboard layout data is partial chess game data, specifically the partial chess game data before each user's move, that is, the chessboard layout data does not include the user's move point; that is to say, for a human-computer interactive game, when re-learning, there is one game data and multiple chessboard layout data. Specifically, when the user plays chess first, there is no chessboard layout data before the user's first move, so the number of chessboard layout data can be one less than the number of user's move points; when the AI plays chess first, the number of chessboard layout data can be the same as the number of user's move points.
[0093] In specific implementation, the selection process of the recommended drop point includes:
[0094] Based on the chessboard layout data pushed before each user places a piece, the pre-trained policy network model is used to search on the Monte Carlo tree according to the predetermined search breadth (i.e. MCTS). After the end condition of the MCTS search is reached (such as the search time exceeds 5s), the search ends. The output of the policy network model is a matrix of the size of the chessboard. The value of each drop point in the matrix represents the recommendation degree of the next drop point on the corresponding position on the chessboard. The sum of all the recommendation degree values reflected in the matrix is 1. The policy network model will rank the recommendation degrees of the multiple drop points, and then recommend at least one drop point with a high recommendation degree ranking as the recommended drop point. Among them, the predetermined search breadth can be set according to actual needs, for example, (analysisWideRootNoise)[0,1], the closer the value is to 0, the narrower the search breadth; the closer the value is to 1, the wider the search breadth. Increasing the predetermined search breadth can expand the search range, thereby increasing the weight of the number of visits to each node on the Monte Carlo tree given by the UCB (upper confidence bound) value during the search, thereby increasing the chance of nodes with fewer visits in the Monte Carlo tree being visited.
[0095] In this embodiment, the chessboard layout data corresponding to each user's placement point in the game data is analyzed to obtain at least one recommended placement point corresponding to the user's placement point, in preparation for subsequent bad move point analysis.
[0096] In some exemplary embodiments, Figure 4 As shown, for each of the user's placement points, determining at least one recommended placement point corresponding to the user's placement point specifically includes:
[0097] S401, labeling each user's move point in the game data according to the order of the game;
[0098] S402, determining the chessboard layout data before each user places a chess piece at a predetermined numbered range;
[0099] S403: Input the chessboard layout data into a pre-trained strategy network model to obtain at least one recommended move point output by the strategy network model.
[0100] Among them, the output of the strategy network model also includes the recommendation degree of the user's placement point and the recommendation degree corresponding to each of the recommended placement points; so as to form a recommendation degree comparison between the user's placement point and the corresponding recommended placement point, and determine whether the user's placement point is a bad move.
[0101] Among them, the preset number range can be input by the user or set by AI according to the user's level to meet the needs of users of different levels. If the user is a beginner, due to the limited level of beginners, they may make mistakes in the first few steps or the last few steps, and the number range can be set wider; if the user is a senior student, his first few steps or the last few steps will basically not make mistakes, then the number range can be set narrower. For example, there are 38 user placement points in a game data. The 38 user placement points are numbered according to the order of appearance in the game, and the numbers are recorded as 1-38. The user inputs the number range of 10-30, then AI only analyzes the chessboard layout data corresponding to the user placement points numbered 10-30, and the subsequent screening of bad points is also based on the user placement points within the preset number range of 10-30.
[0102] In specific implementation, the selection process of the recommended drop point includes:
[0103] Based on the chessboard layout data before each user's move within the preset number range, the pre-trained strategy network model is used to search on the Monte Carlo tree according to the predetermined search breadth (i.e. MCTS). After the end condition of the MCTS search is reached (such as the search time exceeds 5s), the search ends. The output of the strategy network model is a matrix of the size of the chessboard. Each value in the matrix represents the recommendation degree of the next move on the corresponding position on the chessboard. The sum of all the recommendation degree values reflected in the matrix is 1. The strategy network model will rank the recommendation degrees of the multiple move points, and then introduce at least one move point with a high recommendation degree ranking as the recommended move point. Among them, the predetermined search breadth can be set according to actual needs, for example, (analysisWideRootNoise)[0,1], the closer the value is to 0, the narrower the search breadth; the closer the value is to 1, the wider the search breadth. Increasing the predetermined search breadth can expand the search range, thereby increasing the weight of the number of visits to each node on the Monte Carlo tree given by the UCB (upper confidence bound) value during the search, thereby increasing the chance of nodes with fewer visits in the Monte Carlo tree being visited.
[0104] In this embodiment, multiple user placement points in the game data are numbered according to the order of the game, and it can be further limited to analyze the chessboard layout data corresponding to the user placement points within a preset number range to obtain the bad move points within the number range. The setting of the number range limits the neural network in the AI to only calculate the recommended placement points for the user placement points within the relevant number range, without calculating the recommended placement points for all user placement points, reducing the number of neural network operations, improving the pertinence of the replay learning, and then improving the efficiency of the replay learning.
[0105] In some exemplary embodiments, the calculation process of the recommendation value may adopt the following two methods:
[0106] Method 1 is:
[0107] Recommendation value = recommendation degree of the user's placement point / total recommendation degrees of recommended placement points with higher recommendation degrees than the user's placement point.
[0108] In specific implementation, based on the chessboard layout data before a user's chessboard placement, the output of the strategy network model is a chessboard-sized matrix, each value in the matrix represents the recommendation degree of the corresponding position on the chessboard for the next step, and the sum of all the recommendation degree values reflected in the matrix is 1; and the user's chessboard placement point must also be in the matrix, and its corresponding recommendation degree will also be reflected in the matrix; the strategy network model will rank the recommendation degrees of the multiple chessboard placement points, and then introduce the top-ranked chessboard placement points as recommended chessboard placement points. Among them, the recommended chessboard placement points corresponding to the user's chessboard placement point are at least 5. Taking the strategy network model outputting the top 5 recommended chessboard placement points for each user's chessboard placement point as an example, the user's chessboard placement point is recorded as T1, and the 5 recommended chessboard placement points are T2, T3, T4, T5, and T6, respectively. The recommendation degree of each recommended chessboard placement point is T1 = 5%, T2 = 5%, T3 = 5%, T4 = 8%, T5 = 7%, and T6 = 10%.
[0109] The calculation process of the recommended value of the user's placement point is:
[0110] Recommended value = 5% / (8%+7%+10%) = 0.2.
[0111] The calculation process of the recommendation value is the process of comparing the user's placement point with multiple recommended placement points, which is more accurate and convincing than comparing with only one best recommended placement point.
[0112] Method 2 is:
[0113] Recommendation value = recommendation degree of the user's placement point / recommendation degree of the placement point with the highest recommendation degree.
[0114] In specific implementation, based on the chessboard layout data before a user places a move, the output of the strategy network model is a matrix of the size of the chessboard. Each value in the matrix represents the recommendation degree of the corresponding position on the chessboard for the next move, and the sum of all the recommendation degree values reflected in the matrix is 1. The strategy network model will rank the recommendation degrees of the multiple move points, and then deduce the move point with the highest recommendation degree as the recommended move point. The recommendation degree of the user's move point is then compared with the recommendation degree of the move point with the highest recommendation degree. For example, the user's move point is recorded as T1, and the move point with the highest recommendation degree is recorded as T2. The recommendation degrees of T1 and T2 are T1=5% and T2=10%.
[0115] The calculation process of the recommended value of the user's placement point is:
[0116] Recommended value = 5% / (10%) = 0.5.
[0117] The calculation process of the recommendation value is the process of comparing the user's placement point with the most recommended placement point. Compared with comparing with multiple recommended placement points, the accuracy will be lower, but the calculation process will also be faster.
[0118] In some exemplary embodiments, Figure 5 As shown, for each bad move point, determining the winning rate value of each recommended move point corresponding to the bad move point, and determining the recommended move point corresponding to the highest winning rate value as the target move point, includes:
[0119] S501, using the pre-trained valuation network model and the fast move network model, searching for the winning rate value of each of the recommended move points in a fast search manner;
[0120] S502. Sort each of the recommended placement points according to all of the winning rate values, determine the highest of the winning rate values as the highest winning rate value, and determine the recommended placement point corresponding to the highest winning rate value as the target placement point.
[0121] The fast search method is MCTS search.
[0122] In specific implementation, the process of selecting the target drop point includes:
[0123] Based on the chessboard layout data pushed for each recommended move point, the valuation network model will initialize a weight value according to the recommendation degree of the recommended move point given by the strategy network model, perform MCTS search according to the search breadth corresponding to the weight value, and use the fast move network model to play the whole game of black and white chess by itself. The weight value is updated according to the result of the self-game. After the end condition of the MCTS search is met, the search ends and the winning rate value of the recommended move point is determined according to the latest weight value. Repeat the above process to obtain the winning rate values of all recommended move points. Afterwards, sort each of the recommended move points according to all the winning rate values, and determine the recommended move point corresponding to the highest winning rate value as the target move point.
[0124] Take a bad move point corresponding to 5 recommended move points as an example. For these 5 recommended move points, the valuation network model will initialize 5 weight values according to the recommendation degree of the 5 recommended move points given by the policy network model, which are recorded as v1, v2, v3, v4, and v5 respectively. The weight value corresponding to the high recommendation degree is also high, and the weight value corresponding to the low recommendation degree is also low. The weight value corresponds to the search breadth of the recommended move point. Then, MCTS search is performed according to the set search breadth. MCTS search requires the use of the valuation network model and the fast move network model. Specifically, based on each recommended move point, the fast move network model will automatically perform a game of black and white chess alternately (taking black chess representing the user side and white chess representing the AI side as an example). If the result of the game is that the black chess representing the user side loses, the weight value of the recommended move point will decrease; if the result of the game is that the black chess representing the user side wins, the weight value of the recommended move point will increase. That is, after one round of searching, the weight values of the five recommended drop points are updated to v1', v2', v3', v4', and v5', and then the search breadth of the corresponding search branch is updated. The above process is repeated, and when the end condition of the MCTS search is reached, the search ends, and the winning rate value of the corresponding recommended drop point is determined based on the latest five weight values. After that, each of the recommended drop points is sorted according to all the winning rate values, and the recommended drop point corresponding to the highest winning rate value is determined as the target drop point.
[0125] The end conditions of the MCTS search include:
[0126] i) the search time exceeds 5s; or, ii) the highest winning rate value is 50% higher than the second highest winning rate value; wherein the second highest winning rate value is the winning rate value closest to the highest winning rate value.
[0127] In the specific implementation, as long as one of the end conditions is met, the MCTS search ends.
[0128] In some exemplary embodiments, Figure 6 As shown, according to the target drop point, replay demonstration data is generated, and the replay demonstration data is displayed to the user, including:
[0129] S601, determining the chessboard layout data of the target placement point;
[0130] S602, inputting the chessboard layout data of the target chessboard placement point into a pre-trained strategy network model to obtain at least one recommended chessboard placement point for the next step output by the strategy network model;
[0131] S603, using the pre-trained valuation network model and the fast move network model, searching for the winning rate value of each of the recommended next move points in a fast search manner;
[0132] S604, sorting each of the recommended placement points for the next step according to all the winning rate values, determining the highest of the winning rate values as the highest winning rate value, and determining the recommended placement point corresponding to the highest winning rate value as the target placement point for the next step;
[0133] S605, re-determining the chessboard layout data based on the next target placement point;
[0134] S606: In response to determining that the move ending condition is met, generate review demonstration data, and display the review demonstration data to the user.
[0135] By looping through steps S601-S605, the AI can execute the self-playing alternating move process, and then enter step S606 to generate replay demonstration data, and display the replay demonstration data to the user.
[0136] The end conditions of the MCTS search include:
[0137] i) the search time exceeds 5s; or, ii) the highest winning rate value is 50% higher than the second highest winning rate value; wherein the second highest winning rate value is the winning rate value closest to the highest winning rate value.
[0138] In the specific implementation, as long as one of the end conditions is met, the MCTS search ends.
[0139] Among them, the end condition of the move is the number of steps of the pre-set replay demonstration data. For example, the chessboard layout data is PATH0, and the target move point is B1, then the current chess record is PATH0+B1, and the number of steps of the replay demonstration data is set to 5 steps. PolicyNet, ValueNet and FastNet are used to jointly analyze the next target move point W2, and the current chess record is updated to PATH0+B1+W2. The above steps are repeated, and black and white chess pieces are placed alternately, and finally a chess record of PATH0+B1+W2+B3+W4+B5 style is formed. B1+W2+B3+W4+B5 is the 5-step chess game trend for the target move point B1, and PATH0+B1+W2+B3+W4+B5 is the replay demonstration data for the target move point B1. According to the above process, corresponding replay demonstration data can also be generated for the target move points corresponding to other bad hand points for users to study and study independently.
[0140] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a device for learning to capture pieces and review the game.
[0141] refer to Figure 7 The device for learning to capture a piece and then replay the game comprises:
[0142] An acquisition module 701 acquires game data of a user after a human-computer interactive game, and determines a plurality of user placement points of the user in the game according to the game data;
[0143] The calculation module 702 determines, for each of the user's placement points, at least one recommended placement point corresponding to the user's placement point; and determines a recommended value of the user's placement point based on the user's placement point and the recommended placement point corresponding to the user's placement point;
[0144] A bad-hand point module 703 determines the user's placement point that meets a predetermined condition among all the recommended values as a bad-hand point;
[0145] Display module 704 determines, for each bad move point, the winning rate value of each recommended move point corresponding to the bad move point, and determines the recommended move point corresponding to the highest winning rate value as the target move point; generates review demonstration data based on the target move point, and displays the review demonstration data to the user.
[0146] In some optional implementations, the calculation module is specifically configured to determine chessboard layout data before the user places a chess piece; input the chessboard layout data into a pre-trained policy network model to obtain at least one of the recommended chess piece placement points output by the policy network model;
[0147] Among them, at least one of the recommended placement points is the same as the user's placement point.
[0148] The output of the strategy network model also includes the recommendation degree corresponding to each of the recommended placement points; the recommendation degree of the user placement point is equal to the recommendation degree of the recommended placement point that is the same as the user placement point;
[0149] The calculation process of the recommended value is:
[0150] Recommendation value = recommendation degree of the user's placement point / total recommendation degrees of recommended placement points with higher recommendation degrees than the user's placement point.
[0151] In some optional embodiments, the display module is configured to use a pre-trained valuation network model and a fast move network model to search for the winning rate value of each recommended placement point in a fast search manner; sort each recommended placement point according to all the winning rate values, determine the highest winning rate value as the highest winning rate value, and determine the recommended placement point corresponding to the highest winning rate value as the target placement point.
[0152] In some optional embodiments, the display module is further configured to determine the chessboard layout data of the target drop point;
[0153] Inputting the chessboard layout data of the target chessboard placement point into a pre-trained strategy network model to obtain at least one next recommended chessboard placement point output by the strategy network model;
[0154] Using the pre-trained valuation network model and the fast move network model, searching for the winning rate value of each of the recommended next move points in a fast search manner;
[0155] Sort each of the recommended placement points for the next step according to all of the winning rate values, determine the highest of the winning rate values as the highest winning rate value, and determine the recommended placement point corresponding to the highest winning rate value as the target placement point for the next step;
[0156] Based on the next target drop point, re-determine the chessboard layout data;
[0157] In response to determining that the move ending condition is met, replay demonstration data is generated and displayed to the user.
[0158] The fast search method is MCTS search.
[0159] Furthermore, the end condition of the MCTS search includes:
[0160] The search time exceeds 5s;
[0161] or,
[0162] The highest win rate value is 50% higher than the second highest win rate value;
[0163] Among them, the second highest winning rate value is the winning rate value closest to the highest winning rate value.
[0164] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for learning to capture pieces and review the game as described in any of the above embodiments is implemented.
[0165] Figure 8 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040 and a bus 1050. It is characterized in that the processor 1010, the memory 1020, the input / output interface 1030 and the communication interface 1040 realize communication connection with each other inside the device through the bus 1050.
[0166] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0167] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0168] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured as a component in the device (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. It is characterized in that the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0169] The communication interface 1040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0170] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0171] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.
[0172] The electronic device of the above-mentioned embodiment is used to implement the corresponding method for learning to capture pieces and review the game in any of the above-mentioned embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0173] Exemplary Program Products
[0174] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method for learning to capture pieces and review the game as described in any of the above embodiments.
[0175] The above-mentioned non-transitory computer-readable storage medium can be any available medium or data storage device that can be accessed by a computer, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical storage (such as CD, DVD, BD, HVD, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)), etc.
[0176] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the game-taking and review learning method as described in any embodiment in the above exemplary method part, and have the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0177] It is known to those skilled in the art that the embodiments of the present invention may be implemented as a system, method or computer program product. Therefore, the present disclosure may be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit", "module" or "system". In addition, in some embodiments, the present invention may also be implemented in the form of a computer program product in one or more computer-readable media, which contains computer-readable program code.
[0178] Any combination of one or more computer-readable media may be used. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive examples) of computer-readable storage media may include, for example: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device.
[0179] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0180] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0181] Computer program code for performing the operation of the present invention may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0182] It should be understood that each box in the flowchart and / or block diagram and the combination of boxes in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine, and these computer program instructions are executed by a computer or other programmable data processing device to produce a device that implements the functions / operations specified in the boxes in the flowchart and / or block diagram.
[0183] These computer program instructions may also be stored in a computer-readable medium that enables a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable medium produce a product that includes an instruction device that implements the functions / operations specified in the blocks in the flowchart and / or block diagram.
[0184] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby enabling the instructions executed on the computer or other programmable device to provide a process for implementing the functions / operations specified in the blocks in the flowchart and / or block diagram.
[0185] In addition, although the operations of the method of the present invention are described in a specific order in the accompanying drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. On the contrary, the steps depicted in the flow chart can be performed in a different order. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step, and / or one step can be decomposed into multiple steps.
[0186] The use of the verbs "comprise", "include" and their conjugations mentioned in the application documents does not exclude the presence of elements or steps other than those recorded in the application documents. The article "a" or "an" before an element does not exclude the presence of a plurality of such elements.
[0187] Although the spirit and principle of the present invention have been described with reference to several specific embodiments, it should be understood that the present invention is not limited to the disclosed specific embodiments, and the division of various aspects does not mean that the features in these aspects cannot be combined to benefit, and this division is only for the convenience of expression. The present invention is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the attached claims. The scope of the attached claims conforms to the broadest interpretation, thereby including all such modifications and equivalent structures and functions.
[0188] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. In line with the concept of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.
[0189] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present application difficult to understand, the known power supply / ground connection with the integrated circuit (IC) chip and other components may or may not be shown in the provided drawings. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform to be implemented in the embodiments of the present application (that is, these details should be fully within the scope of understanding of those skilled in the art). In the case of elaborating specific details (e.g., circuits) to describe exemplary embodiments of the present application, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.
[0190] Although the present application has been described in conjunction with specific embodiments of the present application, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0191] The embodiments of the present application are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of the present application.
Claims
1. A method for learning to capture pieces by reviewing the game, comprising: Obtaining game data of the user after human-computer interactive chess game, and determining multiple user chess placement points of the user in the game according to the game data, wherein the game data is complete chess game data after human-computer interactive chess game, and the game data includes artificial intelligence chess placement points and multiple user chess placement points; For each of the user's placement points, determine at least one recommended placement point corresponding to the user's placement point; and determine a recommended value for the user's placement point based on the user's placement point and the recommended placement point corresponding to the user's placement point; Determine the user's placement point that meets the predetermined condition among all the recommended values as a bad move point; For each of the bad move points, determine the winning rate value of each of the recommended move points corresponding to the bad move point, determine the highest of the winning rate values as the highest winning rate value, and determine the recommended move point corresponding to the highest winning rate value as the target move point; based on the target move point, generate review demonstration data, and display the review demonstration data to the user, wherein the review demonstration data refers to demonstration data of the self-playing alternating move process executed by artificial intelligence based on the target move point.
2. The method according to claim 1, wherein: The determining of at least one recommended placement point corresponding to the user's placement point specifically includes: Determine the chessboard layout data before the user places the chess piece; The chessboard layout data is input into a pre-trained policy network model to obtain at least one recommended move point output by the policy network model.
3. The method according to claim 2, wherein: The output of the strategy network model also includes the recommendation degree of the user's placement point and the recommendation degree corresponding to each of the recommended placement points; The calculation process of the recommended value is: Recommendation value = recommendation degree of the user's placement point / total recommendation degrees of recommended placement points with higher recommendation degrees than the user's placement point.
4. The method according to claim 1, wherein: The step of determining, for each bad move point, a winning rate value of each of the recommended move points corresponding to the bad move point, and determining the recommended move point corresponding to the highest winning rate value as the target move point includes: Using the pre-trained valuation network model and the fast move network model, searching for the winning rate value of each of the recommended move points in a fast search manner; Each of the recommended placement points is sorted according to all of the winning rate values, the highest of the winning rate values is determined as the highest winning rate value, and the recommended placement point corresponding to the highest winning rate value is determined as the target placement point.
5. The method according to claim 1, wherein: Generate replay demonstration data according to the target drop point, and display the replay demonstration data to the user, including: Determine the chessboard layout data of the target drop point; Inputting the chessboard layout data of the target chessboard placement point into a pre-trained strategy network model to obtain at least one next recommended chessboard placement point output by the strategy network model; Using the pre-trained valuation network model and the fast move network model, searching for the winning rate value of each of the recommended next move points in a fast search manner; Sort each of the recommended move points for the next step according to all the win rate values, determine the highest of the win rate values as the highest win rate value, and determine the recommended move point corresponding to the highest win rate value as the target move point for the next step; Based on the next target drop point, re-determine the chessboard layout data; In response to determining that the move ending condition is met, replay demonstration data is generated and displayed to the user.
6. The method according to claim 4 or 5, wherein: The fast search method is MCTS search.
7. The method according to claim 6, wherein: The end conditions of the MCTS search include: The search time exceeds 5s; or, The highest win rate value is 50% higher than the second highest win rate value; Among them, the second highest winning rate value is the winning rate value closest to the highest winning rate value.
8. A device for learning to capture a piece and review the game, comprising: An acquisition module is used to acquire game data of a user after a human-computer interactive game, and determine multiple user chess placement points of the user in the game according to the game data, wherein the game data is complete chess game data after the human-computer interactive game, and the game data includes artificial intelligence chess placement points and multiple user chess placement points; A calculation module, for each of the user's placement points, determines at least one recommended placement point corresponding to the user's placement point; and determines a recommended value of the user's placement point based on the user's placement point and the recommended placement point corresponding to the user's placement point; A bad-hand point module determines the user's placement point that meets a predetermined condition among all the recommended values as a bad-hand point; A display module determines, for each of the bad move points, the winning rate value of each of the recommended move points corresponding to the bad move point, and determines the recommended move point corresponding to the highest winning rate value as the target move point; generates replay demonstration data based on the target move point, and displays the replay demonstration data to the user, wherein the replay demonstration data refers to demonstration data of the self-play alternating move process executed by artificial intelligence based on the target move point.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Chessboard information processing method and device based on artificial intelligence, equipment and medium
CN111475771A