A multimodal interactive chess teaching method and system based on large models

Through large language models and multimodal interaction methods, the problem of insufficient personalized guidance and real-time feedback in Go teaching has been solved, personalized Go teaching has been achieved, and the learning efficiency and Go skills of Go learners have been improved.

CN119896849BActive Publication Date: 2025-10-03SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510042229.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-10-03
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

Existing Go teaching methods lack systematic scientific analysis and personalized guidance, are difficult to adapt to the learning progress and styles of different Go learners, lack instant feedback and high-level playing partners, have fixed teaching content and lack multimodal interaction capabilities, which affects the popularization of logical thinking and Go education.

Method used

A multimodal interactive teaching method based on a large language model is adopted. By acquiring chess game images and voice interaction, the game can be identified and personalized teaching content and difficulty adjustment can be provided. In combination with the chess learner's chess level, mood and style, multimodal interaction methods such as voice, visual and tactile feedback are used to achieve real-time analysis and feedback.

Benefits of technology

It realizes personalized Go teaching, improves learning efficiency and participation, provides in-depth tactical analysis and real-time feedback, adapts to the needs of different Go learners, and enhances the learning experience and improvement of Go skills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119896849B_ABST
    Figure CN119896849B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology and provides a multimodal interactive chess teaching method and system based on a large language model. The method comprises: obtaining a chess game image and identifying the chess game, then generating a chess game description using a large language model in combination with the chess player's chess skill level, and controlling a voice interaction mechanism to feedback the chess game description; analyzing the chess game, teaching difficulty, and player style to obtain a recommended move for the learner, and controlling a chessboard lighting interaction mechanism and / or a voice interaction mechanism to feedback the recommended move to complete the learner's move; determining whether the game is over, and if not, returning to analyzing the robot's move position; and if so, determining the winner. The system uses the large language model to update the learner's chess player style and skill level based on each move's position, the chess game after the move, the learner's emotional state, and the winner, and adjusts the teaching difficulty for the learner. This ensures that each learner can improve their chess skills at a pace that suits them.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a multimodal interactive chess teaching method and system based on a large model. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] In the field of education, chess education is considered an effective means of quality education, with a positive impact on improving students' overall quality. Go, in particular, a highly strategic and competitive intellectual sport, plays a vital role in cultivating logical thinking, decision-making, and psychological fortitude. The rapid development of artificial intelligence (AI) has led to significant breakthroughs in the field of Go. AlphaGo's victory marked an unprecedented level of AI application in the game and opened up new avenues for Go education. However, existing Go AI applications are primarily focused on high-level games and tournaments, with limited support for widespread education and individual learning.

[0004] Currently, traditional Go teaching methods rely primarily on the personal experience of Go teachers, lacking systematic scientific analysis and personalized guidance. Instructional content is often rigid and difficult to adapt to the learning pace and style of individual learners. Furthermore, while Go learning resources are plentiful, they lack effective integration and intelligent processing, making it difficult for learners to obtain targeted guidance and advice. Furthermore, self-learners often face challenges with a lack of immediate feedback and skilled playing partners. Furthermore, most Go teaching software and online platforms lack natural language interaction capabilities, making it difficult to provide in-depth tactical analysis and strategic advice, and making it difficult to tailor instructional content and difficulty to the learner's specific situation. These limitations not only limit Go's potential in improving logical thinking, concentration, patience, and decision-making skills, but also hinder the development of intelligent and personalized Go teaching experiences and the widespread adoption of the game. Summary of the Invention

[0005] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides a multimodal interactive chess teaching method and system based on a large model. Relying on the technical advantages of the large language model, personalized Go teaching is realized. Through the large language model and multimodal interaction method, the teaching content and difficulty can be flexibly adjusted according to the chess learner's chess level, mood and style, and a learning strategy can be tailored for each chess learner, effectively improving learning efficiency and promoting the widespread dissemination and popularization of Go education.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] A first aspect of the present invention provides a multimodal interactive chess teaching method based on a large model, comprising:

[0008] After acquiring a chessboard image and identifying the Go board, the system uses a large language model to analyze the robot's move position and reasons for its moves, taking into account the difficulty of the teaching and the player's style. The system then uses a specific interaction mode to provide feedback on the robot's move position, completes the robot's move, and provides feedback on the reasons for its move through a voice interaction mechanism. After acquiring a chessboard image and identifying the Go board, the system uses a large language model to generate a description of the board based on the player's skill level, controls the voice interaction mechanism to provide feedback on the description, and waits for the player to complete their move.

[0009] Determine whether the game is over. If not, return to reacquire the chess game image and analyze the robot's chess position. If it is over, give the winner. Based on the position of each move, the Go game after the move, the learner's emotional state and the winner, use the large language model to update the learner's learning style and chess level. Determine the interaction mode based on the learner's chess level. And adjust the teaching difficulty of the learner based on the learner's current winning rate.

[0010] Furthermore, the large language model adopts the multimodal large language model GPT-4o, which is pre-trained and fine-tuned through a Go game database; the pre-training adopts a masked language model and a next sentence prediction task; and the fine-tuning is to fine-tune the large language model for the application task based on the pre-training.

[0011] Furthermore, the application task is to generate a chess game description based on the chess game and the chess player's chess skill level; generate the chess player's style and chess skill level based on the position of each chess move, the chess game after the move, the chess player's emotional state and the winner; or, based on the chess game, teaching difficulty and chess player's style, analyze the next move and the reason for the move.

[0012] Furthermore, if the chess player is a beginner, the interaction modes are chessboard lighting interaction and voice interaction;

[0013] If the player is an amateur, the interaction mode is board lighting interaction or machine piece placement interaction.

[0014] If the chess player is a professional chess player, the interaction mode is machine-moving / moving interaction.

[0015] Furthermore, the method for adjusting the teaching difficulty is:

[0016] D new =D current +α·(P correct -P target )

[0017] Among them, D new is the adjusted teaching difficulty, D current is the teaching difficulty before adjustment, α is the learning rate, P correct is the winning rate of chess learners so far, P target is the target win rate.

[0018] Furthermore, the Go game recognition step includes: graying, adaptive filtering, contrast enhancement, perspective transformation and interpolation of the game image, and then using a convolutional neural network model to identify the chess piece coordinates and chess piece types in the game image to obtain the Go game.

[0019] A second aspect of the present invention provides a multimodal interactive chess teaching system based on a large model, comprising:

[0020] The multimodal interactive teaching module is configured to: obtain a chess game image and identify the Go game position; then, based on the teaching difficulty and the player's style, analyze the robot's chess position and reasons for the move through a large language model; feedback the robot's chess position through a predetermined interaction mode, complete the robot's move, and feedback the reasons for the move through a voice interaction mechanism; obtain a chess game image and identify the Go game position; then, based on the player's skill level, generate a chess game description through a large language model; control the voice interaction mechanism to feedback the chess game description; and wait for the player to complete the move;

[0021] The learner information update module is configured to: determine whether the game is over. If not, return to re-acquire the chess game image and analyze the robot's chess position; if it is over, give the winner, and use the large language model to update the learner's learning style and chess level based on the position of each chess move, the Go game after the move, the learner's emotional state and the winner, and determine the interaction mode based on the learner's chess level; and adjust the teaching difficulty of the learner based on the learner's current winning rate.

[0022] Furthermore, the large language model adopts the multimodal large language model GPT-4o, which is pre-trained and fine-tuned through a Go game database; the pre-training adopts a masked language model and a next sentence prediction task; and the fine-tuning is to fine-tune the large language model for the application task based on the pre-training.

[0023] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multimodal interactive chess teaching method based on a large model as described above.

[0024] A fourth aspect of the present invention provides a computer device comprising a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, wherein when the processor executes the program, the steps of the multimodal interactive chess teaching method based on a large model as described above are implemented.

[0025] Compared with the prior art, the present invention has the following beneficial effects:

[0026] Large-model personalized teaching strategy: This invention uses a large language model to more deeply analyze the behavior, style and chess skills of chess learners. It dynamically adjusts the teaching content and difficulty according to the chess learners' learning progress, level, emotional state and style, and combines multimodal interaction methods to provide a personalized learning experience, ensuring that every chess learner can improve his or her chess skills at a pace that suits him or her.

[0027] Multimodal interaction improves learning engagement: The present invention significantly improves the engagement and learning experience of chess learners through multimodal interaction methods, such as voice, visual and tactile feedback. In contrast, existing technologies only provide a single modality or limited interaction methods, and fail to fully utilize the potential of multimodal interaction to enhance learning motivation.

[0028] Large-model-enhanced intelligent game analysis: This invention uses a large language model to conduct in-depth analysis of Go games, providing players with accurate game descriptions. This is relatively rare in existing technologies. Most existing technologies lack such in-depth analysis capabilities, or require human assistance to provide similar analysis.

[0029] Real-time feedback and adaptability of large-scale model optimization: The large language model in this invention can process the interactive data of chess learners in real time, provide feedback on the learners' chess playing style and chess skills, and dynamically adjust the teaching content and difficulty based on the feedback. This adaptability is not available in existing technologies, which require more manual intervention or cannot respond to the needs of chess learners in real time. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0031] Figure 1 This is a diagram of a multimodal interactive chess teaching system and device architecture based on a large model according to an embodiment of the present disclosure;

[0032] Figure 2 A diagram of a Go teaching language and lighting interaction device according to an embodiment of the present disclosure;

[0033] Figure 3This is a diagram of the interactive device for Go teaching raising / dropping pieces according to an embodiment of the present disclosure;

[0034] Figure 4 Schematic diagram of the hardware platform of a large-scale model-based multimodal interactive chess teaching system according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0035] To make the objectives, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0036] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0037] Example 1

[0038] This embodiment provides a multimodal interactive chess teaching method based on a large model.

[0039] This embodiment provides a multimodal interactive chess teaching method based on a large model. By acquiring and analyzing Go game data, a rich chess game database is constructed and maintained. This database contains a variety of historical chess games and tactics, chess games played by chess learners, and related chess game teaching materials, etc., which are used to train and optimize the pre-trained large language model so that it can understand and generate natural language content related to chess.

[0040] This embodiment provides a large-model-based multimodal interactive chess teaching method, which uses a large multimodal language model, such as GPT-4o, for pre-training and fine-tuning, so that the large language model can understand and process images, text descriptions, and voice commands of Go games; and by designing specific pre-training tasks, such as masked language model (MLM) and next sentence prediction (NSP) tasks, the large language model learns the coherence and strategy analysis of the Go game.

[0041] This embodiment provides a multimodal interactive chess teaching method based on a large model. Through a large language model and multimodal interaction, it dynamically adjusts the teaching content and difficulty according to the chess learner's learning progress, level, emotional state and style, thereby providing a personalized learning experience.

[0042] This embodiment provides a multimodal interactive chess teaching method based on a large model, which integrates an adaptive teaching module and a large language model. This integration not only enhances the understanding and generation capabilities of Go teaching content, but also ensures that the teaching content can meet the needs of different Go learners, thus achieving truly personalized teaching.

[0043] This embodiment provides a multimodal interactive chess teaching method based on a large model, such as Figure 1 As shown, the following steps are included:

[0044] Step 1: Obtain and process chess game history data.

[0045] Step 101: Multi-source data collection of historical Go games: Collect historical Go game data from various Go competition records, online Go game platforms, etc.; at the same time, use OCR technology and computer vision technology to extract historical Go game data from historical Go books and Go explanation videos.

[0046] Among them, some of the chess game history data are historical chess games in SGF format containing chess game comments, and this part of the data does not need to perform step 102.

[0047] Among them, the collected chess game historical data includes: the coordinates of each move, the type of move, the description text of the move position, the reason for the move (reason for the move), the game after the move, the game comments (game description), the difficulty of the game, the emotional state of the players, and the final outcome.

[0048] Among them, the description text of the chess piece placement position can be: coordinate indication type (such as placing the black piece at the chessboard coordinate (A, 4)), relative position type (such as placing the white piece at the empty intersection immediately to the right of the black piece), or based on a fixed pattern or layout description type (such as, according to the star position small flying hanging corner fixed pattern, the black piece is placed at the star position hanging corner in the upper right corner), etc.

[0049] For example, a description of a move might read: "After defending the corner with a small flying piece at the star position, continue to move to the side." The rationale for this move is to expand one's sphere of influence on this side through reasonable spacing, creating a wider pattern with the potential to convert this area into tangible points (territory) in the future. This is similar to the once-popular "three-in-a-row" layout, which involves placing pieces in star positions on one side of the board and then moving to the side to create a grand pattern, luring the opponent to attack and disrupt, and then using the player's strength to attack and profit, creating more space.

[0050] Among them, the chess game records whether there is a chess piece and the type of chess piece at each of the 361 intersections on the chessboard.

[0051] Among them, the example of chess game description: This is a middle stage of the game, with White going first. White has formed a certain solid position in the upper right and lower left corners, while Black is eyeing it with several external positions. The distribution of pieces on both sides is relatively balanced, and they are in the stage of testing each other and looking for breakthrough points.

[0052] Step 102, data conversion: Develop or use existing conversion tools to convert the game information (text describing the position of each move, move coordinates, move type, reason for the move, the game after the move, game comments (game description), game difficulty, player emotional state, and final outcome) into SGF format.

[0053] Among them, the SGF file uses a tree structure to represent the game: a complete game is divided into multiple nodes, a node represents a chess move (a chess step), and each node is associated with a description text of the position of the piece, the coordinates of the piece, the type of piece, the reason for the piece, the game after the piece is placed, the game commentary (game description), the game difficulty (can be used as teaching difficulty), and the emotional state of the chess player; the points are connected by branches, and the branches represent the direction from one chess game to the next; each SGF file is associated with the final win or loss result.

[0054] Among them, data conversion involves writing scripts or using specialized libraries to handle the format conversion of chess game information to ensure data consistency and readability.

[0055] Among them, data conversion converts the chessboard positions from the traditional coordinate system to a numerical range. For example, the 361 intersections of the chessboard are mapped to values ​​​​from 0 to 360. This conversion helps large language models process and learn the relative positions of the chessboard space.

[0056] Step 103, data cleaning and annotation: The converted SGF files are cleaned, including: checking the legitimacy of the moves, the accuracy of the wins and losses, and the appropriateness of the game description, ensuring the accuracy of the move information, and removing records with format errors or incompleteness; then, the SGF files are annotated, including the move type of each node (such as capture, capture, and robbery), the game stage (such as opening, middle game, and closing game), and the styles and skill levels of both players associated with each SGF file, to obtain a dataset.

[0057] Step 104, data augmentation: Enhance the dataset by simulating different chess game variations or using known chess game theories, including randomly changing part of the chess game or generating new chess game variants based on specific strategies to improve the generalization ability of the large language model.

[0058] Step 105, data integration and storage: Store the dataset in a database suitable for large-scale data processing, such as a relational database or a NoSQL database, for subsequent model training and analysis; for large-scale datasets, a distributed file system or cloud database is required to optimize storage and access efficiency; and thus obtain a Go game database.

[0059] Through the above steps, a structured, cleaned, and optimized Go game dataset can be prepared to provide high-quality input for large language model training. This ensures that the large language model can learn effective features from the dataset, thereby providing accurate predictions and personalized teaching content.

[0060] Step 2: Pre-training and fine-tuning of large language models.

[0061] Step 201: Large language model selection: Use an advanced multimodal large language model, such as GPT-4o.

[0062] Among them, GPT-4o can process a variety of data types including text and images, providing a basis for multimodal understanding of Go games.

[0063] Step 202: Pre-training task design.

[0064] Pre-training tasks were designed, including masked language model (MLM) and next sentence prediction (NSP) tasks, which aim to train a large language model to understand Go positions, game descriptions, and the learner's voice commands (piece placement descriptions). The MLM task randomly masks some words in the input text (game descriptions or piece placement descriptions), allowing the large language model to predict these words. The NSP task, on the other hand, constructs a sample consisting of two sentences, allowing the large language model to determine whether the second sentence is the next sentence of the first. Furthermore, the MLM task randomly masks some piece positions in the input game, allowing the large language model to predict these positions. Meanwhile, the NSP task constructs a sample consisting of two games, allowing the large language model to determine whether the second game is the next step of the first. This is very helpful for understanding the coherence of Go games.

[0065] Suppose there is a text description sentence of a Go game: "White has an advantage at [MASK]." In the MLM task, the large language model needs to predict the masked word "corner". The output of the large language model is a probability distribution P(w i |w1,...,w i-1 ,w i+1 w1,...,w N ), where w i is the masked word, and w1,...,w i-1 ,w i+1 w1,...,w N It is the unobscured word in the sentence.

[0066] Suppose two sentences are given. The first sentence is "Black launched an offensive in the center." and the second sentence is "This forced White to strengthen its defense." In the NSP task, the large language model needs to determine whether the second sentence is the next sentence of the first sentence. The large language model outputs a probability P(s2|s1), which represents the probability that the second sentence follows the first sentence.

[0067] During the pre-training process, the large language model sees a large number of chess game images and text pairs as well as the learner's voice instructions. The large language model needs to learn how to extract useful features from this data and make predictions on the MLM and NSP tasks. The training goal of the large language model is to minimize the loss function L = L MLM +L NSP , where L MLM is the loss of MLM mission, L NSP It is the loss of NSP mission.

[0068] Step 203, model fine-tuning: Based on pre-training, a fine-tuning dataset is constructed for specific application tasks, and the large language model is fine-tuned. The parameters of the large language model are adjusted to minimize the loss function of the specific task and improve the performance of the large language model in Go game analysis and teaching.

[0069] In this embodiment, the application tasks include:

[0070] (1) Application Task 1: Generate the chess player's chess skill level and style by analyzing the position of each chess move, the chess situation after the move, the emotional state of the chess player and the winner, and the large language model;

[0071] (2) Application Task 2: Based on the chess game and the chess learner’s chess skill level, the large language model is used to analyze and generate a detailed description of the chess game;

[0072] (3) Application Task 3: Based on the chess game, teaching difficulty, and learner’s style, the large language model analyzes and provides specific recommendations for the next move, chess position information, and reasons for the move.

[0073] When performing application tasks, the large language model performs multimodal data fusion: it integrates Go game images, text descriptions (game descriptions), and the learner's voice commands. This requires the large language model to simultaneously process and understand information from different modalities, improving its generalization and in-depth understanding of Go games. For example, if the game shows White placing a stone in the star position in the upper left corner, and the text description reads "White takes an offensive move in the star position," the large language model fuses information from these two modalities by learning a joint representation R = f(I, T), where I represents the game image features and T represents the text (voice command text or game description) features.

[0074] Step 204: Performance Evaluation and Optimization: Cross-validation is used to evaluate the performance of the large language model on different fine-tuning datasets to ensure its generalization capabilities. Furthermore, based on feedback from players and learning outcomes, the parameters of the large language model are further adjusted to achieve continuous optimization.

[0075] Among them, the data sets used for pre-training and fine-tuning of the large language model all come from the Go game database in step 1.

[0076] Step 3, obtaining the existing information of the chess learner: in response to the user registration instruction, the external display interaction processing module controls the external display interaction mechanism to display the registration interface, receives the interaction information fed back by the chess learner, and obtains the teaching difficulty and interaction mode selected by the user.

[0077] The interactive information fed back by chess learners is their basic information, including: name, age, gender, contact information, competition information, honors obtained, chess learner's style and chess skill level, etc.

[0078] Step 4: Based on the teaching difficulty, interaction mode, learner style, and chess skill level obtained in step 3 or step 5, complete the multimodal interactive AI chess game;

[0079] Step 401, game start: prepare to start the game, the chess learner or AI interaction processing module is in standby state, the chessboard detection camera and multimodal interaction device are ready, and the chess play order is confirmed.

[0080] Step 402, real-time chess game data acquisition: the chessboard detection camera collects chess game images, and converts the chess game images into digital signals and sends them to the computing host.

[0081] Among them, the chessboard detection camera uses a high-resolution camera.

[0082] Step 403: Real-time chess game data processing and chess game status update: After the computing host receives the digital signal from the chessboard detection camera, it updates the chess game to ensure that the current state of the chess game is consistent with the actual chess piece positions.

[0083] (1) Preprocessing: For Go game images, the image preprocessing method enhanced by deep learning is used to improve the quality of the game images and provide clear input for the CNN model.

[0084] Among them, deep learning enhanced image preprocessing methods include: grayscale, adaptive filtering, contrast enhancement, etc.

[0085] (2) Perspective transformation and interpolation: Automated calibration and multi-view processing of chessboard images to achieve more accurate perspective transformation and obtain a bird's-eye view; bilinear interpolation to improve resolution and image quality; exploring deep learning-based methods to automatically correct chessboard image distortion, such as using deep networks to predict transformation parameters and achieve perspective conversion.

[0086] (3) CNN model recognition: Develop a customized convolutional neural network (CNN) model to identify chess piece coordinates and chess piece type (black or white) in chess game images.

[0087] The CNN model accurately identifies chess pieces by learning the mapping between image features and board states. Convolutional neural network training involves a large amount of chess game image data and may employ data augmentation techniques, such as rotation and scaling, to improve the model's generalization capabilities.

[0088] In this embodiment, a high-resolution camera is used to capture chess game images in real time, and a convolutional neural network (CNN) is used to perform image recognition and processing.

[0089] The image is obtained using convolutional neural networks (CNNs) to identify the positions of chess pieces. For each piece identified, the probability of its position on the board is calculated, which can be expressed by the following formula:

[0090]

[0091] Among them, P(position) is the probability that the chess piece is at a specific position, z is the eigenvalue of the position, N is the total number of positions on the chessboard, and e is the base of the natural logarithm.

[0092] In the context of Go, the local perception capabilities of CNN models can capture local patterns and features in Go board images. The Go board is a two-dimensional structure, and the state of each position is related to surrounding positions. Therefore, CNNs are more suitable for processing Go's local structure. Furthermore, CNNs use parameter sharing, which reduces the number of model parameters and improves training efficiency and generalization. When processing Go board images, CNNs can ignore the absolute positions of the pieces and focus on the local structure and relative relationships of the board.

[0093] (4) SGF format conversion: The chess game recognized by the CNN model is converted into SGF format to ensure the compatibility of the data with Go software and database, facilitating further analysis and storage.

[0094] Step 404, multimodal interactive AI chess game: After obtaining the chess game image and identifying the chess game status, combined with the chess learner's chess learning information data (teaching difficulty, chess learner's style), a large language model is used to analyze and generate recommended moves, chess position information and move reasons, and the robot completes the chess round through the corresponding interaction mode.

[0095] For Go beginners, a large language model is used to analyze game and player information, generating basic game descriptions and simple recommended moves. Based on the generated recommended moves and position information, the AI ​​controls the voice interaction and board lighting interaction to provide multimodal prompts for player placement, guiding the player through the AI's turn. The board lighting interaction highlights the recommended moves, providing visual cues for beginners to more intuitively understand the game and recommended moves. The voice interaction provides verbal guidance and explains the rationale for the recommended moves in simple, understandable language, helping beginners understand the importance and logic behind each step. The move rationale begins with teaching the basic rules and gradually introduces fundamental tactics. As players master the fundamental concepts, the difficulty increases to accommodate their progress and needs.

[0096] For amateur Go players: A large language model provides in-depth analysis of the game, providing more complex tactical analysis and a variety of possible moves. Based on personal preference, players can choose to use the board lighting interface to highlight recommended moves, or the machine move / drop interface to simulate their opponent's moves. Voice interaction provides deeper tactical analysis, encouraging players to explore different moves. Reasoning for each move includes discussion of tactical variations and game replays, improving their tactical understanding and application.

[0097] For professional Go players: A large language model provides advanced tactical analysis, including possible opponent responses and predictions of advanced strategies. A machine-based move / place interaction mechanism simulates the opponent's moves, providing a realistic playing experience for professional Go players. A voice interaction mechanism provides real-time tactical advice and opponent analysis, aiding professional Go players in decision-making during games. The move reasoning system focuses on in-depth research into advanced strategies and the application of psychological tactics to improve the competitive level of professional Go players.

[0098] In summary, according to the needs and preferences of chess learners, we can flexibly choose the interaction method to complete multimodal interactive AI chess games and realize AI chess teaching.

[0099] For players of varying skill levels, the AI ​​can adjust the difficulty of its teaching to suit them. By incorporating the difficulty parameter into the large language model, the convolutional neural network model can provide players of varying skill levels with moves of varying difficulty. For example, for beginners, the AI ​​will employ more basic and straightforward moves, simplifying tactics; for amateurs, it will employ intermediate tactics, such as set patterns and strategic positions; and for professionals, it will employ advanced tactics and strategies, including complex set pattern variations and endgame techniques.

[0100] Step 405: Obtain a chess game image, and identify the chess piece coordinates and chess piece type (black or white) in the chess game image using the method of step 403 to obtain the chess game.

[0101] Step 406: Generate a game description based on the game and the player's skill level through a large language model, and provide feedback through a voice interaction mechanism.

[0102] When generating chess game descriptions, the large language model chooses to use different languages ​​and explanation depths based on the player's chess skill level: for beginners, the game description may be simpler and more direct, focusing on explaining the basic rules; for amateur players, the game description includes more tactical analysis, such as fixed pattern selection; for professional players, the game description may include more tactical analysis and discussion of move variations, as well as in-depth analysis of advanced tactics and strategies.

[0103] Step 407, AI recommends moves: Based on the game situation, teaching difficulty, and player style, the current game situation is analyzed through a large language model, and moves are recommended to the learner. Feedback can be provided through the interactive mechanism of the board lighting or through voice, and encouragement or explanation of the strategic significance of the recommended moves (reasons for the moves) are given to help the learner gain a deeper understanding of Go strategies.

[0104] This step is similar to step 404, except that step 404 is to complete the robot's move, while step 407 is to complete the chess learner's move.

[0105] Among them, the large language model adjusts the complexity of recommended moves according to the difficulty of teaching: for beginners, the recommended moves may focus more on basic rules and simple tactics; for amateur players, the recommended moves involve intermediate tactics, such as pattern construction; and for professional players, the recommended moves include deeper tactics and strategies, recommend advanced tactics, and explain subsequent changes; for amateur or professional players, you can choose not to provide feedback on the recommended moves for learners.

[0106] During the game, AI uses speech synthesis technology to make suggestions on where to place the pieces, such as "It is recommended to place the piece at position C10" or "You have made great progress, keep it up! Next, I suggest you make some arrangements in the middle area of ​​the board," thereby enhancing the learning motivation of chess learners.

[0107] It should be noted that on a 19-way Go board, the coordinate system usually consists of letters and numbers to identify specific positions on the board. Letters from A to T (or S) represent horizontal columns, and numbers from 1 to 19 represent vertical rows. Therefore, C10 represents the intersection of the 3rd column (column C) and the 10th row. This coordinate system helps to describe the position of chess pieces in the game, facilitating communication and analysis. For example, in Go teaching or playing, the instructor may instruct the student or opponent to place the chess piece at position C10. This precise coordinate indication is crucial for understanding the game and executing the strategy.

[0108] Step 408: The chess learner completes the move.

[0109] As an implementation method, the chess player can respond with voice "C10" to execute the recommended move, or can respond with voice to the move he or she decides on, such as "C11". Then, based on the chess game and the voice response (the description of the chess position, or the description of the chess position and the chess game description), the chess position is generated through a large language model, and then the machine's lifting / dropping mechanism makes corresponding actions on the chessboard.

[0110] As another implementation method, the chess learner directly moves the chess pieces.

[0111] Step 409: After the player has finished making their move, determine whether the game is over. If so, determine the winner. Otherwise, return to step 402, and the player and the robot continue the AI ​​game, continuously providing real-time analysis and recommendations.

[0112] Among them, the voice interaction processing module uses advanced voice recognition technology to convert the Go learner's voice commands into text, and can provide natural voice feedback, allowing Go learners to communicate through natural language and obtain Go teaching content and strategy suggestions; it can also explain the strategic significance of its recommended moves through voice, enhancing the learner's understanding of Go strategy.

[0113] The interactive board lighting processing module, based on AI analysis, can illuminate the chessboard from the top down to designated locations on the chessboard, such as C10, to provide visual cues to the player. To further enhance the interactive experience, the interactive board lighting mechanism can employ dynamic lighting effects, such as gradients or flashes, to attract the player's attention and enhance retention. This visual cue is crucial for helping players focus and understand the tactics of the game. This module dynamically adjusts the lighting on the board, such as color or brightness gradients, to indicate the urgency or importance of a move, providing intuitive visual feedback. This intuitive visual cue helps players quickly identify AI-recommended moves while also increasing the interactivity and fun of the game. The interactive board lighting processing module can also confirm the effectiveness of a move or provide immediate feedback through lighting changes after the player completes a move.

[0114] The machine's interactive processing module for moving and placing pieces not only enhances players' engagement but also helps them improve their skills through practice by providing real-time feedback and personalized move recommendations. This module takes into account the complexity of Go and the specific needs of teaching, ensuring that every player can improve their skills at a pace that suits them.

[0115] The external display interactive processing module displays real-time game analysis and AI-generated strategic recommendations on an external display. When the AI ​​recommends move C10, the display highlights that position and provides tactical analysis. The external display also provides a visual representation of the game, helping players understand the layout and potential moves from different perspectives. Furthermore, the display can be used to showcase Go-related cultural and historical knowledge, enhancing players' understanding of Go culture.

[0116] Step 5: Real-time information feedback for chess learners.

[0117] During the game, the player's facial expressions and body language (such as personal posture, hand movements, etc.) are captured through the player detection camera, and their emotional state is analyzed in real time using computer vision technology.

[0118] First, by analyzing the position of each chess move in the game, the situation after the move, the emotional state, and the final outcome, a large language model is used to predict the changes in the chess learner's style and chess skill level, and to update the chess learner's style and chess skill level to predict future performance and learning needs, thereby achieving personalized customization of teaching content.

[0119] Then, adjust the interaction mode according to the chess skills of the learners.

[0120] Finally, calculate the winning rate of the chess learner so far and adjust the difficulty of the chess learning content:

[0121] D new =D current +α·(P correct -P target )

[0122] Among them, D new Is the new teaching difficulty, D current is the current difficulty, α is the learning rate, P correct is the proportion of correct answers given by chess learners (the winning rate so far), P target is the target win rate.

[0123] Through the above steps, the large language model for Go teaching can not only analyze the Go game, but also dynamically adjust the difficulty of the teaching content according to the learning progress of the Go learners, providing Go learners with a personalized, interactive and fruitful learning experience, so as to achieve personalized Go teaching and thus improve learning effects.

[0124] In this embodiment, the teaching module can ensure that each chess learner can improve his or her chess skills at a pace that suits him or her while enjoying personalized teaching content. This strategy not only improves learning efficiency, but also enhances the learner's learning motivation and satisfaction.

[0125] In this embodiment, a teaching experience adapted to the needs of chess learners of different levels is provided, thereby improving the learning efficiency and chess skills of the learners. This innovative interactive design can stimulate the interest of chess learners and help them make rapid progress on the road of Go.

[0126] Example 2

[0127] This embodiment provides a multi-modal interactive chess teaching system based on a large model, such as Figure 2 、 Figure 3 and Figure 4 As shown, it specifically includes: a chessboard detection camera 1, a chess learner detection camera 4, a chessboard lighting interaction mechanism 2, a voice interaction mechanism 3, a local controller 5, a standard Go chessboard 6, a machine lifting / dropping interaction mechanism 7, a Go robot body 8, a power supply, a computing host, a cloud server, an external display interaction mechanism, and a chess piece box.

[0128] Chessboard detection camera: captures the position and status of chess pieces on the board, providing clear chess game image input.

[0129] Player's posture detection camera: monitors the player's playing behavior and posture, captures the player's facial expressions and body language, and analyzes the player's playing status in real time. This is used as information input for the adaptive teaching module to predict the player's chess playing style and chess skill level based on the player's feedback and behavior.

[0130] Go Robot Body: Made of high-strength materials, it ensures stability and durability. Its modular design facilitates maintenance and upgrades, adapting to technological developments and changes.

[0131] Standard Go board: The standard size of the Go board meets international standards, providing a standard game platform.

[0132] Piece box: stores black / white Go pieces, divided into black chess box and white chess box.

[0133] The voice interaction mechanism includes: a microphone array: used to capture the user's voice commands and support far-field voice recognition; a speaker: provides clear voice feedback for teaching guidance and interaction.

[0134] The interactive board lighting mechanism includes: an infrared lighting component that provides uniform illumination, ensuring clear visibility of the chess pieces on the board, aiding image recognition and improving accuracy; and a rotating component that creates dynamic lighting effects to enhance the visual experience. The lighting effects also highlight key teaching points, helping students understand the game.

[0135] The robot's interactive mechanism for lifting and placing chess pieces includes: the X, Y, and Z three-axis moving parts of the robotic arm: the three-axis moving parts are used to achieve precise positioning of the robotic arm on the chessboard, enabling flexible movement and positioning of the robotic arm on the chessboard, simulating the lifting and placing operations of the human arm; the lifting and placing parts: a precisely designed clamping mechanism ensures the stable clamping and placement of chess pieces.

[0136] Power supply: Provides stable power supply for the entire teaching system. It can use built-in batteries or external power adapter to support long-term operation of the teaching system, including overcharge, over-discharge, overheating and short-circuit protection to ensure safe power use and extend battery life.

[0137] The local controller is equipped with a high-performance processor for executing algorithms and data processing. It also has ample RAM to ensure smooth processing of large-scale chess game data. High-speed storage devices such as SSDs improve data read and write speeds and reduce latency. It also includes a voice interaction processing module, a chessboard lighting interaction processing module, a machine piece placement interaction processing module, and an external display interaction processing module.

[0138] Computing Host: Equipped with a high-performance processor and large-capacity random access memory (RAM), the system is highly efficient and fluid when executing complex algorithms and processing large amounts of Go game data. The processor's high-speed computing power supports real-time game analysis and dynamic adjustment of teaching strategies, while simultaneously running multiple applications, including image recognition, natural language processing, and machine learning models. The computing host includes a built-in, pre-trained and fine-tuned large language model that can understand and generate natural language content related to Go, generate Go teaching strategies, provide accurate game descriptions, and recommend moves and positions.

[0139] Cloud Server: Dynamically allocates computing resources based on demand, processes large-scale data sets and complex machine learning tasks, synchronizes and backs up chess game data and user learning progress to the cloud, and ensures data security and cross-device access.

[0140] External display interaction mechanism: provides a high-resolution touch screen display for displaying chess game analysis, teaching content and interactive interface.

[0141] The computing host is equipped with a multimodal interactive teaching module and a chess player information update module.

[0142] The multimodal interactive teaching module provides multimodal prompts to the player based on the generated recommended moves and chess position information, controlling the voice interaction mechanism and the chessboard lighting interaction mechanism to guide the player through the AI's chess rounds. Furthermore, based on the generated recommended moves and chess position information, the system sends chess position instructions to the robot through the robot's mechanical lifting / dropping mechanism, which then autonomously completes the AI's rounds. The system flexibly selects interaction methods based on the player's needs and preferences, enabling multimodal interactive AI chess games and AI-based chess teaching. Image recognition technology is used to acquire and analyze chess game images, and advanced image processing techniques such as grayscaling, adaptive filtering, contrast enhancement, and perspective transformation are employed to identify the coordinates and types of chess pieces within the game. A large language model is then used to understand and parse the player's natural language input, including questions and feedback, to generate personalized teaching content. The module generates detailed game descriptions and provides feedback to the learner via a voice interaction mechanism. It also analyzes and provides recommended moves based on the game, the teaching difficulty, and the learner's style. These moves are provided to the learner as visual cues via a board lighting interaction mechanism or as verbal guidance via a voice interaction mechanism. Furthermore, the module provides multimodal feedback and interaction methods, including real-time voice feedback, visual cues, and tactile feedback, simulating a realistic game experience. The teaching difficulty level dynamically adjusts based on the learner's win rate and target win rate, ensuring that the learner consistently learns at an appropriate level of challenge. The module also monitors the learner's emotional state and adjusts teaching strategies accordingly. Finally, the module tracks the learner's learning progress and, based on each move's placement, the resulting game, the learner's emotional state, and the winning side, leverages a large language model to update the learner's skill level and style. This continuously optimizes the teaching content to provide a more effective learning experience.

[0143] The chess learner information update module is configured to: determine whether the game is over. If not, return to re-acquire the chess game image and analyze the robot's chess position; if it is over, give the winner, and use the large language model to update the chess learner's style and chess level based on the position of each chess move, the chess game after the move, the chess learner's emotional state and the winner, and determine the interaction mode based on the chess learner's chess level; and adjust the teaching difficulty of the chess learner based on the chess learner's current winning rate.

[0144] Furthermore, the method of adjusting the teaching difficulty is as follows:

[0145] D new =D current +α·(P correct -P target )

[0146] Among them, D newis the adjusted teaching difficulty, D current is the teaching difficulty before adjustment, α is the learning rate, P correct is the winning rate of chess learners so far, P target is the target win rate.

[0147] Furthermore, the chess game recognition step includes: grayscale conversion, adaptive filtering, contrast enhancement, perspective transformation and interpolation of the chess game image, and then using a convolutional neural network model to identify the chess piece coordinates and chess piece types in the chess game image to obtain the chess game.

[0148] It should be noted here that the various modules in this embodiment correspond one-to-one to the various steps in Example 1, and the specific implementation processes are the same, which will not be repeated here.

[0149] Example 3

[0150] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the multimodal interactive chess teaching method based on a large model as described in the first embodiment above are implemented.

[0151] Example 4

[0152] This embodiment provides a computer device comprising a computer-readable storage medium (volatile memory and non-volatile storage medium), a processor, a communication interface (i.e., a network interface), and a computer program stored on the computer-readable storage medium and executable by the processor. The processor, the communication interface, and the computer-readable storage medium may be connected via a bus or other means. The communication interface is configured to receive and transmit data, and when the processor executes the program, it implements the steps of the large-model-based multimodal interactive chess teaching method described in the first embodiment.

[0153] Among them, any reference to memory, storage, database or other media provided in this application and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSR DRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0154] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0155] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0156] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0157] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A multimodal interactive chess teaching method based on a large language model, characterized by: include: After acquiring a chessboard image and identifying the Go board, the system uses a large language model to analyze the robot's move position and reasons for its moves, taking into account the difficulty of the teaching and the player's style. The system then uses a specific interaction mode to provide feedback on the robot's move position, completes the robot's move, and provides feedback on the reasons for its move through a voice interaction mechanism. After acquiring a chessboard image and identifying the Go board, the system uses a large language model to generate a description of the board based on the player's skill level, controls the voice interaction mechanism to provide feedback on the description, and waits for the player to complete their move. Determine whether the game is over. If not, return to reacquire the chess game image and analyze the robot's chess position. If the game is over, determine the winner. Based on the placement of each move, the chess game after the move, the learner's emotional state, and the winner, use the large language model to update the learner's learning style and chess skill level. Determine the interaction mode based on the learner's chess skill level. Adjust the teaching difficulty of the learner based on the learner's current winning rate. The large language model adopts the modal large language model GPT-4o, which is obtained by pre-training and fine-tuning using a Go game database; the pre-training uses a masked language model and a next sentence prediction task; the fine-tuning is to fine-tune the large language model based on the pre-training and for the application task; The method for adjusting the teaching difficulty is as follows: ;in, is the adjusted teaching difficulty, is the teaching difficulty before adjustment, is the learning rate, is the winning rate of chess learners so far. is the target win rate.

2. A multimodal interactive chess teaching method based on a language large model as claimed in claim 1, characterized in that: The application tasks are: generating a chess game description based on the chess game and the learner's chess skill level; generating the learner's style and chess skill level based on the position of each chess move, the chess game after the move, the learner's emotional state and the winner; or, based on the chess game, teaching difficulty and learner's style, analyzing the next move and the reason for the move.

3. A multimodal interactive chess teaching method based on a language large model as claimed in claim 1, characterized in that: If the player is a beginner, the interaction modes are board lighting interaction and voice interaction. If the player is an amateur, the interaction mode is board lighting interaction or machine piece placement interaction. If the chess player is a professional chess player, the interaction mode is machine-moving / moving interaction.

4. A multimodal interactive chess teaching method based on a language large model as claimed in claim 1, characterized in that: The Go game recognition step includes: graying, adaptive filtering, contrast enhancement, perspective transformation and interpolation of the game image, and then using a convolutional neural network model to identify the coordinates and types of chess pieces in the game image to obtain the Go game.

5. A multimodal interactive chess teaching system based on a large language model, characterized by: include: The multimodal interactive teaching module is configured to: obtain a chess game image and identify the Go game position; then, based on the teaching difficulty and the player's style, analyze the robot's chess position and reasons for the move through a large language model; feedback the robot's chess position through a predetermined interaction mode, complete the robot's move, and feedback the reasons for the move through a voice interaction mechanism; obtain a chess game image and identify the Go game position; then, based on the player's skill level, generate a chess game description through a large language model; control the voice interaction mechanism to feedback the chess game description; and wait for the player to complete the move; The learner information update module is configured to: determine whether the game is over; if not, return to reacquire the chess game image and analyze the robot's chess position; if the game is over, indicate the winner, and use the large language model to update the learner's learning style and chess skill level based on the position of each move, the chess game after the move, the learner's emotional state, and the winner; determine the interaction mode based on the learner's chess skill level; and adjust the learning difficulty of the learner based on the learner's current winning rate; The large language model adopts the modal large language model GPT-4o, which is obtained by pre-training and fine-tuning using a Go game database; the pre-training uses a masked language model and a next sentence prediction task; the fine-tuning is to fine-tune the large language model based on the pre-training and for the application task; The method for adjusting the teaching difficulty is as follows: ;in, is the adjusted teaching difficulty, is the teaching difficulty before adjustment, is the learning rate, is the winning rate of chess learners so far. is the target win rate.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of a multimodal interactive chess teaching method based on a large language model as described in any one of claims 1 to 4 are implemented.

7. A computer device comprising a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, wherein: When the processor executes the program, the steps of a multimodal interactive chess teaching method based on a large language model as described in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Weiqi referee system based on MLP neural network and computer vision

    CN110399888A

  • Vision system for monitoring board games and method thereof

    US20170100661A1