Autonomous game operation (AI) proxy method and device

Through the autonomous game operation AI agent with a multi-model collaborative architecture, the game screen is analyzed in real time and operation instructions are generated, which solves the problem that the player's operation cannot keep up with the consciousness, provides substantial operation assistance, adapts to dynamic changes in the game, and improves the inclusiveness of the game.

CN120706580AActive Publication Date: 2025-09-26HAIMA CLOUD TIANJIN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511211794.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-09-26
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing game assistance systems and automation tools cannot effectively solve the operational difficulties players face in high-precision, fast-response or complex operation sequences. They are particularly prone to errors when the game environment changes dynamically and cannot provide substantial operational assistance.

Method used

It adopts a multi-model collaborative architecture, uses the large language model (LLM) to analyze game images and tasks to generate tactical intent, and generates precise operation instructions through the action execution model. It monitors the completion or failure of tasks in real time and automatically returns the control of game operations to the user.

Benefits of technology

It enables intelligent operations to perform complex tasks in the dynamically changing environment of the game, helps players overcome limitations in operational skills, improves the inclusiveness of the game, and enables users to experience the complete game content without obstacles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706580A_ABST
    Figure CN120706580A_ABST
Patent Text Reader

Abstract

The invention provides an autonomous game operation AI proxy method and device, and the method comprises the steps: S10, capturing a game picture in real time after an AI proxy takes over a game operation control right; s11, game pictures and tasks are analyzed through a large language model LLM, and tactical intentions are generated; s12, an operation instruction is generated through a pre-trained action execution model according to the tactical intention and the game picture, the operation instruction is executed, in the execution process of the scheme, whether a task is completed or failed or not is monitored by analyzing the game picture, and when it is monitored that the task is completed or failed, the game operation control right is returned to the user; according to the method and the system, the pain point that the operation of a player cannot keep up with consciousness is fundamentally solved, the method and the system can adapt to dynamic change and uncertainty in a game, highly complex and nonlinear tasks are executed, and substantive operation help is provided for the user to play the game.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of games, and in particular to an autonomous game operation AI agent method and device. Background Art

[0002] Currently, players can use game assistance systems and automated tools when they encounter difficulties in games. Existing game assistance systems can "understand" the game and "hear" the player's questions, providing strategic suggestions in the form of voice. These systems solve the problem of "not knowing what to do." However, they cannot resolve the dilemma of "knowing what to do, but not being able to execute it." For challenges requiring high precision, rapid response, or complex control sequences, simple voice guidance cannot provide substantial operational assistance to players, resulting in a "knowing is easy, doing is difficult" gap. Existing automated tools, such as keystroke wizards or built-in macro functions, can only record and replay fixed control sequences. They lack real-time perception of the game environment and are unable to intelligently adjust to dynamic changes in the battle situation. If the game scene or enemy behavior deviates slightly from the pre-set script, these tools are prone to errors and lack true intelligence and adaptability. Summary of the Invention

[0003] Therefore, in order to solve the problem of players getting stuck due to insufficient operating skills, the embodiments of the present application provide an autonomous game operation AI agent method and device, which fundamentally solves the pain point of players' "operations not keeping up with consciousness", can adapt to dynamic changes and uncertainties within the game, perform highly complex and nonlinear tasks, and provide users with substantial operational assistance in playing games.

[0004] In a first aspect, an embodiment of the present application provides an autonomous game operation AI agent method, comprising: S10. After the AI ​​agent takes over the control of the game operation, the game screen is captured in real time; S11. Analyze game screens and tasks through the Large Language Model (LLM) to generate tactical intent. S12. Generate operation instructions based on tactical intentions and the game screen through a pre-trained action execution model, and execute the operation instructions. During the execution of the above scheme, the game screen is analyzed to monitor whether the task is completed or failed. When the task is monitored to be completed or failed, the control of the game operation is returned to the user.

[0005] In a second aspect, an embodiment of the present application further provides an autonomous game operation AI agent device, comprising: The capture unit is used to capture the game screen in real time after the AI ​​agent takes over the control of the game operation; The generation unit is used to analyze game images and tasks through the large language model (LLM) to generate tactical intent; The execution unit is used to generate operation instructions based on tactical intentions and game screens through a pre-trained action execution model, and execute the operation instructions. During the execution of the above scheme, the game screen is analyzed to monitor whether the task is completed or failed. When the task is monitored to be completed or failed, the control of the game operation is returned to the user.

[0006] In summary, the autonomous game operation AI agent method and device provided in the embodiments of the present application adopt a multi-model collaborative architecture. After taking over the control of the game operation, the large language model LLM is responsible for understanding the user's intention and the current game screen, decomposing high-level tasks into specific tactical intentions, and the action execution model receives the tactical intentions and real-time game screen, generates and executes precise, low-latency control instruction streams, and automatically returns control to the user after completing the task or the operation fails. The entire solution fundamentally solves the pain point of players' "operations not keeping up with consciousness" and provides substantial operational assistance rather than staying at the suggestion level. In addition, unlike rigid macros and other automation tools, the entire solution makes decisions based on the understanding of real-time images, can adapt to dynamic changes and uncertainties within the game, and perform highly complex and nonlinear tasks, which greatly helps players who have difficulty passing certain game challenges due to physical reasons, reaction speed or operating skills limitations, so that they can also experience the complete game content without obstacles, thereby improving the inclusiveness of the game. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 A flowchart of an embodiment of an autonomous game operation AI agent method provided in an embodiment of the present application; Figure 2 This is a structural diagram of an embodiment of an autonomous game operation AI agent device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0008] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of illustration and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowcharts can be implemented out of sequence, and steps without logical context can be reversed or implemented simultaneously. In addition, those skilled in the art, under the guidance of the contents of this application, can add one or more other operations to the flowchart, or remove one or more operations from the flowchart.

[0009] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.

[0010] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.

[0011] Reference Figure 1 As shown, an embodiment of the present application provides a flow chart of an autonomous game operation AI agent method, the method comprising: S10. After the AI ​​agent takes over the control of the game operation, the game screen is captured in real time; In this embodiment, it should be noted that while the user is playing the game, it is possible to continuously monitor whether the user has triggered a signal requesting takeover. If a signal requesting takeover triggered by the user is detected, the user's physical input signals (such as keyboard, mouse, handle, etc. input signals) can be intercepted through the API (application programming interface) or driver at the operating system level, and at the same time, the permission to simulate input instructions is obtained, and the control of the game operation is handed over to the AI ​​agent, and then the AI ​​agent executes steps S10, S11 and S12.

[0012] S11. Analyze game screens and tasks through the Large Language Model (LLM) to generate tactical intent. In this embodiment, it should be noted that the Large Language Model (LLM) can be a GPT-4o or Llama 3 model. The Large Language Model (LLM) analyzes the game screen and tasks to understand the mission objectives and develop high-level tactics. For example, if the task is "Help me get there," the Large Language Model (LLM) analyzes the current game screen and the task and generates a tactical intent: sprint to avoid the first wave of arrows. After generating the tactical intent, the Large Language Model (LLM) can send it to a pre-trained action execution model.

[0013] S12. Generate operation instructions based on tactical intentions and the game screen through a pre-trained action execution model, and execute the operation instructions. During the execution of the above scheme, the game screen is analyzed to monitor whether the task is completed or failed. When the task is monitored to be completed or failed, the control of the game operation is returned to the user.

[0014] In this embodiment, it should be noted that after receiving tactical intent, the action execution model instantly generates and outputs a specific, low-level, time-sequenced control instruction stream (e.g., instructions for moving the mouse 350 pixels to the right, instructions for pressing the left mouse button, etc.) based on the real-time game screen, and injects these control instructions into the game. The next loop then proceeds to execute steps S10, S11, and S12. During this process, the Large Language Model (LLM) continuously analyzes the game screen and determines, based on the mission objectives, whether the mission is completed (e.g., the boss's health bar disappears, or specific visual features appear when the character reaches the target area, indicating mission completion) or failed (e.g., the appearance of a screen showing the player's death, such as "YOU DIED," indicating mission failure). Once mission completion or failure is detected, control of the physical input signal is immediately released, seamlessly returning game control to the user. This prevents the AI ​​agent from being trapped in endless, fruitless attempts. Instead, control is promptly returned to the user upon mission completion or failure, ensuring ultimate user control and a positive experience. Monitoring mission completion or failure can be performed after steps S11 or S12. In addition, it should be noted that the game can be a game running on a terminal or a cloud game running on a cloud server. If the game is a game running on a terminal, then executing the operation instruction in step S12 refers to executing the operation instruction on the game in the terminal to control the game; if the game is a game running on a cloud server, then executing the operation instruction in step S12 refers to sending the operation instruction to the cloud server and controlling the game on the cloud server based on the operation instruction, such as sending the operation instruction directly to the operating system of the cloud server, or converting the operation instruction into a corresponding operation instruction on the cloud server side and sending it to the operating system of the cloud server.

[0015] The autonomous game operation AI agent method provided in the embodiment of the present application adopts a multi-model collaborative architecture. After taking over the control of the game operation, the large language model LLM is responsible for understanding the user's intention and the current game screen, breaking down high-level tasks into specific tactical intentions, and the action execution model receives the tactical intentions and real-time game screen, generating and executing precise, low-latency control instruction streams, and automatically returning control to the user after completing the task or the operation fails. The entire solution fundamentally solves the pain point of players' "operations not keeping up with consciousness" and provides substantial operational assistance rather than staying at the suggestion level. In addition, unlike rigid automation tools such as macros, the entire solution makes decisions based on the understanding of real-time images, can adapt to dynamic changes and uncertainties within the game, and perform highly complex, nonlinear tasks. It greatly helps players who have difficulty passing certain game challenges due to physical reasons, reaction speed or operational skill limitations, so that they can also experience the complete game content without obstacles, thereby improving the inclusiveness of the game.

[0016] Based on the above method embodiment, before step S12, the following steps may be further included: Collect game video streams and expert operation instructions synchronized with the game video streams (such as keyboard key events, mouse movement trajectory and click data, joystick coordinates and button status data, etc.); The action execution model is trained using game video stream as input and expert operation instructions as output.

[0017] In this embodiment, it should be noted that the action execution model can be a Transformer model, or a Transformer architecture comprising a visual encoder and a recurrent neural network (the output of the Transformer-based visual encoder is used as the input of a sequence decoder based on a recurrent neural network (RNN)). When training the action execution model, the collected dataset can be used for training, with game screen frame sequences as input and corresponding expert operation instructions as output, learning a direct mapping from what is seen to what is done.

[0018] Based on the aforementioned method embodiment, the triggering condition for the AI ​​agent to take over the control of the game operation may include the user pressing a specific key, and the task may be obtained by analyzing the user's voice instructions captured by the microphone.

[0019] In this embodiment, it should be noted that the specific key can be a combination of keys, such as the Ctrl key + the Shift key + the G key. The task can be text converted from the user's voice command. For example, if the user's voice command is "Help me beat him", the task is: Help me beat him. In addition to manipulating specific keys to trigger the AI ​​agent to take over game operation control, the manipulation of specific keys can also be combined with user voice input to trigger the AI ​​agent to take over game operation control. For example, when the user presses the Ctrl key + the Shift key + the G key and voice inputs "Help me beat him", a signal requesting takeover is triggered, transferring game operation control to the AI ​​agent. By triggering the AI ​​agent to take over game operation control through a simple keystroke, or a simple keystroke + voice, the intervention of the AI ​​agent is explicitly authorized and controlled by the user, and the control takeover process is smooth and intuitive, minimizing the disruption to the game immersion.

[0020] Based on the aforementioned method embodiment, returning the game operation control right to the user may further include: The execution results of the task are fed back through voice.

[0021] In this embodiment, it should be noted that when returning control, voice feedback can be generated through text-to-speech technology to clearly report the results of task execution to the user, such as "Task completed, control returned" or "Sorry, I failed, please take over."

[0022] The following is an example to illustrate the solution of the embodiment of the present invention.

[0023] First, we invited several gaming experts to take on the challenge of defeating the challenging boss "Forge Knight" in a specific game multiple times. Using a pre-configured data collection tool, we simultaneously captured their first-person game footage and the corresponding commands for each precise keyboard, mouse, and controller operation (for example, pressing the Shift key -> pressing the W key -> clicking the right mouse button -> releasing the Shift key). We then used this massive amount of collected data (game screen video frame sequence, operation command sequence) to train an action execution model. The goal of the model learning is: when fed a continuous segment of game footage (for example, 24 frames of image over the past second), we can predict the most likely command the expert will enter in the next frame.

[0024] After the action execution model was trained, a user playing the game became frustrated after failing to defeat the boss "Forge Knight" after multiple attempts. The user pressed the hotkeys Ctrl+Shift+G and said into the microphone, "Help me defeat him!" This triggered a takeover request, transferring game control to the AI ​​agent. The AI ​​agent then began intercepting input from the user's physical input devices, such as the keyboard and mouse, and transcribed the audio into text, "Help me defeat him!"

[0025] The Large Language Model (LLM) then analyzes the current game screen (which shows the character "Forge Knight") and the text, understands the mission objective, and begins developing high-level tactics. It first generates the first tactical intent: "Keep your distance, observe its attack pattern, and primarily evade." Based on the mission objective, it determines whether the mission is completed or failed. If the mission is neither completed nor failed, the tactical intent is sent to the action execution model. Based on the real-time game screen, the action execution model generates specific control commands (such as the command for pressing the roll key backward or sideways) and injects them into the game through input. This process continues in a continuous loop.

[0026] When the Large Language Model (LLM) recognizes from the game screen that the Crucible Knight performs a move with a large afterswing (such as a ground stomp), it immediately updates the tactical intent in the next loop to "find the attack window and launch a quick counterattack" and sends this tactical intent to the action execution model. After receiving the new intent, the action execution model immediately adjusts its output.

[0027] This cycle of Large Language Model (LLM) planning, action execution model execution, and Large Language Model (LLM) observation and replanning continues until the Large Language Model (LLM) monitors the Crucible Knight's health bar reaching zero or the player character's death screen appears. There are two final outcomes: Scenario 1 (Mission Success): The Crucible Knight is defeated, and the AI ​​agent immediately releases control of the physical input device and plays a voice message: "Mission Completed, Control Returned," allowing the user to regain control. Scenario 2 (Mission Failure): The player character dies, and the AI ​​agent similarly releases control and plays a voice message: "Sorry, I failed, please take over," allowing the user to regain control and perform other actions.

[0028] Reference Figure 2 FIG. 1 is a schematic diagram of the structure of an autonomous game operation AI agent device provided in an embodiment of the present application, the device comprising: The capture unit 20 is used to capture the game screen in real time after the AI ​​agent takes over the control of the game operation; A generation unit 21 is used to analyze the game screen and tasks through a large language model (LLM) to generate tactical intent; The execution unit 22 is used to generate operation instructions based on tactical intentions and the game screen through a pre-trained action execution model, and execute the operation instructions. During the execution of the above scheme, the game screen is analyzed to monitor whether the task is completed or failed. When the task is monitored to be completed or failed, the control of the game operation is returned to the user.

[0029] The autonomous game operation AI agent device provided in the embodiment of the present application adopts a multi-model collaborative architecture. After taking over the control of the game operation, the large language model LLM is responsible for understanding the user's intention and the current game screen, breaking down high-level tasks into specific tactical intentions, and the action execution model receives the tactical intentions and real-time game screen, generates and executes precise, low-latency control instruction streams, and automatically returns control to the user after completing the task or the operation fails. The entire solution fundamentally solves the pain point of players' "operations not keeping up with consciousness" and provides substantial operational assistance rather than staying at the suggestion level. In addition, unlike rigid automation tools such as macros, the entire solution makes decisions based on the understanding of real-time images, can adapt to dynamic changes and uncertainties within the game, and perform highly complex, nonlinear tasks. It greatly helps players who have difficulty passing certain game challenges due to physical reasons, reaction speed or operating skills limitations, so that they can also experience the complete game content without obstacles, thereby improving the inclusiveness of the game.

[0030] Based on the above device embodiment, the device may further include: The training unit is used to collect the game video stream and the expert operation instructions synchronized with the game video stream before the execution unit works, and use the game video stream as input and the expert operation instructions as output to train the action execution model.

[0031] Based on the aforementioned device embodiment, the triggering condition for the AI ​​agent to take over the control of the game operation may include the user pressing a specific key, and the task can be obtained by analyzing the user's voice instructions captured by the microphone.

[0032] Based on the aforementioned device embodiment, returning the game operation control right to the user may further include: The execution results of the task are fed back through voice.

[0033] The implementation process of the autonomous game operation AI agent device provided in the embodiment of the present application is consistent with the autonomous game operation AI agent method provided in the embodiment of the present application, and the effect that can be achieved is also the same as the autonomous game operation AI agent method provided in the embodiment of the present application, which will not be repeated here.

[0034] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. An autonomous game operation AI agent method, characterized in that: include: S10. After the AI ​​agent takes over the control of the game operation, the game screen is captured in real time; S11. Analyze game screens and tasks through the Large Language Model (LLM) to generate tactical intent. S12. Generate operation instructions based on tactical intentions and the game screen through a pre-trained action execution model, and execute the operation instructions. During the execution of the above scheme, the game screen is analyzed to monitor whether the task is completed or failed. When the task is monitored to be completed or failed, the control of the game operation is returned to the user.

2. The method according to claim 1, wherein Before step S12, the method further includes: Collect game video streams and expert operation instructions synchronized with the game video streams; The action execution model is trained using game video stream as input and expert operation instructions as output.

3. The method according to claim 1 or 2, wherein: The trigger conditions for the AI ​​agent to take over control of game operations include the user pressing a specific button and obtaining the task by analyzing the user's voice commands captured by the microphone.

4. The method according to claim 1, wherein The returning the game operation control right to the user also includes: The execution results of the task are fed back through voice.

5. An autonomous game operation AI agent device, characterized in that: include: The capture unit is used to capture the game screen in real time after the AI ​​agent takes over the control of the game operation; The generation unit is used to analyze game images and tasks through the large language model (LLM) to generate tactical intent; The execution unit is used to generate operation instructions based on tactical intentions and game screens through a pre-trained action execution model, and execute the operation instructions. During the execution of the above scheme, the game screen is analyzed to monitor whether the task is completed or failed. When the task is monitored to be completed or failed, the control of the game operation is returned to the user.

6. The device according to claim 5, characterized in that Also includes: The training unit is used to collect the game video stream and the expert operation instructions synchronized with the game video stream before the execution unit works, and use the game video stream as input and the expert operation instructions as output to train the action execution model.

7. The device according to claim 5 or 6, characterized in that The trigger conditions for the AI ​​agent to take over control of game operations include the user pressing specific buttons, and the tasks are obtained by analyzing the user's voice commands captured by the microphone.

8. The device according to claim 5, wherein The returning the game operation control right to the user also includes: The execution results of the task are fed back through voice.

Citation Information

Patent Citations

  • Game assisting method and device

    CN120532141A

  • Selective Recommendation by Mapping Game Decisions and Behaviors to Predefined Attributes

    US20230102506A1