Apparatus, method, non-transitory machine-readable medium
A machine learning system interprets natural language inputs to configure video games for training tasks, refining scenarios based on user performance feedback, addressing the lack of effective training in open-ended games and enhancing skill development.
Patent Information
- Application Number
- PCT/EP2025/058080
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-25
- Filing Date
- 2025-03-25
- Publication Date
- 2025-10-02
AI Technical Summary
Existing video game training systems, particularly for open-ended games, lack effective training assistance and feedback, with bots often being insufficiently challenging and users struggling to define clear objectives, leading to suboptimal skill development.
A machine learning-based system that interprets natural language inputs to configure video games for specific training tasks, evaluates user performance, and iteratively refines the training scenario based on feedback to enhance skill development.
Provides personalized and adaptive training experiences, improving user skills by offering challenging scenarios and expert feedback, thereby enhancing engagement and performance in competitive gaming.
Smart Images

Figure EP2025058080_02102025_PF_FP_ABST
Abstract
Description
[0001] APPARATUS, METHOD, NON-TRANSITORY MACHINE-READABLE MEDIUM
[0002] Field
[0003] The present disclosure relates to a machine learning based video game training system for video games. In particular, examples of the present disclosure relate to an apparatus, a method and a non-transitory machine-readable medium.
[0004] Background
[0005] E-Sports and competitive computer gaming have been rising in popularity during the past years. Many people participate in ranked games, where they can improve their skills against similarly leveled opponents. This may be, however, a high-stakes environment that may not welcome experimentation or self-paced training. For that, people opt to play against bots, but these can fail to be challenging enough. Guidelines for improving are often inexistent beyond basic tutorials, and players flock to other media, such as YouTube, to find expert advice on how to defeat bosses, perform certain combos or the like.
[0006] Therefore, it may be desirable to improve the training for video games.
[0007] Summary
[0008] According to a first aspect, the present disclosure provides an apparatus comprising circuitry configured to obtain a natural language input comprising a game task with regards to a video game. The circuitry is further configured to determine, by a first machine learning model, a first data structure based on the natural language input. The first data structure comprises first configuration data for the video game corresponding to the game task. The circuitry is further configured to receive a performance evaluation of a user playing the video game configured with the first data structure.
[0009] According to a fourth aspect, the present disclosure provides a method comprising obtaining a natural language input comprising a game task with regards to a video game. The method further comprises determining, by a first machine learning model, a first data structure based on the input. The first data structure comprises first configuration data for the video game corresponding to the game task. The method further comprises receiving a performance evaluation of a user playing the video game configured with the first data structure.
[0010] Further aspects are set forth in the appended set of claims.
[0011] Brief description of the Figures
[0012] Some examples of apparatuses and / or methods will be described in the following by way of example only, and with reference to the accompanying figures, in which
[0013] Fig. 1 illustrates a block diagram of an example of an apparatus;
[0014] Fig. 2 illustrates a block diagram of a machine learning based video game training system for video games; and
[0015] Fig. 3 illustrates a flowchart of an example of a method.
[0016] Detailed Description
[0017] Some examples are now described in more detail with reference to the enclosed figures. However, other possible examples are not limited to the features of these embodiments described in detail. Other examples may include modifications of the features as well as equivalents and alternatives to the features. Furthermore, the terminology used herein to describe certain examples should not be restrictive of further possible examples.
[0018] Throughout the description of the figures same or similar reference numerals refer to same or similar elements and / or features, which may be identical or implemented in a modified form while providing the same or a similar function. The thickness of lines, layers and / or areas in the figures may also be exaggerated for clarification.
[0019] When two elements A and B are combined using an “or”, this is to be understood as disclosing all possible combinations, i.e. only A, only B as well as A and B, unless expressly defined otherwise in the individual case. As an alternative wording for the same combinations, "at least one of A and B" or "A and / or B" may be used. This applies equivalently to combinations of more than two elements.
[0020] If a singular form, such as “a”, “an” and “the” is used and the use of only a single element is not defined as mandatory either explicitly or implicitly, further examples may also use several elements to implement the same function. If a function is described below as implemented using multiple elements, further examples may implement the same function using a single element or a single processing entity. It is further understood that the terms "include", "including", "comprise" and / or "comprising", when used, describe the presence of the specified features, integers, steps, operations, processes, elements, components and / or a group thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, processes, elements, components and / or a group thereof.
[0021] When playing video games, for instance in e-sports or competitive gaming, training may be very important. Users may improve their skills by repeated play, practice against bots, challenge modes, etc. However, bots may in some instances not be challenging enough or with regards to some video games bots may not even be available. This is especially an issue for open ended-video games where no training assistance is available. Further, a challenge in open-ended games may be, that it may be difficult to define a predefined final objective, or such a final objective may not be straightforward and only achievable after mastering a wide range of gaming skills. For instance, in World of Warcraft, a popular massively multiplayer online role-playing game (MMO), a user may want to improve in dungeon crawling (player versus environment) while another user may want to become better at raids (player versus player). In another example, different users may want to do so with a wide range of personalized characters (race, class, gear, abilities, etc.). However, while there may be gaming skills that may transfer across domains within a video game or amongst video games, it may be desirable to provide the relevant exercises and feedback to the user.
[0022] In some previous approaches for chess, there may be a machine directed training provided by a combination of strong heuristics provided by a chess engine (such as Stockfish) and platform graphical user interfaces (GUIs) that may allow players to learn from their actions. However, in this case, there is no training system for alternative chess game modes such as king-of-the- hill. Furthermore, a player who is exercising a video game may be left to practice the exercise on his own, not receiving any feedback.
[0023] In the present disclose systems, apparatuses and methods are provided that provide improve and automate training and coaching for open-ended video games.
[0024] Fig- 1 illustrates a block diagram of an example of an apparatus 100. The apparatus 100 comprises circuitry that is configured to provide the functionality of the apparatus 100. The apparatus 100 comprises a processing circuitry 110. For example, the processing circuitry 110 may be a single dedicated processor, a single shared processor, or a plurality of individual processors, some of which or all of which may be shared, a digital signal processor (DSP) hardware, an application specific integrated circuit (ASIC), a neuromorphic processor or a field programmable gate array (FPGA). The processing circuitry 110 may optionally be coupled to, e.g., memory such as read only memory (ROM) for storing software, random access memory (RAM) and / or non-volatile memory. For example, the apparatus 100 may comprise memory configured to store instructions, which when executed by the processing circuitry 110, cause the processing circuitry 110 to perform the steps and methods described herein.
[0025] The circuitry 110 is configured to obtain a natural language input comprising a game task with regards to a video game. Further, the circuitry 110 is configured to determine, by a first machine learning model, a first data structure based on the natural language input. The first data structure comprises first configuration data for the video game corresponding to the game task.
[0026] For instance, the video game may be an electronic game that involves interaction with a user interface or input device - such as a joystick, controller, keyboard, or motion sensing device - to generate visual feedback for a player. In some examples, the video game may be an open- ended game. An open-ended game may refer to video games that offer players the freedom to explore, create, and make choices within a vast, interactive environment, often without a fixed endpoint or linear storyline. Examples include “Minecraft”, “World of Warcraft”, “Grand Theft Auto”, “League of Legends” and the like. For instance, the game task in the video game may refer to a specific objective or set of objectives that a player performs within the game's environment. The game task may range from simple actions like collecting certain items to complex missions involving multiple steps or challenges. In some embodiments the game task may define an exercising objective, or a training objective with regards to the video game. In some embodiments the game task may define a scenario in the video game. For instance, if the video game is an action game the natural language input may be “I want to train fighting” that is the game task may be “train fighting”. In another example, if the video game is chess, the natural language input may be “I want to become better at openings” and the corresponding game task may be “train chess openings” (that is for instance the first 15 or 30 moves), and the corresponding training objective may be becoming better at or training the chess openings.
[0027] In some examples, the first machine learning model is trained to analyze and process the obtained natural language input and interpret the game task with regards to the video game that is contained in the natural language input. In some examples, the first machine learning model is trained to determine the first data structure based on the analyzed and interpreted game task with regards to the video game. The first data structure comprises the first configuration data for the video game which may comprise configuration data for the video game corresponding to the initial game task described in the natural language input.
[0028] In some embodiments the circuitry 110 is further configured to set up the video game with the first data structure comprising the first configuration data as a video game configuration. For instance, the first configuration data may be comprising commands and setup configurations which may be input to the video game in order for a user being able to perform the game task with regards to the video game that is contained in the natural language input. For instance, the first data structure may be in a computer readable format that may be directly readable by an input interface of the video game. That is, in other words, the circuitry 110 may be configured to translate a natural language input (for instance a user's natural language instructions) into executable first data structure that may be applied to configure the video game tailored to the described game task, i.e., such that the requested game task can be performed and or played within the video game. With regards to the examples given above, based on the natural language input may be “I want to train fighting” the circuitry 110 may determine the first data structure comprising the first configuration data that configures the video game such that a fight (for instance a fist fight and / or a sword fight) may be exercised in the video game. That is the first data structure comprising the first configuration data may comprise setting data, for gaming environment and a gaming character where fighting may be exercised. In yet another example, if the video game is “League of Legends”, a request from the user may be “I want to become better at ganking” and the first machine learning model may determine the first data structure comprising the first configuration data that configures the video game such that it will load an appropriate scenario for the user, with a hero suitable for the game task, necessary equipment, and the right positioning to be able to try it immediately, rather than having to start the game from zero and wait for opportunities to practice this rare event.
[0029] Further, the processing circuitry 110 is configured to receive a performance evaluation of a user playing the video game configured with the first data structure. In some embodiments the apparatus 100 (for instance the circuitry 110) may determine the performance evaluation of the user playing the video game configured with the first data structure. In another embodiment the performance evaluation of the user playing the video game configured with the first data structure may be determined by a second apparatus and transmitted to the circuitry 110. For instance, the, performance evaluation may be determined by a server hosting the video game.
[0030] For instance, the performance of the user playing the video game configured with the first configuration data of the first data structure, is determined based on an internal metric of the video game, i.e., it is measured in a game-internal metric (also referred to as value function). For instance, the game-internal metric may be reaching a certain game-internal level or collecting certain game-internal credits or the like. In some examples, circuitry 110 is further configured to determine the performance evaluation of the user playing the video game configured with the first data structure based on a required time to complete the game task, for instance a time required to successfully complete the game task. In some embodiments the received performance evaluation may be displayed to the user. In another embodiment, the performance evaluation may be fed back into the first machine learning model together with the natural input language (comprising the game task) in order to determine a second data structure comprising second configuration data based thereon. In this case, the second configuration data for the video game may match the intended requested game task better than the first configuration data based on the feedback data. For instance, the machine learning model may retrieve the performance evaluation or a game-internal metric from a database according to the assessment of the machine learning model which produces the query to do so.
[0031] Therefore, user (i.e., players) may improve their video game skills and may further engage with the video game and improve in the competitive sphere due to the disclosed technique. Further, user may specify gaming training scenarios more easily in natural language which may reduce friction. Further, the transformation (translation) of the natural language input comprising the game task into a corresponding configuration for the video game may be, for instance iteratively, improved. That is the performance evaluation enhances the translation bay the first machine learning model of the natural language input into the first data structure.
[0032] The circuitry 110 may be further configured to determine, by the first machine learning model, a second data structure based on the performance evaluation and the natural language input. The second data structure may comprise second configuration data for the video game corresponding to the game task. The second data structure comprising the second configuration data for the video game may therefore correspond to the game task and the performance evaluation. The performance evaluation may be considered as feedback data.
[0033] In some examples, the second machine learning model is trained to analyze and process the obtained natural language input together with the performance evaluation. That is, the game task that is contained in the natural language input may be interpreted by the first machine learning model together with the feedback data (i.e., performance evaluation). Based thereon, the second machine learning model may be trained to determine the second data structure comprising the second configuration data for the video game. The second data structure comprising the second configuration data may comprise configuration data for the video game which is corresponding to the initial game task described in the natural language input and is further refined by the feedback data (i.e., the performance evaluation). That is the second configuration data may comprise commands and setup configurations which may be input to the video game in order for the user being able to perform the game task with regards to the video game that is contained in the natural language input. However, the second configuration data for the video game may match the intended requested game task better than the first configuration data based on the feedback data (performance evaluation). That is, the second data structure comprising the second configuration data and the corresponding setup configurations for the video game result in a set up for the video game that provides an improved training for the user compared to the first data structure comprising the first configuration data. In other words, the first machine learning model may use previously received instructions from the user together with the feedback to close the loop and generate better suited training scenarios, query more appropriate agents, etc.
[0034] The circuitry 110 may be configured to improve, based on the performance evaluation, the translation of the natural language input into the executable second data structure. The second data structure may be applied to configure the video game, which may better match the described game task better than the first data structure. The transformation (translation) of the natural language input comprising the game task into a corresponding configuration for the video game may be improved based on the performance evaluation. This process may be carried out iteratively several times to further improve and refined the determined data structure comprising the configuration data for the video game as described above.
[0035] In some embodiments the circuitry 100 may display the received performance evaluation to the user, for instance on a display screen (such as monitor, TV, smartphone). In some embodiments the circuitry 110 may convey the performance evaluation to the user with natural language, for instance by using the first machine learning model or another machine learning model to translate the text into speech.
[0036] In some embodiments the game-internal feedback may be provided during performing the game task to provide feedback to the user, for example, by displaying a rating of the current performance of the user. For instance, the input from the controller (such as gamepad, mouse & keyboard, or other embodiments) and the game-internal metric may be used to determine mistakes and opportunities for further improvement.
[0037] With regards to the examples given above, based on the natural language input “I want to train fighting” the circuitry 110 may have determined the first data structure comprising the first configuration data that configures the video game such that a fight, for instance a fist fight and a sword fight, was exercised in the video game by the user. For instance, the performance evaluation determines that the player scored 50 game-internal credits with sword fighting and 70 game-internal credits with fist fighting. This performance evaluation, i.e., feedback data, together with the natural language input “I want to train fighting” is input into the first machine learning model which is trained to analyze and process the natural language input together with the feedback data. The first machine learning model may then determine be trained to determine the second data structure comprising the second configuration data that configures the video game such that a fight may be exercised, however with special focus on the sword fight which has scored lower game-internal credits in the feedback data. That is the second data structure comprising the second configuration data may comprise setting data, for gaming environment and a gaming character where especially sword fighting may be exercised.
[0038] In some embodiments the circuitry 110 may be further configured to determine a reference performance value of playing the video game configured with the first data structure. In some examples the reference performance value may correspond to a performance evaluation of playing the video game configured with the first data structure by another player, having skills on a predetermined skill level. For instance, the reference performance value may correspond to a performance evaluation of playing the video game configured with the first data structure by a professional player, an amateur player or a rookie. In some examples the reference performance value may be measured in the same metric as the performance evaluation of a user playing the video game configured with the first data structure. For instance, the reference performance value of playing the video game configured with the first data structure may be determined based on game-internal metrics, like reaching a certain game-internal level or collecting certain number of game-internal credits or the like. In some examples, the reference performance is evaluated in an evaluation metric such as time required to successfully complete the game task. The reference performance value may indicate a highest possible gameinternal level that is achieved or a maximum number of scored game-internal credits or the like or a minimum required time to successfully complete the game task or the like. In another embodiment the reference performance value may be determined by a second apparatus, for instance by a server hosting the video game.
[0039] In some embodiments the circuitry 110 is further configured to determine a difference between the reference performance value and the performance evaluation. In some examples the performance and the performance evaluation may each be a single number or a vector of the same dimension or the like. For instance, as described above, the reference performance value and the performance evaluation may both be measured in game-internal metrics, such as a reached game-internal level, or number of collected game-internal credits. The difference of the reference performance value and the performance evaluation in these examples may be a difference of reached game-internal levels, or a difference in the number of scored game-internal credits or the time difference required to successfully complete the game task or the like. With regards to the examples given above, based on the natural language input may be “I want to train fighting”, the reference performance value may be 100 game-internal credits with sword fighting are scorable and a maximum of 100 credits with fist fighting are scorable collectable. The difference value in this case, which may be provide to the first machine learning model as feedback data may then be 30 game-internal credit points for fist fighting and 50 game-internal credit points for sword fighting.
[0040] In some examples, the difference between the reference performance value and the performance evaluation may be considered as feedback data as described above and fed back to first machine learning model together with the natural language input. The first machine learning model may then determine be trained to determine the second data structure comprising the second configuration data that configures the video game such that a fight may be exercised, however with special focus on the sword fight which had the bigger difference between the reference performance value and the performance evaluation. That is the second data structure comprising the second configuration data may comprise setting data, for gaming environment and a gaming character where especially sword fighting may be exercised.
[0041] In some embodiments the first machine learning model may be trained to communicate with the user and ask questions to gather more information about the game task and thereby increase the specificity of it.
[0042] In some embodiments the circuitry 110 may be further configured to display the virtual agent executing the video game set up with the first configuration data for the video game corresponding to the game task. A virtual agent may be a software generated entity that is participating in the game or interacting with another player, executing tasks, providing guidance, or demonstrating gameplay mechanics within the video game's world. For example, the virtual agent may be a pre-trained virtual agent, for instance pre-trained on the game task at hand. For instance, different pre-trained virtual agents may be available for the same game task with different playstyles and proficiency levels. Therefore, the virtual agent may be personalized to the user to provide the different experiences. However, the virtual agent may be configurable from the user side. For instance, the virtual agent may be configured based in the first data structure comprising the first configuration data for the video game corresponding to the game task. In other words, if the user requests a demonstration, after the circuitry has obtained the natural language input, the virtual agent is deployed according to the game task in the configured game. The virtual agent may take charge of the input controls typically generated by the user (player) and execute the game task. In some embodiments this may be done in a slow- motion and / or with highlights for important parts of the demonstration. The user may take control of the game task carried out by the virtual agent from in the demonstration at any point in time and still receive feedback from the analysis.
[0043] In other words, the virtual agent may be considered as a training coach, which demonstrates a proper execution of the game task as well as provide corrections and suggestions to the user. This may further improve the training experience of the video game based on the natural language input comprising a game task with regards to a video game. Therefore, the users may learn from expert demonstrations in-game. Further, training scenarios may adapt to user’s performance and may produce challenges that the user may learn form.
[0044] In some embodiments the circuitry 110 may be configured to determine the virtual agent. In some embodiments the circuitry 110 may be configured to receive the virtual agent from a second apparatus. For example, the virtual agent may be retrieved from a database or from the server hosting the video game according to the assessment of the first machine learning model, which produces the query to do so.
[0045] In some examples, the multiple game-internal metrics may be applied functions to find a closest or most transferable to the player’s play style, as well as deploying multiple virtual agents to control other characters besides the user’s. In some embodiments the game-internal metric (also referred to as value function) may be derived together with the virtual agent’s policy.
[0046] Machine learning model A machine learning (such as the first machine-learning model and / or the second machine learning model, see below) may refer to algorithms and statistical models that computer systems (for instance the apparatus 100 or the processing circuitry 110) may use to perform a specific task without using explicit instructions, instead relying on models and inference. For example, in machine-learning, instead of a rule-based transformation of data, a transformation of data may be used, that is inferred from an analysis of historical and / or training data.
[0047] The machine-learning model may denote a data structure and / or set of rules that represents learned knowledge (for instance learned weights of an artificial neural network), e.g. based on machine learning training performed by a machine learning algorithm. The usage of a machine-learning model may imply that the machine-learning model and / or the data structure / set of rules that is the machine-learning model is trained by a machine-learning algorithm. That is a machine learning algorithm may be a method or procedure used to process training data and learn from it, while a machine learning model is the output or the result of that algorithm trained on training data, capable of making predictions or decisions (in some cases in the art, the terms machine learning model and machine learning algorithm are used interchangeably).
[0048] Machine-learning models may be trained to perform their specific task using training input data. There are several different training methods for machine learning models to learn from data. For instance, the methods are supervised learning, unsupervised learning, semi-super- vised learning, reinforcement learning or the like. In some embodiments the one or more different training method may be applied for training the machine learning model. For instance, in supervised learning, the machine learning model learns from labeled data, meaning each input training sample is associated with an output label. The machine learning model makes predictions during the training and adjusts based on the error of its predictions compared to the actual labels. In unsupervised learning the machine learning model learns from unlabeled data, identifying inherent patterns or structures in the input training data without explicit feedback on the accuracy of its predictions. Self-supervised learning may a subset of unsupervised learning where the machine learning model is trained to predict part of the input from another part of the input, using the data itself to generate labels for training. In semi-supervised learning elements of supervised and unsupervised learning are combined, where the machine learning model learns from a mix of labeled and unlabeled training data, which may be useful when acquiring labeled data is costly or time-consuming. Further, in reinforcement learning, one or more software actors (called “software agents”) are trained to take actions in an environment. Based on the taken actions, a reward is calculated. Reinforcement learning is based on training the one or more software agents to choose the actions such, that the cumulative reward is increased, leading to software agents that become better at the task they are given (as evidenced by increasing rewards).
[0049] There may be several different machine learning algorithms that may be used to train the machine learning model. For instance, decision trees may be used. A decision tree as a predictive model. In other words, the machine-learning model may be based on a decision tree. In a decision tree, observations about an item (e.g. a set of input values) may be represented by the branches of the decision tree, and an output value corresponding to the item may be represented by the leaves of the decision tree. Decision trees may support both discrete values and continuous values as output values. If discrete values are used, the decision tree may be denoted a classification tree, if continuous values are used, the decision tree may be denoted a regression tree.
[0050] Alternatively, the machine-learning algorithm may be a support vector machine, a random forest model or a gradient boosting model. Support vector machines (i.e. support vector networks) are supervised learning models with associated learning algorithms that may be used to analyze data (e.g. in classification or regression analysis). Support vector machines may be trained by providing an input with a plurality of training input values that belong to one of two categories. The support vector machine may be trained to assign a new input value to one of the two categories. Alternatively, the machine-learning model may be a Bayesian network, which is a probabilistic directed acyclic graphical model. A Bayesian network may represent a set of random variables and their conditional dependencies using a directed acyclic graph. Alternatively, the machine-learning model may be based on a genetic algorithm, which is a search algorithm and heuristic technique that mimics the process of natural selection.
[0051] Association rules are a further technique that may be used in machine-learning algorithms. In other words, the machine-learning model may be based on one or more association rules. Association rules are created by identifying relationships between variables in large amounts of data. The machine-learning algorithm may identify and / or utilize one or more relational rules that represent the knowledge that is derived from the data. The rules may e.g. be used to store, manipulate, or apply the knowledge.
[0052] In some examples, a k-Nearest Neighbors (KNN) algorithm may be used which classifies data points based on the majority vote of their neighbors, being simple yet effective for classification and regression. In yet another example, linear regression may be used which predicts a continuous outcome by establishing a linear relationship between input variables, while logistic regression extends this to binary classification tasks.
[0053] In some examples, clustering algorithms like K-Means may be used to organize data into clusters based on similarity, useful in unsupervised learning scenarios. In this regard, also, anomaly detection (i.e. outlier detection) may be used, which is aimed at providing an identification of input values that raise suspicions by differing significantly from the majority of input or training data. In other words, the machine-learning model may at least partially be trained using anomaly detection, and / or the machine-learning algorithm may comprise an anomaly detection component.
[0054] For example, the machine-learning model may be an artificial neural network (ANN). ANNs are systems that are inspired by biological neural networks, such as can be found in a retina or a brain. ANNs comprise a plurality of interconnected nodes and a plurality of connections, so- called edges, between the nodes. There are usually three types of nodes, input nodes that receiving input values, hidden nodes that are (only) connected to other nodes, and output nodes that provide output values. Each node may represent an artificial neuron. Each edge may transmit information, from one node to another. The output of a node may be defined as a (nonlinear) function of its inputs (e.g. of the sum of its inputs). The inputs of a node may be used in the function based on a “weight” of the edge or of the node that provides the input. The weight of nodes and / or of edges may be adjusted in the learning process. In other words, the training of an artificial neural network may comprise adjusting the weights of the nodes and / or edges of the artificial neural network, i.e. to achieve a desired output for a given input. There are several types of ANNs. For instance, multilayer perceptron’s (MLPs), the simplest form of ANNs, consist of input, hidden, and output layers and are fundamental for basic classification and regression tasks. Convolutional neural networks (CNNs) are designed to process data with a grid-like topology, such as images, leveraging convolutional layers to detect spatial hierarchies and patterns, making them ideal for image and video analysis. Recurrent Neural Networks (RNNs) are a class of neural networks where connections between nodes form a directed graph along a temporal sequence, allowing them to exhibit temporal dynamic behavior and memory of past inputs, ideal for processing sequential data like speech or text. Deep Neural Networks are advanced forms of neural networks with multiple hidden layers that allow for modeling of complex data patterns through a deeper and more intricate computational structure.
[0055] Further, feature learning, also known as representation learning (also related to latent space learning), is a fundamental aspect in the realm of machine learning where machine learning models / algorithms autonomously identify and extract relevant features from raw data. It's particularly integral to deep learning, where neural networks, through layers of processing, learn increasingly complex representations. Furthermore, some techniques may be applied to some of the machine-learning algorithms. For example, feature learning may be used. In other words, the machine-learning model may at least partially be trained using feature learning, and / or the machine-learning algorithm may comprise a feature learning component. Feature learning algorithms, which may be called representation learning algorithms, may preserve the information in their input but also transform it in a way that makes it useful, often as a preprocessing step before performing classification or predictions. Feature learning may be based on principal components analysis or cluster analysis, for example.
[0056] Further, the ANNs may be organized into different architectures which define a specific structure and design of a network, determining how it processes and learns from data. One architecture is a autoencoder, designed for learning efficient encodings by compressing input into a latent-space representation and then reconstructing it, often used for dimensionality reduction and feature learning. Another architecture is a generative adversarial network (GAN) involving two competing networks, a generator, and a discriminator, which learn to generate new data with the same statistics as the training set. Another architecture is RNN-based Long Short-Term Memory (LSTM) networks and Gated Recurrent Units (GRUs), offer architectures tailored for sequential data processing, capable of maintaining information across long sequences, which is crucial fortasks like language modeling and time series analysis. Another architecture are encoder-decoder models, often built with RNNs, LSTMs, or transformers, constitute a specific architecture for transforming sequences, where the encoder processes the input sequence and the decoder generates the output, essential for applications like machine translation and summarization.
[0057] Another architecture is a so called transformer model. These transformer models may be used in so called large language models (LLMs). Transformer models may be a type of deep learning model that utilizes self-attention mechanisms to process sequential data without relying on recurrent or convolutional layers. These models are for instance based on the scientific paper “Attention is all you need.”, Vaswani, Ashish, et al, published in Advances in neural information processing systems 30 (2017). Transformer models allow each element in a sequence to interact with every other element directly, enabling the model to capture complex dependencies and relationships within the data. This architecture may consist of an encoder and decoder, or just an encoder or just a decoder, each composed of multiple layers of attention and feed-forward networks, facilitating parallel processing and significantly improving efficiency and effectiveness in tasks like machine translation, text generation, and language understanding. Some examples of transformer models are for instance, BERT (Bidirectional Encoder Representations from Transformers) which utilizes an encoder-only architecture, making significant strides in tasks like sentence classification, question answering, and language understanding by pre-training on a large corpus of text and then fine-tuning for specific tasks. Another example are decoder-only architectures like GPT (Generative Pre-trained Transformer) focus on generating text. GPT models are trained on a large dataset to predict the next word in a sequence, enabling applications in text generation, chatbots, and even creative writing. Another example is T5 (Text-to-Text Transfer Transformer), which frames all NLP tasks as a text-to-text problem, using a modified encoder-decoder architecture to perform a variety of tasks with a single model, from translation to summarization.
[0058] Transformers may be trained using a two-stage process that combines unsupervised (or self- supervised) and supervised learning techniques. For instance, in a first phase, referred to pretraining the transformer models are trained based in an unsupervised or self-supervised manner on a large corpus of input text data. This text does not have to be specific to the task that the transformer network should carry out. This phase doesn't rely on labeled data. Instead, the transformer model generates its own labels from the provided data. For example, GPT models use autoregressive language modeling where they predict the next word in a sequence given the previous words, learning to understand context, grammar, and language structure. BERT, on the other hand, uses masked language modeling, where some percentages of the input tokens are masked, and the goal is for the model to predict the original token at each masked position.
[0059] In a second optional stage, the pre-trained transformer model may be fine-tuned on a specific task, which may involve supervised learning. During this stage, the model may be trained on a smaller, labeled dataset specific to the task at hand. The fine-tuning may adjust the weights of the pre-trained transformer model to optimize its performance for the specific task, whether it's text classification, sentiment analysis, question-answering, or any other NLP task. This phase leverages the general language understanding the model has gained during pre-training to achieve high performance with relatively little task-specific data.
[0060] The first machine learning model (and the second machine learning model) may be based on any of the above mentioned machine learning algorithms trained by any of the above machine learning training methods.
[0061] In some examples, the first machine learning model is ANN. In some examples the first machine learning model is an LLM.
[0062] In some examples, the circuitry 110 may be further configured to train the first machine learning model based on video game specific training data. For example, the first machine learning model may be trained in a supervised learning method based on labelled training data. The labelled training data may comprise pairs of natural language inputs comprising a game task, and configuration data for the video game corresponding to the game task. Furthermore, the first machine learning model may be trained with labelled training data comprising triplets of natural language inputs (comprising a game task), configuration data for the video game corresponding to the game task, and feedback data as described above (for example a performance evaluation).
[0063] In some embodiment the training or of the training of the first machine learning model may be performed by another apparatus equipped with hardware specialized for ANN training, for instance with graphical processing units (GPUs) or the like. The trained ANN, i.e. the first machine learning model, may be transmitted to the apparatus. In some examples, the natural language input may be text input or voice input or both, in a natural language. The natural language input may be in a natural language used by humans, such as German, Japanese, or English or the like.
[0064] In some embodiments the circuitry 110 is further configured to perform automatic speech recognition (ASR), when the natural language input is speech input, and forward the text output of the automatic speech recognition to the first machine learning model. In some examples, the automatic speech recognition may be performed by a second machine learning model. ASR may be able to interpret and transcribe natural language speech into text. For instance, the ASR systems may be based on ANNs, including DNNs, CNNs, RNNs or the like. An ASR system may be pre-trained. For example, the pre-trained ASR “DeepSpeech”, as described in the scientific paper “Deep speech: Scaling up end-to-end speech recognition.”, by Hannun, Awni, et al., published on arXiv preprint arXiv: 1412.5567 (2014) may be used.
[0065] In some examples, the first data structure (and the second data structure) may have a predetermined format compatible with an application programming interface (API) of the video game. That is the first data structure may be in a computer readable format that is defined by the video game API and may be directly readable by an input interface of the video game. In some embodiments the predetermined data structure of the first data structure may be obtained possible by prompting with one or more examples. In some examples, the first machine learning model may be integrated with a guided programming paradigm (also referred to as constraintbased generation), where the machine learning model is trained to output data structures in a predetermined format. For instance, “guidance-ai” programming paradigm may be used in this regard.
[0066] In some examples, the machine learning model may comprise a data formatting tool which an unstructured input from a for module of the first machine learning model and formats that input into the predetermined first data structure. For instance, the data formatter JSON formatter may be used to format the first data structure into a JSON format, without altering the underlying data.
[0067] Example of the first machine learning model being an LLM In the following example, the first machine learning model may be an ANN in a transformer architecture. For instance, it may be a decoder-only pre-trained transformer model, for instance one of the examples as disclosed in the scientific paper “Opt: Open pre-trained transformer language models ”, by Zhang, Susan, et al, published on arXiv preprint arXiv:2205.01068 (2022).
[0068] For example, the decoder-only pre-trained transformer model may be further trained by unsupervised learning based on un-labelled training data, which is specific to the video game. For instance, this un-labelled training data may be collected from an online forum dedicated to the video game or to videos, or books or journals dedicated to the video game, or a recording of a conference dedicated to the video game or similar. The corresponding data may be converted from video or audio form into text form by well-known software and thus supplied to the transformer as training data.
[0069] For example, this pre-trained transformer model may be further trained with a supervised learning method based on labelled training data. The labelled training data may comprise pairs of natural language inputs comprising a game task, and configuration data for the video game corresponding to the game task. Furthermore, the pre-trained transformer model may be trained with labelled training data comprising triplets of natural language inputs (comprising a game task), configuration data for the video game corresponding to the game task, and feedback data as described above (for example a performance evaluation). The training data (the unsupervised training data as well as the labelled pairs and triplets) may be generated manually or automatically. For instance, labelled training data may be labelled automatically. This may yield a transformer model (first machine learning model), trained specifically for the video game.
[0070] Further, the second machine learning model for the ASR system may be a pre-trained ANN as described in the “Deep speech: Scaling up end-to-end speech recognition.”, by Hannun, Awni, et al., published on arXiv preprint arXiv: 1412.5567 (2014). This may be used to transform the natural language input of the user into speech and forward it to the trained transformer model. Fig- 2 illustrates a block diagram of a machine learning based video game training system 200 for open-ended games. The first machine learning model 204 (large language model LLM), running on the circuitry 110, obtains the natural language input comprising a game task with regards to a video game 212 from the user 202. The first machine learning model 202 determines the first data structure 206 based on the natural language input. The first data structure 206 comprises first configuration data (structured exercise configuration) for the video game 212 corresponding to the game task. That is, in order to understand the goals of the user 202 and translate them into a suitable training exercise, the LLM 204 is utilized, providing a natural language interface to the user 202 and generate the first data structured for the video 212 game.
[0071] Further, the circuitry 110 sets up the video game 212 with the first data structure 206 as a video game configuration and loads the exercise scenario. The user 202 plays the video game 212 configured with the structure exercise configuration of the first data structure 206. The user 202 may therefore use the display screen 202a and the keyboard 202b. Further, an evaluator 214 determines a performance evaluation (feedback for player) 216 of the user 202 playing the video game 212 configured with the first data structure 206. The performance may be displayed as feedback information to the user 202 on the display screen 202a. The performance may also be fed back to the first machine learning model 204 together with the natural language input to determine a second data structure. Further, optionally, the circuitry 110 (for instance the first machine learning model 204) may query from a database 208 a virtual agent (expert agent) and a game-internal metric (value function) 210 as described above. The virtual agent may be configured and then demonstrate the game task in the video game 212 to the user 202. The game-internal metric (value function) 210 may be transmitted to the evaluator 212.
[0072] In other words, it is provided a system 200 to deliver automated coaching for open-ended games that players can use to improve their skills towards participation in the competitive sphere. The LLM 204 is utilized to understand the intended practice and / or objective of the player 202. Proficient virtual agents 210 may be used to infer optimal policies to use as reference. Further comparative analysis and feedback mechanisms may be used to communicate suggestions to the player 202. Further, structured generation of training exercises may be generated that may adapt to a player’s goals and performance. Fig- 3 illustrates a flowchart of an example of a method 300. The method 300 may, for instance, be performed by an apparatus as described herein, such as apparatus 100. The method 300 comprises obtaining 310 a natural language input comprising a game task with regards to a video game. The method 300 further comprises determining 320, by a first machine learning model, a first data structure based on the input. The first data structure comprises first configuration data for the video game corresponding to the game task. The method 300 further comprises receiving 330 a performance evaluation of a user playing the video game configured with the first data structure.
[0073] More details and aspects of the method 300 are explained in connection with the proposed technique or one or more examples described above, e.g., with reference to Fig. 1. The method 300 may comprise one or more additional optional features corresponding to one or more aspects of the proposed technique, or one or more examples described above.
[0074] The aspects and features described in relation to a particular one of the previous examples may also be combined with one or more of the further examples to replace an identical or similar feature of that further example or to additionally introduce the features into the further example.
[0075] In the following, some examples of the proposed concept are presented:
[0076] An example (e.g., example 1) relates to apparatus comprising circuitry configured to obtain a natural language input comprising a game task with regards to a video game, determine, by a first machine learning model, a first data structure based on the natural language input, wherein the first data structure comprises first configuration data for the video game corresponding to the game task, and receive a performance evaluation of a user playing the video game configured with the first data structure.
[0077] Another example (e.g., example 2) relates to a previous example (e.g., example 1) or to any other example, further comprising that the circuitry is further configured to determine, by the first machine learning model, a second data structure based on the performance evaluation and the natural language input, wherein the second data structure comprises second configuration data for the video game corresponding to the game task. Another example (e.g., example 3) relates to a previous example (e.g., one of the examples 1 to 2) or to any other example, further comprising that the circuitry is further configured to determine the performance evaluation of the user playing the video game configured with the first data structure.
[0078] Another example (e.g., example 4) relates to a previous example (e.g., example 3) or to any other example, further comprising that the circuitry is further configured to determine the performance evaluation of the user playing the video game configured with the first data structure based on an internal metric of the video game.
[0079] Another example (e.g., example 5) relates to a previous example (e.g., one of the examples 3 or 4) or to any other example, further comprising that the circuitry is further configured to determine the performance evaluation of the user playing the video game configured with the first data structure based on a required time to complete the game task.
[0080] Another example (e.g., example 6) relates to a previous example (e.g., one of the examples 1 to 5) or to any other example, further comprising that the circuitry is further configured to determine a reference performance value of playing the video game configured with the first data structure.
[0081] Another example (e.g., example 7) relates to a previous example (e.g., example 7) or to any other example, further comprising that the circuitry is further configured to determine a difference between the reference performance value and the performance evaluation.
[0082] Another example (e.g., example 8) relates to a previous example (e.g., one of the examples 1 to 7) or to any other example, further comprising that the circuitry is further configured to display a virtual agent executing the video game set up with the first configuration data for the video game corresponding to the game task.
[0083] Another example (e.g., example 9) relates to a previous example (e.g., one of the examples 1 to 8) or to any other example, further comprising that the circuitry is further configured to set up the video game with the first data structure as a video game configuration. Another example (e.g., example 10) relates to a previous example (e.g., one of the examples 1 to 9) or to any other example, further comprising that the first data structure has a predetermined format.
[0084] Another example (e.g., example 11) relates to a previous example (e.g., one of the examples 1 to 10) or to any other example, further comprising that the first data structure has a predetermined format compatible with an application programming interface, API, of the video game.
[0085] Another example (e.g., example 12) relates to a previous example (e.g., one of the examples 1 to 11) or to any other example, further comprising that the circuitry is further configured to train the first machine learning model, based on video game specific training data.
[0086] Another example (e.g., example 13) relates to a previous example (e.g., one of the examples 1 to 12) or to any other example, further comprising that first machine learning model is an artificial neural network, ANN, for example a large language model.
[0087] Another example (e.g., example 14) relates to a previous example (e.g., one of the examples 1 to 13) or to any other example, further comprising that the natural language input is a text input and / or a speech input.
[0088] Another example (e.g., example 15) relates to a previous example (e.g., one of the examples 1 to 14) or to any other example, further comprising that the circuitry is further configured to perform automatic speech recognition when the natural language input is speech input and forward the output of the automatic speech recognition to the first machine learning model.
[0089] Another example (e.g., example 16) relates to a previous example (e.g., example 15) or to any other example, further comprising that the automatic speech recognition is performed by a second machine learning model.
[0090] Another example (e.g., example 17) relates to a previous example (e.g., one of the examples 1 to 16) or to any other example, further comprising that the video game is open-ended game. Another example (e.g., example 18) relates to a previous example (e.g., one of the examples 1 to 17) or to any other example, further comprising that the game task defines an exercising objective with regards to the video game.
[0091] An example (e.g., example 19) relates to a method comprising obtaining a natural language input comprising a game task with regards to a video game, determining, by a first machine learning model, a first data structure based on the input, wherein the first data structure comprises first configuration data for the video game corresponding to the game task, and receiving a performance evaluation of a user playing the video game configured with the first data structure.
[0092] Another example (e.g., example 20) relates to a non-transitory machine-readable medium having stored thereon a program having a program code for performing any one of the methods according to example 19, when the program is executed on a processor or a programmable hardware.
[0093] Another example (e.g., example 21) relates to a program having a program code for performing any one of the methods according to example 19, when the program is executed on a processor or a programmable hardware.
[0094] Examples may further be or relate to a (computer) program including a program code to execute one or more of the above methods when the program is executed on a computer, processor or other programmable hardware component. Thus, steps, operations or processes of different ones of the methods described above may also be executed by programmed computers, processors or other programmable hardware components. Examples may also cover program storage devices, such as digital data storage media, which are machine-, processor- or computer- readable and encode and / or contain machine-executable, processor-executable or computerexecutable programs and instructions. Program storage devices may include or be digital storage devices, magnetic storage media such as magnetic disks and magnetic tapes, hard disk drives, or optically readable digital data storage media, for example. Other examples may also include computers, processors, control units, (field) programmable logic arrays ((F)PLAs), (field) programmable gate arrays ((F)PGAs), graphics processor units (GPU), application- specific integrated circuits (ASICs), integrated circuits (ICs) or system-on-a-chip (SoCs) systems programmed to execute the steps of the methods described above.
[0095] It is further understood that the disclosure of several steps, processes, operations or functions disclosed in the description or claims shall not be construed to imply that these operations are necessarily dependent on the order described, unless explicitly stated in the individual case or necessary for technical reasons. Therefore, the previous description does not limit the execution of several steps or functions to a certain order. Furthermore, in further examples, a single step, function, process or operation may include and / or be broken up into several sub-steps, - functions, -processes or -operations.
[0096] If some aspects have been described in relation to a device or system, these aspects should also be understood as a description of the corresponding method. For example, a block, device or functional aspect of the device or system may correspond to a feature, such as a method step, of the corresponding method. Accordingly, aspects described in relation to a method shall also be understood as a description of a corresponding block, a corresponding element, a property or a functional feature of a corresponding device or a corresponding system.
[0097] The following claims are hereby incorporated in the detailed description, wherein each claim may stand on its own as a separate example. It should also be noted that although in the claims a dependent claim refers to a particular combination with one or more other claims, other examples may also include a combination of the dependent claim with the subject matter of any other dependent or independent claim. Such combinations are hereby explicitly proposed, unless it is stated in the individual case that a particular combination is not intended. Furthermore, features of a claim should also be included for any other independent claim, even if that claim is not directly defined as dependent on that other independent claim.
Claims
ClaimsWhat is claimed is:
1. Apparatus comprising circuitry configured to: obtain a natural language input comprising a game task with regards to a video game; determine, by a first machine learning model, a first data structure based on the natural language input, wherein the first data structure comprises first configuration data for the video game corresponding to the game task; and receive a performance evaluation of a user playing the video game configured with the first data structure.
2. The apparatus of claim 1, wherein the circuitry is further configured to determine, by the first machine learning model, a second data structure based on the performance evaluation and the natural language input, wherein the second data structure comprises second configuration data for the video game corresponding to the game task.
3. The apparatus of claim 1, wherein the circuitry is further configured to determine the performance evaluation of the user playing the video game configured with the first data structure.
4. The apparatus of claim 3, wherein the circuitry is further configured to determine the performance evaluation of the user playing the video game configured with the first data structure based on an internal metric of the video game.
5. The apparatus of claim 3, wherein the circuitry is further configured to determine the performance evaluation of the user playing the video game configured with the first data structure based on a required time to complete the game task.
6. The apparatus of claim 1, wherein the circuitry is further configured to determine a reference performance value of playing the video game configured with the first data structure.
7. The apparatus of claim 7, wherein the circuitry is further configured to determine a difference between the reference performance value and the performance evaluation.
8. The apparatus of claim 1, wherein the circuitry is further configured to display a virtual agent executing the video game set up with the first configuration data for the video game corresponding to the game task.
9. The apparatus of claim 1, wherein the circuitry is further configured to set up the video game with the first data structure as a video game configuration.
10. The apparatus of claim 1, wherein the first data structure has a predetermined format.
11. The apparatus of claim 1, wherein the first data structure has a predetermined format compatible with an application programming interface, API, of the video game.
12. The apparatus of claim 1, wherein the circuitry is further configured to train the first machine learning model, based on video game specific training data.
13. The apparatus of claim 1, wherein first machine learning model is an artificial neural network, ANN, for example a large language model.
14. The apparatus of claim 1, wherein the natural language input is a text input and / or a speech input.
15. The apparatus of claim 1, wherein the circuitry is further configured to perform automatic speech recognition when the natural language input is speech input and forward the output of the automatic speech recognition to the first machine learning model.
16. The apparatus of claim 1, wherein the video game is open-ended game.
17. The apparatus of claim 1, wherein the game task defines an exercising objective with regards to the video game.
18. A method comprising: obtaining a natural language input comprising a game task with regards to a video game;determining, by a first machine learning model, a first data structure based on the input, wherein the first data structure comprises first configuration data for the video game corresponding to the game task; and receiving a performance evaluation of a user playing the video game configured with the first data structure.
19. A non-transitory machine-readable medium having stored thereon a program having a program code for performing any one of the methods according to claim 18, when the program is executed on a processor or a programmable hardware.
20. A program having a program code for performing any one of the methods according to claim 18, when the program is executed on a processor or a programmable hardware.