Appartus, system and method

A context-aware encoder and decoder system for esports games addresses latency issues in gaming communication by encoding and decoding game-specific audio data, ensuring efficient and low-latency communication and translation.

WO2025202182A1PCT designated stage Publication Date: 2025-10-02SONY GROUP CORP +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/058096
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-25
Filing Date
2025-03-25
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

In the realm of esports and competitive video gaming, effective and fast communication between team members is crucial for success, but existing technologies struggle to achieve ultra-low latency verbal communication, particularly in gaming environments.

Method used

A context-aware, game-specific encoder and decoder system utilizing neural networks to encode and decode gaming data, enabling ultra-low latency audio communication and translation, leveraging contextual information about the game state to compress and reproduce audio signals efficiently.

Benefits of technology

The system achieves ultra-low latency communication and translation, enhancing team coordination and expanding the pool of potential esports athletes by improving language-independent communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025058096_02102025_PF_FP_ABST
    Figure EP2025058096_02102025_PF_FP_ABST
Patent Text Reader

Abstract

It is provided an apparatus comprising circuitry configured to obtain natural language data comprising content relating to a video game. Further, the circuitry is configured to obtain state information data relating to a current state of the video game. Further, the circuitry is configured to determine, by an encoder, compressed encoded input data based on the natural language data and on the state information data. The encoder is adapted for the video game.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] APP ARTUS, SYSTEM AND METHOD

[0002] Field

[0003] The present disclosure relates to a context-aware, game-specific, audio encoder and decoder. In particular, examples of the present disclosure relate to apparatuses, a system, and methods.

[0004] Background

[0005] In the realm of esport and competitive video gaming, effective and fast communication, for instance between team members, during a competitive video game is important to success. Latency of the communication transmission may therefore be an important factor as it may influence the subsequent reaction time of the players in the video game. The delay in relaying essential information, whether it be strategic cues, enemy positions, or tactical instructions, can directly impact players' ability to react promptly and make informed decisions.

[0006] Therefore, faster communication of players in video games may be desirable.

[0007] Summary

[0008] According to a first aspect, the present disclosure provides an apparatus comprising circuitry configured to obtain natural language data comprising content relating to a video game. The circuitry is further configured to obtain state information data relating to a current state of the video game. The circuitry is further configured to determine, by an encoder, compressed encoded input data based on the natural language data and on the state information data, wherein the encoder is adapted for the video game.

[0009] According to a second aspect, the present disclosure provides an apparatus comprising circuitry configured to obtain compressed encoded data, the compressed encoded data is encoding natural language data, the natural language data comprising content relating to a video game. The circuitry is further configured to obtain state information data relating to a state of the video game. The circuitry is further configured to determine, by a decoder, reconstructed data based on the compressed encoded data and on the state information data, wherein the decoder is adapted for the video game.

[0010] According to a third aspect, the present disclosure provides a system comprising a first apparatus comprising circuitry configured to obtain natural language data comprising content relating to a video game. The circuitry is further configured to obtain state information data relating to a current state of the video game. The circuitry is further configured to determine, by an encoder, compressed encoded input data based on the content relating to the video game of the natural language data and on the current state information data, wherein the encoder is adapted for the video game. The circuitry is further configured to transmit the compressed encoded input data to a second apparatus. The system further comprises a second apparatus comprising circuitry configured to obtain the compressed encoded input data. The circuitry is further configured to obtain the current state information data. The circuitry is further configured to determine, by a decoder, reconstructed data based on the encoded input data and on the state information data, wherein the decoder is adapted for the video game.

[0011] According to a fourth aspect, the present disclosure provides a method comprising obtaining natural language data comprising content relating to a video game. The method further comprises obtaining state information data relating to a current state of the video game. The method further comprises determining, by an encoder, compressed encoded input data based on the natural language data and on the state information data, wherein the encoder is adapted for the video game.

[0012] According to a fourth aspect, the present disclosure provides a method obtaining compressed encoded data, the compressed encoded data is encoding natural language data, the natural language data comprising content relating to a video game. The method further comprises obtaining state information data relating to a state of the video game. The method further comprises determining, by a decoder, reconstructed data based on the compressed encoded data and on the state information data, wherein the decoder is adapted for the video game.

[0013] Further aspects are set forth in the appended set of claims.

[0014] Brief description of the Figures Some examples of apparatuses and / or methods will be described in the following by way of example only, and with reference to the accompanying figures, in which

[0015] Fig. 1 illustrates a block diagram of an example of an apparatus;

[0016] Fig. 2 illustrates a block diagram of an example of an apparatus;

[0017] Fig. 3 illustrates a block diagram of an example of an apparatus;

[0018] Fig. 4 illustrates a block diagram of an example of system;

[0019] Fig. 5 illustrates an example of an artificial neural network (ANN)-based encoder-decoder system;

[0020] Fig. 6 illustrates an example of encoder-decoder system;

[0021] Fig. 7 illustrates a flowchart of an example of a method; and

[0022] Fig. 8 illustrates a flowchart of an example of a method.

[0023] Detailed Description

[0024] Some examples are now described in more detail with reference to the enclosed figures. However, other possible examples are not limited to the features of these embodiments described in detail. Other examples may include modifications of the features as well as equivalents and alternatives to the features. Furthermore, the terminology used herein to describe certain examples should not be restrictive of further possible examples.

[0025] Throughout the description of the figures same or similar reference numerals refer to same or similar elements and / or features, which may be identical or implemented in a modified form while providing the same or a similar function. The thickness of lines, layers and / or areas in the figures may also be exaggerated for clarification. When two elements A and B are combined using an “or”, this is to be understood as disclosing all possible combinations, i.e. only A, only B as well as A and B, unless expressly defined otherwise in the individual case. As an alternative wording for the same combinations, "at least one of A and B" or "A and / or B" may be used. This applies equivalently to combinations of more than two elements.

[0026] If a singular form, such as “a”, “an” and “the” is used and the use of only a single element is not defined as mandatory either explicitly or implicitly, further examples may also use several elements to implement the same function. If a function is described below as implemented using multiple elements, further examples may implement the same function using a single element or a single processing entity. It is further understood that the terms "include", "including", "comprise" and / or "comprising", when used, describe the presence of the specified features, integers, steps, operations, processes, elements, components and / or a group thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, processes, elements, components and / or a group thereof.

[0027] In order to realizes low-latencies, especially in the realm of esport and competitive video gaming, in previous approaches specialized hardware such as gaming mice or the like may have been used. Further, in order to realize low latency, techniques such as HDMI’s Auto Low Latency Mode (ALLM) were used. ALLM is a feature designed to enhance gaming experiences by minimizing input lag between gaming consoles and displays. When enabled, ALLM may automatically switch the display into a low-latency mode when it detects that a compatible gaming device is connected. However, it may be also desirable to improve the latency of communication, for instance among team members. Since, fast and effective communication especially in certain languages may be a lacking skill for some otherwise talented esport athletes. An improved low-latency communication may widen the pool of available esport athlete candidates, particularly with regards to a player's spoken language.

[0028] The disclosed technique proposes a context-aware, game-specific, audio encoder and decoder and encoder-decoder framework to achieve ultra-low latency verbal communication in esport. The proposed solution also enables translations. That is the disclosed technique proposes an encoder / decoder which is designed to deliver ultra-low latency audio communication between team members in a video game and / or translations of said communication. This may be based on exploiting contextual information about the video game. In this regard the disclosed technique may not only allow to effectively compress and reproduce the original audio signal but also allows to transmit the same message with a minimized number of bytes.

[0029] Encoder

[0030] Fig- 1 illustrates a block diagram of an example of an apparatus 100. The apparatus 100 comprises circuitry that is configured to provide the functionality of the apparatus 100. The apparatus 100 comprises a processing circuitry 110. For example, the processing circuitry 110 may be a single dedicated processor, a single shared processor, or a plurality of individual processors, some of which or all of which may be shared, a digital signal processor (DSP) hardware, an application specific integrated circuit (ASIC), a neuromorphic processor or a field programmable gate array (FPGA). The processing circuitry 110 may optionally be coupled to, e.g., memory such as read only memory (ROM) for storing software, random access memory (RAM) and / or non-volatile memory. For example, the apparatus 100 may comprise memory configured to store instructions, which when executed by the processing circuitry 110, cause the processing circuitry 110 to perform the steps and methods described herein.

[0031] The circuitry 110 is configured to obtain natural language data comprising content relating to a video game. For instance, the video game may be an electronic game that involves interaction with a user interface or input device - such as a joystick, controller, keyboard, or motion sensing device - to generate visual feedback for a player. For example, a game like chess or tic tac toe, or a racing game or an action game or strategy game like “World of Warcraft” or “The Legend of Zelda” or the like.

[0032] The natural language data may be data obtained from a natural language utterance, for instance utter by a player of playing the video game. The natural language data may comprise content and / or acoustic features. Content may refer to the intelligible elements, like spoken words and meaning. Acoustic features may refer to technical aspects such as frequency, pitch, timbre, and amplitude of the utterance. The natural language data may be voice input which is available in a common audio format such as Waveform Audio File Format (WAV), MPEG Layer 3 (MP3), Free Lossless Audio Codec (FLAC), Advanced Audio Coding (AAC) or the like. For instance, the encoder may receive the natural language data in any one of the mentioned or any other format.

[0033] In some examples, the natural language data may be text input and / or voice input or both, in a natural language. The natural language data may be in a natural language used by humans, such as German, Japanese, or English or the like.

[0034] The natural language data comprises content relating to a video game. The content relating to the video game may comprise information or data that describes, discusses, and / or reports on the current events, status, actions, and / or circumstances within the video games environment. For example, the content relating to the video game may be real-time data, interactions, and narrative elements that reflect the immediate context, actions, and events occurring within the game's environment, alongside details of the game's mechanics, design, and structure relevant to the current gameplay situation. Additionally, content relating to the video game may comprise updates on the game state, such as changes in the environment, progression in the storyline, or alterations in the player's status (health, resources, etc.), constitute vital content for the player.

[0035] For example, the content relating to the video game represents the dynamic aspects of video game, such as a player's decisions, actions, or position within the game's environment. Further, the content relating to the video game may be content that is communicated and shared among players of the video game or pertain to an individual’ s engagement with the game. When players interact with each other, the content relating to the video game that is exchanged may comprise strategic decisions, such as planning the next move in a strategy game or sharing tactical information during a cooperative mission.

[0036] For example, in chess or tic tac toe video game the natural language data comprising content relating to the video game may be a player discussing his potential next move or a tic tac toe player declaring his strategy to a team member. In a strategy video game “World of Warcraft” the natural language data comprising content relating to the video game may be a player describing his current mission or location and asking a team member to help him with an attack or the like. Similarly, in racing or action games, the natural language data comprising content relating to the video game cover a player describing his current position or strategy with team members.

[0037] The circuitry 110 is further configured to obtain state information data relating to a current state of the video game. State information data relating to a current state of the video game may be a set of data points that partially or completely describes or quantifies the current state of the video game. The state information data comprises current values of variables and / or elements within the video game. The state information data may include enough data points describing the relevant elements within the game to make a first game state at a first point in time distinguishable from a second game state at a second point in time. In other words, the state information data relating to a current state of the video game captures the context of the natural language input.

[0038] In some examples, state information data may be obtained from a server running the video game, which is determining the state information data. In some examples, the state information data is obtained may be in a stateless representation. For instance, language settings may also be extracted from the game settings.

[0039] In some examples, the circuitry 110 may be further configured to determine the state information data based on current values of variables representing the video game. For instance, the state information data may comprise the current values of variables and / or elements within the video game. In some examples, determining the state information data may involve directly accessing the video game (that is the apparatus and memory storing and running the video game) to read out the values representing relevant variables and / or elements within the video game at a given point in time. In some examples, external tools may be utilized to capture the state information data by interpreting the video game's output. For example, based on availbe image recognition tools the current state of the video game may be analyzed to determine the positions of players and objects in a game, translating visual data back into a structured format.

[0040] For example, in a video game, the state information data may include the positions of all units on the map, their health levels, available resources, and the current game time. For instance, may capture the player's location, movement velocity, score, and the statuses of on-screen enemies and items. For instance, capturing the state information may involve directly accessing the game's memory to read values representing the player positions, health, and inventory in a structured format like an array or object in programming languages. For instance, the game state may be represented in JSON format, where player data, game environment settings, and other variables are stored in key-value pairs, offering a clear, structured, and readable snapshot of the game's current state.

[0041] For example, in chess the state information data may include positions of all 32 pieces on the board, each identified by its type and color, along with additional information like the player's turn, castling rights etc. This may be represented in an 8x8 matrix, where each cell contains information about the piece occupying that square, if any. For example, in tic-tac-toe, the state information data may include the markings of each of the 9 cells in the 3x3 grid, identified by either an “X”, “O”, or blank, representing the moves made by the two players. Additionally, the state information data may capture which player's turn it is and whether the game has ended in a win or a draw. This may be represented in a 3x3 matrix, where each cell contains information about the mark in that particular cell, if any.

[0042] In the example of an action video game like “The Legend of Zelda”, the state information data may comprise the protagonist's location on the map, their health, inventory items, quest progress, and the status of key game-world elements. Each of these may be stored in a detailed, structured format, for example as a complex object with nested arrays or dictionaries, to capture the game's multifaceted environment and mechanics accurately.

[0043] In a racing video game, the game state may capture each racer's position, lap time, current speed, and car condition, along with the track conditions and the positions of various obstacles or power-ups. This information may be structured in a way that each car and track element is an object with its attributes, facilitating a detailed representation of the race's current dynamics.

[0044] The circuitry 110 is configured to determine, by an encoder, compressed encoded input data based on natural language data and on the state information data. The encoder is adapted for the video game. The circuitry 110 may be configured to execute the encoder, which determines the compressed encoded input data. The encoder may be a system or algorithm designed to transform the natural language data from its original format, in its original size, into a more compact format with reduced size. This transformation is aimed at reducing the size of the natural langue input data, making it more efficient to store and / or transmit.

[0045] When the encoder (executed by the circuitry 110) processes the natural language data, it may analyze the natural language data, and may identify and eliminate redundancies and represent the same information more compactly. The encoding process is aimed at capturing the essential information of the input while most often reducing its dimensionality or size, which may be particularly beneficial in applications like data compression, language translation, and image processing. In the encoding stage, the encoder compresses or transforms the input data into a condensed representation, often called a latent space or a feature space. For example, common phrases or repeated words in the natural language data may be substituted with shorter placeholders. The encoder may leverage patterns or structures inherent in the language to further compress the data. This process is further leveraged by including the state information data. The content relating to the video game may have a direct contextual connection to the state information data relating to the current state of the video game. In other words, what a player is saying (or typing) in natural language while playing the video game may be contextually directly related the current state of the video game, which is captured in the state information data. The encoder may use this contextual connection to further compress the natural language data and determine the compressed encoded input data. Therefore, the encoder is adapted to the specific computer game. That is the encoder is specifically adapted the specific video game, for instance trained with regards to the video game (see below), to take the direct contextual connection between the content relating to this specific video game and the state information data of the specific video game into account when determining the compressed encoded data to further compress the natural language data without losing information.

[0046] In some examples, the encoder be adapted to compress the natural language data by discarding the all audio features of the natural language data (such as the speaker’s language) and only keep the content of the natural language data. The encoder may be used together with a decoder, and the two of them may form an encoderdecoder system (see below). The decoder may obtain the compressed encoded data and reverses the compression process, reconstructing the natural language data back to its original format and original size or almost original size or even extend it (see below)

[0047] In some embodiments the circuitry 110 is further configured to transmit the compressed encoded input data to the decoder. The decoder may as well receive the state information data relating to the current state of the video game and take its direct contextual connection into account when determining the reconstructed data.

[0048] In some examples, the compressed encoded input data may be sent to a video game server from where it is transmitted to another player with a decoder. For example, the compressed encoded input data is transmitted together with mouse / keyboard / controller input (and later to the game server), instead of relying on an external voice server such as Discord. In another example, a dedicated message broker may be implemented to handle the communication. This may be particularly useful when the publisher and receiver are in the same location, enabling point-to-point communication without networking. In some examples communication to an audience through streaming platforms such as Twitch may be used.

[0049] This allows for low latency transmission of the natural language data and for low-latency communication in video games. Low latency contextual translation for player-to-player audio communication in games may especially be useful in in competitive esport setting. The encoder may limit the vocabulary to the specifics of the game, which may enable content moderation and censorship.

[0050] Encoder / Decoder Techniques

[0051] The encoder and also the decoder (see below) may each be used alone or together in an encoder-decoder system. There are several different encoders, decoders and / or encoder-decoder systems utilizing a variety of techniques to transform, compress, and / or reconstruct data. In some examples, the encoder and decoder may be adapted to each other, as the effectiveness of the encoder in capturing the key features of the input directly influences the quality of the output produced by the decoder. Well-known encoder / decoder techniques comprise Discrete Cosine Transform (DCT), wavelet transforms, Modified Discrete Cosine Transform (MDCT), statistical models, such as Markov models or the like.

[0052] In some examples the encoder (and the decoder, see below) may be based on Artificial Neural Networks (ANNs). For instance, the encoder comprises an ANN which compresses the input data into a latent representation, and a decoder may comprise an ANN that reconstructs the output from this latent representation. Theses ANNs may be trained iteratively by adjusting the network's parameters to minimize a defined loss function (see below), which may quantify the difference between the input data and its reconstructed output or similar differences.

[0053] Further, the ANNs may be organized into different architectures which define a specific structure and design of a network, determining how it processes and learns from data.

[0054] One architecture are Autoencoders. These models are a type of neural network architecture designed for unsupervised learning, consisting of an encoder and a decoder. The encoder compresses the input data into a lower-dimensional latent space representation, while the decoder reconstructs the original input from this compressed representation. The network is trained to minimize the reconstruction error between the input and the output, effectively learning to compress and decompress data efficiently.

[0055] Another architecture that may be used as encoder-decoder system are Variational Autoencoders (VAEs). These models extend the concept of traditional autoencoders by incorporating probabilistic modeling. In addition to encoding the input into a latent space, VAEs also learn the distribution of the latent space, allowing for sampling and generating new data points. VAEs impose a regularization term, typically based on Kullback-Leibler divergence, on the latent space distribution during training, encouraging it to resemble a predefined prior distribution (often Gaussian). This regularization term ensures that the latent space remains smooth and continuous, facilitating meaningful interpolation between data points.

[0056] Another architecture that may be used as encoder-decoder system are Conditional Variational Autoencoders (CVAEs). These models extend the variational autoencoder (VAE) as described above. CVAEs introduce conditioning variables into the generative process, enabling the model to generate outputs based on given conditions. The encoder in a CVAE learns to represent the input data as a distribution in latent space, conditioned on auxiliary information. The decoder then uses samples from this distribution, along with the conditioning variable, to generate diverse and targeted outputs.

[0057] Another architecture is a convolutional autoencoders. These networks employ convolutional layers to encode the input and deconvolutional layers for decoding. In the encoder, convolutional layers extract hierarchical spatial features, while pooling layers reduce dimensionality, creating a compact latent representation. The decoder reversely reconstructs the original input from this latent space, using deconvolutional layers to upscale the representation and convolutional layers to refine the details.

[0058] Another architecture that may be used as encoder-decoder system are RNN Encoder-Decoders. These models leverage the temporal processing capabilities of Recurrent Neural Networks (RNNs). The encoder summarizes the input sequence into a fixed-size vector, capturing its temporal dynamics. The decoder, often another RNN, generates the output sequence from this vector, preserving the sequential information.

[0059] Another architecture that may be used as encoder-decoder system are transformer models. These models abandon recurrent layers for self-attention mechanisms, enabling the model to weigh the importance of different parts of the input data regardless of their positions. The transformer models allow parallel processing and effectively captures long-range dependencies. The transformer consists of an encoder stacking self-attention and feed-forward layers to encode the input into a series of representations, which the decoder, using similar layers and additional cross-attention with the encoder's output, decodes into the target sequence. The paper “Attention is all you need.”, Vaswani, Ashish, et al, published in Advances in neural information processing systems 30 (2017) a transformer model is described.

[0060] Another architecture that may be used as encoder-decoder system are U-Nets. This architecture features an encoder-decoder structure with skip connections that directly concatenate feature maps from the encoder to the decoder. This design enables precise localization by combining high-level contextual information and low-level detail features, essential for tasks like segmentation. In the paper ”U-net: Convolutional networks for biomedical image segmentation ”, by Ronneberger, Olaf, Philipp Fischer, and Thomas Brox, published in Medical image computing and computer-assisted intervention-MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18. Springer International Publishing, 2015, a U-Net model is described.

[0061] Another architecture that may be used as encoder-decoder system or may be used together with another ANN-based encoder-decoder system is Generative Adversarial Networks (GANs). This architecture is class of Al algorithms involving two neural networks, a generator, and a discriminator, which are trained simultaneously through adversarial processes in unsupervised learning. The generator aims to produce data resembling the training set, while the discriminator evaluates this against real data, trying to differentiate between actual and generated samples. Through iterative training, the generator improves its ability to create realistic outputs, aiming to fool the discriminator, which, in turn, gets better at distinguishing real data from fakes. This dynamic training continues until the generator produces highly realistic data, and the discriminator is essentially guessing, marking an equilibrium. GANs are renowned for their prowess in generating high-quality, lifelike images and are utilized across various domains, from art generation to medical imagery, demonstrating the versatile applications of this powerful Al architecture.

[0062] All the above mentioned encoders, decoders and encoder-decoder systems and the different architectures or combinations thereof may be used for separately for the encoder and the decoder and / or encoder-decoder system combined.

[0063] For example, the ANN-based encoder (and decoder, see below) may further use vector quantization. In this technique a high-dimensional input vector is mapped to a finite set of points in a lower-dimensional space, which is particularly prevalent in speech encoding and image compression. Central to this technique is the concept of a codebook (also referred to as a dictionary, specifically in case of speech encoding), which provides a predefined set of codewords or vectors in the lower-dimensional space. Each high-dimensional input vector is compared against the codebook, and the closest codeword in the codebook is selected as its representation, effectively reducing the data's dimensionality. This process is important for diminishing the bit rate. The use of a carefully constructed codebook ensures that the essential characteristics of the original signal or image are preserved, allowing for a faithful reconstruction of the original data from its compressed form. In some examples, the encoder as described above may comprise a trained ANN. The encoder ANN may be trained (see Fig. 3 below) based on training data related to the video game. For instance, the encoder may be trained based on training data comprising natural language data comprising content relating to the video game and corresponding state information data relating to the current state of the video game. The ANN-based encoder may be trained to determine the compressed encoded input data based on the natural language data and based on the state information data, such that the compressed encoded input data has a reduced size. The trained ANN-based encoder may receive as an input the natural language data and the state information data and determine the compressed encoded input data. This ANN-based encoder is able to further reduce the size of the compressed encoded input data without losing information compared to an ANN-based encoder that is only trained to determine the compressed encoded input data based on the natural language data without using the state information data. An implementation example is given Fig. 5 below.

[0064] As described above the encoder may be adapted to a specific video game. That is the encoder is adapted to be explicitly effective in compressing natural language data comprising content relating to the specific video game based on the state information data of the specific video game. However, there may be a plurality of encoders, each being adapted, for instance a trained ANN, to explicitly effectively compressing natural language data comprising content relating to the respective video game based on the state information data of the respective video game.

[0065] In this regard, the circuitry 110 may be further configured to identify the video game. For example, the circuitry 110 may identify the video game to which the content of the natural language data is relating. The circuitry 110 may be further configured to load the encoder adapted for the identified video game among a plurality of encoders adapted for a respective plurality of the video games. In some examples, the encoder may be adapted to a plurality of different video games.

[0066] In some examples, the circuitry 110 may identify the video game among a plurality of video games based on an entry in the data structure comprising the natural language data comprising the content relating to the video game. For instance, the data structure may comprise a header file, which comprises an entry indicating the specific video game. In another example the circuitry 110 may identify the video game among a plurality of video games based on a natural language data processing tool that analyzes the content of the natural language data relating to the video game and / or the state information data.

[0067] For example, if the apparatus 100 and the circuitry 110 are part of an input device such as a headphone, the circuity 110 may download the specific encoder adapted for the identified video game among a plurality of encoders adapted for a respective plurality of the video games in advance. For example, the different encoders may be distributed through a platform where users can download the specific one for the game they want and push it as a firmware update to their device.

[0068] In some examples, the circuitry 110 is further configured to determine the compressed encoded input data by the encoder wherein the encoder comprises a video game-specific codebook. The compressed encoded input data may be represented as data based on the video gamespecific codebook.

[0069] For instance, the encoder may utilize vector quantization as described above. In this technique the high-dimensional data structure representing the natural language data may be mapped onto a finite set of points in a lower-dimensional space codebook. The finite set of points in the lower-dimensional space may be referred to as codebook (also referred to as a dictionary). The codebook provides a predefined set of quantized vectors (referred to as codewords) in the lower-dimensional space. Each high-dimensional input vector is compared against the codebook, and the closest quantized vector in the codebook is selected as its representation, effectively reducing the data's dimensionality. In this regard, the compressed encoded input data may be represented by one or more quantized vectors (i.e., codewords) of the game-specific codebook.

[0070] The codebook of the encoder may be especially generated to represent natural language data relating to a specific video game. For instance, the codebook may be generated manually. In some examples, the codebook may be generated by iteratively selecting codewords that optimally represent the input data (a batch of input data). In some examples, the codebook may be generated during a training process of an ANN-based encoder (see below). There may be a plurality of codebooks, wherein each of the plurality of codebooks is specifically adapted for a respective video game. For example, the codewords of each of the plurality of codebooks is generated specifically for a respective video game.

[0071] In some examples, the circuitry 110 is further configured to identify the video game (as described above). In some examples, the circuitry 110 is configured to load the game-specific codebook for the identified video game among a plurality of game-specific codebooks each being adapted to a respective video game.

[0072] In some examples, the codebooks may be interpretable by designing an understandable yet minimal language for a specific video game (see for example Fig. 6). This allows to train the encoder and decoder independently and / or use cascading systems for Automatic Speech Recognition (ASR), Text-to-Text Translation (T2TT) and Text-to- Speech Synthesis (TTS), with the game-specific vocabulary and grammar as a new language.

[0073] The encoder as described above may be used as stand-alone application or together with a decoder as described below as an encoder-decoder system.

[0074] Further details and aspects are mentioned in connection with the examples described below. The example shown in Fig. 1 may include one or more optional additional features corresponding to one or more aspects mentioned in connection with the proposed concept or one or more examples described below (e.g., Figs. 2 - 8).

[0075] Decoder

[0076] Fig- 2 illustrates a block diagram of an example of an apparatus 200. The apparatus 100 comprises circuitry that is configured to provide the functionality of the apparatus 200. The apparatus 200 comprises a processing circuitry 210. The circuitry 210 may be similar or different to the circuitry 110.

[0077] The circuitry 210 is configured to obtain compressed encoded data. The compressed encoded data is encoding natural language data. The natural language data comprising content relating to a video game. Further, circuitry 210 is further configured to obtain state information data relating to a state of the video game. Further, the circuitry 210 is configured to determine, by a decoder, reconstructed data based on the compressed encoded data and on the state information data. The decoder is adapted for the video game.

[0078] The compressed encoded data, encoding natural language data may be a compressed and encoded natural language utterance, for instance uttered by a player who is playing the video game. The natural language data may have comprised content and / or acoustic features before being compressed and encoded. Content may refer to the intelligible elements, like spoken words. Acoustic features may refer technical aspects such as frequency, pitch, timbre, and amplitude of the utterance. The compressed encoded data may comprise parts (or all) of the content and / or acoustic features of the natural language data.

[0079] The state information data may be state information data comprising a state of the video game, that was acquired at a point in time when the natural language data was compressed and encoded and / or when the natural language utterance that is compressed and encoded in the natural language data was uttered.

[0080] For example, the compressed encoded data and / or the state information data are obtained by the circuitry 210 from the first apparatus 110. For example, they are transmitted together from the circuitry 110 to the circuitry 210. In another example the compressed encoded data and / or the state information data are obtained from another apparatus for example from the server running the video game.

[0081] In some examples, the decoder obtains the compressed encoded data and reverses the compression process, reconstructing the natural language data back to its original format and original size or almost original size or even expand it. The decoder takes the condensed representation and may reconstruct or generate output that may either be a replica of the original input or a transformation of the input into a new domain. The decoder may be used together with an encoder, which may form an encoder-decoder system.

[0082] In some examples, the reconstructed data may comprise the content relating to the video game of the natural language data. The reconstructed data may comprise some or all of the content and / or acoustic features of the natural language data. The decoder utilizes state information data during reconstruction to enhance the reconstruction process. The state information data may provide additional information that helps interpret the compressed encoded data more accurately. The state information data may provide additional contextual information that the decoder may use to make more informed decisions when reconstructing the original data from its compressed form.

[0083] This allows the compressed encoded data to have a small size and being highly compressed without losing information during decompression. This allows for low latency transmission and reception of the natural language data and for low-latency communication in video games. This may especially be useful in in competitive esport setting. Low latency contextual translation for player-to-player audio communication in games. Since the encoder limits the vocabulary to the specifics of the game, content moderation and censorship may be easily applied.

[0084] In some examples, the decoder may be based on a trained ANN. The decoder ANN may be trained (see Fig. 3 below) based on training data related to the video game. For instance, the decoder may be trained together with the encoder as described above. For example, the decoder may be trained based on training data comprising natural language data comprising content relating to the video game and corresponding state information data relating to the current state of the video game. The ANN-based decoder may be trained to determine the reconstructed data. The trained ANN-based decoder may receive as an input the compressed encoded data and based on the state information data and be trained to determine the reconstructed data. The reconstructed data may comprise the content relating to the video game of the natural language data.

[0085] In some examples the decoder may comprise a look-up table. For example, the look-up table may be a pre-defined, stored set of values or data points that the decoder uses to quickly decode and decompressed the encoded data back to their original (or approximate) data during the decoding process. The vector quantiziation may be an example of a look-up table. Instead of performing complex calculations in real-time, the decoder may consult look-up table to find the corresponding output for a given input value, significantly speeding up the decoding process. In some examples an indexing approach is used in this regard. Therefore, latency may be reduced even further by filtering the top-most likely results based on the state information data (i.e., game context) before the compressed encoded message arrives or before decoding starts.

[0086] In some examples, the circuitry 210 may be further configured to determine, by the decoder, extended reconstructed data based on the compressed encoded data and on the state information data. The extended reconstructed data may comprise content extending the content relating to the video game of the natural language data based on the state information data. For example, the extended reconstructed data comprises additional information or details that go beyond the content relating to the video game of the natural language data. For example, the extended reconstructed data is enhanced by insights of the video game derived from the state information data. In other words, the decoder is enriching the content relating to the video game of the natural language data with extra context or details that are relevant to the current state or scenario in the video game, as provided by the state information data (see also Fig. 6 below).

[0087] For example, an ANN-based decoder may be trained to determine the extended reconstructed data based on supervised training with labelled training data triplets, comprising the compressed encoded data, the state information data, and the extended reconstructed data. This process results in a more comprehensive or enriched output that incorporates and reflects the current in-game context, thereby providing a more detailed or nuanced interpretation or extension of the original natural language content.

[0088] In some examples, the reconstructed data comprises the content relating to the video game of the natural language data in a language different than the language of the natural language data. That is the reconstructed data is a translation of the natural language data. For example, the ANN-based decoder is trained to determine the reconstructed data in a language that is different from the language of natural language input. For example, the encoder as described above and the decoder are trained together, wherein the encoder is trained to obtain the natural language input in a first language and the decoder is trained to determine the reconstructed data in a second language.

[0089] In some examples, the circuitry 210 is further configured to generate a voice output based on the reconstructed data. For example, a Text-to-Speech (TTS) engine may be used in this regard. A TTS engine may be designed to read written text aloud, often using synthesized voices. For example, the circuity 210 may transmit the determined reconstructed data or a generate voice output based thereon to a speaker or headphones or earphones or the like. For example, the circuitry 210 may be part of a speaker or headphones or earphones or the like. For instance, the decoder may utilize vector quantization as described above. The compressed encoded data in this case may comprise indices or identifiers for codewords of a codebook. These indices may be used to look up the corresponding quantized vectors in the codebook. Each index may point to a specific vector in the codebook that represents an approximation of the original high-dimensional data. The decoder then determines the reconstructed data by sequentially replacing each codeword with its high-dimensional representation from the codebook. This may be an example of a look-up table.

[0090] In some examples, the circuitry 210 may further configured to decode the compressed encoded data by the decoder, wherein the decoder comprising a video game-specific codebook. The compressed encoded data is represented as data based on the video game-specific codebook. In some examples if an encoder and a decoder may be used together, they may use the same codebook.

[0091] As described above, the codebook of the decoder may be especially generated to represent natural language data relating to the specific video game. For instance, the codebook may be generated manually. In some examples, the codebook may be generated by iteratively selecting codewords that optimally represent the input data. In some examples, the codebook may be generated during a training process of an ANN-based encoder. There may be a plurality of codebooks, wherein each of the plurality of codebooks is specifically adapted for a respective video game. For example, the codewords of each of the plurality of codebooks is generated specifically for a respective video game.

[0092] In some examples, the decoder comprises a plurality of game-specific dictionaries each being adapted to a respective video game. In another example, the circuitry 210 may be further configured to identify the video game. Further, the circuitry 210 may be configured to load the game-specific codebook for the identified video game among a plurality of game-specific codebooks each being adapted to a respective video game. In some examples, the circuitry 210 may identify the video game among a plurality of video games based on an entry in the compressed encoded data. For instance, the compressed encoded data may comprise a header file, which comprises an entry indicating the specific video game. In another example the circuitry 210 may identify the video game among a plurality of video games based on natural language data processing tool that analyzes the state information data.

[0093] For example, if the apparatus 200 and the circuitry 210 are part of an output device such as speakers the circuity 210 may download the specific decoder adapted for the identified video game among a plurality of decoders adapted for a respective plurality of the video games in advance. For example, the different decoders may be distributed through a platform where users can download the specific one for the game they want and push it as a firmware update to their device.

[0094] Further details and aspects are mentioned in connection with the examples described above or below. The example shown in Fig. 2 may include one or more optional additional features corresponding to one or more aspects mentioned in connection with the proposed concept or one or more examples described above (e.g., Fig. 1) or below (e.g., Figs. 3 - 8).

[0095] Training of an ANN-based encoder / decoder

[0096] Fig- 3 illustrates a block diagram of an example of an apparatus 300. The apparatus 300 comprises circuitry that is configured to provide the functionality of the apparatus 300. The apparatus 300 comprises a processing circuitry 310. The circuitry 310 may be similar or different to the circuitry 110.

[0097] The circuitry 310 is configured to train a first ANN, of an encoder and a second ANN of a decoder based on training data. The training data comprises natural language data. The natural language data comprises content relating to a video game. Further, the first ANN is trained to encode natural language data comprising content relating to the video game and the second ANN is trained to reconstruct the content of the natural language data. The encoder and the decoder may each comprise an ANN. For example, during training the encoder ANN receives the natural language data and the state information data as input and determines as output of the encoder ANN the compressed encoded input data. The compressed encoded input data and the state information data is input into the decoder which determines the reconstructed data. Then a loss function may be determined which quantifies the difference between the natural language data and the reconstructed data and adapts the weights of the encoder ANN and the decoder ANN to minimize this loss function. Further, the training may be based on backpropagation to effectively learn to encode and decode the data efficiently. Backpropagation, coupled with optimization algorithms such as stochastic gradient descent may typically be used to update the network weights iteratively. During training, the network learns from a dataset representative of the task at hand, which could include images, text, audio, or any other relevant data format. Data augmentation techniques may be applied to increase dataset diversity and improve model generalization. The training process aims to teach the network to efficiently encode input data into a latent representation and decode it back to the original form accurately, ensuring that the reconstructed output closely matches the input data across various examples.

[0098] In some examples, the training may be based on training data pairs with similar content and in different languages. That is the encoder will receive the natural language input in a first language and the loss function may determine the difference between the natural language data in a second language and the reconstructed data. That is the encoder and decoder learn to compress / decompress the data and further translate it from the first language to the second language.

[0099] The encoder and decoder may be trained with regards to a specific video game. That is the training data may comprise natural language data comprising content relating to the specific video game and corresponding state information data relating to the current state of the specific video game.

[0100] The training data may be collected from an online forum dedicated to the video game or to videos, or books or journals dedicated to the video game, or a recording of a conference dedicated to the video game or similar. In some embodiments the decoder is trained during training to determine an extended reconstructed data that comprises additional information or details that go beyond the content relating to the video game of the natural language data. In this case a loss function may be used that quantifies the difference between extended reconstructed data and labeled training data that comprises the additional information or details that go beyond the content relating to the video game.

[0101] In some examples, the circuitry 310 is further configured to adapt, while training the first ANN and the second ANN, a game-specific codebook. The encoder ANN and the decoder ANN may be used together with a vector quantization as described above. The entries of the codebook may be adapted during training, such that the output of the encoder may be minimized for the generated codewords.

[0102] In some examples, the codebook may be interpretable by designing an understandable yet minimal language for a specific video game (see for example Fig. 6). This allows to train the encoder and decoder independently and / or use cascading systems for Automatic Speech Recognition (ASR), Text-to-Text Translation (T2TT) and Text-to- Speech Synthesis (TTS), with the game-specific vocabulary and grammar as a new language.

[0103] For example, the trained encoder ANN may be used as an encoder as described with regards to Fig. 1 and the trained decoder ANN may be used as a decoder as described with regards to Fig. 2.

[0104] Further details and aspects are mentioned in connection with the examples described above or below. The example shown in Fig. 3 may include one or more optional additional features corresponding to one or more aspects mentioned in connection with the proposed concept or one or more examples described above (e.g., Figs. 1-2) or below (e.g., Figs. 4 - 8).

[0105] Fig- 4 illustrates a block diagram of an example of system 400. The system 400 comprise a first apparatus 410 which comprises circuitry that is configured to provide the functionality of the first apparatus 410. The first apparatus 410 comprises a processing circuitry 420. The circuitry 420 may be similar or different to the circuitry 110. The system 400 further comprise a second apparatus 430 which comprises circuitry that is configured to provide the functionality of the second apparatus 430. The second apparatus 430 comprises a processing circuitry 440.

[0106] The circuitry 440 may be similar or different to the circuitry 110.

[0107] The first apparatus 410 comprises circuitry 410 configured to obtain natural language data comprising content relating to a video game. The circuitry 410 is further configured to obtain state information data relating to a current state of the video game. The circuitry 410 is further configured to determine, by an encoder, compressed encoded input data based on the content relating to the video game of the natural language data and on the current state information data. The encoder is adapted for the video game. The circuitry 410 is further configured to transmit the compressed encoded input data to a second apparatus 430.

[0108] The second apparatus comprises circuitry 440 configured to obtain the compressed encoded input data. The circuitry 440 is further configured to obtain the current state information data. The circuitry 440 is further configured to determine, by a decoder, reconstructed data based on the encoded input data and on the state information data. The decoder is adapted for the video game.

[0109] The first apparatus 410, the circuitry 420 and the encoder may be implemented as described above (e.g., with regards to Fig. 1). The second apparatus 430, the circuitry 440 and the decoder may be implemented as described above (e.g., with regards to Fig. 2).

[0110] In one example, apparatus 100 and circuitry 110 running the encoder and apparatus 200 and circuitry 1210 running the decoder may be part of a headphone device as input / output device, for example each connected to a computer or console, with either networking or direct communication between the devices.

[0111] In another example, apparatus 100 and circuitry 110 running the encoder may part of an input such as microphone and the and apparatus 200 and circuitry 1210 running the decoder may be part of an output device such as speaker and preset each for the communication with its pair on the other side (microphone_player_l, speaker_player_2).

[0112] Further details and aspects are mentioned in connection with the examples described above or below. The example shown in Fig. 4 may include one or more optional additional features corresponding to one or more aspects mentioned in connection with the proposed concept or one or more examples described above (e.g., Figs. 1-3) or below (e.g., Figs. 5 - 8).

[0113] Example of ANN-based Encoder-Decoder System

[0114] The encoder and decoder as described above may be based on an ANN, such as disclosed in the scientific paper “High fidelity neural audio compression ”, by Defossez, A., Copet, J., Synnaeve, G., & Adi, Y. (2022), arXiv preprint arXiv:2210.13438.

[0115] The encoder-decoder architecture of the paper together with a vector quantization may be as illustrated in Fig. 5. Fig. 5 illustrates an example of ANN-based encoder-decoder system 500. The encoder 510 may receive the natural language data 512 comprising content relating to a video game and the state information data 514 relating to the current state of the video game as input. The encoder 510 determines output data 516. The output data 516 is input into a vector quantizer 520, comprising a codebook, and determining the encoded compressed input data 522 (i.e., the quantized representation). The vector quantizer 520 may be considered as a part of the encoder 510. The encoded compressed input data 522 and the state information data 514 relating to the current state of the video game is obtained as input by the decoder 530. The decoder 530 determines the reconstructed data (or expanded reconstructed data) 532. The context-aware encoding / decoding is achieved by providing the state information data 514 relating to the current state of the video game (comprising a game state and configuration) as input to encoder 510 and the decoder 530.

[0116] During training of the encoder-decoder system 500, the message-centric representation may be enabled, by replacing the reconstruction losses “Is” and “It” from the “High fidelity neural audio compression”-paper by comparing audio pairs comprising the same content relating to the video game but in different languages, for example from different speakers, es descriebd above with regards to Fig. 3. The training data and the training may be as described above (e.g., with regards to Fig. 3). Further explanations mining training data and training this system is stated in the book by Barrault, Loic, et al. “SeamlessM4T-Massively Multilingual & Multimodal Machine Translation.” arXiv preprint arXiv:2308.11596 (2023). In some examples, multiple game-specific codebooks (also referred to as vocabularies) may be trained and used. These may be implemented by learning different dictionaries for the vector quantization (one for each game).

[0117] Further details and aspects are mentioned in connection with the examples described above or below. The example shown in Fig. 5 may include one or more optional additional features corresponding to one or more aspects mentioned in connection with the proposed concept or one or more examples described above (e.g., Figs. 1-4) or below (e.g., Figs. 6 - 8).

[0118] Fig- 6 illustrates an example of encoder-decoder system 600. A first player 602 and a second player 604 play together in the same team the video game tic-tac-toe, for example against a third player or a bot. The first player utters the natural language utterance 603 “Defend bottom line” in English. This natural language utterance may be available in a computer readable format as natural language data, for example as wav file. The natural language data is obtained by the encoder 602. Further, the encoder obtains the state information data comprising the current state 608 of the video game. For the in tic-tac-toe video game the state information data 608 includes the markings of each of the 9 cells in the 3x3 grid, identified by either an “X”, “O”, or blank as shown in Fig. 6. The encoder 606 determines the encoded compress input data as described above and transmits to the decoder 610. The encoded compress input data may be a minimal game-dependent representation (such as “[3,2]X”), possibly devoid of any audio features, such as the speaker’s language. The decoder further obtains the state information data comprising the current state 608 of the video game. Based on the encoded compress input data and the state information data comprising the current state 608 of the video game the decoder determines expanded reconstructed data 605 “Put X on bottom center”. In yet another example the decoder 610 may determine translated expanded reconstructed data, for example in Spanish, like “Pon X abajo en el centro”.

[0119] Further details and aspects are mentioned in connection with the examples described above or below. The example shown in Fig. 6 may include one or more optional additional features corresponding to one or more aspects mentioned in connection with the proposed concept or one or more examples described above (e.g., Figs. 1-5) or below (e.g., Figs. 7 - 8). Fig- 7 illustrates a flowchart of an example of a method 700. The method 700 may, for instance, be performed by an apparatus as described herein, such as apparatus 100. The method 700 comprises obtaining 710 natural language data comprising content relating to a video game. The method 700 further comprises obtaining 720 state information data relating to a current state of the video game. The method 700 further comprises determining 730, by an encoder, compressed encoded input data based on the natural language data and on the state information data, wherein the encoder is adapted for the video game.

[0120] More details and aspects of the method 700 are explained in connection with the proposed technique or one or more examples described above, e.g., with reference to Fig. 1. The method 700 may comprise one or more additional optional features corresponding to one or more aspects of the proposed technique, or one or more examples described above.

[0121] Further details and aspects are mentioned in connection with the examples described above or below. The example shown in Fig. 7 may include one or more optional additional features corresponding to one or more aspects mentioned in connection with the proposed concept or one or more examples described above (e.g., Figs. 1-6) or below (e.g., Fig. 8).

[0122] Fig- 8 illustrates a flowchart of an example of a method 800. The method 800 may, for instance, be performed by an apparatus as described herein, such as apparatus 200. The method 800 comprises obtaining 810 compressed encoded data, the compressed encoded data is encoding natural language data, the natural language data comprising content relating to a video game. The method 800 further comprises obtaining 820 state information data relating to a state of the video game. The method 800 further comprises determining 830, by a decoder, reconstructed data based on the compressed encoded data and on the state information data, wherein the decoder is adapted for the video game.

[0123] More details and aspects of the method 800 are explained in connection with the proposed technique or one or more examples described above, e.g., with reference to Fig. 2. The method 800 may comprise one or more additional optional features corresponding to one or more aspects of the proposed technique, or one or more examples described above. Further details and aspects are mentioned in connection with the examples described above. The example shown in Fig. 8 may include one or more optional additional features corresponding to one or more aspects mentioned in connection with the proposed concept or one or more examples described above (e.g., Figs. 1-7).

[0124] In the following, some examples of the proposed concept are presented:

[0125] An example (e.g., example 1) relates to an apparatus comprising circuitry configured to obtain natural language data comprising content relating to a video game, obtain state information data relating to a current state of the video game, determine, by an encoder, compressed encoded input data based on the natural language data and on the state information data, wherein the encoder is adapted for the video game.

[0126] Another example (e.g., example 2) relates to a previous example (e.g., example 1) or to any other example, further comprising that the circuitry is further configured to transmit the compressed encoded input data to a decoder.

[0127] Another example (e.g., example 3) relates to a previous example (e.g., one of the examples 1 to 2) or to any other example, further comprising that the circuitry is further configured to identify the video game, and load the encoder adapted for the identified video game among a plurality of encoders adapted for a respective plurality of the video games.

[0128] Another example (e.g., example 4) relates to a previous example (e.g., one of the examples 1 to 3) or to any other example, further comprising that the circuitry is further configured to determine the state information data based on current values of variables representing the video game.

[0129] Another example (e.g., example 5) relates to a previous example (e.g., one of the examples 1 to 4) or to any other example, further comprising that the encoder comprises an artificial neural network, ANN, wherein the ANN is trained based on training data related to the video game.

[0130] Another example (e.g., example 6) relates to a previous example (e.g., one of the examples 1 to 4) or to any other example, further comprising that the circuitry is further configured to determine the compressed encoded input data by the encoder, the encoder comprising a video game-specific codebook, wherein the compressed encoded input data is represented as data based on the video game-specific codebook.

[0131] Another example (e.g., example 7) relates to a previous example (e.g., example 6) or to any other example, further comprising that the compressed encoded data is represented by one or more quantized vector based on a video game-specific codebook.

[0132] Another example (e.g., example 8) relates to a previous example (e.g., one of the examples 6 to 7) or to any other example, further comprising that the circuitry is further configured to identify the video game, and load the game-specific codebook for the identified video game among a plurality of game-specific codebooks each being adapted to a respective video game.

[0133] Another example (e.g., example 9) relates to a previous example (e.g., one of the examples 1 to 8) or to any other example, further comprising that the natural language data is a text input and / or a voice input.

[0134] Another example (e.g., example 10) relates to a previous example (e.g., one of the examples 1 to 9) or to any other example, further comprising that the natural language data is a voice input and audio features of the voice input are discarded in the encoded signal.

[0135] An example (e.g., example 11) relates to an apparatus comprising circuitry configured to obtain compressed encoded data, the compressed encoded data is encoding natural language data, the natural language data comprising content relating to a video game, obtain state information data relating to a state of the video game, determine, by a decoder, reconstructed data based on the compressed encoded data and on the state information data, wherein the decoder is adapted for the video game.

[0136] Another example (e.g., example 12) relates to a previous example (e.g., example 11) or to any other example, further comprising that the reconstructed data comprises the content relating to the video game of the natural language data. Another example (e.g., example 13) relates to a previous example (e.g., one of the examples 11 to 12) or to any other example, further comprising that the circuitry is further configured to determine, by the decoder, extended reconstructed data based on the compressed encoded data and on the state information data, wherein the extended reconstructed data comprises content extending the content relating to the video game of the natural language data based on the state information data.

[0137] Another example (e.g., example 14) relates to a previous example (e.g., one of the examples 11 to 13) or to any other example, further comprising that the decoder comprising a look-up table.

[0138] Another example (e.g., example 15) relates to a previous example (e.g., one of the examples 11 to 14) or to any other example, further comprising that the reconstructed data comprises the content relating to the video game of the natural language data in a language different than the language of the natural language data.

[0139] Another example (e.g., example 16) relates to a previous example (e.g., one of the examples 11 to 15) or to any other example, further comprising that the circuitry is further configured to generate a voice output based on the reconstructed data.

[0140] Another example (e.g., example 17) relates to a previous example (e.g., one of the examples 11 to 16) or to any other example, further comprising that the circuitry is further configured to decode the compressed encoded data by the decoder, the decoder comprising a video gamespecific codebook, wherein the compressed encoded data is represented as data based on the video game-specific codebook.

[0141] Another example (e.g., example 18) relates to a previous example (e.g., example 17) or to any other example, further comprising that the decoder comprises a plurality of game-specific dictionaries each being adapted to a respective video game.

[0142] Another example (e.g., example 19) relates to a previous example (e.g., one of the examples 11 to 18) or to any other example, further comprising that the circuitry is further configured to identify the video game, and load the game-specific codebook for the identified video game among a plurality of game-specific codebooks each being adapted to a respective video game.

[0143] An example (e.g., example 20) relates to a system comprising a first apparatus comprising circuitry configured to obtain natural language data comprising content relating to a video game, obtain state information data relating to a current state of the video game, determine, by an encoder, compressed encoded input data based on the content relating to the video game of the natural language data and on the current state information data, wherein the encoder is adapted for the video game, transmit the compressed encoded input data to a second apparatus, the second apparatus comprising circuitry configured to obtain the compressed encoded input data, obtain the current state information data, determine, by a decoder, reconstructed data based on the encoded input data and on the state information data, wherein the decoder is adapted for the video game.

[0144] An example (e.g., example 21) relates to an apparatus comprising circuitry configured to train a first artificial neural network, ANN, of an encoder and a second ANN of a decoder based on training data comprising natural language data, the natural language data comprising content relating to a video game, wherein the first ANN is trained to encode natural language data comprising content relating to the video game and the second ANN is trained to reconstruct the content of the natural language data.

[0145] Another example (e.g., example 22) relates to a previous example (e.g., example 21) or to any other example, further comprising that the circuitry is further configured to adapt, while training the first ANN and the second ANN, a game-specific codebook.

[0146] Another example (e.g., example 23) relates to a previous example (e.g., one of the examples 21 to 22) or to any other example, further comprising that the training is based on training data pairs with similar content and in different languages.

[0147] An example (e.g., example 24) relates to a method comprising obtaining natural language data comprising content relating to a video game, obtaining state information data relating to a current state of the video game, determining, by an encoder, compressed encoded input data based on the natural language data and on the state information data, wherein the encoder is adapted for the video game.

[0148] An example (e.g., example 25) relates to a method comprising obtaining compressed encoded data, the compressed encoded data is encoding natural language data, the natural language data comprising content relating to a video game, obtaining state information data relating to a state of the video game, determining, by a decoder, reconstructed data based on the compressed encoded data and on the state information data, wherein the decoder is adapted for the video game.

[0149] Another example (e.g., example 26) relates to a non-transitory machine-readable medium having stored thereon a program having a program code for performing any one of the methods according to examples 24 or 25, when the program is executed on a processor or a programmable hardware.

[0150] Another example (e.g., example 27) relates to a program having a program code for performing any one of the methods according to examples 24 or 25, when the program is executed on a processor or a programmable hardware.

[0151] Examples may further be or relate to a (computer) program including a program code to execute one or more of the above methods when the program is executed on a computer, processor or other programmable hardware component. Thus, steps, operations or processes of different ones of the methods described above may also be executed by programmed computers, processors or other programmable hardware components. Examples may also cover program storage devices, such as digital data storage media, which are machine-, processor- or computer- readable and encode and / or contain machine-executable, processor-executable or computerexecutable programs and instructions. Program storage devices may include or be digital storage devices, magnetic storage media such as magnetic disks and magnetic tapes, hard disk drives, or optically readable digital data storage media, for example. Other examples may also include computers, processors, control units, (field) programmable logic arrays ((F)PLAs), (field) programmable gate arrays ((F)PGAs), graphics processor units (GPU), application-specific integrated circuits (ASICs), integrated circuits (ICs) or system-on-a-chip (SoCs) systems programmed to execute the steps of the methods described above. It is further understood that the disclosure of several steps, processes, operations or functions disclosed in the description or claims shall not be construed to imply that these operations are necessarily dependent on the order described, unless explicitly stated in the individual case or necessary for technical reasons. Therefore, the previous description does not limit the execution of several steps or functions to a certain order. Furthermore, in further examples, a single step, function, process or operation may include and / or be broken up into several sub-steps, - functions, -processes or -operations.

[0152] If some aspects have been described in relation to a device or system, these aspects should also be understood as a description of the corresponding method. For example, a block, device or functional aspect of the device or system may correspond to a feature, such as a method step, of the corresponding method. Accordingly, aspects described in relation to a method shall also be understood as a description of a corresponding block, a corresponding element, a property or a functional feature of a corresponding device or a corresponding system.

[0153] The following claims are hereby incorporated in the detailed description, wherein each claim may stand on its own as a separate example. It should also be noted that although in the claims a dependent claim refers to a particular combination with one or more other claims, other examples may also include a combination of the dependent claim with the subject matter of any other dependent or independent claim. Such combinations are hereby explicitly proposed, unless it is stated in the individual case that a particular combination is not intended. Furthermore, features of a claim should also be included for any other independent claim, even if that claim is not directly defined as dependent on that other independent claim.

Claims

ClaimsWhat is claimed is:

1. An apparatus comprising circuitry configured to: obtain natural language data comprising content relating to a video game; obtain state information data relating to a current state of the video game; determine, by an encoder, compressed encoded input data based on the natural language data and on the state information data, wherein the encoder is adapted for the video game.

2. The apparatus of claim 1, wherein the circuitry is further configured to transmit the compressed encoded input data to a decoder.

3. The apparatus of claim 1, wherein the circuitry is further configured to identify the video game; and load the encoder adapted for the identified video game among a plurality of encoders adapted for a respective plurality of the video games.

4. The apparatus of claim 1, wherein the circuitry is further configured to determine the state information data based on current values of variables representing the video game.

5. The apparatus of claim 1, wherein the encoder comprises an artificial neural network, ANN, wherein the ANN is trained based on training data related to the video game.

6. The apparatus of claim 1, wherein the circuitry is further configured to determine the compressed encoded input data by the encoder, the encoder comprising a video game-specific codebook, wherein the compressed encoded input data is represented as data based on the video game-specific codebook.

7. The apparatus of claim 6, wherein the compressed encoded data is represented by one or more quantized vector based on a video game-specific codebook.

8. The apparatus of claim 6, wherein the circuitry is further configured to identify the video game; andload the game-specific codebook for the identified video game among a plurality of gamespecific codebooks each being adapted to a respective video game.

9. The apparatus of claim 1, wherein the natural language data is a text input and / or a voice input.

10. The apparatus of claim 1, wherein the natural language data is a voice input and audio features of the voice input are discarded in the encoded signal.

11. An apparatus comprising circuitry configured to: obtain compressed encoded data, the compressed encoded data is encoding natural language data, the natural language data comprising content relating to a video game; obtain state information data relating to a state of the video game; determine, by a decoder, reconstructed data based on the compressed encoded data and on the state information data, wherein the decoder is adapted for the video game.

12. The apparatus of claim 11, wherein the reconstructed data comprises the content relating to the video game of the natural language data.

13. The apparatus of claim 11, wherein the circuitry is further configured to determine, by the decoder, extended reconstructed data based on the compressed encoded data and on the state information data, wherein the extended reconstructed data comprises content extending the content relating to the video game of the natural language data based on the state information data.

14. The apparatus of claim 11, wherein the reconstructed data comprises the content relating to the video game of the natural language data in a language different than the language of the natural language data.

15. The apparatus of claim 11, wherein the circuitry is further configured to generate a voice output based on the reconstructed data.

16. The apparatus of claim 11, wherein the circuitry is further configured to decode the compressed encoded data by the decoder, the decoder comprising a video game-specificcodebook, wherein the compressed encoded data is represented as data based on the video game-specific codebook.

17. The apparatus of claim 16, wherein the decoder comprises a plurality of game-specific dictionaries each being adapted to a respective video game.

18. A system comprising: a first apparatus comprising circuitry configured to obtain natural language data comprising content relating to a video game; obtain state information data relating to a current state of the video game; determine, by an encoder, compressed encoded input data based on the content relating to the video game of the natural language data and on the current state information data, wherein the encoder is adapted for the video game; transmit the compressed encoded input data to a second apparatus; the second apparatus comprising circuitry configured to: obtain the compressed encoded input data; obtain the current state information data; determine, by a decoder, reconstructed data based on the encoded input data and on the state information data, wherein the decoder is adapted for the video game.

19. A method comprising: obtaining natural language data comprising content relating to a video game; obtaining state information data relating to a current state of the video game; determining, by an encoder, compressed encoded input data based on the natural language data and on the state information data, wherein the encoder is adapted for the video game.

20. A method comprising:obtaining compressed encoded data, the compressed encoded data is encoding natural language data, the natural language data comprising content relating to a video game; obtaining state information data relating to a state of the video game; determining, by a decoder, reconstructed data based on the compressed encoded data and on the state information data, wherein the decoder is adapted for the video game.

Citation Information

Patent Citations

  • Speaker conversion for video games

    US11605388B1

  • Generating speech in the voice of a player of a video game

    US11790884B1

  • Compressing audio waveforms using neural networks and vector quantizers

    US20230186927A1