Voice chat translation
The method and system address the loss of context and emotion in voice chat translation by converting speech to text, translating with context and emotion data, and generating output speech, enhancing user experience in virtual environments.
Patent Information
- Application Number
- JP2024572201
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-08
- Filing Date
- 2023-06-07
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2043-06-07
AI Technical Summary
Existing voice chat translation systems in virtual environments fail to preserve the context and emotions of user speech, relying on text-based translations that lose the nuances of communication.
A method and system that converts user speech to text, translates it into a different language while retaining context and emotion data, and generates output speech using a speech waveform modulator to maintain the original characteristics.
Enables accurate understanding of context and emotion across language barriers, providing a more immersive and enjoyable experience in virtual environments.
Smart Images

Figure 2025523410000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application is an international application that claims the benefit of priority under 35 U.S.C. § 119(e) of U.S. Provisional Application No. 63 / 350,154, filed on June 8, 2022, entitled "VOICE CHAT TRANSLATION", the entire content of which is incorporated herein by reference.
[0002] Embodiments generally relate to audio input and output via a computer device, and more particularly, to methods, systems, and computer - readable media for providing voice chat translation that preserves the characteristics, context, and emotions of a user's voice in a virtual environment such as a metaverse place in a virtual metaverse.
Background Art
[0003] Computer audio (e.g., chat between users of a computer device) often consists of providing monaural or stereo audio when it is received from a listening device or microphone. When the audio should be translated for various users speaking different languages, most solutions rely on text - based translation that provides only a simple function including a word - for - word or phrase - by - phrase translation presented in text. Thus, much of the context and / or emotion associated with a user's directed chat may be lost in translation.
[0004] The description of the background art provided herein is for the purpose of presenting the context of the present disclosure. The scope described in this background art section, the inventors' research named herein, and aspects of the description that may not be considered prior art at the time of filing are not admitted as prior art to the present disclosure, either expressly or implicitly.
Summary of the Invention
Means for Solving the Problem
[0005] The implementation of this application relates to providing voice chat translation in a virtual metaverse.
[0006] According to one aspect, a method implemented by a computer includes receiving a request to translate audio related to a chat function of a metaverse place in a virtual metaverse, the audio being received from a first user among a plurality of users, the plurality of users being associated with the metaverse place; retrieving translation data related to a second user among the plurality of users, the translation data including at least a language preference related to the second user, the second user being associated with a user device; converting the audio received from the first user into text, the audio including input speech in a first language spoken by the first user; translating the text into a second language, the second language being defined by the language preference, the translated text including context data from the audio; converting the translated text into output speech including the context data; and providing the output speech to the user device.
[0007] In some implementations, the request identifies the first user and the second user, and the request originates from a computing device associated with the first user.
[0008] In some implementations, the method further includes retrieving a voice output preference related to the first user, the voice output preference at least partially overriding the translation data related to the second user.
[0009] In some implementations, the method further includes moderating the text to remove words based on a text moderation filter before translating the text into a second language.
[0010] In some implementations, the translating step includes providing the text as an input to a trained machine learning model and receiving the translated text as an output from the trained machine learning model.
[0011] In some implementations, the context data includes sentiment data extracted from audio.
[0012] In some implementations, the method further includes preprocessing the audio to extract sentiment data.
[0013] In some implementations, the step of converting the translated text to output speech includes using a speech waveform modulator to create a moderated speech waveform that at least partially includes the context data.
[0014] In some implementations, the method further includes translating the text into a plurality of different languages to create a plurality of different translated texts and converting the plurality of different translated texts to a plurality of different output speeches.
[0015] In some implementations, the method further includes providing the plurality of different output speeches to a plurality of user devices, wherein each output speech includes utterances in a language associated with each respective user device of the plurality of user devices.
[0016] According to another example, in response to execution by a processing device, the processing device receives a request to translate audio related to a chat function of a metaverse place in a virtual metaverse, the audio being received from a first user among a plurality of users, the plurality of users being associated with the metaverse place, retrieving translation data related to a second user among the plurality of users, the translation data including at least a language preference related to the second user, the second user being associated with a user device, converting the audio received from the first user into text, the audio including input speech in a first language spoken by the first user, translating the text into a second language, the second language being defined by the language preference, the translated text including context data from the audio, translating the translated text into output speech including the context data, and providing the output speech to the user device. A non-transitory computer-readable medium storing instructions for causing the above-described operations to be performed is disclosed.
[0017] In some implementations, the request identifies the first user and the second user, and the request originates from a computing device associated with the first user.
[0018] In some implementations, the operations further include retrieving a voice output preference related to the first user, the voice output preference at least partially overriding the translation data related to the second user.
[0019] In some implementations, the operations further include moderating the text to remove words based on a text moderation filter before translating the text into the second language.
[0020] In some implementations, translating includes providing text as input to a trained machine learning model and receiving translated text as output from the trained machine learning model.
[0021] In some implementations, the context data includes sentiment data extracted from audio.
[0022] In some implementations, the operation further includes preprocessing the audio to extract sentiment data.
[0023] In some implementations, converting the translated text to output speech includes using a speech waveform modulator to create an adjusted speech waveform that at least partially includes the context data.
[0024] In some implementations, the operation further includes translating the text into a plurality of different languages to create a plurality of different translated texts, converting the plurality of different translated texts into a plurality of different output speeches, and providing the plurality of different output speeches to a plurality of user devices, wherein each output speech includes utterances in a language associated with each respective user device of the plurality of user devices.
[0025] According to yet another aspect, there is provided a system including a memory storing instructions and a processing device coupled to the memory and operable to access the memory, wherein when the instructions are executed by the processing device, the processing device is caused to receive a request to translate audio related to a chat function of a metaverse place in a virtual metaverse, the audio being received from a first user among a plurality of users, the plurality of users being associated with the metaverse place, retrieve translation data related to a second user among the plurality of users, the translation data including at least a language preference related to the second user, the second user being associated with a user device, convert the audio received from the first user into text, the audio including input speech in a first language spoken by the first user, translate the text into a second language, the second language being defined by the language preference, the translated text including context data from the audio, convert the translated text into output speech including the context data, and provide the output speech to the user device.
[0026] According to another aspect, details of some, features, and implementations of the systems, devices, methods, and non-transitory computer-readable media disclosed herein may be combined to form additional aspects including omitting and / or modifying some or a portion of individual components or features, and including additional components or features, and / or other modifications, and all such modifications are within the scope of the present disclosure.
Brief Description of the Drawings
[0027]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 4C
Figure 4D
Figure 5
Figure 6
Figure 7
Figure 8
Mode for Carrying Out the Invention
[0028] One or more implementations described herein relate to voice chat translation related to an online virtual experience platform. Features may include automatically converting speech to text while retaining context and / or emotion data in a metaverse place of a virtual metaverse, translating the text into a different language, and automatically generating speech from the translated text using the context and / or emotion data. The generated speech retains at least a portion of the context and / or emotion from the source speech.
[0029] The features described herein provide automatic translation of audio for output in a client device coupled to an online platform, such as an online virtual experience platform or an online game platform. The online platform may provide a virtual metaverse having a plurality of metaverse places associated therewith. Virtual avatars associated with users can move through, participate in, and interact with items, characters, other avatars, and objects within the metaverse place. The avatar can move from one metaverse place to another while engaging in a voice chat that provides an immersive and enjoyable experience by enabling communication with users who speak different languages. Different audio streams from multiple users (avatars associated with multiple users) can be translated based on language preferences defined by the user and provided to other users.
[0030] Through automatic translation that retains context and / or emotion, users can accurately understand context and / or emotion through chat despite language barriers. This may provide a more immersive and enjoyable experience for users of the virtual experience platform.
[0031] Online virtual experience platforms and online game platforms (also referred to as "user-generated content platforms" or "user-generated content systems") provide various ways for users to interact with each other. For example, users of an online virtual experience platform may create games or other content or resources (such as characters, graphics, items for use in gameplay and / or within a virtual metaverse, etc.) within the online platform.
[0032] Users of an online virtual experience platform may engage in activities such as sharing various virtual items (such as inventory items, game items, etc.) that cooperate towards a common goal in a metaverse place, game, or game creation, participating in audio chats (such as audio chats with automatic translation), sending electronic messages to each other. Users of an online virtual experience platform may interact with other users, for example, by playing games that include characters (avatars) or other game objects and mechanisms. An online virtual experience platform may also enable users of the platform to communicate with each other. For example, users of an online virtual experience platform may communicate with each other using voice messages or live voice interactions (such as via voice chat with automatic translation), text messaging, video messaging (such as including audio translation), or a combination of the above. Some online virtual experience platforms can provide a virtual three-dimensional environment or multiple linked environments within a metaverse, within which users can interact with each other or play online games.
[0033] To help enhance the entertainment value of an online virtual experience platform, the platform can provide rich audio for playback on user devices. The audio can include, for example, different audio streams from different users and background audio. According to various implementations described herein, different audio streams can be captured and automatically translated based on the user listening. For example, a first user may request to engage in a voice chat with a second user with automatic translation. Thereafter, the audio stream from the first user may be translated before being provided to the second user. Additionally, the audio stream may be provided to other users with or without automatic translation, based on, for example, user settings, language settings, override settings, and / or other settings.
[0034] Figures 1-3: Exemplary System Architectures Figure 1 shows an exemplary network environment 100 according to some implementations of the present disclosure. The network environment 100 (also referred to herein as a "system") includes an online virtual experience platform 102, a first client device 110, and a second client device 116 (collectively referred to herein as "client devices 110 / 116"), all of which are connected via a network 122. The online virtual experience platform 102 can include, among other things, a virtual experience engine 104, one or more virtual experiences 105, a voice chat translation component 106, and a data store 108.
[0035] Client device 110 can include a virtual experience application 112, and client device 116 can include a virtual experience application 118. Users 114 and 120 can each use client devices 110 and 116 to interact with online virtual experience platform 102 and other users who utilize online virtual experience platform 102.
[0036] Network environment 100 is provided for illustration purposes. In some implementations, network environment 100 may include the same, fewer, more, or different elements configured in the same or different manner as shown in FIG. 1.
[0037] In some implementations, network 122 may include a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or wide area network (WAN)), a wired network (e.g., an Ethernet network), a wireless network (e.g., an 802.11 network, a Wi-Fi® network, or a wireless local area network (WLAN)), a cellular network (e.g., a long term evolution (LTE) network), routers, hubs, switches, server computers, or combinations thereof.
[0038] In some implementations, data store 108 may be a non-transitory computer-readable memory (e.g., random access memory), a cache, a drive (e.g., a hard drive), a flash drive, a database system, or another type of component or device capable of storing data. Also, data store 108 may include multiple storage components (e.g., multiple drives or multiple databases) that may span multiple computing devices (e.g., multiple server computers).
[0039] In some implementations, the online virtual experience platform 102 may include a server having one or more computing devices (e.g., a cloud computing system, a rack-mounted server, a server computer, a cluster of physical servers, a virtual server, etc.). In some implementations, the server may be included in the online virtual experience platform 102, be an independent system, or be part of another system or platform.
[0040] In some implementations, the online virtual experience platform 102 may include one or more computing devices (such as a rack-mounted server, a router computer, a server computer, a personal computer, a mainframe computer, a laptop computer, a tablet computer, a desktop computer, etc.), a data store (e.g., a hard disk, memory, database), a network, software components, and / or hardware components that execute operations on the online virtual experience platform 102 and may be used to provide users with access to the online virtual experience platform 102. Also, the online virtual experience platform 102 may include a website (e.g., one or more web pages) or application backend software that may be used to provide users with access to the content provided by the online virtual experience platform 102. For example, users 114 / 120 may access the online virtual experience platform 102 using virtual experience applications 112 / 118 on client devices 110 / 116, respectively.
[0041] In some implementations, the online virtual experience platform 102 may include a social network that provides connections between users, or a user-generated content system that enables users (e.g., end users or consumers) to communicate with other users via the online virtual experience platform 102. The communication may include voice chat (e.g., synchronous and / or asynchronous voice communication with or without automatic translation), video chat (e.g., synchronous and / or asynchronous video communication with or without automatic audio translation), or text chat (e.g., synchronous and / or asynchronous text-based communication with or without automatic text translation).
[0042] In some implementations of the present disclosure, a "user" may be represented as an individual person. However, other implementations of the present disclosure include the case where a "user" (e.g., a creating user) is an entity controlled by a set of users or an automated source. For example, a set of individual users associated as a community or group within a user-generated content system may be considered a "user".
[0043] In some implementations, the online virtual experience platform 102 may be a virtual game platform. For example, the game platform may provide single-player or multiplayer games to a community of users who may access or interact with games (e.g., user-generated games or other games) using the client devices 110 / 116 via the network 122. In some implementations, the games (also referred to herein as "video games", "online games", "metaverse places", or "virtual experiences") may be, for example, two-dimensional (2D) games, three-dimensional (3D) games (e.g., 3D user-generated games), virtual reality (VR) games, or augmented reality (AR) games. In some implementations, users may search for games and game items and participate in gameplay with other users in one or more games. In some implementations, the games may be played in real time with other users of the game. Similarly, some users may engage in real-time voice or video chat with other users of the game. As described herein, the real-time voice or video chat may include automatic translation.
[0044] In some implementations, instead of or in addition to the online virtual experience platform 102 and / or the voice chat translation component 106, other collaboration platforms may be used with the features described herein. For example, social networking platforms, purchasing platforms, messaging platforms, creation platforms, etc. may be used with the automatic translation feature such that translated audio is provided to users outside of games and / or virtual experiences.
[0045] In some implementations, gameplay may refer to the interaction of one or more players using client devices (e.g., 110 and / or 116) within a game (e.g., virtual experience 105), or the representation of interactions on the display of client device 110 or 116 or other output devices. In some implementations, gameplay may instead refer to interactions within a virtual experience or metaverse place, which may include different, or the same purposes, and may not be similar to some games. Further, although referred to as "players", the terms "avatar", "user", and / or other terms may be used to refer to users who participate in and / or interact with an online virtual experience.
[0046] One or more virtual experiences 105 are provided by an online virtual experience platform. In some implementations, virtual experience 105 may include an electronic file that can be executed or loaded using software, firmware, or hardware configured to present virtual content (e.g., digital media items) to entities. In some implementations, virtual experience applications 112 / 118 may be executed in relation to virtual experience engine 104, and virtual experience 105 may be rendered in relation to virtual experience engine 104. In some implementations, virtual experience 105 may have a common set of rules or common goals, and the virtual environments of virtual experience 105 share a common set of rules or common goals. In some implementations, different virtual experiences may have different rules or goals from each other.
[0047] In some implementations, a game and / or virtual experience may have one or more environments (also referred to herein as "game environments", "metaverse places", or "virtual environments"), and the multiple environments may be linked. An example of an environment may be a three-dimensional (3D) environment. The virtual experience 105 or one or more environments of the virtual experience may be collectively referred to herein as the "world", "game world", "virtual world", "universe", or "metaverse". An example of a world may be the 3D metaverse place of the virtual experience 105. For example, a user may construct a metaverse place that is linked to another metaverse place created by a different user different from the first user. A character of the virtual experience may cross a virtual boundary to enter an adjacent metaverse place. Further, sounds, theme music, and / or background music may also cross the virtual boundary so that an avatar standing near the virtual boundary may hear audio including at least a portion of the sound emitted from the adjacent metaverse place.
[0048] It may be noted that 3D environments or 3D worlds use graphics that use a three-dimensional representation of geometric data to represent content (or at least present the content so that it appears as 3D content regardless of whether a 3D representation of the geometric data is used). 2D environments or 2D worlds use graphics that use a two-dimensional representation of geometric data to represent game content.
[0049] In some implementations, the online virtual experience platform 102 can host one or more virtual experiences 105 and allow users to interact with the virtual experiences 105 (e.g., experience, play games, game-related content, virtual content, or search for other content) using the virtual experience applications 112 / 118 of the client devices 110 / 116. Users (e.g., 114 and / or 120) of the online virtual experience platform 102 may play, create, interact with, or build the virtual experiences 105, search for the virtual experiences 105, communicate with other users, create, build, and / or search for objects (e.g., also referred to herein as "items" or "game objects" or "virtual game items") of the virtual experiences 105. For example, when generating user-generated virtual items, the user may, among other things, create characters, decorations for the characters, one or more virtual environments for interactive games, or build structures used within the virtual experience 105.
[0050] In some implementations, a user may buy, sell, or trade virtual game objects, such as in-platform currency (e.g., virtual currency), with other users of the online virtual experience platform 102. In some implementations, the online virtual experience platform 102 may send game content to a game application (e.g., the virtual experience application is 112). In some implementations, game content (also referred to herein as "content") may refer to any data or software instructions related to the online virtual experience platform 102 or the game application (e.g., game objects, games, user information, videos, images, commands, media items, etc.).
[0051] In some implementations, a game object (e.g., also referred to herein as an "item" or an "object" or a "virtual game item") may refer to an object that is used, created, shared, or otherwise depicted in the virtual experience 105 of the online virtual experience platform 102 or in the virtual experience applications 112 or 118 of the client devices 110 / 116. For example, game objects may include parts, models, characters, tools, weapons, clothing, buildings, vehicles, currency, flora, fauna, components of the foregoing (e.g., windows of a building), and the like.
[0052] It should be noted that the online virtual experience platform 102 hosting the virtual experience 105 is provided for purposes of illustration and not limitation. In some implementations, the online virtual experience platform 102 may host one or more media items that may include communication messages from one user to one or more other users. Media items may include, but are not limited to, digital videos, digital movies, digital photos, digital music, audio content, melodies, website content, social media updates, e-books, e-magazines, digital newspapers, digital audiobooks, e-journals, web blogs, real simple syndication (RSS) feeds, digital comics, software applications, and the like. In some implementations, a media item may be an electronic file that can be executed or loaded using software, firmware, or hardware configured to present the digital media item to an entity.
[0053] In some implementations, the virtual experience 105 may be associated with a particular user or a particular group of users (e.g., a private game), or may be made widely available to users of the online virtual experience platform 102 (e.g., a published game). In some implementations where the online virtual experience platform 102 associates one or more virtual experiences 105 with a particular user or group of users, the online virtual experience platform 102 may use user account information (e.g., user account identifiers such as a username and password) to associate a particular user with the virtual experience 105. Similarly, in some implementations, the online virtual experience platform 102 may use developer account information (e.g., developer account identifiers such as a username and password) to associate a particular developer or group of developers with the virtual experience 105.
[0054] In some implementations, the online virtual experience platform 102 or the client devices 110 / 116 may include a virtual experience engine 104 or virtual experience applications 112 / 118. The virtual experience engine 104 may include virtual experience applications similar to the virtual experience applications 112 / 118. In some implementations, the virtual experience engine 104 may be used for the deployment or execution of the virtual experience 105. For example, the virtual experience engine 104 may include, among other features, a rendering engine (a "renderer") for 2D, 3D, VR, or AR graphics, a physics engine, a collision detection engine (and collision response), a sound engine, a machine learning model, a translation component, a spatialized audio manager / engine, an audio mixer, an audio subscription exchange, an audio subscription logic, an audio subscription prioritizer, a real-time communication engine, a scripting function, an animation engine, an artificial intelligence engine, a networking function, a streaming function, a memory management function, a threading function, a scene graph function, or video support for cinematics. The components of the virtual experience engine 104 may generate commands (such as rendering commands, collision commands, physics commands, etc.) that assist in calculating and rendering the virtual experience and may translate audio (such as converting audio to text, translating text, converting the translated text to speech, etc.). In some implementations, the virtual experience applications 112 / 118 of the client devices 110 / 116 may each independently cooperate with the virtual experience engine 104 of the online virtual experience platform 102 or work in combination with both of them.
[0055] In some implementations, both the online virtual experience platform 102 and the client devices 110 / 116 execute a virtual experience engine (104, 112, and 118 respectively). The online virtual experience platform 102 that uses the virtual experience engine 104 may execute some or all of the functions of the virtual experience engine (e.g., generate physical commands, rendering commands, spatial audio commands, etc.), or may offload some or all of the functions of the virtual experience engine to the virtual experience engine 104 of the client device 110. In some implementations, each virtual experience 105 may have a different ratio between the functions of the virtual experience engine executed on the online virtual experience platform 102 and the functions of the virtual experience engine executed on the client devices 110 and 116.
[0056] For example, the virtual experience engine 104 of the online virtual experience platform 102 may be used to generate physical commands when there is a collision between at least two virtual objects, while additional functions of the virtual experience engine (e.g., generate rendering commands, or combine spatial audio streams) may be offloaded to the client device 110. In some implementations, the ratio between the functions of the virtual experience engine executed on the online virtual experience platform 102 and the functions of the virtual experience engine executed on the client device 110 may be changed (e.g., dynamically) based on the conditions of the gameplay. For example, if the number of users participating in the gameplay of the virtual experience 105 exceeds a threshold number, the online virtual experience platform 102 may execute one or more functions of the virtual experience engine that were previously executed by the client device 110 or 116.
[0057] For example, a user may interact with virtual experience 105 on client devices 110 and 116 and may send control instructions (such as user inputs like right, left, up, down, user selections, or character position and speed information, etc.) to online virtual experience platform 102. After receiving the control instructions from client devices 110 and 116, online virtual experience platform 102 may send game play instructions (such as position and speed information of characters participating in group game play, or commands such as rendering commands, collision commands, spatial audio commands, etc.) to client devices 110 and 116 based on the control instructions. For example, online virtual experience platform 102 may perform one or more logical operations (such as using virtual experience engine 104) according to the control instructions to generate game play instructions for client devices 110 and 116. In other cases, online virtual experience platform 102 may pass one or more of the control instructions from one client device 110 to other client devices (such as 116) participating in virtual experience 105. Client devices 110 and 116 may use the game play instructions and render the game play for presentation on the displays of client devices 110 and 116.
[0058] In some implementations, the control instruction may refer to an instruction indicating an in - character action of the user. For example, the control instruction may include user input, user selection, gyroscope position and orientation data, force sensor data, etc. for controlling in - character actions such as right, left, up, down. The control instruction may include character position and velocity information. In some implementations, the control instruction is directly sent to the online virtual experience platform 102. In other implementations, the control instruction may be sent from the client device 110 to another client device (e.g., 116), and the other client device generates game play instructions using the local virtual experience engine 104. The control instruction may include an instruction to play a voice communication message or other sound from another user on an audio device (e.g., speaker, headphones, etc.).
[0059] In some implementations, the game play instruction may refer to an instruction that enables the client device 110 (or 116) to render the game play of a virtual experience such as a multiplayer game or a virtual experience. The game play instruction may include one or more of user input (e.g., control instruction), character position and velocity information, or commands (e.g., physical commands, rendering commands, collision commands, etc.).
[0060] In some implementations, a character (or more generally, a virtual object) is constructed from components that automatically bind together to assist a user in editing, and one or more of the components may be selected by the user. One or more characters (also referred to herein as "avatars" or "models") may be associated with a user, and the user may control the character to facilitate the user's interaction with game 105. In some implementations, a character may include components such as body parts (e.g., head, hair, arms, legs, etc.) and accessories (e.g., T-shirts, glasses, decorative images, tools, etc.). In some implementations, customizable character body parts may include, among other things, head types, body part types (arms, legs, torso, and hands), face types, hair types, and skin types. In some implementations, customizable accessories may include clothing (e.g., shirts, pants, hats, shoes, glasses, etc.), weapons, or other tools.
[0061] In some implementations, the user may also control the scale of the character (e.g., height, width, or depth) or the scale of the components of the character. In some implementations, the user may control the proportions of the character (e.g., blocky, anatomical, etc.). It should be noted that in some implementations, a character may not include character game objects (e.g., body parts), but the user may control the character (without character game objects) to facilitate the user's interaction with the virtual experience (e.g., a puzzle game where there are no rendered character game objects, but the user still controls the character to control in-game actions).
[0062] In some implementations, components such as body parts may be in basic geometric shapes such as blocks, cylinders, spheres, or some other basic shapes such as wedges, rings, tubes, channels. In some implementations, the creator module may publish the user's character for viewing or use by other users of the online virtual experience platform 102. In some implementations, creating, modifying, or customizing a character, other virtual object, virtual experience 105, or virtual environment may be performed by the user using a user interface (e.g., a developer interface), either by scripting or without scripting (or by an application programming interface (API) or without an API). It should be noted that, for purposes of illustration and not limitation, the character is described as having the form of a humanoid robot. It should further be noted that the character may have any form such as a vehicle, an animal, an inanimate object, or some other creative form.
[0063] In some implementations, the online virtual experience platform 102 may store characters created by a user in the data store 108. In some implementations, the online virtual experience platform 102 may hold a character catalog and a virtual experience catalog that may be presented to the user via the virtual experience engine 104, the virtual experience 105, and / or the client devices 110 / 116. In some implementations, the virtual experience catalog includes images of virtual experiences stored in the online virtual experience platform 102. Additionally, the user may select a character (e.g., a character created by the user or another user) from the character catalog to participate in a selected experience. The character catalog includes images of characters stored in the online virtual experience platform 102. In some implementations, one or more of the characters in the character catalog may have been created or customized by the user. In some implementations, the selected character may have a character setting that defines one or more of the components of the character.
[0064] In some implementations, a user's character can include a component configuration, and the component configuration, appearance, and more broadly the character's appearance may be defined by a character setting. In some implementations, the character setting of a user's character may be at least partially selected by the user. In other implementations, the user may select a character having a default character setting or other character setting selected by another user. For example, the user may select a default character having a predefined character setting from a character catalog, and further, the user may customize the default character by changing a part of the character setting (for example, adding a customized shirt with a logo). The character setting may be associated with a specific character by the online virtual experience platform 102.
[0065] In some implementations, the client device 110 or 116 may each include a computing device such as a personal computer (PC), a mobile device (for example, a laptop, a mobile phone, a smartphone, a tablet computer, or a netbook computer), a network-connected television, a game console, etc. In some implementations, the client device 110 or 116 may sometimes be referred to as a "user device". In some implementations, one or more client devices 110 or 116 may be connected to the online virtual experience platform 102 at any time. It should be noted that the number of client devices 110 or 116 is given by way of example and not limitation. In some implementations, any number of client devices 110 or 116 may be used.
[0066] In some implementations, each client device 110 or 116 may respectively include an instance of a virtual experience application 112 or 118. In one implementation, the virtual experience application 112 or 118 may search for virtual experiences, virtual items, or other content, control virtual characters of virtual experiences hosted by the online virtual experience platform 102, or enable a user to use and interact with the online virtual experience platform 102, such as viewing or uploading content such as virtual experiences 105, images, video items, web pages, documents, etc. In one example, the virtual experience application may be a web application (e.g., an application that operates in conjunction with a web browser) that can access, retrieve, present, or navigate content provided by a web server (such as virtual characters in a virtual environment). In another example, the virtual experience application may be a native application (e.g., a mobile application, app, or game program) installed and executed locally on the client device 110 or 116 that enables a user to interact with the online virtual experience platform 102. The virtual experience application may sometimes render, display, or present content to the user (such as web pages, user interfaces, media viewers, audio streams). In an implementation, the virtual experience application may also include an embedded media player embedded in a web page.
[0067] According to aspects of the present disclosure, the virtual experience application 112 / 118 may be an online virtual experience platform application for a user to build, create, edit content, upload it to the online virtual experience platform 102, and interact with the online virtual experience platform 102 (e.g., play a virtual experience 105 hosted by the online virtual experience platform 102). Thus, the virtual experience application 112 / 118 may be provided to the client device 110 or 116 by the online virtual experience platform 102. In another example, the virtual experience application 112 / 118 may be an application downloaded from a server.
[0068] In some implementations, a user may log in to the online virtual experience platform 102 via the virtual experience application. The user may access the user account by providing user account information (e.g., username and password), and the user account is associated with one or more characters available to participate in one or more virtual experiences 105 of the online virtual experience platform 102.
[0069] Generally, the functions described as being performed by the online virtual experience platform 102 can also be performed by the client device 110 or 116 or the server as appropriate in other implementations. Additionally, the functions attributable to a particular component can be performed by different or multiple components operating together. Also, the online virtual experience platform 102 can be accessed as a service provided to other systems or devices through an appropriate application programming interface (API), and thus is not limited to use on a website.
[0070] In some implementations, the online virtual experience platform 102 may include a voice chat translation component 106.
[0071] In some implementations, the voice chat translation component 106 may include an application programming interface (API) that includes a set of computer-executable code that provides functionality to users and / or developers in the form of function calls that enable software components to communicate and / or provide / receive data. The API includes a plurality of defined software functions related to voice chat translation, and those software functions can be used by developers to enable audio translation functions for voice chat and video chat, and may include any functions related to audio playback on user devices.
[0072] In some implementations, the voice chat translation component 106 is a software component that provides an automatic voice chat translation function based on user settings. For example, in some implementations, the voice chat translation component may include one or more machine learning models, one or more text translation components, one or more audio conversion components, one or more text-to-speech components, one or more plugins for communicating with multiple third-party services, and / or any other suitable components. FIGS. 3 and 4 show different sub-components that may be included as part of the voice chat translation component 106 in some implementations.
[0073] The operation of the online virtual experience platform 102 regarding providing automatic voice chat translation will be more fully described with reference to FIG. 2 hereinafter.
[0074] FIG. 2 is a diagram of an exemplary network environment 200 (e.g., a subset of network environment 100) for providing automatic voice chat translation in a virtual metaverse according to some implementations. Network environment 200 is provided for illustrative purposes. In some implementations, network environment 200 may include the same, fewer, more, or different elements configured in the same or different ways as shown in FIG. 2.
[0075] As shown in FIG. 2, online virtual experience platform 102 may communicate with client device 110 and client device 116 via network 122 such that user audio stream 232 is received from client device 110 and translated audio stream 234 is provided for output at client device 116. Online virtual experience platform 102 may also communicate with communication server 202 and relay server 210 via network 122.
[0076] In addition to the components shown in FIG. 1, online virtual experience platform 102 may include a voice chat plugin 208 for communicating with communication server 202. Voice chat plugin 208 may perform separation of the audio stream and / or identification of the audio stream to be translated by voice chat translation component 106. In this way, audio stream 232 may be sent in its native form to any client device, but voice chat plugin 208 may indicate to media server 204 that a translated version of audio stream 232 should be provided to other client devices. Thus, voice chat plugin 208 may enable both native and translated communications to occur at substantially the same or similar times.
[0077] The communication server 202 may be a third-party communication server and / or a separate server existing within the online virtual experience platform 102. The communication server 202 may include a media server 204 that communicates operably with the chat service 206.
[0078] The media server 204 is a server configured to connect components of the network environment 100 and transmit audio streams (or other data) between those components. The media server 204 may, for example, facilitate real-time communication between various client devices and between each client device and the online virtual experience server 102.
[0079] The chat service 206 may be a software service configured to enable voice chat and / or video chat (with audio) between the client device and the online virtual experience server 102.
[0080] The relay server 210 may be a third-party relay server and / or a separate server existing on the online virtual experience platform 102. The relay server 210 may include a TURN server 212 that communicates operably with the TURN management component 214.
[0081] The TURN server 212 may implement the Traversal Using Relay NAT (TURN) protocol. The TURN server 212 may relay network traffic. For example, the TURN server 212 may support communication between the client device 110 and the client device 116 via the network 122.
[0082] In addition to other functions, the TURN management component 214 may implement the communication protocol and control messaging with the TURN server 212.
[0083] Hereinafter, the translation of audio for chat using the voice chat translation component 106 and available translation data will be more fully described with reference to FIG. 3.
[0084] FIG. 3 is a diagram of an exemplary voice translation pipeline 300 for automatically translating voice chat (or video chat with audio) in a virtual metaverse according to some implementations. Pipeline 300 is provided for purposes of illustration. In some implementations, pipeline 300 may include the same, fewer, more, or different elements configured in the same or different ways as shown in FIG. 3.
[0085] As shown in FIG. 3, pipeline 300 begins with the reception of source audio from a voice chat (or video chat) at stage 302. The source audio may be associated with translation data obtained at stage 304. The translation data may include user settings for translation, language settings, and other user settings.
[0086] Upon obtaining the translation data and receiving the audio, the voice chat translation component 106 may begin the translation (as shown, for example, in the dashed box 306).
[0087] At stage 308, the source audio may be converted from the form received from the media server 204 to another format suitable for text extraction. For example, if the media server 204 uses a first format (e.g., OPUS), stage 308 may include a conversion from the first format to a second format (e.g., WAV).
[0088] In stage 310, the (optionally) converted audio is converted into text. For example, the converted audio may be processed to extract phonemes or other audio cues, and those phonemes or other audio cues may be used to estimate the text. In some implementations, a trained machine learning model is used to convert the audio into text.
[0089] In stage 312, machine translation of the text is performed to translate the text from a first language to a second language. The machine translation may hold context data and / or sentiment data. For example, the context data and / or sentiment data may be identified using a trained machine learning model that identifies context and / or sentiment from phrases, phonemes, audio cues, accentuation, emphasis, etc. within the received audio stream. In some implementations, the context data may include specific emphasis, accentuation, and other attributes from a first speaker. In these and other implementations, the sentiment data may be extracted from the context data (such as by identifying stronger sentiment with a stronger accentuation / emphasis, for example). In some implementations, a trained machine learning model or sub-model may preprocess the audio to identify context data and / or sentiment data for use in the translated speech waveform (such as by adjusting to enhance or suppress sentiment in the synthesized speech).
[0090] In some implementations, the context data and / or sentiment data may be further identified based on an analysis of video or animation included in the received data (when the chat is a video chat). Such analysis may be performed by a trained machine learning model or other techniques configured to identify sentiment from one or more frames of the video or animation.
[0091] In stage 314, the translated text is converted to speech by a speech synthesizer or TTS (Text-to-Speech). The speech synthesizer may utilize context data and / or emotion data to modify the generated speech waveform (or directly in the generation process) in order to output speech that conveys the same context and / or emotion. For example, the speech synthesizer may receive input emotion data and provide in the output speech an accented pronunciation that reflects the emotion indicated by the emotion data. For example, the speech synthesizer may provide a varying speech pattern that reflects the context indicated by the context data. For example, if the received audio is from an indoor context with reverberation or background noise, the output waveform may be generated to include reverberation or background noise.
[0092] Other techniques for improving the translation of emotion and context may also be utilized. For example, a speech-to-speech translation system that gives expressiveness to the translation may be phoneme-based. The voice can be decomposed into phonemes such that there are some differences between different languages and their dialects. The probability matrix may be based on the person speaking and the language characteristics of the person's source language. Using this probability matrix, it is possible to effectively render and / or probabilistically identify the most likely phonemes that follow other phonemes.
[0093] In stage 316, the speech waveform is optionally converted back from the second format to the first format. In this way, the audio output stage 318 may provide an audio stream that can be input by the media server 204 and directed to the chat recipient in a manner similar to an untranslated voice chat.
[0094] Subsequently, a brief description of some exemplary methods of performing automatic voice chat translation is presented with reference to FIGS. 4A - 4D.
[0095] Figures 4A - 4D: Exemplary Method for Automatic Voice Translation Figure 4A is a diagram showing an exemplary per - user voice machine learning model training method 400 according to some implementations. As shown, method 400 may include training one or more machine learning models per user. Method 400 may be trained to provide enhanced translation accuracy and may also include storing models associated with a particular user. Multiple trained models may also be used for multilingual translation. Method 400 may, in some implementations, use transfer learning techniques.
[0096] As shown in Figure 4A, user 414 may provide input voice chat audio 402 to voice chat server 404 (or server 102). Voice processing plugin 406 and voice pre - processing stage 408 may filter and / or remove noise or other artifacts from audio 402. Thereafter, training data injection system 422 may generate training records for training a machine learning model and store the training records in training data store 434. In some implementations, training data cleanup processor 438 may adjust the stored training records.
[0097] Thereafter, or substantially simultaneously, machine learning model evaluator processor 424 and machine learning model generator processor 432 may generate a data model representing the machine learning model being trained and store the data model in model data store 426. Similarly, different versions of the machine learning model may be stored in data store 428. A reference model and / or a base model may be stored from and / or retrieved from data store 436 for use in training to create the models stored in 426 / 428.
[0098] As shown in FIG. 4A, in some implementations, the machine learning model may be generated, trained, and adjusted for each user. In this way, faster voice chat translation with improved context and / or sentiment may be achieved. Other variations, including machine learning models based on specific dialects, specific languages, etc., may be implemented in some implementations instead of or in combination with the per-user model.
[0099] FIG. 4B is a diagram showing an exemplary method 410 for moderating and adjusting a voice chat according to some implementations. As shown, method 410 may include extending the voice translation system to include trust and safety features that allow any listener (e.g., a young listener for whom certain content may be inappropriate or not allowed) to safely participate in the voice chat. Method 410 may filter inappropriate voice chat messages, block non-voice audio, and allow the user to adjust how the voice sounds to the recipient of the voice chat.
[0100] As shown in FIG. 4B, user 414 initiates a chat request to chat with user 416. Voice chat audio 402 is sent from the user device associated with user 414 to a voice chat server 404 (or server 102). A voice processing plugin 406 sends the processed audio waveform to a speech-to-text (STT) system 444 to create text. The text may then undergo text moderation and / or filtering 448 to remove offensive or moderated content.
[0101] In some implementations, the steps of voice adjustment pre-processor 442 and voice adjustment 446 may be performed to adjust the synthesized speech waveform to mimic the emotion in the translated language. Further, in some implementations where direct phoneme translation may be used, pre-processing 442 may also include moderation activities based on phonemes related to the moderated content.
[0102] FIG. 4C is a diagram showing an exemplary method 420 for player control of voice chat output according to some implementations. As shown, method 420 may include a voice adjustment system 458 to enable a user to control how their voice sounds to other users.
[0103] As shown in FIG. 4C, user 414 requests to chat with user 416. Voice chat audio 402 is provided to voice chat server 404 (or server 102) as described above and undergoes voice processing 406. In some implementations, voice output preferences 452 (i.e., including override preferences and other preferences related to user 414) are sent to voice preference service 454 for storage in data store 462.
[0104] Voice output generation system 456 may retrieve a per-user trained machine learning model from data store 464 and / or a voice pack model from data store 466 for use in synthesizing a speech waveform from the translated text. Thereafter, voice adjustment system 458 may create the desired voice chat output audio that is sent to voice chat server 404 and routed to user 416.
[0105] FIG. 4D is a diagram showing an exemplary voice generation method 430 according to some implementations. As shown, method 430 may include a voice generation application programming interface (API) that is exposed to developers. The developers may use the exposed API to add voices to non-player characters (NPCs) using text input, and may also include localized voices based on the speech translation pipeline 300. For example, multilingual output may be provided for different texts input by developers for output in different regions.
[0106] As shown in FIG. 4D, native language information 470, NPC text information 472, and voice characteristic definitions 474 may be provided to the platform 102, and then they may be routed to the audio data store 490, the machine translation system 312, and the text-to-speech (TTS) voice generation system 314. The machine translation system may generate multiple different translated texts 482 for the generation of multiple different output speeches 486 in different languages respectively associated with the target computing device and / or user device. Additionally, although shown as being associated with NPCs, they may vary to include multiple translations of user chats such that multiple different recipients each speaking a different language may engage in the chat.
[0107] FIG. 5 is a diagram showing an exemplary voice machine learning model training method 500 according to some implementations. As shown, method 500 may include training a machine learning model with higher reliability through the use of backpropagation and text-to-speech services.
[0108] For example, user audio 502 may be used by the STT component 508 to create text 510. The model management system 504 may implement a per-user model training method 506 using the translated user audio 512 and the translated speech text 516 by backpropagation to improve accuracy.
[0109] Other variations of the exemplary method of FIG. 5 may be applicable in some implementations, and all such variations are considered to be within the scope of the exemplary embodiments.
[0110] Figure 6: Exemplary Method for Translating Voice Chat FIG. 6 is a flowchart of an exemplary method 600 for automatically translating voice chat in a metaverse place according to some implementations. In some implementations, method 600 may be implemented, for example, on a server system, such as the online virtual experience platform 102 shown in FIG. 1. In some implementations, some or all of method 600 may be implemented in a system such as one or more of the client devices 110 and 116 shown in FIG. 1, and / or in both a server system and one or more client systems. In the example being described, the system implementing may include one or more processors or processing circuits and one or more storage devices such as a database or other accessible storage. In some implementations, different components of one or more servers and / or clients may execute different blocks or other portions of method 600. Method 600 may begin at block 602.
[0111] At block 602, a request to translate audio is received. The audio is associated with a chat function of a metaverse place in the virtual metaverse from a first user among a plurality of users. The audio is received from the first user. The plurality of users are associated with the metaverse place and / or the chat with the first user. Further, the first user may be associated with a first user device (such as client device 110). Block 602 is followed by block 604.
[0112] In block 604, translation data related to a second user among a plurality of users is retrieved. The translation data includes at least a language setting related to the second user. Further, the second user is associated with a second user device (such as client device 116 for example). Block 604 is followed by block 606.
[0113] In block 606, the audio received from the first user is converted into text. The audio includes the input speech in the first language spoken by the first user. For example, a machine learning model may be trained to extract phonemes from the audio, use the extracted phonemes to reproduce the text related to the input speech, and extract context from the audio. Block 606 is followed by block 608.
[0114] In block 608, the text is translated into a second language. The second language is defined by the language preference, and the translated text includes context data and / or sentiment data. For example, a machine learning model may be used to extract context data based on the speech of the first user, and may also be used to extract sentiment data based on the speech of the first user. The context data and / or sentiment data may be encoded in any suitable format including accents, variations, and other notations that may be embedded in the text and / or included separately from the text with appropriate timestamps or synchronization marks. Block 608 is followed by block 610.
[0115] In block 610, the translated text is converted into output speech that includes context data and / or sentiment data. For example, a text-to-speech model specific to or unique to a user may be provided with the text, context data, and / or sentiment data. The text-to-speech model may also be referred to as a speech synthesizer or speech synthesis model. Using the context data and / or sentiment data together with the speech data, the speech synthesizer may generate a speech-based speech waveform in a second language that includes at least one or more of the context data and / or sentiment data. Block 610 is followed by block 612.
[0116] In block 612, the output speech is provided to a second user device. For example, the speech waveform may be converted to a specific audio format for routing using relay server 210 and / or processing by media server 204. Thereafter, a second user device (e.g., client device 116) may receive and output the audio for playback to the second user.
[0117] As described above, a system, method, and computer-readable medium may provide automatic translation of voice chat in a virtual experience. Variations of the above-described techniques may include additional features that result in an improved user experience and reduced latency in translation.
[0118] For example, the following improvements may be implemented in method 600 of FIG. 6 and / or pipeline 300 of FIG. 3.
[0119] Different voice models: Each output model may be trained for a specific voice. Thus, a special language model may be implemented so that any user can speak as virtually any character voice available as an output model. Noise removal / training for different age groups and dialects may generate different output models, and those output models may be used to change the apparent age of the voice to match the user or user settings. Different dialects may also generate different output models, and those output models may be used to modify the speech waveform to more closely match the local dialect.
[0120] Safety: Since the expressiveness of the voice can be tracked, the system may mute the voice output when the player is speaking aggressively. Additionally, speech-to-text (STT) may be applied to the output voice to check for inappropriate content and / or context before providing the audio to be output on a second user device. Further, the voice output may be modified to anonymize the voice of the original speaker without losing the expressiveness of the voice.
[0121] Latency: A predictive model of the phoneme mapping to other phonemes in other languages may be used to reduce latency. For example, English has approximately 42 different phonemes, Spanish has approximately 24 phonemes, and a probability model that maps phonemes as a stream to their mappings in other languages using audio data and splits based on translation may be used to reduce latency. Thus, in some scenarios, the latency inherent in word-to-word translation may be avoided, and the focus may be on how sounds and voices are planned. For example, the stream may be translated as soon as it arrives without waiting for an entire block or sentence to be fully spoken.
[0122] In these examples, for languages with different structures, low-latency voice translation may be possible. Take, for example, a language that has the adjective-noun-verb order as compared to verb-noun-adjective. The sounds are probabilistically mapped and can ignore the structure of the sentence itself. To obtain reliable results, audio translation may be terminated before the person finishes speaking.
[0123] In these examples, there may still be some structuring mismatches, but in many cases, they can be detected early before the audio stream ends. As a step towards per-phoneme processing, the final audio can be reconstructed during processing according to the language rules of each language.
[0124] Figure 7: Exemplary method for translating voice chat based on phoneme prediction Figure 7 is a flowchart of an exemplary method 700 for automatically translating voice chat in a metaverse place according to some implementations. In some implementations, method 700 may be implemented, for example, on a server system, such as the online virtual experience platform 102 shown in FIG. 1. In some implementations, some or all of method 700 may be implemented in a system such as one or more of the client devices 110 and 116 shown in FIG. 1, and / or in both the server system and one or more client systems. In the example being described, the system implementing may include one or more processors or processing circuits and one or more storage devices such as a database or other accessible storage. In some implementations, different components of one or more servers and / or clients may execute different blocks or other parts of method 700. Method 700 may begin at block 702.
[0125] In block 702, a request to translate audio is received. The audio is associated with the chat function of a metaverse place in the virtual metaverse from a first user among a plurality of users. The audio is received from the first user. The plurality of users are associated with the metaverse place and / or the chat with the first user. Further, the first user may be associated with a first user device (such as client device 110). Block 702 is followed by block 704.
[0126] In block 704, translation data related to a second user among the plurality of users is retrieved. The translation data includes at least the language settings related to the second user. Further, the second user is associated with a second user device (such as client device 116). Block 704 is followed by block 706.
[0127] In block 706, the audio received from the first user is converted into phonemes. The audio includes the input speech in a first language spoken by the first user. For example, a machine learning model may be trained to extract phonemes from the audio and further extract context from the audio using the extracted phonemes. Block 706 is followed by block 708.
[0128] In block 708, the phonemes from block 706 are processed for each phoneme to determine a high-confidence match of phonemes between the first language and the second language. For example, the per-phoneme processing may include predictive phoneme processing based on the probabilistically generated translated audio. The final result may be selected based on the confidence level such that a higher confidence level is selected first.
[0129] Phoneme-by-phoneme processing may also include reconstructing the audio based on a probabilistically generated translated audio that varies based on the target language. In this way, a streaming structure can be adjusted for different languages. Further, in this way, a method implemented by a computer may include predictive translation during speech-to-speech synthesis, translation during speech-to-speech synthesis, and other speech-to-speech synthesis methods. Block 708 is followed by block 710.
[0130] In block 710, output speech including context data and / or emotion data is generated based on highly reliable phoneme predictions. For example, a per-user or user-specific speech model may be provided with phonemes, context data, and / or emotion data. The speech model may sometimes also be referred to as a speech synthesizer or speech synthesis model. Using the context data and / or emotion data together with the phoneme data, the speech synthesizer may generate a phoneme-based speech waveform in a second language that includes at least one or more of the context data and / or emotion data. Block 710 is followed by block 712.
[0131] In block 712, the output speech is provided to a second user device. For example, the speech waveform may be converted to a specific audio format for routing using relay server 210 and / or processing by media server 204. Thereafter, a second user device (e.g., client device 116) may receive and output the audio for playback to a second user.
[0132] Subsequently, a more detailed description of various computing devices that may be used to implement the different devices shown in FIGS. 1-6 is provided with reference to FIG. 8.
[0133] Figure 8: Exemplary Computing Device Figure 8 is a block diagram of an exemplary computing device 800 that may be used to implement one or more features described herein according to some implementations. In one example, device 800 implements a computer device (e.g., 102, 110, and / or 116 of FIG. 1) and may be used to execute an implementation of an appropriate method described herein. The computing device 800 can be any suitable computer system, server, or other electronic or hardware device. For example, the computing device 800 can be a mainframe computer, a desktop computer, a workstation, a portable computer, or an electronic device (portable device, mobile device, cell phone, smartphone, tablet computer, television, TV set-top box, personal digital assistant (PDA), media player, game device, wearable device, etc.). In some implementations, device 800 includes a processor 802, a memory 804, an input / output (I / O) interface 806, and an audio / video input / output device 814 (e.g., a display screen, a touch screen, display goggles or glasses, audio speakers, headphones, a microphone, etc.).
[0134] Processor 802 can be one or more processors and / or processing circuits for executing program code and controlling the basic operations of device 800. A "processor" includes any suitable hardware and / or software system, mechanism, or component that processes data, signals, or other information. The processor may include a general-purpose central processing unit (CPU), multiple processing units, a system with dedicated circuits for realizing functions, or other systems. Processing is not necessarily limited to a specific geographical location or have time limitations. For example, the processor may execute the functions of the processor in "real time", "offline", "batch mode", etc. Some parts of the processing may be executed at different times and in different locations by different (or the same) processing systems. A computer may be any processor that communicates with a memory.
[0135] Memory 804 is generally provided within device 800 for access by processor 802, is suitable for storing instructions for execution by the processor, and can be any suitable processor-readable storage medium that is separate from and / or integrated with processor 802, such as random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, etc. Memory 804 can store software that operates on server device 800 by processor 802, including operating system 808, application 810, and related data 812. In some implementations, application 810 may include instructions that enable processor 802 to execute some or all of the functions described herein, such as some or all of the methods of FIGS. 6 and / or 7.
[0136] For example, memory 804 may include software instructions for automatically translating voice chats in a metaverse place. Any of the software within memory 804 may alternatively be stored in any other suitable storage location or computer-readable medium. Further, memory 804 (and / or other connected storage devices) can store the instructions and data used in the features described herein. Memory 804 and any other type of storage (magnetic disk, optical disk, magnetic tape, or other tangible media) may be considered a "storage" or "storage device".
[0137] I / O interface 806 can provide functionality to enable the server device 800 to interface with other systems and devices. For example, network communication devices, storage devices (e.g., memory and / or data store 108), and input / output devices can communicate via interface 806. In some implementations, the I / O interface may connect to an interface device that includes input devices (such as a keyboard, pointing device, touch screen, microphone, camera, scanner, etc.) and / or output devices (such as a display device, speaker device, printer, monitor, etc.).
[0138] To facilitate illustration, FIG. 8 shows one block for each of processor 802, memory 804, I / O interface 806, software blocks 808 and 810, and database 812. These blocks may represent one or more processors or processing circuits, operating systems, memories, I / O interfaces, applications, and / or software modules. In other implementations, device 800 may not have all of the components shown, and / or may have other elements including other types of elements instead of or in addition to the elements shown herein. Although online virtual experience platform 102 is described as performing operations as described in some implementations herein, online virtual experience platform 102 or any suitable component or combination of components of a similar system, or any suitable one or more processors associated with such a system, may perform the described operations.
[0139] The user device can also implement and / or be used with the features described herein. An exemplary user device can be a computer device that includes some components similar to device 800, such as processor 802, memory 804, and I / O interface 806. An operating system, software, and applications suitable for the client device can be provided in the memory and used by the processor. The I / O interface for the client device can be connected to a network communication device as well as input and output devices, such as a microphone for capturing sound, a camera for capturing images or video, an audio speaker device for outputting sound, a display device for outputting images or video, or other output devices. For example, the display device within the audio / video input / output device 814 can be connected to (or included in) device 800 to display pre- and post-processing of images as described herein, and such display device can include any suitable display device, such as an LCD, LED, or plasma display screen, CRT, television, monitor, touch screen, 3-D display screen, projector, or other visual display device. Some implementations can provide an audio output device, such as voice output or synthesis for speaking text.
[0140] The methods, blocks, and / or operations described herein may be executed in an order different from that shown or described, as appropriate, and / or may be executed simultaneously (partially or fully) with other blocks or operations. Some blocks or operations may be executed for some portions of data and, for example, may be executed again later for other portions of data. Not all of the described blocks and operations are necessarily executed in various implementations. In some implementations, the blocks and operations may be executed multiple times in different orders and / or at different times within the method.
[0141] In some implementations, some or all of the method may be implemented on a system such as one or more client devices. In some implementations, one or more of the methods described herein may be implemented, for example, on a server system and / or on both a server system and a client system. In some implementations, different components of one or more servers and / or clients may execute different blocks, operations, or other portions of the method.
[0142] One or more of the methods described herein (e.g., methods 400 - 700) can be implemented by computer program instructions or code that can be executed on a computer. For example, the code can be implemented by one or more digital processors (e.g., a microprocessor or other processing circuitry), and can be stored in a computer program product including a non - transitory computer - readable medium (e.g., a storage medium), such as a magnetic, optical, electromagnetic, or semiconductor storage medium including semiconductor or solid - state memory, magnetic tape, removable computer diskette, random access memory (RAM), read - only memory (ROM), flash memory, hard magnetic disk, optical disk, solid - state memory drive, etc. The program instructions can be included in an electronic signal and provided as a software - as - a - service (SaaS) form of a service distributed from a server (e.g., a distributed system and / or a cloud computing system). Alternatively, one or more of the methods can be implemented in hardware (such as logic gates) or in a combination of hardware and software. Exemplary hardware can be a programmable processor (e.g., a field - programmable gate array (FPGA), a complex programmable logic device), a general - purpose processor, a graphics processor, an application - specific integrated circuit (ASIC), etc. One or more of the methods can be executed as part of an application or component running on a system, or as an application or software that runs in cooperation with other applications and an operating system.
[0143] One or more of the methods described herein can be executed as a stand-alone program that can be executed on any type of computing device, a program executed on a web browser, or within a mobile application (a "mobile app") executed on a mobile computing device (e.g., a mobile phone, smartphone, tablet computer, wearable device (such as a wristwatch, bracelet, jewelry, hat, goggles, glasses, etc.), laptop computer, etc.). In one example, a client / server architecture can be used, for example, a mobile computing device (as a client device) can send user input data to a server device and receive final output data (e.g., for display) from the server for output. In another example, all calculations can be executed within a mobile app (and / or other apps) on the mobile computing device. In another example, the calculations can be divided between the mobile computing device and one or more server devices.
[0144] Exemplary terms Item 1. A method implemented by a computer for voice chat translation in a virtual metaverse, the method comprising: receiving a request to translate audio related to a chat function of a metaverse place in the virtual metaverse, the audio being received from a first user among a plurality of users, the plurality of users being associated with the metaverse place; retrieving translation data related to a second user among the plurality of users, the translation data including at least a language preference related to the second user, the second user being associated with a user device; converting the audio received from the first user into text, the audio including input speech in a first language spoken by the first user; translating the text into a second language, the second language being defined by the language preference, the translated text including context data from the audio; converting the translated text into output speech including the context data; and providing the output speech to the user device.
[0145] Item 2. The subject matter of any of the preceding items, wherein the request identifies the first user and the second user, and the request originates from a computing device associated with the first user.
[0146] Item 3. The subject matter of any of the preceding items, further comprising retrieving a voice output preference related to the first user, the voice output preference at least partially overriding the translation data related to the second user.
[0147] Item 4. The subject matter of any of the preceding items, further comprising moderating the text to remove words based on a text moderation filter before translating the text into the second language.
[0148] Item 5. The subject matter of any of the preceding items, wherein the translation step includes providing text as an input to a trained machine learning model and receiving the translated text as an output from the trained machine learning model.
[0149] Item 6. The subject matter of any of the preceding items, wherein the context data includes sentiment data extracted from audio.
[0150] Item 7. The subject matter of any of the preceding items, further including the step of preprocessing the audio to extract sentiment data.
[0151] Item 8. The subject matter of any of the preceding items, wherein the step of converting the translated text into output speech includes using a speech waveform modulator to create an adjusted speech waveform that at least partially includes the context data.
[0152] Item 9. The subject matter of any of the preceding items, further including the step of translating the text into a plurality of different languages to create a plurality of different translated texts, and the step of converting the plurality of different translated texts into a plurality of different output speeches.
[0153] Item 10. The subject matter of any of the preceding items, further including the step of providing a plurality of different output speeches to a plurality of user devices, wherein each output speech includes utterances in a language associated with each respective user device of the plurality of user devices.
[0154] Item 11. In response to execution by a processing device, receiving, by the processing device, a request to translate audio related to a chat function of a metaverse place in a virtual metaverse, wherein the audio is received from a first user among a plurality of users, the plurality of users being associated with the metaverse place; retrieving translation data related to a second user among the plurality of users, the translation data including at least a language preference related to the second user, the second user being associated with a user device; converting the audio received from the first user into text, the audio including input speech in a first language spoken by the first user; translating the text into a second language, the second language being defined by the language preference, the translated text including context data from the audio; converting the translated text into output speech including the context data; and storing instructions for causing the user device to perform operations including providing the output speech to the user device in a non-transitory computer-readable medium.
[0155] Item 12. The subject matter of any of the preceding items, wherein the request identifies a first user and a second user, and the request originates from a computing device associated with the first user.
[0156] Item 13. The subject matter of any of the preceding items, further including retrieving a voice output preference related to the first user, the voice output preference at least partially overriding the translation data related to the second user.
[0157] Item 14. The subject matter of any of the preceding items, further including moderating the text to remove words based on a text moderation filter before translating the text into the second language.
[0158] Item 15. The subject matter of any of the preceding items, wherein translating comprises providing text as an input to a trained machine learning model and receiving translated text as an output from the trained machine learning model.
[0159] Item 16. The subject matter of any of the preceding items, wherein the context data includes sentiment data extracted from audio.
[0160] Item 17. The subject matter of any of the preceding items, wherein the operation further includes preprocessing the audio to extract sentiment data.
[0161] Item 18. The subject matter of any of the preceding items, wherein converting the translated text to output speech includes using a speech waveform modulator to create an adjusted speech waveform that at least partially includes the context data.
[0162] Item 19. The subject matter of any of the preceding items, wherein the operation further includes translating the text into a plurality of different languages to create a plurality of different translated texts, converting the plurality of different translated texts into a plurality of different output speeches, and providing the plurality of different output speeches to a plurality of user devices, wherein each output speech includes utterances in a language associated with each respective user device of the plurality of user devices.
[0163] Item 20. A system including a memory storing instructions and a processing device coupled to the memory and operable to access the memory, the instructions, when executed by the processing device, causing the processing device to receive a request to translate audio related to a chat function of a metaverse place in a virtual metaverse, the audio being received from a first user among a plurality of users, the plurality of users being associated with the metaverse place, retrieve translation data related to a second user among the plurality of users, the translation data including at least a language preference related to the second user, the second user being associated with a user device, convert the audio received from the first user into text, the audio including input speech in a first language spoken by the first user, translate the text into a second language, the second language being defined by the language preference, the translated text including context data from the audio, convert the translated text into output speech including the context data, and provide the output speech to the user device.
[0164] Conclusion The description has been presented in relation to specific implementations, but these specific implementations are merely exemplary and not limiting. The concepts shown in the examples may apply to other examples and implementations.
[0165] In situations where a particular implementation considered in this specification may obtain or use user data (e.g., user demographics, user behavior data on the platform, user search history, purchased and / or viewed items, user friendships on the platform, etc.), the user is provided with options to control whether such information is collected, stored, or used, and how it is collected, stored, or used. That is, the implementations considered in this specification collect, store, and / or use user information upon receiving explicit user authorization and complying with applicable regulations.
[0166] The user is enabled to control whether a program or feature collects user information about that particular user or other users related to the program or feature. For each user for whom information is to be collected, options (e.g., via a user interface) are presented that enable the user to exercise control over the collection of information related to that user, to give permission or authorization regarding whether the information is to be collected and which parts of the information are to be collected. Further, certain data may be modified in one or more ways before being stored or used such that information that can identify an individual is removed. As an example, a user's identity may be modified (e.g., by substitution using a pseudonym, numerical value, etc.) such that information that can identify an individual cannot be determined. As another example, a user's geographical location may be generalized to a larger area (e.g., city, postal code, state, country, etc.).
[0167] It should be noted that the functional blocks, operations, features, methods, devices, and systems described in this disclosure may be integrated or separated into different combinations of systems, devices, and functional blocks as known to those skilled in the art. Any suitable programming language and programming technology may be used to implement the routines of a particular implementation. Different programming technologies, such as procedural or object-oriented programming technologies, may be used. The routines may be executed on a single processing device or multiple processors. Steps, operations, or calculations may be presented in a particular order, but the order may be changed in different particular implementations. In some implementations, multiple steps or operations shown as sequential herein may be executed simultaneously.
Description of the Reference Numerals
[0168] 100 Network environment 102 Online virtual experience platform, server 104 Virtual experience engine 105 Game, virtual experience 106 Voice chat translation component 108 Data store 110 First client device 112 Virtual experience application 114 User 116 Second client device 118 Virtual experience application 120 User 122 Network 200 Network environment 202 Communication server 204 Media server 206 Chat service 208 Voice chat plugin 210 Relay server 212 TURN server 214 TURN management component 232 User audio stream 234 Translated audio stream 300 Voice translation pipeline 312 Machine translation system 314 TTS voice generation system 400 Voice machine learning model training method per user 402 Input voice chat audio 404 Voice chat server 406 Voice processing plugin, voice processing 408 Voice preprocessing stage 410 Method for moderation and adjustment of voice chat 414 User 416 User 418 Developer user 420 Method for player control of voice chat output 422 Training data injection system 424 Machine learning model evaluator processor 426 Model data store 428 Data store 430 Voice generation method 432 Machine learning model generator processor 434 Training data store 436 Data store 438 Training data cleanup processor 442 Voice adjustment preprocessor, preprocessing 444 STT system 446 Voice adjustment 448 Text moderation and / or filtering 452 Voice output preference 454 Voice preference service 456 Voice output generation system 458 Voice adjustment system 462 Data store 464 Data store 466 Data store 470 Native language information 472 NPC text information 474 Voice characteristic definition 482 Translated text 486 Multiple output speeches 490 Audio data store 500 Method for training voice machine learning model 502 User audio 504 Model management system 506 Method for training model for each user 508 STT component 510 Text 512 Translated user audio 516 Translated speech text 600 Method 700 Method 800 Computing device 802 Processor 804 Memory 806 I / O interface 808 Operating system 810 Application 812 Related data 814 Audio / video input / output device
Claims
1. A computer-implemented method for voice chat translation in a virtual metaverse, comprising: receiving a request to translate audio related to a chat function of a metaverse place in the virtual metaverse, wherein the audio is received from a first user among a plurality of users, and the plurality of users are associated with the metaverse place; retrieving translation data related to a second user among the plurality of users, wherein the translation data includes at least a language preference related to the second user, and the second user is associated with a user device; converting the audio received from the first user into text, wherein the audio includes input speech in a first language spoken by the first user; translating the text into a second language, wherein the second language is defined by the language preference, and the translated text includes context data from the audio; converting the translated text into output speech including the context data; and providing the output speech to the user device.
2. The computer-implemented method according to claim 1, wherein the request identifies the first user and the second user, and the request originates from a computing device associated with the first user.
3. The computer-implemented method according to claim 1, further comprising retrieving a voice output preference related to the first user, wherein the voice output preference at least partially overrides the translation data related to the second user.
4. The computer-implemented method according to claim 1, further comprising moderating the text to remove words based on a text moderation filter before translating the text into the second language.
5. The method according to claim 1, wherein the step of translating includes providing the text as an input to a trained machine learning model and receiving the translated text as an output from the trained machine learning model.
6. The method according to claim 1, wherein the context data includes sentiment data extracted from the audio.
7. The method according to claim 6, further comprising a step of preprocessing the audio to extract the sentiment data.
8. The method according to claim 1, wherein the step of converting the translated text into output speech includes using a speech waveform modulator to create an adjusted speech waveform that at least partially includes the context data.
9. The method according to claim 1, further comprising translating the text into a plurality of different languages to create a plurality of different translated texts, and converting the plurality of different translated texts into a plurality of different output speeches.
10. The method according to claim 9, further comprising a step of providing the plurality of different output speeches to a plurality of user devices, wherein each output speech includes utterances in a language associated with each respective user device of the plurality of user devices.
11. In response to execution by a processing device, the processing device receives a request to translate audio related to a chat function of a metaverse place in a virtual metaverse, the audio being received from a first user among a plurality of users, the plurality of users being associated with the metaverse place; retrieves translation data related to a second user among the plurality of users, the translation data including at least a language preference related to the second user, the second user being associated with a user device; converts the audio received from the first user into text, the audio including input speech in a first language spoken by the first user. Translating the text into a second language, where the second language is defined by the language preference and the translated text includes context data from the audio; Converting the translated text into output speech including the context data; Storing instructions for causing a user device to perform operations including providing the output speech to the user device, in a non - transitory computer - readable medium. **Claim 12** The non - transitory computer - readable medium of claim 11, wherein the request identifies the first user and the second user, and the request originates from a computing device associated with the first user. **Claim 13** The non - transitory computer - readable medium of claim 11, wherein the operations further include retrieving a voice output preference associated with the first user, and the voice output preference at least partially overrides the translation data associated with the second user. **Claim 14** The non - transitory computer - readable medium of claim 11, wherein the operations further include moderating the text to remove words based on a text moderation filter before translating the text into the second language. **Claim 15** The non - transitory computer - readable medium of claim 11, wherein translating includes providing the text as an input to a trained machine - learning model and receiving the translated text as an output from the trained machine - learning model. **Claim 16** The non - transitory computer - readable medium of claim 11, wherein the context data includes sentiment data extracted from the audio. **Claim 17** The non - transitory computer - readable medium of claim 16, wherein the operations further include pre - processing the audio to extract the sentiment data. **Claim 18** The non - transitory computer - readable medium of claim 11, wherein converting the translated text into output speech includes using a speech waveform modulator to create a modulated speech waveform that at least partially includes the context data. **Claim 19** The operations include translating the text into a plurality of different languages to create a plurality of different translated texts; converting the plurality of different translated texts into a plurality of different output speeches; providing the plurality of different output speeches to a plurality of user devices, each output speech including utterances in a language associated with each respective user device of the plurality of user devices, the non-transitory computer-readable medium of claim 11 further comprising providing.
20. A system, a memory storing instructions, a processing device coupled to the memory and operable to access the memory, the instructions causing the processing device, when executed by the processing device, receiving a request to translate audio related to a chat function of a metaverse place in a virtual metaverse, the audio being received from a first user among a plurality of users, the plurality of users being associated with the metaverse place; retrieving translation data related to a second user among the plurality of users, the translation data including at least a language preference related to the second user, the second user being associated with a user device; converting the audio received from the first user into text, the audio including input speech in a first language spoken by the first user; translating the text into a second language, the second language being defined by the language preference, the translated text including context data from the audio; converting the translated text into an output speech including the context data; providing the output speech to the user device; A system that executes operations including.
Citation Information
Patent Citations
Voice translation device, method and program
JP2012073941A
Multilingual text-to-speech synthesis method
JP2021511536A
Natural Language Translation in AR
JP2022510752A
Full-duplex speech translation system, full-duplex speech translation method, and program
WO2019111346A1