Adjusting the audible area of ​​the avatar's voice.

A generative model in mixed reality systems calculates voice propagation distance to define an audible area, addressing noise and overhearing issues, enhancing communication and immersion by allowing intuitive voice control.

JP2026516939APending Publication Date: 2026-05-27INTERNATIONAL BUSINESS MACHINE CORPORATION

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
INTERNATIONAL BUSINESS MACHINE CORPORATION
Filing Date
2024-02-12
Publication Date
2026-05-27

Smart Images

  • Figure 2026516939000001_ABST
    Figure 2026516939000001_ABST
Patent Text Reader

Abstract

According to one embodiment, a method, computer system, and computer program product are provided for adjusting the audible area of ​​an avatar's voice. The present invention may comprise the steps of: receiving source audio in a microphone; generating the received audio; calculating the user's voice propagation distance based on a generation model, the source audio, the received audio, and template text sentences describing categories of mixed reality environments experienced by the user; drawing a virtual circle within the mixed reality environment with a radius equal to the voice propagation distance, centered on a user avatar representing the user; and transmitting the source audio to one or more participants within the mixed reality environment, represented by one or more participant avatars located within the virtual circle.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The present invention generally relates to the field of computing, and more specifically, to mixed reality. Mixed reality is a field related to merging the real world and the virtual world so that physical objects and digital objects coexist and interact in real time. Mixed reality is not exclusively performed in either the physical or virtual world, but is a hybrid of reality and virtual reality; thus, mixed reality describes anything in the reality-virtual continuum, excluding the two extremes, namely, a purely physical environment and a purely virtual environment. Accordingly, mixed reality includes augmented virtuality (AV), augmented reality (AR), and virtual reality (VR). Mixed reality has found practical applications in areas such as remote work, architecture, gaming, military, academic and commercial training, and in social networking.

[0002] Mixed reality systems use software to generate images, sounds, tactile feedback, and other sensations to augment the real-world environment. The generation of this augmented environment can be realized using a mobile device such as a mobile phone or a tablet, while more specialized devices in the form of glasses or headsets are also used, where elements typically generated by a computer are projected or mapped onto lenses in front of the user's eyes, overlaying the scenery of the real environment. Using computer augmentation, information about the world around the user, as well as other digital elements overlaid on the world, becomes interactive and digitally manipulable.

[0003] One emerging application of mixed reality is that of mixed reality social networks. In such applications, multiple individuals may be placed in a virtual reality environment, where each is represented by a virtual avatar, and they may be able to see the avatars of other participants and interact freely. The goal of mixed reality social networks is to virtually recreate real-world social gatherings in a mixed reality environment, enabling individuals who cannot or cannot physically exist with each other to interact and socialize in a way that recreates a face-to-face physical presence as faithfully as possible within a mixed reality space that can be accessed from their own homes simply by wearing a mixed reality device. For this purpose, mixed reality social networks must be able to receive spoken voices of users represented by avatars and transmit recorded voices to other avatars in the mixed reality environment in real time, in order to facilitate natural voice communication and conversation between users.

[0004] Because mixed reality environments are virtual, and users are not physically near each other or not all physically located, sound cannot propagate naturally through the virtual environment when a user speaks. Instead, whether a user can hear another participant's voice usually depends on whether the user and participants have definitively and mutually opted in to persistent shared voice channels, friend lists, single-instance calls, etc., through intentional selection. However, mixed reality social networks are designed to mimic social gatherings in real environments; for this purpose, a user's avatar may be placed in the virtual environment alongside those of individuals from all over the world, most of whom may be unknown to the user. To facilitate natural conversation, the mixed reality environment must allow the user to interact with participants who were previously unknown. To address this problem, attempts have been made in the art, for example, by requiring the user to select those participants that the user can hear and / or be heard by, or by requiring the user to filter what the user hears or can be heard by the user based on different topics, intentions, group membership, or any other criteria. This results in a situation where users must manually select or modify rules before communicating with individuals whose voices they are unable to hear. This places an obstacle to natural and effortless communication between users in a virtual reality environment.

[0005] Attempts have been made in the art to eliminate the need for user action before enabling a user to interact with other participants; currently, in some solutions attempted in the art, virtual characters existing in the virtual space are audible to all persons in the virtual environment. This is problematic in that in large chat rooms with many users, all users may be able to hear each other's voices, potentially resulting in an overwhelming amount of sound that can interfere with conversations, or more specifically, create an audible situation, thereby making audible communication as a whole impossible or even damaging the user's hearing. Another solution attempted in the art transmits the user's voice to all participants in the virtual environment whose avatars are located at a predetermined distance from the user's avatar. However, the size of the radius presents certain challenges: if the voice is transmitted too far within the virtual reality environment, the user may inadvertently overhear or be overheard by conversations of participants with whom they do not wish to communicate. If the voice is not transmitted far enough within the virtual reality environment, the user may be unable to communicate with users who are far away. To address such issues, users may have no other option but to manually adjust the radius, thereby reintroducing the requirement for user actions that make it appear as though the voice transmission opt-in within the radius has been removed. [Overview of the project]

[0006] According to one embodiment, a method, computer system, and computer program product are provided for adjusting the audible area of ​​a user's voice within a mixed reality environment. The present invention may comprise the steps of: receiving source audio in a microphone; calculating the user's voice propagation distance based on a generative model, the source audio, the received audio, and template text sentences describing categories of mixed reality environments experienced by the user; drawing a virtual circle within the mixed reality environment with a radius equal to the voice propagation distance, centered on a user avatar representing the user; and transmitting the source audio to one or more participants within the mixed reality environment, represented by one or more participant avatars located within the virtual circle.

[0007] Such embodiments address the challenges faced by the art by enabling the user to see the radius of a circle that describes the range within which the user's voice is audible to other participants in the mixed reality environment, and by enabling the automatic adjustment of the radius of the circle by simply changing the volume of the user's voice. This, as a result, enables the user to intuitively see and control who can hear their voice, thereby improving user privacy, ease of communication between users, and user immersion, and reducing the emotional and physical friction that may arise from participating in a mixed reality environment over a long period of time; in addition, the volume-based voice propagation method eliminates instances of overwhelming noise resulting from hearing the voices of all participants in the mixed reality environment, and improves the interface between the user and the system by eliminating, for example, the intermediate selection step between encountering participants and interacting with them in the mixed reality environment.

[0008] Furthermore, the present invention may optionally include the steps of: converting the source audio, the received audio, and the template text sentence into a plurality of source audio tokens, received audio tokens, and text tokens using an audio tokenization module and a text tokenizer; generating an input sequence including a first separator token, a source audio token, a second separator token, a third separator token, a received audio token, a fourth separator token, a fifth separator token, a text token, and a sixth separator token; and providing the input sequence as input to the generative model. Such an audio tokenization module may further include a chunk-by-chunk image-based audio tokenizer using a VQ-VAE model.

[0009] Aspects of the present invention may preferentially include a step of dynamically updating the circle to graphically represent the sound propagation distance of the source audio in real time while the user is speaking, and enabling the user to modify the user's volume within a sentence in order to provide a more granular and intuitive control to those who can hear the user's voice.

[0010] Aspects of the present invention may preferentially include the step of multiplying the speech propagation distance by a predetermined scaling factor in order to normalize any discrepancies between distance in the real environment and distance within the mixed reality environment.

[0011] Aspects of the present invention may optionally include a generative model trained via autoregressive language modeling.

[0012] Furthermore, embodiments may take the form of related computer program products accessible from a computer-enabled or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For the purposes of this description, the computer-enabled or computer-readable medium may be any device that includes means for storing, communicating, propagating or transferring programs for use by or in connection with an instruction execution system, apparatus or device. [Brief explanation of the drawing]

[0013] These and other objects, features and advantages of the present invention will become apparent from the following detailed description of their exemplary embodiments, which will be read in conjunction with the accompanying drawings. The examples are for clarity to facilitate understanding of the invention by those skilled in the art in conjunction with the detailed description, and various features of the drawings are not to scale. The drawings are as follows:

[0014] [Figure 1] An exemplary networked computer environment according to at least one embodiment is shown.

[0015] [Figure 2] This is an operation flowchart illustrating a dynamic voice adjustment process according to at least one embodiment.

[0016] [Figure 3] This document provides an exemplary use case of a system that implements a dynamic voice adjustment process according to at least one embodiment.

[0017] [Figure 4] This diagram illustrates an exemplary generative model of a system implementing a dynamic voice adjustment process according to at least one embodiment.

[0018] [Figure 5]A diagram showing an exemplary audio tokenization module of a generative model according to at least one embodiment.

[0019] [Figure 6] A diagram showing an exemplary per-chunk image-based tokenizer of an audio tokenization module of a generative model according to at least one embodiment.

[0020] [Figure 7] A diagram showing a training mode of a generative model according to at least one embodiment. **DETAILED DESCRIPTION OF THE INVENTION**

[0021] Detailed embodiments of the claimed structures and methods are disclosed herein; however, it is to be understood that the disclosed embodiments are merely exemplary of the claimed structures and methods that may be embodied in various forms. However, the present invention may be embodied in many different forms and should not be construed as limited to the exemplary embodiments set forth herein. In the description, details of well-known features and techniques may be omitted in order to avoid unnecessarily obscuring the presented embodiments.

[0022] Embodiments of the present invention relate to the field of computing, and more specifically, to mixed reality. The exemplary embodiments described below provide, among other things, a system, method, and program product for dynamically adjusting the audible area of a virtual avatar within a virtual environment by utilizing a text generative model that infers sound propagation distances from pairs of source and received audio signals.

[0023] According to one or more embodiments, the present invention is a method for dynamically and automatically adjusting the audible area of the voice of a virtual character in a mixed reality environment by using a generation model that receives two audio inputs and a template text sentence comprising categories of the mixed reality environment, calculating the farthest propagation distance of the user's voice based on the output template text sentence from the generation model, and drawing a virtual circle around the user's avatar based on the voice propagation distance.

[0024] In one or more embodiments of the present invention, the mixed reality environment can be a hybrid environment comprising both physical and virtual elements. The mixed reality environment can comprise a hybrid physical-virtual world in which one or more users can enter, view, move around, and interact through the medium of a mixed reality device. The mixed reality environment can include an extended reality environment that generates a hybrid augmented reality environment in which generated images, sounds, tactile feedback, and other sensations are integrated into the real-world environment to comprise both virtual and real-world elements. The mixed reality environment can include a virtual reality environment that completely replaces the physical environment with virtual elements, such that a user experiencing the virtual reality environment cannot see any object or element of the physical world; however, the virtual reality environment is fixed to a real-world location, such that the movement of the user, virtual objects, virtual environmental effects, and elements all occur relative to corresponding locations in the physical environment. All users in a single mixed reality environment can potentially see and / or interact with the same virtual objects and virtual elements and interact with each other's virtual representations, or avatars.

[0025] In some embodiments of the present invention, the mixed reality device may be any device or combination of devices capable of recording real-world information that a mixed reality program can overlay with computer-generated perceptual elements to generate a mixed reality environment; the mixed reality device may further record user actions, location, movement, etc., to track the user's movement within the mixed reality environment and their interactions with it. The mixed reality device may display the mixed reality environment to the user. The mixed reality device may be equipped with or comprise a number of sensors such as cameras, microphones, accelerometers, etc., and / or may be equipped with or comprise a number of user interface devices such as displays, touchscreens, speakers, etc. In some embodiments, the mixed reality device may be a headset worn by the user.

[0026] In one or more embodiments of the present invention, a user may be an individual who interacts with a mixed reality environment through the use of a mixed reality device. A participant may be a non-user individual who interacts with the mixed reality environment in a similar manner through the use of a mixed reality device.

[0027] Various aspects of this disclosure are illustrated by narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of mechanical logic included in embodiments of computer program products (CPPs). With respect to any flowchart, depending on the technology involved, operations may be performed in a different order than those shown in a given flowchart. For example, again depending on the technology, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated stage, simultaneously, or in a manner that at least partially overlaps in time.

[0028] Embodiments of a computer program product ("CPP Embodiment" or "CPP") are terms used in this disclosure to describe any set of one or more storage media ("Multiple Media") that are collectively comprised of a set of one or more storage devices that collectively contain machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A "storage device" is any tangible device capable of holding and storing instructions for use by a computer processor. Computer-readable storage media may be, but are not limited to, electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, mechanical storage media, or any preferred combination of those described above. Some known types of storage devices, including these media, include diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices (such as pits / lands formed on the main surface of a punch card or disk), or any suitable combination of the foregoing. When the term "computer-readable storage medium" is used in this disclosure, it shall not be interpreted as storage in the form of a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, optical pulses passing through optical fiber cables, electrical signals communicated through wires, and / or other transmission media.As will be understood by those skilled in the art, data is typically moved at several intermittent points during the normal operation of a storage device, such as during access, defragmentation, or garbage collection; however, data is not temporary while it is stored, and therefore the storage device is not considered temporary.

[0029] The exemplary embodiments described below provide a system, method, and program product for dynamically adjusting the audible area of ​​a virtual avatar within a virtual environment, utilizing a text generation model that infers the sound propagation distance from a pair of source and received audio signals.

[0030] Referring here to Figure 1, the computing environment 100 includes an example of an environment for executing at least some computer code involved in performing the method of the invention, such as a code block 145, which may comprise a mixed reality social network 107 and a dynamic voice adjustment program 108. In addition to the code block 145, the computing environment 100 includes, for example, a computer 101, a wide area network (WAN) 102, an end user device (EUD) 103, a remote server 104, a public cloud 105, and a private cloud 106. In this embodiment, the computer 101 includes a processor set 110 (including processing circuits 120 and a cache 121), a communication fabric 111, volatile memory 112, persistent storage 113 (including an operating system 122 and a code block 145 as identified above), a peripheral device set 114 (including a user interface (UI), a device set 123, storage 124, and an Internet of Things (IoT) sensor set 125), and a network module 115. The remote server 104 includes the remote database 130. The public cloud 105 includes the gateway 140, the cloud orchestration module 141, the host physical machine set 142, the virtual machine set 143, and the container set 144.

[0031] Computer 101 may take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device currently known or to be developed in the future, capable of running programs, accessing networks, or querying databases such as remote database 130. As is well understood in the field of computer technology, and depending on the technology, the execution of a computer implementation may be distributed among multiple computers and / or across multiple locations. On the other hand, in this presentation of the computing environment 100, in order to keep the presentation as concise as possible, the detailed discussion focuses on a single computer, specifically computer 101. Computer 101 may be located in the cloud, even if not shown in the cloud in Figure 1. On the other hand, it is not necessary for computer 101 to be in the cloud unless it may be shown in some definitive way.

[0032] The processor set 110 includes one or more computer processors of any type currently known or to be developed in the future. The processing circuitry 120 may be distributed across multiple packages, for example, multiple coordinated integrated circuit chips. The processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. The cache 121 is memory located within the processor chip package and is typically used for data or code that should be available for high-speed access by threads or cores running on the processor set 110. The cache memory is typically organized into multiple levels depending on its relative proximity to the processing circuitry. Alternatively, some or all of the cache for the processor set may be located "off-chip". In some computing environments, the processor set 110 may operate with qubits and be designed to perform quantum computing.

[0033] Computer-readable program instructions are typically loaded into computer 101, causing the processor set 110 of computer 101 to execute a series of operational steps, thereby enabling the computer implementation method. As a result, the instructions thus executed instantiate the method specified in the flowcharts and / or narrative descriptions of the computer implementation method contained herein (collectively referred to as the "Method of the Invention"). These computer-readable program instructions are stored in various types of computer-readable storage media, such as the cache 121 and other storage media described later. The program instructions and associated data are accessed by the processor set 110 to control and direct the execution of the Method of the Invention. In the computing environment 100, at least some of the instructions for executing the Method of the Invention may be stored in code blocks 145 in persistent storage 113.

[0034] The communication fabric 111 is a signal-conducting path that enables various components of the computer 101 to communicate with one another. Typically, this fabric is made up of switches and conductive paths, such as buses, bridges, physical input / output ports, and similar switches and conductive paths. Other types of signal-conducting paths, such as optical fiber communication paths and / or wireless communication paths, may be used.

[0035] The volatile memory 112 is any type of volatile memory currently known or to be developed in the future. Examples include dynamic random-access memory (RAM) or static RAM. Typically, volatile memory is characterized by random access, but this is not required unless explicitly stated. In computer 101, the volatile memory 112 is located in a single package and resides inside computer 101, but alternatively or in addition, the volatile memory may be distributed across multiple packages and / or located externally to computer 101.

[0036] The persistent storage 113 is any form of non-volatile storage for a computer, currently known or to be developed in the future. The non-volatility of this storage means that the stored data is maintained regardless of whether the computer 101 and / or the persistent storage 113 are directly powered. The persistent storage 113 may be read-only memory (ROM), but typically, at least a portion of the persistent storage allows for writing, deleting, and rewriting of data. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. The operating system 122 can take several forms, such as various known proprietary operating systems employing a kernel or open-source portable operating system interface type operating systems. The code contained in the code block 145 typically includes at least some computer code involved in performing the method of the invention.

[0037] The peripheral device set 114 includes a set of peripheral devices for the computer 101. Data communication connections between the peripheral devices and other components of the computer 101 can be implemented in various ways, including Bluetooth connections, Near-Field Communication (NFC) connections, connections made by cables (such as universal serial bus (USB) type cables), insertable connections (e.g., secure digital (SD) cards), connections made via local area communication networks, and even connections made via wide area networks such as the internet. In various embodiments, the UI device set 123 may include components such as a display screen, speakers, microphones, wearable devices (such as mixed reality headsets, goggles, and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. Storage 124 is external storage such as an external hard drive, or insertable storage such as an SD card. Storage 124 may be persistent and / or volatile. In some embodiments, storage 124 may take the form of a quantum computing memory device that stores data in the form of qubits. In embodiments where computer 101 needs to have large-capacity storage (for example, computer 101 locally stores and manages a large database), this storage may be provided by peripheral storage devices designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. The IoT sensor set 125 consists of sensors that can be used in Internet of Things applications. For example, one sensor may be a first microphone and another sensor may be a second microphone, where either the first or second microphone is integrated into a mixed reality device.Other examples of sensors may include gyroscopes, cameras, LiDAR, accelerometers, etc., for head tracking, motion tracking, tilt detection, and other such functions.

[0038] The network module 115 is a collection of computer software, hardware, and firmware that enables computer 101 to communicate with other computers via the WAN 102. The network module 115 may include hardware such as a modem or Wi-Fi® signal transceiver, software for packetizing and / or depackaging data for communication network transmission, and / or web browser software for exchanging data over the internet. In some embodiments, the network control and network forwarding functions of the network module 115 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing software-defined networking (SDN)), the control and forwarding functions of the network module 115 are performed on physically separate devices, thereby allowing the control function to manage multiple different network hardware devices. Computer-readable program instructions for performing the method of the present invention can typically be downloaded from an external computer or external storage device to computer 101 via a network adapter card or network interface included in the network module 115.

[0039] WAN102 is any wide area network (e.g., the Internet) that can transmit computer data over non-local distances using any currently known or future-developed technology for communicating computer data. In some embodiments, the WAN may be replaced and / or complemented by a local area network (LAN), such as a Wi-Fi network, designed to exchange data between devices located in a local area. The WAN and / or LAN typically include computer hardware such as copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.

[0040] The end-user device (EUD) 103 is any computer system used and controlled by an end-user (e.g., a customer of the company running computer 101) and can take any of the forms described above in relation to computer 101. Typically, EUD 103 receives useful and valuable data from the operation of computer 101. For example, in a hypothetical case where computer 101 is designed to provide recommendations to the end-user, these recommendations would typically be communicated from the network module 115 of computer 101 to EUD 103 via WAN 102. In this way, EUD 103 can display or otherwise present the recommendations to the end-user. In some embodiments, EUD 103 may be a client device such as a thin client, heavy client, mainframe computer, or desktop computer.

[0041] The remote server 104 is any computer system that provides at least some data and / or functionality to computer 101. The remote server 104 may be controlled and used by the same entity that operates computer 101. The remote server 104 represents a machine that collects and stores useful and valuable data for use by other computers, such as computer 101. For example, in the hypothetical case where computer 101 is designed and programmed to provide recommendations based on historical data, this historical data may be provided to computer 101 from the remote database 130 of the remote server 104.

[0042] The public cloud 105 is any computer system available for use by multiple entities, providing on-demand availability of computer system resources and / or other computer functions, particularly data storage (cloud storage) and computing power, without direct active management by the user. Cloud computing typically leverages resource sharing to achieve coherence and economies of scale. Direct active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented by virtual computing environments running on various computers that make up the host physical machine set 142, which is the universe of physical computers available in and / or to the public cloud 105. The virtual computing environment (VCE) typically takes the form of virtual machines from the virtual machine set 143 and / or containers from the container set 144. These VCEs may be stored as images and are understood to be transportable either as images or after instantiation of the VCEs, in and between various physical machine hosts. The cloud orchestration module 141 manages the transfer and storage of images, deploys new instances of VCE, and manages active instances of VCE deployments. The gateway 140 is a collection of computer software, hardware, and firmware that enables the public cloud 105 to communicate over the WAN 102.

[0043] Here, some further explanation of virtualized computing environments (VCEs) is provided. A VCE can be stored as an "image." From this image, a new active instance of the VCE can be instantiated. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses virtualization at the operating system level. This refers to an operating system feature where the kernel enables the existence of multiple isolated user-space instances called containers. These isolated user-space instances typically behave like actual computers from the perspective of the programs running within them. Computer programs running on a normal operating system can utilize all the resources of that computer, including connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, programs running inside a container can only use the contents of the container and the devices allocated to the container; this feature is known as containerization.

[0044] The private cloud 106 is similar to the public cloud 105, except that its computing resources are available only for use by a single enterprise. While the private cloud 106 is shown interacting with the WAN 102, in other embodiments, the private cloud may be completely disconnected from the internet and accessible only via a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types), often implemented by different vendors. Each of the multiple clouds remains a separate discrete entity, but the larger hybrid cloud architecture is coupled by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the multiple configuration clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.

[0045] According to this embodiment, the mixed reality social network 107 may be a program capable of generating and maintaining a mixed reality environment and enabling individuals to connect with and interact with the mixed reality environment and with one another through the use of avatars. An avatar may be a three-dimensional virtual character representing an individual. The individual to whom the avatar corresponds may be constrained to a single viewpoint and location of the avatar, but may navigate and interact with the mixed reality environment and other avatars, objects, or other virtual elements within the mixed reality environment by controlling the movement and actions of the avatar. In addition, the individual to whom the avatar corresponds may control the appearance, dimensions, and other characteristics of the avatar. In some embodiments, parts of the avatar, such as hands and / or head, may be directly mapped to their physical counterparts on the individual to whom the avatar corresponds, so that when the individual moves the mapped body part, the avatar performs the same movement with the corresponding virtual body part. The mixed reality social network 107 enables individuals, such as users and other participants, to interact with each other via audible speech. The mixed reality social network 107 may include, be integrated with, or otherwise be configured to interoperate with a dynamic voice adjustment program 108.

[0046] According to this embodiment, the dynamic voice adjustment program 108 may be a program capable of dynamically adjusting the audible area of ​​a virtual avatar within a virtual environment by utilizing a text generation model that infers the sound propagation distance from a pair of source and received audio signals. When executed, the dynamic voice adjustment program 108 may cause the computing environment 100 to execute a dynamic voice adjustment process 200. The dynamic voice adjustment process 200 can be described in more detail below with reference to Figure 2. In embodiments of the present invention, the dynamic voice adjustment program 108 may be stored and / or executed within or by any number of devices or combinations thereof, including the computer 101, the end-user device 103, the remote server 104, the private cloud 106, and / or the public cloud 105, the peripheral device set 114, and any other devices connected to the server 112 and / or the WAN 102. Furthermore, the dynamic voice adjustment program 108 may be distributed in its operation by any number of the aforementioned devices or combinations thereof. The dynamic voice adjustment program 108 may, in an embodiment, be a module or subcomponent of the mixed reality social network 107 and operate as an independent application that is invoked by or communicating with the mixed reality social network 107, or interoperates with the mixed reality social network 107 in a different manner.

[0047] Referring here to Figure 2, an operational flowchart illustrating the dynamic voice adjustment process 200 is shown according to at least one embodiment. In 202, the dynamic voice adjustment program 108 may record source audio using a microphone. The dynamic voice adjustment program 108 may operate, be integrated with, or otherwise communicate with a microphone that may be integrated into or on the user's mixed reality device, and as a result, the audio recorded by the microphone, i.e., the “source audio,” has volume characteristics as close as possible to those of the original user's speech created by the user and recorded by the microphone as source audio.

[0048] In 204, the dynamic audio adjustment program 108 may generate received audio with the same duration as the source audio. During the inference or non-training phase of the trained model, the received audio may have zero-decibel audio of the same length as the source audio, facilitating audible area adjustment. By setting the received audio to zero decibels, the dynamic audio adjustment program 108 may predict the furthest distance the source audio will propagate, in other words, the distance at which the source audio will attenuate to zero decibels. In some embodiments, zero-decibel audio clips with very short durations, such as 1 millisecond, may be pre-stored by the dynamic audio adjustment program 108. If the duration of the input source audio is 5 milliseconds, a 1-millisecond 0-decibel audio clip may be duplicated five times, and the five 0-decibel audio clips may be concatenated to generate a 0-decibel audio clip with a duration of 5 milliseconds as the input received audio. In this way, the input source sound is simulated to be attenuated to 0 decibels, and the trained model can predict the furthest relevant sound propagation distance.

[0049] In 206, the dynamic speech adjustment program 108 may receive a template text sentence indicating a category of mixed reality environment the user is experiencing. The template text sentence may be a text sentence converted by a text tokenizer into a sequence of text tokens. A text token comprises a subcomponent of text corresponding to individual syllables, phonemes, and / or even words that together form a complete sentence; for example, in the sentence "the frog hopped over the log," the text tokenizer may divide the sentence into tokenized words in the following order: ["the", "frog", "hopped", "over", "the", "log"]. All tokenized words form a vocabulary, where each word is assigned a unique index. A text token is a one-hot encoding embedding such as a vector. For example, the vocabulary may contain only three words, and the assigned index for the word "frog" may be 2. Next, the text token for the word "frog" is the one-hot encoded vector [0,0,1]. One-hot encoding is the process of defining data variables as vectors of integers, where one integer is high and the rest are low. One-hot encoding is best suited to machine learning paradigms where the features of the object being defined are nominal and lack any kind of natural order / ordinal relationship; in such scenarios, integer or label encoding models can lead machine learning models to infer ordinal relationships where they do not exist, resulting in errors; for example, a machine learning model might mistakenly assume that higher numbers are more important. One-hot learning models avoid such errors.

[0050] In embodiments, a text statement may comprise a template that follows a consistent structure and has template parameters; for example, a test statement might read "Sound propagates in [site_category]". The term "[site_category]" is an example of a template parameter and may be filled in by a user, a human programmer, a mixed reality social network 107, or several other agents that have categories that broadly describe a mixed reality environment in such a way that it enables the dynamic voice adjustment program 108 to infer general acoustic characteristics of a mixed reality environment, such as "meeting room," "wooded clearing," "cave," or "open-plan house." If the template parameter is not filled in, the dynamic voice adjustment program 108 may automatically infer the template parameter based on the attributes of the mixed reality environment, for example, through a classification technique such as a trained deep learning classification model. For example, a visual tagline displayed in a virtual environment that reads "Annual All-Attendance Meeting of a Company" may suggest a large virtual meeting room. A virtual environment filled with virtual avatars that look like doctors and patients may suggest a virtual hospital. A virtual room with a large virtual movie screen may suggest a virtual cinema. Several keywords in content that other virtual avatars are talking about, such as "project," "development," and "bug," may suggest a virtual project meeting room. The dynamic voice adjustment program 108 may analyze such visual, auditory, and / or text details of the mixed reality environment to infer the characteristics of the mixed reality environment and select template parameters that cause, suggest, and / or describe such characteristics. Once obtained, the dynamic voice adjustment program 108 may tokenize the filled parameters as text tokens.

[0051] In 208, the dynamic speech adjustment program 108 may calculate the user's speech propagation distance based on source audio, received audio, and template text sentences using a generative model. The generative model may be a machine learning algorithm that takes as input a dual-modal sequence which may consist sequentially of a category of the environment in which the user's utterance propagates, source audio, and template text sentences which include the received audio. The model may provide a template text sentence as output, where the output template text sentence includes a template parameter with a predicted value multiplied by, for example, 0.5, which represents the number of meters the sound from the source propagates to the location where the received audio is acquired. For example, in an embodiment where the distance between the first microphone and the second microphone may be 5 meters when collecting training data, the predicted value may be multiplied by 0.5. The relevant template text sentence that defines the propagation distance as the target output is "The sound propagation distance is

[10] × 0.5 meters". Thus, for 5 meters, the model predicts an integer value such as 10, i.e., a multiple; multiplying 10 by 0.5 results in 5. In embodiments of the present invention, "0.5 meters" may be considered a hyperparameter. Other embodiments may set alternative values ​​for the hyperparameter.

[0052] In some embodiments, the generative model may be trained via left-to-right token prediction, also known as autoregressive language modeling; left-to-right token prediction may be most suitable for applications where the output is a sequence of text tokens that form a phrase or sentence to be parsed from left to right by a human user or software program. During the training stage, the generative model may be provided with training data. Each element of the training data may consist of manually annotated template text sentences defining site categories such as shopping stores, source audio, and received audio, and template text sentences defining propagation distance as the target output. The generative model may be trained until it converges. The training process of the generative model can be described in more detail with respect to Figure 7.

[0053] In 210, the dynamic speech adjustment program 108 may draw a virtual circle within the mixed reality environment, centered on the user's avatar and having a radius equal to the speech propagation distance. Here, the dynamic speech adjustment program 108 may generate a virtual element within the mixed reality environment that describes the circle, where the circle represents the maximum distance from the user's avatar at which the user's speech, emitted from the user, propagates in the real-world analogue of the mixed reality environment before attenuating to zero decibels, and describes the region of the mixed reality environment. Here, the user's speech becomes audible to others, and as a result, participants represented by avatars located inside the circle will be able to hear the user's speech, while participants represented by avatars outside the circle will not be able to hear the user's speech. The circle may be generated to be placed on the top of terrain and / or visible through obstacles, and may be visible only to the user, only to participants selected or otherwise designated by the user as belonging to a group that can see the user and the circle, or to all persons. The circle may be continuously redrawn to reflect the position of the user's avatar, and as a result, it remains centered on the user's avatar even when the avatar is moving. In some embodiments, the default setting may automatically adjust the audible area based on the volume of the user's voice emanating from the user's throat; however, in some embodiments, the dynamic voice adjustment program 108 may allow the user to manually adjust the circle, for example, by increasing the radius of the circle, thereby reducing the volume at which a user speaking in a naturally gentle voice may need to speak to be audible at a certain distance. Volume-based sound propagation provides a simulation of improving individual immersion within the mixed reality social network 107 by dynamically adjusting the sound propagation distance in a real environment, creating a convenient, natural, and intuitive user experience, and reducing the consumption of time spent within the mixed reality environment.

[0054] In some embodiments of the present invention, circles may be associated with individual user utterances and may be drawn in real time to reflect the volume of the user utterance as spoken by the user and / or transmitted within the mixed reality environment. The circles may disappear once the user utterance has finished being spoken and / or transmitted, then remain for a predetermined amount of time before disappearing, and / or remain until the next user utterance is recognized before disappearing, etc. A user utterance can be any sound emitted by the user within a range from individual phonemes to any location between an entire sentence. In some embodiments of the present invention, circles may be redrawn in real time to reflect changes in volume within the user utterance. For example, by default, circles may be drawn to reflect the volume of a sentence; upon receiving source audio from a user containing a sentence, the dynamic speech adjustment program 108 may determine the propagation distance and redraw the circle for that sentence. However, the granularity can be adjusted either automatically by the dynamic speech adjustment program 108 or manually via user selection, so that the dynamic speech adjustment program 108 redraws the circles to indicate the furthest propagation distance for individual phrases, words, or even syllables. In some embodiments, the dynamic speech adjustment program 108 may indicate the volume of the user's utterance in real time at each consecutive interval by redrawing the circles at regular intervals as the user's utterance is received.

[0055] In 212, the dynamic voice adjustment program 108 can transmit source audio to all participants within the mixed reality environment where the corresponding avatar is located inside the circle. In some embodiments, the dynamic voice adjustment program 108 may measure only the horizontal distance from the user's avatar when determining the boundary of the sound propagation distance, and as a result, the circle describes a cylindrical volume that extends upward and downward relative to the boundary of the mixed reality environment. In some embodiments, the dynamic voice adjustment program 108 may measure both the horizontal and vertical distance from the user's avatar, and as a result, the circle describes a spherical volume at its center located on the user's avatar. In either case, individuals can still hear the user's voice as long as they remain within the volume range described by the circle.

[0056] In some embodiments, the dynamic audio adjustment program 108 may modify the source audio before the avatar transmits it to one or more participants within a volume range described by a circle. For example, the dynamic audio adjustment program 108 may attenuate the volume of the source audio based on the horizontal or overall distance from the user's avatar to the participant's avatar, so that the farther away the participant's avatar is, the quieter the user's sound will sound to that participant, thereby mimicking real-world sound propagation. In another example, the dynamic audio adjustment program 108 may reduce the volume of the source audio to a predetermined volume or by a predetermined amount if the source audio exceeds a predetermined volume threshold, where the volume threshold represents a volume at which the source audio may be unnecessarily loud to the participant.

[0057] Referring here to Figure 3, an exemplary use case 300 of a system implementing the dynamic voice adjustment process 200 is shown according to at least one embodiment. Here, user 302 is located inside a physical environment 304 equipped with a microphone 306 located near and / or above user 302. User 302 is wearing a mixed reality headset 308 through which user 302 experiences a mixed reality environment 310, through which user 302 is represented by an avatar 314. The microphone 306 may be integrated into user 302's mixed reality headset 308. User 302 has just called out a friend's name and produced user utterance 312. User utterance 312 is recorded by microphone 306 as source audio 318. The dynamic voice adjustment program 108 determines the sound propagation distance of user utterance 312 and draws a circle 316 inside the mixed reality environment 310. Avatar 320 is located inside circle 316, and therefore the dynamic voice adjustment program 108 transmits source audio 318 to the user represented by avatar 320. Second avatar 322 and third avatar 324 are also presented in the mixed reality environment, but they are not located inside circle 316, so the dynamic voice adjustment program 108 does not transmit source audio 318 to the second avatar 322 and third avatar 324. In three-dimensional space, the vertical cylinder enclosed by the drawn circle 316 is the only area where user 302's voice can be heard by others. In order to dynamically expand or contract the audible area of ​​user 302's voice by adjusting the radius of the drawn circle 316, user 302 simply needs to change the volume of user 302's voice, which is generated in real time from user 302's physical throat.

[0058] Referring here to Figure 4, a diagram illustrating the operation 400 of an exemplary generative model 402 implementing a dynamic audio adjustment process 200 is shown according to at least one embodiment. Here, the generative model 402 is provided with three inputs: a template text statement from a database 404, source audio 406, and received audio 408. The template text statement is provided to a text tokenizer 412 to generate a text token 414, the source audio 406 is provided to an instance of the audio tokenization module 416 to generate an audio token 418A, and the received audio 408 is provided to another instance of the audio tokenization module 416 to generate an audio token 418B. The audio tokens 418A and 418B may comprise a series of one-hot encoding vectors, e.g., [0,0,…,1,0,0], which together represent individual millisecond-length chunks of audio comprising the source audio. The dynamic speech adjustment program 108 generates an input sequence from text tokens 414 and audio tokens 418A and 418B by sequentially ordering the tokens so that they have separator tokens between them. Separator tokens may be placeholder tokens that do not describe sound or text snippets, but they serve as boundaries between text tokens 414 and audio tokens 418A and 418B, allowing the generation model 402 to more easily parse and distinguish each input token. An exemplary input sequence may have the following order of tokens: first text separator token, text token 414, second text separator token, first audio separator token, second audio separator token, source audio token 418A, third audio separator token, received audio token 418B, and fourth audio separator token. The separator tokens serve as boundaries, enclosing text token 414, audio token 418A, and audio token 418B in parentheses, respectively.All text tokens 414, separator tokens, and audio tokens 418A and B share the same one-hot encoding embedding space; for example, text tokens 414 and separator tokens in the range of 0 to M, and audio tokens 418A and B in the range of M+1 to M+K. This implies a text / separator / audio token dimension, i.e., the vector dimension is the same for all tokens. By utilizing different tokens that use the same embedding space, the dynamic speech adjustment program 108 can distinguish between text or separator tokens and audio tokens. For example, in an embodiment using two very small embedding spaces, the token indices of the first embedding space may comprise text tokens from 0 to 1, with the corresponding one-hot encoding text embeddings being (1,0) and (0,1). In the second embedding space, the token indices of audio tokens may also be from 0 to 1, with the corresponding one-hot encoding audio embeddings being (1,0) and (0,1). If the probability distribution P(u) output from the model, which in this case represents the last transformer decoder layer, is (0.2, 0.8), then the model would not know to map the output to a text embedding (0, 1) or an audio embedding (0, 1).

[0059] In one embodiment, the token input sequence is sequentially fed into a transformer decoder comprising multiple transformer decoder layers 420, each layer comprising a masked self-attention 422 and a feedforward neural network 424. The masked self-attention 422 may be a component of the transformer decoder layer 420 used in the decoder and a multi-head attention block, which can leverage self-attention causality to force predictions to pay attention only to tokens at previous positions. The generative model 402 employs left-to-right token prediction, and the masked self-attention 422 makes such predictions possible when using the transformer decoder layer 420. The feedforward neural network (FNN) 424 may be an artificial neural network in which data and computations flow in a single direction, from input data to output. The role and purpose of the FNN 424 is to process the output from one attention layer in a way that better fits the input for the next attention layer.

[0060] The final transformer decoder layer 420 then outputs output text tokens 426 of a template text sentence, which contain predicted values ​​for the template parameters. The generative model 402 then calculates the furthest distance the sound of the source audio 406 can propagate by multiplying the predicted values ​​for the template parameters from the output template text sentence by 0.5. The distance is multiplied by a preset scaling factor, for example, in embodiments where the lengths of one meter in the virtual and real environments are not equal, thereby normalizing any discrepancies between the real and virtual distances and preventing inaccurate predictions of speech propagation distance that may otherwise result. The generative model 402 then outputs the normalized speech propagation distance 428.

[0061] Referring here to Figure 5, a diagram illustrating the operation 500 of an exemplary audio tokenization module 416 of the generative model 402 is shown according to at least one embodiment. The audio tokenization module 416 converts the input audio into a spectrogram 504, which is a 2D image of the sound and graphically represents the frequencies comprising the sound. Here, the input sound comprises source audio 406 representing a recorded version of the original user utterance which may comprise a sentence. In some embodiments, for example, during a training phase where the received audio 408 is greater than zero decibels, the input sound may comprise the received audio 408. The audio tokenization module 416 then divides the spectrogram 504 into chunks 506, each chunk 506 representing a short segment of the original source audio 406 and having a duration on a millisecond scale. Here, the source audio 406 and / or received audio 408 are converted into a sequence of audio tokens 418A, B via the operation of the audio tokenization module 416. The audio tokenization module 416 first converts audio (either source audio 406 or received audio 408) into a spectrogram 504 via the audio-spectrogram conversion module 502; secondly, the audio tokenization module 416 divides the resulting spectrogram 504 into smaller chunks 506 of the same duration, with size 257xtx3, where t is in milliseconds. The audio tokenization module 416 divides the spectrogram 504 into smaller chunks 506 because, otherwise, the spectrogram 504 converted from the source audio would be excessively broad to tokenize. The chunk-by-chunk audio tokenizer 508 is image-based, and therefore the input image, i.e., the spectrogram chunk 506, must not be excessively broad.Finally, the audio tokenization module 416 converts each spectrogram chunk into an audio token using a chunk-by-chunk image-based audio tokenizer 508, outputting audio tokens 418A and 418B. The chunk-by-chunk image-based audio tokenizer 508 can be described in more detail below with reference to Figure 6.

[0062] Referring here to Figure 6, a diagram illustrating the operation 600 of an exemplary chunk-by-chunk image-based audio tokenizer 508 of the audio tokenization module 416 of the generative model 402 is shown according to at least one embodiment. Here, the chunk-by-chunk image-based audio tokenizer 508 uses a VQ-VAE model as an image-based tokenizer to convert a single spectrogram chunk 602 selected from a spectrogram chunk 506 comprising source audio 406 or received audio 408 into audio tokens 418A, B, resulting in audio tokens 418A, B, each corresponding to an audio token 418A, B comprising spectrogram 506. In other words, the chunk-by-chunk image-based audio tokenizer 508 here converts a spectrogram chunk 602 into a one-hot encoding vector of dimension (M+K), where M is the total number of text and separator tokens in this embodiment and K is the total number of embeddings in the learnable codebook 610. The VQ-VAE model is an automated image encoder model whose training is independent of labels annotated as ground truth. Using the trained encoder 606 and codebook 610, feature vectors from encoder 606 can be quantized to integers, e.g., indices of specific codebook embeddings. These integers can be converted to one-hot encoding vectors or audio tokens 418A, B. Encoder 606 is a multilayer convolutional neural network with 1024 hidden units and ReLU activation in each layer. Each layer has 4 receptive fields and a stride of half the width and height of the input image, such as the spectrogram chunk image 604 reconstructed from 2. Decoder 612 has the same architecture as encoder 606, but replaces convolution with deconvolution.The decoder 612 functions to reconstruct an image into an output image 614 that is the same as the input image received by the encoder 606, i.e., the reconstructed image 614 is expected to be the same as the input image 604, thereby ensuring that the chunk-by-chunk image-based audio tokenizer 508 is working correctly. The decoder 612 is used only during training; after training, only the encoder 606, mapping module 608, and codebook 610 are used to tokenize the spectrogram chunks 602 into audio tokens 418A and B. The encoder 606, decoder 612, and codebook 610 are trained in exactly the same way as the VQ-VAE, using all spectrogram chunks 602 that can be converted from the available audio in the pre-collected training data.

[0063] Here, the chunk-based image-based audio tokenizer 508 reshapes the input spectrogram chunk 602, which has a size of 257 × t × 3 (where t represents milliseconds), into an image 604 with a size of P × P × 3. Next, the chunk-based image-based audio tokenizer 508 encodes the reshaped chunk 604 into a 1024-dimensional feature vector via the trained encoder 606. Then, the mapping module 608 searches for the codebook embedding in the codebook 610 that is most similar to the feature vector via nearest neighbor mapping. The indices of the mapped embeddings in the learnable codebook 610, ranging from 1 to K, are converted by the chunk-based image-based audio tokenizer 508 into a one-hot encoded vector of dimension (M + K) as the final audio tokens 418A, B of the original input spectrogram chunk 602.

[0064] Referring here to Figure 7, a diagram illustrating the training process 700 for the generative model 402 is shown according to at least one embodiment. Here, the training process 700 employs left-to-right token prediction, also known as autoregressive language modeling, to train the generative model 402 on pre-collected training data. Each element of the pre-collected training data includes a manually annotated site category 702 (e.g., shopping store), source audio 406, received audio 408, and a target output sound propagation distance (expressed here as a multiple of 0.5 meters). The received audio 408 may be the same length as the source audio 406 and may be zero decibels to train the generative model 402 to predict the furthest propagation distance, or it may be at some non-zero decibel level to train the generative model to predict the propagation distance at which the received audio attenuates to a non-zero decibel level. Training the generative model 402 to recognize the propagation distance at which the received audio is attenuated to a non-zero decibel level enhances the flexibility and functionality of the dynamic audio adjustment program 108, which may enable embodiments in which, for example, transmitted source audio 318 in a mixed reality environment 310 is attenuated based on the distance between the participant 320 and the user's avatar 314.

[0065] In some embodiments, the dynamic voice adjustment program 108 may record the training data itself, in which case the dynamic voice adjustment program 108 may operate with or otherwise communicate with a second microphone; when collecting training data for a generative model, the second microphone may be located anywhere between immediately next to the first microphone and the furthest propagation distance of the source utterance, but the second microphone may be located, for example, only a few meters away from the first microphone. The audio recorded in the second microphone may be used as received audio. The source audio and the received audio are two different recordings of the same user's utterance produced by the user, and therefore they may be of the same length, differing only in volume, with the source audio being at a higher volume than the received audio. For the sake of simplicity in implementation, the distance between the first and second microphones can be an integer multiple of 0.5 meters, such as 0.5, 1, 1.5, ..., K*0.5 (where K is a very large integer), where K*0.5 meters can be considered the maximum distance over which human speech can propagate in the physical world where the volume of the received audio would be attenuated to zero decibels.

[0066] To construct a training dataset, the dynamic speech adjustment program 108 can collect four components of each training data element by playing training audio of human speech at different volume levels, recording the training audio as source audio with a first microphone, and recording the training audio as received audio with a second microphone placed at different propagation distances in a real-world site.

[0067] The dynamic speech adjustment program 108 can perform the training process 700 by randomly selecting elements of the training data. For each randomly selected element of the training data, the model 402 takes a template text sentence as input, such as "The sound propagates in [site_category]", where [site_category] is assigned a site category from the training data, source audio 406, and received audio 408. The generating model 402 then generates a sentence where "The sound propagation distance is 0.5 meters."

number

number

[0068] In some embodiments of the present invention, the generative model 402 may operate during the training process 700 as follows: In this example, the input token sequence is U={u1,…,u m It is referred to as}, where u1 is a text / audio / separator token, and u m is the total length of the input token sequence. In this case, training uses an autoregressive language modeling target to maximize the following likelihood:

number

number

number

number

number

[0069] It should be understood that Figures 2 through 7 only provide illustrations of individual implementations and do not imply any limitation regarding how different embodiments may be implemented.

[0070] The descriptions of various embodiments of the present invention are presented for illustrative purposes only and are not intended to be comprehensive or limitless to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terms used herein have been selected to best describe the principles, practical applications, or technological improvements over the technologies available on the market of the embodiments, or to enable other those skilled in the art to understand the embodiments disclosed herein.

Claims

1. A processor implementation method, wherein the method is: The stage in which the microphone receives source audio; A step of generating a received audio with zero decibels and the same duration as the aforementioned source audio; A step in which the generative model calculates the user's voice propagation distance based on the source audio, the received audio, and template text sentences describing categories of mixed reality environments experienced by the user; A step of drawing a virtual circle within the mixed reality environment with a radius equal to the sound propagation distance, centered on the user avatar representing the user; and The step of transmitting the source audio to one or more participants within the mixed reality environment, which is represented by one or more participant avatars located inside the virtual circle. A method that includes [a certain feature].

2. The step of converting the source audio, the received audio, and the template text sentence into multiple source audio tokens, received audio tokens, and text tokens using a text tokenizer and an audio tokenization module. The method according to claim 1, further comprising:

3. A step of generating an input sequence including a first separator token, the source audio token, a second separator token, a third separator token, the received audio token, a fourth separator token, a fifth separator token, the text token, and a sixth separator token; and Step of providing the aforementioned input sequence as input to the generative model. The method according to claim 2, further comprising:

4. The method according to claim 2, wherein the audio tokenization module comprises a chunk-by-chunk image-based tokenizer using a VQ-VAE model.

5. The method according to claim 1, wherein the circle is dynamically updated in real time to graphically represent the sound propagation distance of the source audio while the user is speaking.

6. The method according to claim 1, wherein the sound propagation distance is multiplied by a predetermined scaling factor.

7. The method according to claim 1, wherein the generative model is trained via autoregressive language modeling.

8. One or more processors, one or more computer-readable memories, and one or more computer-readable storage media; Program instructions for receiving source audio in a microphone, stored in at least one of the one or more storage media for execution by at least one of the one or more processors via at least one of the one or more memories; Program instructions for generating zero-decibel received audio of the same duration as the source audio, stored in at least one of the one or more storage media for execution by at least one of the one or more processors via at least one of the one or more memories; Program instructions for calculating the user's voice propagation distance, stored in at least one of the one or more storage media for execution by at least one of the one or more processors via at least one of the one or more memories, based on a generative model describing the source audio, the received audio, and template text sentences describing categories of mixed reality environments experienced by the user; Program instructions for drawing a virtual circle within the mixed reality environment with a radius equal to the sound propagation distance, centered on the user avatar representing the user, stored in at least one of the one or more storage media for execution by at least one of the one or more processors via at least one of the one or more memories; and Program instructions for transmitting the source audio to one or more participants in the mixed reality environment, represented by one or more participant avatars located inside the virtual circle, stored in at least one of the one or more storage media for execution by at least one of the one or more processors via at least one of the one or more memories. A computer system equipped with the following features.

9. Program instructions for converting the source audio, the received audio, and the template text statement into multiple source audio tokens, received audio tokens, and text tokens by an audio tokenization module stored in at least one of the one or more storage media for execution by at least one of the one or more processors via at least one of the one or more memories. The computer system according to claim 8, further comprising:

10. A program instruction for generating an input sequence including a first separator token, the source audio token, a second separator token, a third separator token, the received audio token, a fourth separator token, a fifth separator token, the text token, and a sixth separator token, stored in at least one of the one or more storage media for execution by at least one of the one or more processors via at least one of the one or more memories; and Program instructions for providing the input sequence as input to the generative model, stored in at least one of the one or more storage media for execution by at least one of the one or more processors via at least one of the one or more memories. The computer system according to claim 9, further comprising:

11. The computer system according to claim 9, wherein the audio tokenization module comprises a chunk-by-chunk image-based tokenizer using a VQ-VAE model.

12. Program instructions for dynamically updating the circle to graphically represent the sound propagation distance of the source audio in real time while the user is speaking, stored in at least one of the one or more storage media for execution by at least one of the one or more processors via at least one of the one or more memories The computer system according to claim 8, further comprising:

13. The computer system according to claim 8, wherein the sound propagation distance is multiplied by a predetermined scaling factor.

14. The computer system according to claim 8, wherein the generative model is trained via autoregressive language modeling.

15. One or more computer-readable storage media; Program instructions for receiving source audio in a microphone, stored in at least one of the one or more storage media; Program instructions for generating zero-decibel received audio of the same duration as the source audio, stored in at least one of the one or more storage media; Program instructions for calculating the user's voice propagation distance, based on a generative model stored in at least one of the one or more storage media, the source audio, the received audio, and template text sentences describing categories of mixed reality environments experienced by the user; Program instructions stored in at least one of the one or more storage media for drawing a virtual circle within the mixed reality environment with a radius equal to the sound propagation distance, centered on the user avatar representing the user; and Program instructions for transmitting the source audio to one or more participants in the mixed reality environment, which is represented by one or more participant avatars located inside the virtual circle, and which are stored in at least one of the one or more storage media mentioned above. A computer program product that includes the following features.

16. The computer program product according to claim 15, wherein the program instructions stored in at least one of the one or more storage media are converted by an audio tokenization module into a plurality of source audio tokens, received audio tokens, and text tokens, the source audio, the received audio, and the template text statement.

17. A computer program product according to claim 16, further comprising: a program instruction for generating an input sequence including a first separator token, the source audio token, a second separator token, a third separator token, the received audio token, a fourth separator token, a fifth separator token, the text token, and a sixth separator token, stored in at least one of the one or more storage media; and a program instruction for providing the input sequence as input to the generation model, stored in at least one of the one or more storage media.

18. The computer program product according to claim 16, wherein the audio tokenization module comprises a chunk-by-chunk image-based tokenizer using a VQ-VAE model.

19. The computer program product according to claim 15, further comprising program instructions stored in at least one of the one or more storage media for dynamically updating the circle to graphically represent the sound propagation distance of the source audio in real time while the user is speaking.

20. The computer program product according to claim 15, wherein the sound propagation distance is multiplied by a predetermined scaling factor.