Natural Language Translation in AR
By eliminating the original foreign language voice in real time in augmented reality devices and translating it into user-understandable voice, users' understanding difficulties in communication in different languages are solved, and an instant and seamless native language communication experience is achieved.
Patent Information
- Application Number
- CN201880100503.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-10-25
- Filing Date
- 2018-12-20
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2038-12-20
AI Technical Summary
In the prior art, when users communicate with people who speak different languages, they need to process the original voice and translated version of foreign language speakers at the same time, which leads to difficulties in understanding, and traditional methods need to wait for translation output, affecting communication fluency.
By implementing active noise cancellation technology in augmented reality devices, the original voice of foreign language speakers is eliminated in real time and translated into a language that is understandable by the user, and the translated voice is directly played back in the earpiece, realizing personalized speech processing and spatial playback.
It improves users' understanding of foreign language communication, making communication more smooth and efficient. Users can instantly hear the translated voice without being disturbed by the original voice, achieving a seamless native language communication experience.
Smart Images

Figure CN113228029B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Non - Provisional Application No. 16 / 170,639, filed on October 25, 2018, the disclosure of which is hereby incorporated by reference in its entirety.
[0003] Background
[0004] Modern smartphones and other electronic devices are capable of performing a wide variety of functions. Many of these functions are provided by the phone's core operating system, and many additional functions can be added via applications. One function that is now built into most modern smartphones is a function known as "text - to - speech" or TTS.
[0005] TTS allows a user to type words or phrases into an electronic device, and the electronic device will present a computerized voice to speak the written words. The TTS function can also be used to read a document or book to a user. The reverse of TTS is speech - to - text (STT), which is also a function commonly provided by most modern smartphones.
[0006] In addition, many smartphones can run applications that perform language translation. For example, in some cases, a user can launch an application that listens for voice input in one language, translates the words into another language, and then plays back the words of the translated language to the user. In other cases, the application can translate the words and present the words in written form to the user.
[0007] Summary
[0008] As will be described in more detail below, the present disclosure describes methods for communicating with a person speaking another language. However, contrary to conventional techniques, the embodiments herein implement active noise cancellation to mute the person speaking in a foreign language and play back a translation of the words of the foreign - language speaker to the listening user. Thus, when the listening user will see the lips of the foreign - language speaker moving, the listening user will only hear the translated version of the words of the foreign - language speaker. By removing the words of the foreign - language speaker and replacing them with words that the listener understands, the listener will more easily understand the speaker. The systems herein operate in real - time such that the listener hears the translated version of the words of the foreign - language speaker substantially as the foreign - language speaker speaks the words, rather than hearing both the foreign - language speaker and the translation simultaneously, or having to wait while the foreign - language speaker speaks and then output the translated version. In addition, due to the implementation of active noise cancellation, the listening user will only hear the translated words, rather than hearing the words of the foreign - language speaker and the translated words. This will greatly enhance the listening user's understanding of the conversation and enable people to communicate more easily and with a higher level of understanding.
[0009] In some cases, active noise cancellation and translation capabilities can be provided on an augmented reality (AR) or virtual reality (VR) device. In fact, in one example, a listening user wearing an AR headset may be conversing with a foreign language speaker who speaks a language the listening user does not understand. When the foreign language speaker speaks, the listening user's AR headset can apply active noise cancellation to the foreign language speaker's words. Then, the translated words of the foreign language speaker are played back to the listening user through the earpiece or other auditory device via the AR headset. This can occur in real time, so that the listening user can clearly and accurately understand the foreign language speaker's words. In such an embodiment, the listening user will only hear the translated version of the foreign language speaker's words and will not have to attempt to filter or ignore the spoken words of the foreign language speaker. If the foreign language speaker is also wearing such an AR headset, the two people can converse back and forth, each speaking in their native language and each hearing a response in their native language, without being hindered by the actual words of the speaker (which, in any case, the listener cannot understand). Additionally, in some embodiments, the voice of the translated words spoken to the listener can be personalized to sound as if it is coming from the foreign language-speaking user.
[0010] In one example, a computer-implemented method for performing natural language translation in AR can include accessing an audio input stream received from a speaking user. The audio input stream can include words spoken by the speaking user in a first language. The method can then include performing active noise cancellation on the words in the audio input stream received from the speaking user such that the spoken words are suppressed before reaching the listening user. Additionally, the method can include processing the audio input stream to identify the words spoken by the speaking user and translating the identified words spoken by the speaking user into a different second language. The method can also include generating spoken words in the different second language using the translated words and playing back the generated spoken words in the second language to the listening user.
[0011] In some examples, the generated spoken words can be personalized for the speaking user such that the generated spoken words in the second language sound as if they are spoken by the speaking user. In some examples, personalizing the generated spoken words can further include processing the audio input stream to determine how the speaking user pronounces various words or syllables and applying the determined pronunciations to the generated spoken words. The personalization can be dynamically applied to the words being played back when the computer determines how the speaking user pronounces a word or syllable during playback of the generated spoken words. In some examples, the speaking user can provide voice samples. These voice samples can be used to determine how the speaking user pronounces words or syllables before receiving the audio input stream.
[0012] In some examples, playing back the generated spoken words to the listening user may further include determining from which direction the speaking user is speaking and spatializing the playback of the generated spoken words to sound as if the spoken words are coming from the direction of the determined speaking user. Determining from which direction the speaking user is speaking may include receiving location data of a device associated with the speaking user, determining from which direction the speaking user is speaking based on the received location data, and spatializing the playback of the generated spoken words to sound as if the spoken words are coming from the direction of the determined speaking user.
[0013] In some examples, determining the direction from which the speaking user is speaking may also include calculating the direction of arrival of sound waves from the speaking user, determining from which direction the speaking user is speaking based on the calculated direction of arrival, and spatializing the playback of the generated spoken words to sound as if the spoken words are coming from the direction of the determined speaking user.
[0014] In some examples, determining from which direction the speaking user is speaking may further include tracking the movement of the listening user's eyes, determining from which direction the speaking user is speaking based on the tracked movement of the listening user's eyes, and spatializing the playback of the generated spoken words to sound as if the spoken words are coming from the direction of the determined speaking user.
[0015] In some examples, processing an audio input stream to identify words spoken by the speaking user may include implementing a speech-to-text (STT) program to identify words spoken by the speaking user and implementing a text-to-speech (TTS) program to generate translated spoken words. The method may also include downloading a voice profile associated with the speaking user and using the downloaded voice profile associated with the speaking user to personalize the generated spoken words such that the generated spoken words in the second language when played back sound as if they were spoken by the speaking user.
[0016] In some examples, the method may further include accessing stored audio data associated with the speaking user and then using the accessed stored audio data to personalize the generated spoken words. In this way, the generated spoken words played back in the second language may sound as if they were spoken by the speaking user. In some examples, the method may further include parsing words spoken by the speaking user, determining that at least one of the words is spoken in a language understood by the listening user, and pausing active noise cancellation for words spoken in a language understood by the listening user.
[0017] In some examples, the audio input stream includes words spoken by at least two different speaking users. The method can then include differentiating the two speaking users based on different voice patterns, generating spoken words for the first speaking user while performing active noise cancellation on both speaking users. Additionally, in some examples, the method can include storing the spoken words generated for the second speaking user until the first user has stopped speaking for a specified amount of time, and then playing back the spoken words generated for the second speaking user.
[0018] In some examples, the method further includes creating a voice model for the second speaking user while the first speaking user is speaking. The method can also include personalizing the generated spoken words for each of the two speaking users such that the generated spoken words in the second language sound as if they are from the voice of each speaking user.
[0019] Additionally, a corresponding system for performing natural language translation in AR can include several modules stored in a memory, including an audio access module that accesses an audio input stream that includes words spoken by a speaking user in a first language. The system can also include a noise cancellation module that performs active noise cancellation on the words in the audio input stream so that the spoken words are suppressed before reaching a listening user. The system can also include an audio processing module that processes the audio input stream to identify the words spoken by the speaking user. A translation module can translate the identified words spoken by the speaking user into a different second language, and a speech generator can use the translated words to generate spoken words in the different second language. Then, a playback module can play back the generated spoken words in the second language to the listening user.
[0020] In some examples, the above method can be encoded as computer-readable instructions on a computer-readable medium. For example, the computer-readable medium can include one or more computer-executable instructions that, when executed by at least one processor of a computing device, can cause the computing device to access an audio input stream that includes words spoken by a speaking user in a first language, perform active noise cancellation on the words in the audio input stream so that the spoken words are suppressed before reaching a listening user, process the audio input stream to identify the words spoken by the speaking user, translate the identified words spoken by the speaking user into a different second language, use the translated words to generate spoken words in the different second language, and play back the generated spoken words in the second language to the listening user.
[0021] In accordance with the general principles described herein, features from any of the above-mentioned embodiments can be used in combination with one another. These and other embodiments, features, and advantages will be more fully understood when the following detailed description is read in conjunction with the accompanying drawings and the claims.
[0022] In particular, embodiments in accordance with the present invention are disclosed in the appended claims, which relate to methods, systems, and storage media, wherein any feature mentioned in one claim category (e.g., method) can also be claimed in another claim category (e.g., system, storage media, and computer program product). The dependencies or back-references in the appended claims are chosen only for formal reasons. However, any subject matter resulting from an intentional back-reference (in particular, a multiple reference) to any preceding claim can also be claimed, such that any combination of the claims and their features is disclosed and can be claimed, regardless of the dependencies chosen in the appended claims. The subject matter that can be claimed includes not only combinations of features as set forth in the appended claims, but also any other combination of features in the claims, where each feature mentioned in the claims can be combined with any other feature or combination of features in the claims. In addition, any of the embodiments and features described or depicted herein can be claimed in a separate claim and / or in any combination with any of the embodiments or features described or depicted herein or in any combination with any of the features of the appended claims.
[0023] In an embodiment in accordance with the present invention, one or more computer-readable non-transitory storage media can embody software that, when executed, is operable to perform a method in accordance with the present invention or any of the above-mentioned embodiments.
[0024] In an embodiment in accordance with the present invention, a system can include: one or more processors; and at least one memory coupled to the processors and including instructions executable by the processors, the processors being operable to perform a method in accordance with the present invention or any of the above-mentioned embodiments when executing the instructions.
[0025] In an embodiment in accordance with the present invention, a computer program product preferably including a computer-readable non-transitory storage medium can be operable to perform a method in accordance with the present invention or any of the above-mentioned embodiments when executed on a data processing system. Brief Description of the Drawings
[0027] The drawings illustrate numerous exemplary embodiments and are part of the specification. These drawings, together with the following description, demonstrate and explain various principles of the present disclosure.
[0028] Figure 1Shows an embodiment of an artificial reality headset.
[0029] Figure 2 Shows an embodiment of an augmented reality headset and a corresponding neckband.
[0030] Figure 3 Shows an embodiment of a virtual reality headset.
[0031] Figure 4 Shows a computing architecture in which the embodiments described herein may operate, including performing natural language translation in augmented reality (AR).
[0032] Figure 5 Shows a flowchart of an exemplary method for performing natural language translation in AR.
[0033] Figure 6 Shows a computing architecture in which natural language translation in AR can be personalized for a user.
[0034] Figure 7 Shows an alternative computing architecture in which natural language translation in AR can be personalized for a user.
[0035] Figure 8 Shows an alternative computing architecture in which natural language translation in AR can be personalized for a user.
[0036] Figure 9 Shows an alternative computing architecture in which natural language translation in AR can be personalized for a user.
[0037] Figure 10 Shows a computing architecture in which a speech-to-text module and a text-to-speech module are implemented during the process of performing natural language translation in AR.
[0038] Figure 11 Shows a computing architecture in which the voices of different users are distinguished to prepare for performing natural language translation in AR.
[0039] In all the figures, the same reference symbols and descriptions indicate similar but not necessarily identical elements. Although the exemplary embodiments described herein admit of various modifications and alternative forms, specific embodiments are shown by way of example in the figures and will be described in detail herein. However, the exemplary embodiments described herein are not intended to be limited to the particular forms disclosed. Rather, the present disclosure covers all modifications, equivalents, and alternatives falling within the scope of the appended claims.
[0040] Detailed Description of Exemplary Embodiments
[0041] The present disclosure generally relates to performing natural language translation in augmented reality (AR) or virtual reality (VR). As will be explained in more detail below, embodiments of the present disclosure may include performing noise cancellation on the voice of a speaking user. For example, if the speaking user is speaking a language that the listening user does not understand, the listening user will not be able to understand the speaking user when the speaking user is speaking. Thus, embodiments herein may perform noise cancellation on the voice of the speaking user such that the listening user does not hear the speaking user. When the voice of the speaking user is muted via noise cancellation, the systems described herein may determine what words the speaking user is saying and may translate those words into a language that the listening user understands. The systems herein may also convert the translated words into speech that is played back to the user's ear via a speaker or other sound transducer. In this way, the ease of understanding of the speaking user by the listening user may be significantly improved. Instead of having one user speak into an electronic device and wait for translation, embodiments herein may operate while the speaking user is speaking. Thus, when the speaking user speaks in one language, the listening user hears in real time the generated voice that speaks the translated words to the listening user. This process may be seamless and automatic. Users may talk to each other without delay, each speaking and hearing in their native language.
[0042] Embodiments of the present disclosure may include or be implemented in conjunction with various types of artificial reality systems. Artificial reality is a form of reality that has been adjusted in some manner before being presented to a user and may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or content generated in combination with captured (e.g., real-world) content. Artificial reality content may include video, audio, haptic feedback, or some combination thereof, any of which may be presented in a single channel or in multiple channels (e.g., stereoscopic video that produces a three-dimensional effect for a viewer). Additionally, in some embodiments, artificial reality may also be associated with applications, products, accessories, services, or some combination thereof that are used, for example, to create content in artificial reality and / or otherwise be used in artificial reality (e.g., to perform an activity in artificial reality).
[0043] Artificial reality systems may be implemented in a variety of different form factors and configurations. Some artificial reality systems may be designed to operate without a near-eye display (NED), examples of which are Figure 1The AR system 100 therein. Other artificial reality systems may include NEDs that also provide visibility of the real world (e.g., Figure 2 the AR system 200 therein) or NEDs that visually immerse the user in an artificial reality (e.g., Figure 3 the VR system 300 therein). Although some artificial reality devices may be autonomous systems, other artificial reality devices may communicate and / or cooperate with external devices to provide an artificial reality experience to the user. Examples of such external devices include handheld controllers, mobile devices, desktop computers, devices worn by the user, devices worn by one or more other users, and / or any other suitable external system.
[0044] Turning to Figure 1 , the AR system 100 generally represents a wearable device that is designed to be sized to fit around a user's body part (e.g., the head). As Figure 1 shown, the system 100 may include a frame 102 and a camera assembly 104, the camera assembly 104 being coupled to the frame 102 and configured to collect information about the local environment by observing the local environment. The AR system 100 may also include one or more audio devices, such as output audio transducers 108(A) and 108(B) and an input audio transducer 110. The output audio transducers 108(A) and 108(B) may provide audio feedback and / or content to the user, and the input audio transducer 110 may capture audio in the user's environment.
[0045] As shown, the AR system 100 may not necessarily include a NED in front of the user's eyes. An AR system without a NED may take various forms, such as a headband, hat, headband, belt, watch, wristband, ankle band, ring, neckband, necklace, chest band, spectacle frame, and / or any other suitable type or form of device. Although the AR system 100 may not include a NED, the AR system 100 may include other types of screens or visual feedback devices (e.g., a display screen integrated into one side of the frame 102).
[0046] The embodiments discussed in this disclosure may also be implemented in AR systems that include one or more NEDs. For example, as Figure 2 shown, the AR system 200 may include a glasses device 202 having a frame 210, the frame 210 being configured to hold a left display device 215(A) and a right display device 215(B) in front of the user's eyes. The display devices 215(A) and 215(B) may act together or independently to present an image or a series of images to the user. Although the AR system 200 includes two displays, the embodiments of this disclosure may be implemented in AR systems having a single NED or more than two NEDs.
[0047] In some embodiments, the AR system 200 may include one or more sensors, such as sensor 240. The sensor 240 may generate a measurement signal in response to the movement of the AR system 200 and may be substantially located on any part of the frame 210. The sensor 240 may include a position sensor, an inertial measurement unit (IMU), a depth camera assembly, or any combination thereof. In some embodiments, the AR system 200 may or may not include the sensor 240, or may include more than one sensor. In embodiments where the sensor 240 includes an IMU, the IMU may generate calibration data based on the measurement signal from the sensor 240. Examples of the sensor 240 may include, but are not limited to, an accelerometer, a gyroscope, a magnetometer, other suitable types of sensors for detecting motion, sensors for error correction of the IMU, or some combination thereof.
[0048] The AR system 200 may also include a microphone array having a plurality of acoustic sensors 220(A)-220(J) (collectively referred to as acoustic sensors 220). The acoustic sensors 220 may be transducers that detect changes in air pressure caused by sound waves. Each acoustic sensor 220 may be configured to detect sound and convert the detected sound into an electronic format (e.g., analog or digital format). Figure 2 The microphone array in may include, for example, ten acoustic sensors: 220(A) and 220(B), which may be designed to be placed in the respective ears of the user; acoustic sensors 220(C), 220(D), 220(E), 220(F), 220(G), and 220(H), which may be positioned at different locations on the frame 210; and / or acoustic sensors 220(I) and 220(J), which may be positioned on the respective neckbands 205.
[0049] The configuration of the acoustic sensors 220 of the microphone array may vary. Although the AR system 200 is shown in Figure 2 as having ten acoustic sensors 220, the number of acoustic sensors 220 may be greater than or less than ten. In some embodiments, using a higher number of acoustic sensors 220 may increase the amount of audio information collected and / or the sensitivity and accuracy of the audio information. Conversely, using a lower number of acoustic sensors 220 may reduce the computational power required by the controller 250 to process the collected audio information. Additionally, the position of each acoustic sensor 220 of the microphone array may vary. For example, the position of the acoustic sensors 220 may include defined positions on the user, defined coordinates on the frame 210, orientations associated with each acoustic sensor, or some combination thereof.
[0050] The acoustic sensors 220(A) and 220(B) can be located at different parts of the user's ear, such as behind the pinna or within the auricle or fossa. Alternatively, in addition to the acoustic sensor 220 inside the ear canal, there can be additional acoustic sensors on or around the ear. Positioning the acoustic sensors beside the user's ear canal enables the microphone array to collect information on how sound reaches the ear canal. By positioning at least two of the acoustic sensors 220 on opposite sides of the user's head (e.g., as a binaural microphone), the AR device 200 can simulate binaural hearing and capture a 3D stereo field around the user's head. In some embodiments, the acoustic sensors 220(A) and 220(B) can be connected to the AR system 200 via a wired connection, and in other embodiments, the acoustic sensors 220(A) and 220(B) can be connected to the AR system 200 via a wireless connection (e.g., a Bluetooth connection). In still other embodiments, the acoustic sensors 220(A) and 220(B) can be used without being combined with the AR system 200 at all.
[0051] The acoustic sensors 220 on the frame 210 can be positioned along the length of the temple, across the bridge, above or below the display devices 215(A) and 215(B), or some combination thereof. The acoustic sensors 220 can be oriented such that the microphone array can detect sound in a wide range of directions around the user wearing the AR system 200. In some embodiments, an optimization process can be performed during the manufacture of the AR system 200 to determine the relative positions of each acoustic sensor 220 in the microphone array.
[0052] The AR system 200 can also include or be connected to an external device (e.g., a paired device), such as a neckband 205. As shown, the neckband 205 can be coupled to the glasses device 202 via one or more connectors 230. The connectors 230 can be wired or wireless connectors and can include electrical and / or non-electrical (e.g., structural) components. In some cases, the glasses device 202 and the neckband 205 can operate independently without any wired or wireless connection between them. While Figure 2Shows components of the eyewear device 202 and the neckband 205 in example locations on the eyewear device 202 and the neckband 205, but these components can be located elsewhere on and / or differently distributed on the eyewear device 202 and / or the neckband 205. In some embodiments, components of the eyewear device 202 and the neckband 205 can be located on one or more additional peripheral devices paired with the eyewear device 202, the neckband 205, or some combination thereof. Additionally, the neckband 205 generally represents any type or form of paired device. Thus, the following discussion of the neckband 205 can also apply to various other paired devices, such as smartwatches, smartphones, wristbands, other wearable devices, handheld controllers, tablet computers, laptop computers, etc.
[0053] Pairing an external device such as the neckband 205 with the AR eyewear device can enable the eyewear device to achieve the form factor of a pair of glasses while still being able to provide sufficient battery and computing power for extended capabilities. Some or all of the battery power, computing resources, and / or additional features of the AR system 200 can be provided by the paired device or shared between the paired device and the eyewear device, thus generally reducing the weight, heat distribution, and form factor of the eyewear device while still maintaining the desired functionality. For example, the neckband 205 can allow components that would otherwise be included on the eyewear device to be included in the neckband 205 because users can tolerate a heavier weight load on their shoulders than they would on their heads. The neckband 205 can also have a larger surface area on which to spread and dissipate heat into the surrounding environment. Thus, the neckband 205 can allow for a greater battery and computing capacity than might otherwise be possible on a stand-alone eyewear device. Because the weight carried in the neckband 205 may be less intrusive to the user than the weight carried in the eyewear device 202, users can tolerate wearing a lighter eyewear device and carrying or wearing the paired device for a longer period of time compared to tolerating wearing a heavy stand-alone eyewear device, thus enabling the artificial reality environment to more fully integrate into the user's daily activities.
[0054] The neckband 205 can be communicatively coupled to the eyewear device 202 and / or other devices. The other devices can provide certain functions to the AR system 200 (e.g., tracking, positioning, depth mapping, processing, storage, etc.). In Figure 2 embodiments, the neckband 205 can include two acoustic sensors (e.g., 220(I) and 220(J)), which are part of a microphone array (or potentially form their own microphone sub-array). The neckband 205 can also include a controller 225 and a power supply 235.
[0055] The acoustic sensors 220(I) and 220(J) of the neckband 205 can be configured to detect sound and convert the detected sound into an electronic format (analog or digital). In Figure 2 an embodiment, the acoustic sensors 220(I) and 220(J) can be positioned on the neckband 205 to increase the distance between the neckband acoustic sensors 220(I) and 220(J) and other acoustic sensors 220 located on the glasses device 202. In some cases, increasing the distance between the acoustic sensors 220 of the microphone array can improve the accuracy of beamforming performed via the microphone array. For example, if sound is detected by acoustic sensors 220(C) and 220(D), and the distance between acoustic sensors 220(C) and 220(D) is greater than, for example, the distance between acoustic sensors 220(D) and 220(E), the determined source location of the detected sound can be more accurate than if the sound were detected by acoustic sensors 220(D) and 220(E).
[0056] The controller 225 of the neckband 205 can process information generated by sensors on the neckband 205 and / or the AR system 200. For example, the controller 225 can process information from the microphone array that describes the sound detected by the microphone array. For each detected sound, the controller 225 can perform DoA estimation to estimate the direction from which the detected sound arrives at the microphone array. When the microphone array detects sound, the controller 225 can populate the audio data set with this information. In an embodiment where the AR system 200 includes an inertial measurement unit, the controller 225 can compute all inertial and spatial calculations from the IMU located on the glasses device 202. The connector 230 can transfer information between the AR system 200 and the neckband 205 and between the AR system 200 and the controller 225. The information can be in the form of optical data, electrical data, wireless data, or any other form of transmittable data. Moving the processing of information generated by the AR system 200 to the neckband 205 can reduce the weight and heat in the glasses device 202, making it more comfortable for the user.
[0057] The power source 235 in the neckband 205 can supply power to the glasses device 202 and / or the neckband 205. The power source 235 can include, but is not limited to, a lithium-ion battery, a lithium polymer battery, a primary lithium battery, an alkaline battery, or any other form of electrical energy storage device. In some cases, the power source 235 can be a wired power source. Including the power source 235 on the neckband 205 rather than on the glasses device 202 can help better distribute the weight and heat generated by the power source 235.
[0058] As mentioned, some artificial reality systems can substantially replace a user's one or more sensory perceptions of the real world with virtual experiences, rather than mixing artificial reality with physical reality. An example of this type of system is a head-worn display system (e.g., Figure 3 the VR system 300 in Figure 3 ), which primarily or completely covers the user's field of view. The VR system 300 can include a front rigid body 302 and a strap 304 shaped to fit around the user's head. The VR system 300 can also include output audio transducers 306(A) and 306(B). Additionally, although not shown in
[0059] ), the front rigid body 302 can include one or more electronic components, which include one or more electronic displays, one or more inertial measurement units (IMUs), one or more tracking transmitters or detectors, and / or any other suitable devices or systems for creating an artificial reality experience.
[0060] In addition to or instead of using a display screen, some artificial reality systems can also include one or more projection systems. For example, the display device in the AR system 200 and / or VR system 300 can include a micro-LED (micro-light emitting diode) projector (using, for example, a waveguide) that projects light into the display device, such as a transparent combination lens that allows ambient light to pass through. The display device can refract the projected light towards the user's pupil and enable the user to view artificial reality content and the real world simultaneously. The artificial reality system can also be configured with any other suitable type or form of image projection system.
[0061] An artificial reality system may also include various types of computer vision components and subsystems. For example, AR system 100, AR system 200, and / or VR system 300 may include one or more optical sensors, such as two-dimensional (2D) or three-dimensional (3D) cameras, time-of-flight depth sensors, single-beam or swept-frequency laser rangefinders, 3D LiDAR sensors, and / or any other suitable type or form of optical sensor. The artificial reality system may process data from one or more of these sensors to identify the user's location, map the real world, provide the user with context about the real-world surroundings, and / or perform various other functions.
[0062] The artificial reality system may also include one or more input and / or output audio transducers. In Figure 1 and Figure 3 the example shown, the output audio transducers 108(A), 108(B), 306(A), and 306(B) may include voice coil speakers, ribbon speakers, electrostatic speakers, piezoelectric speakers, bone conduction transducers, cartilage conduction transducers, and / or any other suitable type or form of audio transducer. Similarly, the input audio transducer 110 may include a capacitive microphone, a dynamic microphone, a ribbon microphone, and / or any other type or form of input transducer. In some embodiments, a single transducer may be used for both audio input and audio output.
[0063] Although not shown in Figures 1 - 3 , the artificial reality system may include a tactile (i.e., haptic) feedback system, which may be incorporated into a headset, gloves, a bodysuit, a handheld controller, environmental devices (e.g., chairs, floor mats, etc.), and / or any other type of device or system. The tactile feedback system may provide various types of cutaneous feedback, including vibration, force, traction, texture, and / or temperature. The tactile feedback system may also provide various types of kinesthetic feedback, such as motion and compliance. Motors, piezoelectric actuators, jet systems, and / or various other types of feedback mechanisms may be used to implement the tactile feedback. The tactile feedback system may be implemented independently of other artificial reality devices, within other artificial reality devices, and / or in combination with other artificial reality devices.
[0064] By providing tactile sensations, audible content, and / or visual content, an artificial reality system can create an entire virtual experience or enhance a user's real-world experience in various contexts and environments. For example, an artificial reality system can assist or augment a user's perception, memory, or cognition within a particular environment. Some systems can enhance a user's interaction with others in the real world, or can enable a more immersive interaction of a user with others in a virtual world. Artificial reality systems can also be used for educational purposes (e.g., for teaching or training in schools, hospitals, government organizations, military organizations, commercial enterprises, etc.), entertainment purposes (e.g., for playing video games, listening to music, watching video content, etc.), and / or for accessibility purposes (e.g., as hearing aids, visual aids, etc.). The embodiments disclosed herein can implement or enhance a user's artificial reality experience in one or more of these contexts and environments and / or in other contexts and environments.
[0065] Some AR systems can use a technique known as "Simultaneous Localization and Mapping" (SLAM) to map a user's environment. SLAM mapping and position recognition techniques can involve various hardware and software tools that can create or update a map of the environment while simultaneously keeping track of the user's position within the mapped environment. SLAM can use many different types of sensors to create the map and determine the user's position within the map.
[0066] SLAM techniques can, for example, implement optical sensors to determine the user's position. Radio devices including WiFi, Bluetooth, Global Positioning System (GPS), cellular, or other communication devices can also be used to determine the user's position relative to a radio transceiver or group of transceivers (e.g., a WiFi router or a group of GPS satellites). Acoustic sensors such as microphone arrays or 2D or 3D sonar sensors can also be used to determine the user's position within the environment. AR and VR devices (e.g., systems 100, 200, or 300 such as Figure 1 , Figure 2 or Figure 3 respectively) can incorporate any or all of these types of sensors to perform SLAM operations, such as creating and continuously updating a map of the user's current environment. In at least some of the embodiments described herein, the SLAM data generated by these sensors can be referred to as "environmental data" and can indicate the user's current environment. This data can be stored in a local or remote data storage (e.g., cloud data storage) and can be provided to the user's AR / VR device on demand.
[0067] When a user wears an AR headset or a VR headset in a given environment, the user may be interacting with other users or other electronic devices acting as audio sources. In some cases, it may be desirable to determine where the audio sources are located relative to the user and then present the audio sources to the user as if they were coming from the location of the audio sources. The process of determining where the audio sources are located relative to the user may be referred to herein as "localization", and the process of reproducing the playback of an audio source signal to appear as if it is coming from a particular direction may be referred to herein as "spatialization".
[0068] The localization of audio sources can be performed in a variety of different ways. In some cases, an AR or VR headset may initiate direction-of-arrival (DOA) analysis to determine the location of the sound source. The DOA analysis may include analyzing the intensity, spectrum, and / or arrival time of each sound at the AR / VR device to determine the direction from which the sound originated. In some cases, the DOA analysis may include any suitable algorithm for analyzing the surrounding acoustic environment in which the artificial reality device is located.
[0069] For example, the DOA analysis can be designed to receive an input signal from a microphone and apply a digital signal processing algorithm to the input signal to estimate the direction of arrival. These algorithms can include, for example, the delay-and-sum algorithm, where the input signal is sampled and weighted and delayed versions of the resulting sampled signals are averaged together to determine the direction of arrival. The least mean square (LMS) algorithm can also be implemented to create an adaptive filter. This adaptive filter can then be used, for example, to identify differences in signal strength or differences in arrival time. Then, these differences can be used to estimate the direction of arrival. In another embodiment, the DOA can be determined by transforming the input signal into the frequency domain and selecting specific bins in the time-frequency (TF) domain for processing. Each selected TF bin can be processed to determine whether the bin includes a portion of the audio spectrum with a direct-path audio signal. Then, those bins with a portion of the direct-path signal can be analyzed to identify the angle at which the microphone array receives the direct-path audio signal. Then, the determined angle can be used to identify the direction of arrival of the received input signal. Other algorithms not listed above can also be used, either alone or in combination with the algorithms above, to determine the DOA.
[0070] In some embodiments, different users may perceive a sound source as coming from slightly different locations. This can be a result of each user having a unique head-related transfer function (HRTF), which can be determined by the user's anatomy including the length of the ear canal and the positioning of the eardrum. The artificial reality device can provide alignment and orientation guidelines that the user can follow to customize the sound signals presented to the user based on their unique HRTF. In some embodiments, the artificial reality device can implement one or more microphones to listen for sounds within the user's environment. An AR or VR headset can use various different array transfer functions (e.g., any of the DOA algorithms identified above) to estimate the direction of arrival of a sound. Once the direction of arrival is determined, the artificial reality device can playback the sound to the user according to the user's unique HRTF. Thus, the DOA estimate generated using the array transfer function (ATF) can be used to determine the direction from which the sound will be played. The playback of the sound can be further improved based on how a particular user hears the sound according to the HRTF.
[0071] In addition to or as an alternative to performing DOA estimation, the artificial reality device can perform localization based on information received from other types of sensors. These sensors can include cameras, IR sensors, thermal sensors, motion sensors, GPS receivers, or in some cases sensors that detect the eye movements of the user. For example, as mentioned above, the artificial reality device can include an eye tracker or gaze detector that determines where the user is looking. A user's eyes often look towards a sound source, even briefly. Such a cue provided by the user's eyes can further help to determine the location of the sound source. Other sensors such as cameras, thermal sensors, and IR sensors can also indicate the location of the user, the location of the electronic device, or the location of another sound source. Any or all of the above methods can be used individually or in combination to determine the location of the sound source, and can also be used to update the location of the sound source over time.
[0072] Some embodiments may implement the determined DOA to generate a more customized output audio signal for a user. For example, an "acoustic transfer function" may characterize or define how sound is received from a given location. More specifically, the acoustic transfer function may define the relationship between the parameters of the sound at its source location and the parameters of the sound signal by which it is detected (e.g., detected by a microphone array or detected by the user's ear). The artificial reality device may include one or more acoustic sensors that detect sound within the device's range. A controller of the artificial reality device may estimate the DOA of the detected sound (e.g., using any of the methods identified above), and based on the parameters of the detected sound, may generate an acoustic transfer function specific to the location of the device. Thus, this customized acoustic transfer function may be used to generate a spatialized output audio signal, where the sound is perceived as coming from a specific location.
[0073] In fact, once the location of one or more sound sources is known, the artificial reality device may reproduce (i.e., spatialize) the sound signal to sound as if it is coming from the direction of the sound source. The artificial reality device may apply filters or other digital signal processing that alters the intensity, spectrum, or arrival time of the sound signal. The digital signal processing may be applied in such a way that the sound signal is perceived as originating from the determined location. The artificial reality device may amplify or suppress certain frequencies or alter the time the signal arrives at each ear. In some cases, the artificial reality device may create an acoustic transfer function specific to the location of the device and the direction of arrival of the detected sound signal. In some embodiments, the artificial reality device may reproduce the source signal in a stereo device or a multi-speaker device (e.g., a surround sound device). In such a case, separate and different audio signals may be sent to each speaker. Each of these audio signals may be altered based on the user's HRTF and based on measurements of the user's location and the location of the sound source to sound as if they are coming from the determined location of the sound source. Thus, in this way, the artificial reality device (or the speakers associated with the device) may reproduce the audio signal to sound as if it is originating from a specific location.
[0074] A detailed description of how natural language translation is performed in augmented reality will be provided below with reference to Figures 4 - 11 For example, Figure 4Illustrated is a computing architecture 400 in which many of the embodiments described herein may operate. The computing architecture 400 may include a computer system 401. The computer system 401 may include at least one processor 402 and at least some system memory 403. The computer system 401 may be any type of local or distributed computer system (including a cloud computer system). The computer system 401 may include program modules for performing various different functions. The program modules may be hardware-based, software-based, or may include a combination of hardware and software. Each program module may use or represent computing hardware and / or software to perform a specified function (including those functions described herein below).
[0075] For example, a communication module 404 may be configured to communicate with other computer systems. The communication module 404 may include any wired or wireless communication device capable of receiving data from other computer systems and / or sending data to other computer systems. These communication devices may include radio devices, such as a hardware-based receiver 405, a hardware-based transmitter 406, or a hardware-based transceiver capable of receiving and sending data in combination. The radio device may be a WIFI radio device, a cellular radio device, a Bluetooth radio device, a Global Positioning System (GPS) radio device, or other types of radio devices. The communication module 404 may be configured to interact with databases, mobile computing devices (such as mobile phones or tablets), embedded systems, or other types of computing systems.
[0076] The computer system 401 may also include other modules, including an audio access module 407. The audio access module 407 may be configured to access a real-time (or stored) audio input stream 409. The audio input stream 409 may include one or more words 410 spoken by a speaking user 408. The words spoken by the speaking user 408 may be in a language that the listening user 413 does not (partially or fully) understand. A noise cancellation module 411 of the computer system 401 may generate a noise cancellation signal 412 that is designed to cancel the audio input stream 409 received from the speaking user (i.e., the audio input stream received at a microphone on the computer system 401 or at a microphone on an electronic device associated with the speaking user 408). Thus, in this way, when the speaking user 408 speaks, the listening user may hear the noise cancellation signal 412 that cancels out the words of the speaking user and substantially mutes them.
[0077] The audio processing module 414 of the computer system 401 can be configured to recognize one or more words 410 or phrases spoken by the speaking user 408. It will be appreciated that the words 410 can include single words, word phrases, or complete sentences. These words or groups of words can be recognized and translated individually, or can be recognized and translated together as a phrase or sentence. Thus, although word recognition and translation are mainly described herein in the singular form, it will be understood that these words 410 can be phrases or complete sentences.
[0078] Each word 410 can be recognized in the language spoken by the speaking user 408. Once the audio processing module 414 has recognized one or more of the words of the speaking user, the recognized words 415 can be fed to the translation module 416. The translation module 416 can use a dictionary, database, or other local or online resources to translate the recognized words 415 into a specified language (e.g., the language spoken by the listening user 413). The translated words 417 can then be fed to the speech generator 418. The speech generator 418 can generate spoken words 419 that convey the meaning of the words 410 of the speaking user. The spoken words 419 can be spoken via computer-generated voice, or in some embodiments, the spoken words 419 can be personalized to sound as if they were spoken by the speaking user 408 himself. These spoken words 420 are provided to the playback module 420 of the computer system 401, where they are played back to the listening user 413. Thus, in this way, active noise cancellation and language translation are combined to allow the speaking user to speak in their native language, while the listening user only hears the translated version of the words of the speaking user. These embodiments will be described in more detail below with reference to Figure 5 method 500.
[0079] Figure 5 is a flowchart of an exemplary computer-implemented method 500 for performing natural language translation in AR. Figure 5 The steps shown can be performed by any suitable computer-executable code and / or computing system (including Figure 4 the system shown). In one example, Figure 5 each step shown can represent an algorithm whose structure includes multiple sub-steps and / or is represented by multiple sub-steps, examples of which will be provided in more detail below.
[0080] As Figure 5As shown, at step 510, one or more systems described herein can access an audio input stream that includes words spoken by a user in a first language. For example, the audio access module 407 can access the audio input stream 409. The audio input stream 409 can include one or more words 410 spoken by the user 408 who is speaking. The audio input stream 409 can be real-time or pre-recorded. The words 410 can be spoken in any language.
[0081] Method 500 can then include performing active noise cancellation on the words 410 in the audio input stream 409 so that the spoken words are suppressed before reaching the listening user (step 520). For example, the noise cancellation module 411 of the computer system 401 can generate a noise cancellation signal 412 that is designed to suppress or reduce the intensity of the voice of the user who is speaking or to completely cancel out the voice of the user who is speaking. Thus, when the user 408 who is speaking is talking, the listening user 413 may not hear the words of the user who is speaking, or may only hear a muffled or muted version of the words. When providing a playback to the listening user, the noise cancellation signal 412 can be used within the computer system 401, or the noise cancellation signal 412 can be sent to a device such as a headset or headphones, where the noise cancellation signal is used to mute the voice of the user who is speaking. If desired, the intensity of the noise cancellation signal 412 can be turned up or down, or can be turned off completely.
[0082] Method 500 can further include processing the audio input stream 409 to identify the words 410 spoken by the user 408 who is speaking (step 530), and translating the identified words spoken by the user into a different second language (step 540). The audio processing module 414 can process the audio input stream 409 to identify which words 410 the user 408 who is speaking has spoken. When identifying the words spoken by the user 408 who is speaking, the audio processing module 414 can use speech-to-text (STT) algorithms, dictionaries, databases, machine learning techniques, or other programs or resources. These identified words 415 are then provided to the translation module 416. The translation module 416 translates the identified words 415 into another language. This new language can be a language spoken by or at least understood by the listening user 413. These translated words 417 can then be provided to the speech generator 418 to generate spoken words.
[0083] Figure 5Method 500 may then include using the translated words to generate spoken words in a different second language (step 550), and playing back the generated spoken words in the second language to a listening user (step 560). For example, a speech generator 418 may receive the translated words 417 (e.g., as a digital text string), and may generate spoken words 419 corresponding to the translated words. The speech generator 418 may use a text-to-speech (TTS) algorithm or other resources (including databases, dictionaries, machine learning techniques, or other applications or programs) to generate the spoken words 419 from the translated words 417. The spoken words may sound as if they are being spoken by a computer-generated voice, or may be personalized (as will be further explained below) to sound as if they are being spoken by the user 408 who is speaking. Once the spoken words 419 have been generated, they may be passed to a playback module 420 for playback to the listening user 413. The spoken words may be sent to speakers that are part of the computer system 401 or that are connected to the computer system 401 via a wired or wireless connection. In this way, the listening user 413 will hear the translated spoken words that represent the words 410 of the speaking user, while the noise cancellation module 411 simultaneously ensures that the only thing the listening user hears is the translated spoken words 419.
[0084] In some embodiments, the spoken words 419 may be played back via speakers on an augmented reality (AR), virtual reality (VR), or mixed reality (MR) head-mounted device (e.g., any one of the head-mounted devices 100, 200, or 300 of Figure 1 , 2 or 3, respectively). Although any of these forms of altered reality may be used in any of the embodiments described herein, the embodiments described below will primarily deal with augmented reality. The AR head-mounted device (e.g., Figure 6The head-mounted device 630A worn by the listening user 603 or the head-mounted device 630B worn by the speaking user 606) may include a transparent lens that allows the user to look out and see the external world while also having an internal reflective surface that allows images to be projected and reflected into the user's eyes. Thus, the user can see everything around them but can also see the virtual elements generated by the AR head-mounted device. Additionally, the AR head-mounted device may provide built-in speakers or may have wired or wireless earphones that fit into the user's ears. These speakers or earphones provide audio to the user's ears, whether the audio is music, video game content, movie or video content, speech, or other forms of audio content. Thus, in at least some embodiments herein, the computer system 401 or at least some modules of the computer system 401 may be incorporated into the AR head-mounted device. Thus, the AR head-mounted device can perform noise cancellation, audio processing, translation, speech generation, and playback through the speakers or earphones of the AR head-mounted device.
[0085] As mentioned above, the generated spoken words 419 may be personalized for the speaking user 408 such that the generated spoken words in the second language sound as if they were spoken by the speaking user 408. In many cases, it may be preferred to make the translated spoken words 419 sound as if they were spoken by the speaking user 408 even if the speaking user is unable to speak the language. This personalization provides a familiar tone and feel to the user's words. Personalization makes the words sound less mechanical and robotic and more familiar and personal. The embodiments herein are designed to make the spoken words 419 sound as if they were pronounced and spoken by the speaking user 408.
[0086] In some embodiments, personalizing the generated spoken words 419 may include processing the audio input stream 409 to determine how the speaking user pronounces various words or syllables. For example, each user may pronounce certain words or syllables in a slightly different way. Figure 6The personalization engine 600 in the computing environment 650 can receive audio input 605 from the speaking user 606 and can activate the pronunciation module 601 to determine how the speaking user pronounces their words. A vocal characteristics analyzer 602 can analyze the audio input 605 to determine the speaking user's pitch, word spacing, and other vocal characteristics. The personalization engine 600 can then apply the determined pronunciation, vocal tone, and other vocal characteristics to the spoken words generated in the personalized audio output signal 604. The personalized audio output 604 is then provided to the listening user 603. During playback of the generated spoken words (e.g., via the AR headset 630A), personalization can be dynamically applied to the words being played back when the personalization engine 600 or the computer system 401 determines how the speaking user 606 pronounces the words or syllables.
[0087] In some cases, as Figure 7 shown in the computing environment 700, the speaking user 606 can provide a voice sample or a voice model. For example, the speaking user 606 can provide a voice model 608 that includes the user's pronunciation, pitch, and other vocal characteristics, which can be used to form a voice profile. In such an embodiment, the personalization engine 600 can forego real-time analysis of the speaking user's voice and can use the characteristics and pronunciation in the voice model 608 to personalize the audio output 604. The voice model 608 can include voice samples that can be used to determine how the speaking user pronounces words or syllables before receiving an audio input stream 605 from the speaking user. A voice model interpreter 607 can interpret the data in the voice model and use it when personalizing the spoken words sent to the listening user. In some embodiments, instead of foregoing real-time analysis of the speaking user's voice, the personalization engine 600 can use data from the voice model 608 in combination with real-time analysis of the speaking user's words to further refine the personalization. In such cases, the refinement from the real-time analysis can be added to the user's voice model or can be used to update the user's voice model 608.
[0088] In some cases, the personalization engine 600 can access stored audio data 613 associated with the speaking user 606 and then use the accessed stored audio data to personalize the generated spoken words. The stored audio data 613 can include, for example, pre-recorded words spoken by the speaking user 606. These pre-recorded words can be used to create a voice model or voice profile associated with the user. Then, the voice model can be used to personalize the voice of the speaking user to obtain an audio output 604 sent to the listening user 603. In this way, the generated spoken words played back in a new (translated) language sound as if they were spoken by the speaking user 606.
[0089] In some cases, the personalization engine 600 can parse the words spoken by the speaking user 606. In some examples, the speaking user 606 and the listening user 603 may speak different languages, but these languages share some similar words. For example, some languages may share similar terms related to computing technology that are borrowed directly from English. In such a case, the personalization engine 600 can parse the words spoken by the speaking user 606 and determine that at least one word is spoken in a language understood by the listening user 603. If such a determination is made, the personalization engine can cause the active noise cancellation for the words spoken in the language understood by the listening user to be temporarily suspended. In this way, those words can be heard by the listening user without noise cancellation and without translation.
[0090] Playing back the generated spoken words 604 to the listening user 603 can additionally or alternatively include determining from which direction the speaking user is speaking and spatializing the playback of the generated spoken words to sound as if the spoken words are coming from the determined direction of the speaking user. For example, as Figure 8 shown in the computing environment 800, the speaking user 606 can provide location information 612 associated with the user. The location data can indicate where the user is based on Global Positioning System (GPS) coordinates or can indicate the user's position within a given room, ballroom, stadium, or other venue. The direction identification module 610 of the personalization engine 600 can use the location data 612 to determine from which direction the speaking user is speaking. Then, the spatialization module 611 can spatialize the playback of the generated spoken words in the audio output 604 to sound as if the spoken words are coming from the determined direction of the speaking user 606. The spatialization module 611 can apply various acoustic processing techniques to make the voice of the speaking user sound as if it is behind the listening user 603, or to the right or left of the listening user, or in front of or away from the listening user. Thus, the personalized voice speaking the translated words can not only sound as if it were spoken by the speaking user, but also sound as if it is coming from the exact position of the speaking user relative to the position of the listening user.
[0091] In some embodiments, determining from which direction a speaking user is speaking may include calculating the direction of arrival of sound waves from the speaking user. For example, Figure 8 a direction recognition module 610 may calculate the direction of arrival of sound waves in an audio input 605 from a speaking user 606. The personalization engine 600 may then determine from which direction the speaking user 606 is speaking based on the calculated direction of arrival, and a spatialization module 611 may spatialize the playback of the spoken words generated in the personalized audio output 604 to sound as if the spoken words are coming from the direction of the determined speaking user. In some cases, in addition to receiving location data 612, this direction of arrival calculation may be performed to further refine the location of the speaking user 606. In other cases, the direction of arrival calculation may be performed without receiving location data 612. In this way, the location of the speaking user can be determined without the user sending specific data indicating their current location. For example, a listening user may implement a mobile device with a camera, or may wear an AR headset with a camera. The direction recognition module 610 may analyze the video feed from the camera to determine the direction of the speaking user, and then spatialize the audio based on the determined direction.
[0092] Additionally or alternatively, determining the direction from which a speaking user 606 is speaking may include tracking the movement of the listening user's eyes. For example, the personalization engine 600 (which may be part of or communicate with an AR headset) may include an eye movement tracker. For example, as Figure 9 shown in a computing environment 900, the personalization engine 600 may include an eye movement tracker 615 that generates eye movement data 616. The eye movement tracker 615 may be part of an AR headset 630A and may be configured to track the eyes of a user (e.g., the eyes of a listening user 603) and determine where the user is looking. In most cases, if a speaking user is speaking to a listening user, the listening user will turn to look at the speaking user in order to actively listen to them. In this way, tracking the eye movements of the listening user can provide a clue as to where the speaking user 606 is speaking from. The direction recognition module 610 may then use the eye movement data 616 to determine from which direction the speaking user 606 is speaking based on the tracked eye movements of the listening user. The spatialization module 611 may then spatialize the playback of the generated spoken words in the manner described above to sound as if the spoken words are coming from the direction of the determined speaking user.
[0093] In some embodiments, processing an audio input stream to identify words spoken by a speaking user may include implementing a speech-to-text (STT) program to identify words spoken by the speaking user, and may also include implementing a text-to-speech (TTS) program to generate translated spoken words. As Figure 10 shown in the computing environment 1000 of , words of the speaking user (e.g., in the audio input 1006 from the speaking user 1007) may be fed into the speech-to-text module 1005, where the words are converted into text or some other digital representation of the words. The translation module 1004 may then use the words in text form to perform a translation from one language to another. Once the translation is performed, the text-to-speech module 1003 may convert the written words into speech. This speech may be included in the audio output 1002. The audio output 1002 may then be sent to the listening user 1001. Thus, some embodiments may use STT and TTS to perform conversions between speech and text and the conversion of text back to speech.
[0094] Figure 11 The computing environment 1100 of shows an example where multiple speaking users are speaking simultaneously. Each speaking user (e.g., the speaking user 1105 or 1107) may provide an audio input stream (e.g., the audio input streams 1104 or 1106, respectively), which includes words spoken by two different users. The voice discrimination module 1103 (which may be Figure 6 part of the personalization engine 600 of and / or Figure 4 part of the computer system 401 of ) may then distinguish between the two speaking users 1105 and 1107 based on different voice patterns or other voice characteristics. The voice discrimination module 1103 may then generate spoken words in the voice output 1102 for one speaking user (e.g., 1105), while performing active noise cancellation for the two speaking users. In this way, the listening user 1101 (who does not understand the languages spoken by the two users) can still receive a translated version of the words of the speaking user 1105.
[0095] In some embodiments, an audio input stream 1106 from another speaking user 1107 can be stored in a data storage. The stored audio stream can then be parsed and translated. Then, when the speaking user 1105 finishes speaking, the voice discrimination module 1103 can cause the stored and translated words of the speaking user 1107 to be converted into spoken words. In some cases, the voice discrimination module 1103 can operate according to a policy that indicates that if two (or more) speaking users are speaking, one speaker (possibly based on eye tracking information to see which speaking user the listening user is looking at) will be selected and the words from the other speaking users will be stored. Then, once the voice discrimination module 1103 determines that the first speaking user has stopped speaking for a specified amount of time, the spoken words generated for the other speaking users will be sequentially played back to the listening user. In certain cases, the policy may be biased towards certain speaking users based on the identity of the speaking users. Thus, even if the listening user 1101 is in a group of people, the system here can focus on a single speaking user (e.g., based on the voice characteristics of that user) or a group of users and record the audio from these users. Then, the audio can be converted to text, translated, converted back to speech and played back to the listening user 1101.
[0096] In some cases, when multiple users are speaking, Figure 4 the personalization engine 400 can create a voice model for each speaking user, or can create a voice model of a second speaking user while the first speaking user is speaking. The personalization engine 400 can also personalize the spoken words generated for each of the users speaking simultaneously. In this way, the spoken words generated in the new (translated) language can sound as if they are from the voices of each different speaking user. Thus, whether two people are having a one-on-one conversation or chatting in a large group, the embodiments here can operate to mute the audio from the speaker, translate the speaker's words, and play back to the listening user a personalized oral translation of the speaker's words.
[0097] In addition, a corresponding system for performing natural language translation in AR can include several modules stored in a memory, including an audio access module that accesses an audio input stream that includes words spoken by a user in a first language. The system can also include a noise cancellation module that performs active noise cancellation on the words in the audio input stream such that the spoken words are suppressed or are substantially inaudible to a listening user. The system can also include an audio processing module that processes the audio input stream to identify the words spoken by the user. A translation module can translate the identified words spoken by the user into a different second language, and a speech generator can generate spoken words in the different second language using the translated words. A playback module can then playback the generated spoken words in the second language to the listening user.
[0098] In some examples, the above method can be encoded as computer-readable instructions on a computer-readable medium. For example, the computer-readable medium can include one or more computer-executable instructions that, when executed by at least one processor of a computing device, can cause the computing device to access an audio input stream that includes words spoken by a user in a first language, perform active noise cancellation on the words in the audio input stream such that the spoken words are suppressed or are substantially inaudible to a listening user, process the audio input stream to identify the words spoken by the user, translate the identified words spoken by the user into a different second language, generate spoken words in the different second language using the translated words, and playback the generated spoken words in the second language to the listening user.
[0099] Thus, two (or more) users can talk to each other, each speaking their own language. The speech of each user is muted to the other user and is translated and spoken back to the listening user in the voice of the speaking user. Thus, users speaking different languages can talk to each other freely, only hearing personalized translated speech. This can greatly assist users in communicating with each other, especially when they do not speak the same language.
[0100] As detailed above, the computing devices and systems described and / or illustrated herein broadly represent any type or form of computing device or system capable of executing computer-readable instructions (such as those included in the modules described herein). In their most basic configuration, these computing devices can each include at least one memory device and at least one physical processor.
[0101] In some examples, the term "memory device" generally refers to any type or form of volatile or non-volatile storage device or medium capable of storing data and / or computer-readable instructions. In one example, a memory device may store, load, and / or maintain one or more of the modules described herein. Examples of memory devices include, but are not limited to, random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid state drive (SSD), optical disk drive, cache, variations or combinations of one or more of these components, or any other suitable storage memory.
[0102] In some examples, the term "physical processor" generally refers to any type or form of hardware-implemented processing unit capable of interpreting and / or executing computer-readable instructions. In one example, a physical processor may access and / or modify one or more of the modules stored in the memory device described above. Examples of physical processors include, but are not limited to, microprocessors, microcontrollers, central processing units (CPUs), field-programmable gate arrays (FPGAs) implementing soft-core processors, application-specific integrated circuits (ASICs), portions of one or more of these components, variations or combinations of one or more of these components, or any other suitable physical processor.
[0103] Although shown as separate elements, the modules described and / or illustrated herein may represent parts of a single module or application. Additionally, in certain embodiments, one or more of these modules may represent one or more software applications or programs that, when executed by a computing device, may cause the computing device to perform one or more tasks. For example, one or more of the modules described and / or illustrated herein may represent modules stored and configured to run on one or more of the computing devices or systems described and / or illustrated herein. One or more of these modules may also represent all or part of one or more dedicated computers configured to perform one or more tasks.
[0104] Furthermore, one or more of the modules described herein may convert data, physical devices, and / or representations of physical devices from one form to another. For example, one or more of the modules described herein may receive data to be converted, convert the data, output the result of the conversion to perform a function, use the result of the conversion to perform a function, and store the result of the conversion to perform a function. Additionally or alternatively, one or more of the modules described herein may convert any other part of a processor, volatile memory, non-volatile memory, and / or physical computing device from one form to another by executing on a computing device, storing data on a computing device, and / or otherwise interacting with a computing device.
[0105] In some embodiments, the term "computer-readable medium" generally refers to any form of device, carrier, or medium that can store or carry computer-readable instructions. Examples of computer-readable media include, but are not limited to, transmission-type media (e.g., carrier waves) and non-transitory media such as magnetic storage media (e.g., hard disk drives, tape drives, and floppy disks), optical storage media (e.g., compact discs (CDs), digital video discs (DVDs), and Blu-ray discs), electronic storage media (e.g., solid-state drives and flash media), and other distribution systems.
[0106] Embodiments of the present disclosure may include or be implemented in conjunction with an artificial reality system. Artificial reality is a form of reality that has been adjusted in some way before being presented to a user, and may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or content generated in combination with captured (e.g., real-world) content. Artificial reality content may include video, audio, tactile feedback, or some combination thereof, any of which may be presented in a single channel or in multiple channels (e.g., stereoscopic video that produces a three-dimensional effect for a viewer). Additionally, in some embodiments, artificial reality may also be associated with applications, products, accessories, services, or some combination thereof for, e.g., creating content in artificial reality and / or being otherwise used in artificial reality (e.g., performing activities in artificial reality). An artificial reality system that provides artificial reality content may be implemented on various platforms, including a head-mounted display (HMD) connected to a host computer system, a standalone HMD, a mobile device or computing system, or any other hardware platform capable of providing artificial reality content to one or more viewers.
[0107] The order of process parameters and steps described and / or illustrated herein is given only as an example and may vary as needed. For example, although the steps shown and / or described herein may be shown or discussed in a particular order, these steps do not necessarily need to be performed in the order shown or discussed. The various exemplary methods described and / or illustrated herein may also omit one or more steps described or illustrated herein, or include additional steps other than those disclosed.
[0108] The foregoing description is provided to enable other technicians in the art to best utilize the various aspects of the exemplary embodiments disclosed herein. The exemplary description is not intended to be exhaustive or limited to any precise form. Many modifications and variations are possible without departing from the spirit and scope of the present disclosure. The embodiments disclosed herein should be considered illustrative in all respects and not restrictive. When determining the scope of the present disclosure, reference should be made to the appended claims and their equivalents.
[0109] Unless otherwise indicated, as used in the specification and claims, the terms "connected to" and "coupled to" (and their derivatives) shall be construed to allow both direct and indirect (i.e., via other elements or components) connection. Further, as used in the specification and claims, the term "a" or "an" shall be construed to mean "at least one of...". Finally, for ease of use, the terms "including" and "having" (and their derivatives) as used in the specification and claims are interchangeable with the word "comprising" and have the same meaning as the word "comprising".
Claims
1. A computer-implemented method, comprising: Accessing an audio input stream that includes one or more words spoken by a speaking user in a first language and one or more words spoken by another speaking user; Performing active noise cancellation on one or more words in the audio input stream, including generating a noise cancellation signal configured to substantially cancel out the audio input stream received from the speaking user; Applying the generated noise cancellation signal to the audio input stream to suppress the spoken words of the speaking user; Processing the audio input stream to identify one or more words spoken by the speaking user; Determining that a first plurality of words spoken by the speaking user are in a language not understood by a listening user and that a second plurality of words spoken by the speaking user are in a language understood by the listening user; Pausing active noise cancellation for the second plurality of words spoken in the language understood by the listening user; Translating the identified first plurality of words spoken by the speaking user into a different second language; Generating spoken words in the different second language using the translated words while performing active noise cancellation on the speaking user and the other speaking user; Generating spoken words for words spoken by the other speaking user suppressed by the active noise cancellation; Storing the generated spoken words for the other speaking user until the speaking user has stopped speaking for a specified amount of time; And Playing back the generated spoken words in the second language to the listening user, wherein the audio input stream provided to the listening user includes a mixture of the original speech of the speaking user, the mixture including the second plurality of words during which active noise cancellation was paused and the generated spoken words played back in the second language, and Wherein, after the speaking user has stopped speaking for a specified amount of time, the stored spoken words generated for the other speaking user are played back to the listening user in sequence.
2. The computer-implemented method according to claim 1, wherein, The generated spoken words are personalized for the speaking user such that the generated spoken words in the second language sound as if they were spoken by the speaking user.
3. The computer-implemented method according to claim 2, wherein, Making the generated spoken words personalized further includes: Processing the audio input stream to determine how the speaking user pronounces one or more words or syllables; and Applying the determined pronunciation to the generated spoken words.
4. The computer-implemented method according to claim 3, wherein, When the computer determines how the speaking user pronounces the words or syllables, the personalization is dynamically applied to the words being played back during the playback of the generated spoken words.
5. The computer-implemented method according to claim 3, wherein, The speaking user provides one or more voice samples, and before receiving the audio input stream, the computer uses the voice samples to determine how the speaking user pronounces one or more of the words or syllables.
6. The computer-implemented method according to claim 1, wherein, Playing back the generated spoken words to the listening user further includes: Determining from which direction the speaking user is speaking; and Spatialize the playback of the generated spoken words to sound as if the spoken words are coming from the direction of the identified speaking user.
7. The computer-implemented method according to claim 6, wherein, Determining from which direction the speaking user is speaking further includes: Receiving location data of a device associated with the speaking user; Determining from which direction the speaking user is speaking based on the received location data; and Spatialize the playback of the generated spoken words to sound as if the spoken words are coming from the direction of the identified speaking user.
8. The computer-implemented method according to claim 6, wherein, Determining from which direction the speaking user is speaking further includes: Calculating the direction of arrival of sound waves from the speaking user; Determining from which direction the speaking user is speaking based on the calculated direction of arrival; and Spatialize the playback of the generated spoken words to sound as if the spoken words are coming from the direction of the identified speaking user.
9. The computer-implemented method according to claim 6, wherein, Determining from which direction the speaking user is speaking further includes: Tracking the movement of the listening user's eyes; Determining from which direction the speaking user is speaking based on the tracked movement of the listening user's eyes; and Spatialize the playback of the generated spoken words to sound as if the spoken words are coming from the direction of the identified speaking user.
10. The computer-implemented method according to claim 1, wherein, Processing the audio input stream to identify one or more words spoken by the speaking user includes implementing a speech-to-text (STT) program to identify the words spoken by the speaking user, and implementing a text-to-speech (TTS) program to generate translated spoken words.
11. A system, comprising: At least one physical processor; A physical memory including computer-executable instructions that, when executed by the physical processor, cause the physical processor to: Access an audio input stream that includes one or more words spoken by a speaking user in a first language and one or more words spoken by another speaking user; Perform active noise cancellation on one or more words in the audio input stream, including generating a noise cancellation signal configured to substantially cancel out the audio input stream received from the speaking user; Apply the generated noise cancellation signal to the audio input stream to suppress the spoken words of the speaking user; Process the audio input stream to identify one or more words spoken by the speaking user; Determine that a first plurality of words spoken by the speaking user are spoken in a language not understood by a listening user, and that a second plurality of words spoken by the speaking user are spoken in a language understood by the listening user; Pause active noise cancellation on the second plurality of words spoken in the language understood by the listening user; Translate the identified first plurality of words spoken by the speaking user into a different second language; Generate spoken words in the different second language using the translated words while performing active noise cancellation on the speaking user and the other speaking user; Generate spoken words for words spoken by the other speaking user that are suppressed by the active noise cancellation; Store the spoken words generated by the additional speaking user until the speaking user has stopped speaking for a specified amount of time; and Play back the generated spoken words in the second language to the listening user, wherein the audio input stream provided to the listening user includes a mixture of the original speech of the speaking user, the mixture including the second plurality of words during which the active noise cancellation is paused and the generated spoken words played back in the second language, and wherein, after the speaking user has stopped speaking for a specified amount of time, the stored spoken words generated for the additional speaking user are sequentially played back to the listening user.
12. The system according to claim 11, further comprising: Download a voice profile associated with the speaking user; and Use the downloaded voice profile associated with the speaking user to personalize the generated spoken words such that the generated spoken words in the second language played back sound as if spoken by the speaking user.
13. The system according to claim 11, further comprising: Access one or more portions of the stored audio data associated with the speaking user; and Use the accessed stored audio data to personalize the generated spoken words such that the generated spoken words played back in the second language sound as if spoken by the speaking user.
14. The system according to claim 11, further comprising: Parse the words spoken by the speaking user; Determine that at least one of the words is spoken in a language understood by the listening user; and Pause the active noise cancellation for the words spoken in a language understood by the listening user.
15. The system according to claim 11, further comprising: Determine that the audio input stream includes words spoken by at least two different speaking users; Distinguish the two speaking users based on one or more voice patterns; and Generate spoken words for the first speaking user while performing active noise cancellation for both speaking users.
16. The system according to claim 15, further comprising: Store the spoken words generated for the second speaking user until the first speaking user has stopped speaking for a specified amount of time; and Play back the spoken words generated for the second speaking user.
17. The system according to claim 16, further comprising personalizing the generated spoken words for each of the two speaking users such that the generated spoken words in the second language sound as if from the voice of each speaking user.
18. The system according to claim 11, wherein, At least a portion of the computer-executable instructions stored on the physical memory is processed by at least one remote physical processor separate from the system.
19. The system according to claim 18, wherein, One or more policies indicate when and which portions of the computer-executable instructions will be processed on the at least one remote physical processor separate from the system.
20. A non - transitory computer - readable medium that includes one or more computer - executable instructions that, when executed by at least one processor of a computing device, cause the computing device to: Access an audio input stream that includes one or more words spoken by a speaking user in a first language and one or more words spoken by another speaking user; Perform active noise cancellation on one or more words in the audio input stream, including generating a noise cancellation signal that is configured to substantially cancel out the audio input stream received from the speaking user; Apply the generated noise cancellation signal to the audio input stream to suppress the spoken words of the speaking user; Process the audio input stream to identify one or more words spoken by the speaking user; Determine that a first plurality of words spoken by the speaking user are spoken in a language that the listening user does not understand and that a second plurality of words spoken by the speaking user are spoken in a language that the listening user understands; Pause active noise cancellation for the second plurality of words spoken in the language that the listening user understands; Translate the first plurality of words identified as being spoken by the speaking user into a different second language; Generate spoken words in the different second language using the translated words while performing active noise cancellation on the speaking user and the other speaking user; Generate spoken words for words spoken by the other speaking user that are suppressed by the active noise cancellation; Store the spoken words generated for the other speaking user until the speaking user has stopped speaking for a specified amount of time; And Play back the generated spoken words in the second language to the listening user, wherein the audio input stream provided to the listening user includes a mixture of the original speech of the speaking user, the mixture including the second plurality of words during which active noise cancellation is paused and the generated spoken words played back in the second language, and Wherein, after the speaking user has stopped speaking for a specified amount of time, the stored spoken words generated for the other speaking user are played back to the listening user in sequence.
21. A computer - implemented method, comprising: Access an audio input stream that includes one or more words spoken by a speaking user in a first language and one or more words spoken by another speaking user; Perform active noise cancellation on one or more words in the audio input stream, including generating a noise cancellation signal that is configured to substantially cancel out the audio input stream received from the speaking user; Apply the generated noise cancellation signal to the audio input stream to suppress the spoken words of the speaking user; Process the audio input stream to identify one or more words spoken by the speaking user; Determine that the first plurality of words spoken by the speaking user are spoken in a language that the listening user does not understand, and that the second plurality of words spoken by the speaking user are spoken in a language that the listening user understands; Pause active noise cancellation for the second plurality of words spoken in the language that the listening user understands; Translate the identified first plurality of words spoken by the speaking user into a different second language; Generate spoken words in the different second language using the translated words, while performing active noise cancellation on the speaking user and the other speaking user; Generate spoken words for the words spoken by the other speaking user that are suppressed by the active noise cancellation; Store the spoken words generated for the other speaking user until the speaking user has stopped speaking for a specified amount of time; and Play back the generated spoken words in the second language to the listening user, wherein the audio input stream provided to the listening user includes a mixture of the original speech of the speaking user, the mixture including the second plurality of words during which the active noise cancellation was paused and the generated spoken words played back in the second language, and wherein, after the speaking user has stopped speaking for a specified amount of time, the stored spoken words generated for the other speaking user are played back to the listening user in sequence.
22. The computer-implemented method according to claim 21, wherein, The generated spoken words are personalized for the speaking user such that the generated spoken words in the second language sound as if they were spoken by the speaking user.
23. The computer-implemented method according to claim 22, wherein, Making the generated spoken words personalized further includes: Processing the audio input stream to determine how the speaking user pronounces one or more words or syllables; and Applying the determined pronunciation to the generated spoken words.
24. The computer-implemented method according to claim 23, wherein, When the computer determines how the speaking user pronounces the word or syllable, during playback of the generated spoken word, the personalization is dynamically applied to the word being played back; and / or wherein, the speaking user provides one or more voice samples, and before receiving the audio input stream, the computer uses the one or more voice samples to determine how the speaking user pronounces one or more of the words or syllables.
25. The computer-implemented method according to any one of claims 21 to 24, wherein, Playing back the generated spoken words to the listening user further includes: Determining from which direction the speaking user is speaking; and Spatializing the playback of the generated spoken words to sound as if the spoken words are coming from the determined direction of the speaking user.
26. The computer-implemented method according to claim 25, wherein, Determining from which direction the speaking user is speaking further includes: Receiving position data of a device associated with the speaking user; Based on the received position data, determining from which direction the speaking user is speaking; and Spatializing the playback of the generated spoken words to sound as if the spoken words are coming from the determined direction of the speaking user; and / or wherein, determining from which direction the speaking user is speaking further includes: Calculating the direction of arrival of sound waves from the speaking user; Determine from which direction the speaking user is speaking based on the calculated direction of arrival; and Spatialize the playback of the generated spoken words to sound as if the spoken words are coming from the determined direction of the speaking user; and / or wherein determining from which direction the speaking user is speaking further comprises: Tracking the movement of the listening user's eyes; Determine from which direction the speaking user is speaking based on the tracked movement of the listening user's eyes; and Spatialize the playback of the generated spoken words to sound as if the spoken words are coming from the determined direction of the speaking user.
27. The computer-implemented method according to any one of claims 21 to 24 and 26, wherein, Processing the audio input stream to identify one or more words spoken by the speaking user includes implementing a speech-to-text (STT) program to identify the words spoken by the speaking user, and implementing a text-to-speech (TTS) program to generate translated spoken words.
28. A system, comprising: At least one physical processor; A physical memory including computer-executable instructions that, when executed by the physical processor, cause the physical processor to: Access an audio input stream that includes one or more words spoken by a speaking user in a first language and one or more words spoken by another speaking user; Perform active noise cancellation on one or more words in the audio input stream, including generating a noise cancellation signal configured to substantially cancel out the audio input stream received from the speaking user; Process the audio input stream to identify one or more words spoken by the speaking user; Determine that a first plurality of words spoken by the speaking user are in a language not understood by the listening user, and that a second plurality of words spoken by the speaking user are in a language understood by the listening user; Pause active noise cancellation for the second plurality of words spoken in the language understood by the listening user; Translate the identified first plurality of words spoken by the speaking user into a different second language; Generate spoken words in the different second language using the translated words while performing active noise cancellation on the speaking user and the other speaking user; Generate spoken words for words spoken by the other speaking user that are suppressed by the active noise cancellation; Store the spoken words generated for the other speaking user until the speaking user has stopped speaking for a specified amount of time; and Playback the generated spoken words in the second language to the listening user, wherein the audio input stream provided to the listening user includes a mixture of the original speech of the speaking user, the mixture including the second plurality of words for which active noise cancellation was paused during that time and the generated spoken words played back in the second language, and wherein, after the speaking user has stopped speaking for a specified amount of time, the stored spoken words generated for the other speaking user are played back to the listening user in sequence.
29. The system according to claim 28, further comprising: Download a voice profile associated with the speaking user; and Use the downloaded voice profile associated with the speaking user to personalize the generated spoken words such that the generated spoken words in the second language when played back sound as if spoken by the speaking user.
30. The system according to claim 28 or 29, further comprising: Access one or more portions of stored audio data associated with the speaking user; and Use the accessed stored audio data to personalize the generated spoken words such that the generated spoken words in the second language when played back sound as if spoken by the speaking user.
31. The system according to claim 28 or 29, further comprising: Parse the words spoken by the speaking user; Determine that at least one of the words is spoken in a language understood by the listening user; and Pause active noise cancellation for the words spoken in a language understood by the listening user.
32. The system according to claim 28 or 29, further comprising: Determine that the audio input stream includes words spoken by at least two different speaking users; Distinguish the two speaking users according to one or more voice patterns; and Generate spoken words for the first speaking user while performing active noise cancellation for both speaking users; Optionally, further comprising: Store the spoken words generated for the second speaking user until the first speaking user has stopped speaking for a specified amount of time; and Play back the spoken words generated for the second speaking user; Optionally, further comprising personalizing the generated spoken words for each of the two speaking users such that the generated spoken words in the second language sound as if from the voice of each speaking user.
33. The system according to claim 28 or 29, wherein, At least a portion of the computer-executable instructions stored on the physical memory is processed by at least one remote physical processor separate from the system; Optionally, wherein one or more policies indicate when and which portions of the computer-executable instructions will be processed on the at least one remote physical processor separate from the system.
34. A non-transitory computer-readable medium comprising one or more computer-executable instructions that, when executed by at least one processor of a computing device, cause the computing device to perform the method according to any one of claims 21 to 27 or: Access an audio input stream that includes one or more words spoken by a speaking user in a first language and one or more words spoken by another speaking user; Perform active noise cancellation on one or more words in the audio input stream, including generating a noise cancellation signal configured to substantially cancel out the audio input stream received from the speaking user; Apply the generated noise cancellation signal to the audio input stream to suppress the spoken words of the speaking user; Process the audio input stream to identify one or more words spoken by the speaking user; Determine that the first plurality of words spoken by the speaking user are in a language not understood by the listening user, and that the second plurality of words spoken by the speaking user are in a language understood by the listening user; Pause active noise cancellation for the second plurality of words spoken in the language understood by the listening user; Translate the identified first plurality of words spoken by the speaking user into a different second language; Generate spoken words in the different second language using the translated words, while performing active noise cancellation for the speaking user and the other speaking user; Generate spoken words for words spoken by the other speaking user that are suppressed by the active noise cancellation; Store the spoken words generated for the other speaking user until the speaking user has stopped speaking for a specified amount of time; and Play back the generated spoken words in the second language to the listening user, wherein the audio input stream provided to the listening user includes a mixture of the original speech of the speaking user, the mixture including the second plurality of words during which the active noise cancellation is paused and the generated spoken words played back in the second language, and wherein, after the speaking user has stopped speaking for a specified amount of time, the stored spoken words generated for the other speaking user are sequentially played back to the listening user.
Citation Information
Patent Citations
System for computer-assisted communication and / or computer-assisted human analysis
GB201603610D0
Control method of interpretation apparatus, control method of interpretation server, control method of interpretation system and user terminal
US20140303958A1
Audio source spatialization
US20160165350A1
Hearing assistance with automated speech transcription
WO2017142775A1