Methods, apparatuses, and computer program products for providing large language models that facilitate real-time translation

CN122514760APending Publication Date: 2026-08-04META PLATFORMS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
META PLATFORMS INC
Filing Date
2025-01-08
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

取决于回复,该过程可能需要时间,并且在一些示例中,翻译可能效率低下

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122514760A_ABST
    Figure CN122514760A_ABST
Patent Text Reader

Abstract

A system and method for translating speech are provided. The system can detect audio signals associated with speech data of one or more first users and determine that the speech data is associated with a first language. The system can translate words in the speech data associated with the first language into other words associated with a second language. The system can present the other words translated into the second language as text items to a display of a device of one or more second users, or output the other words in the second language as audio to one or more second users. In response to the speech data, the system can generate a first set of responses in the first language presented via the display and a second set of corresponding other responses in the second language. The other responses may include translations of the responses in the first language into the second language.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims priority to U.S. Provisional Application No. 63 / 619,136, entitled “Large Language Models in LiveTranslation”, filed on January 9, 2024. Technical Field

[0002] The various examples disclosed herein generally relate to methods, apparatus, and computer program products for speech translation systems. Background Technology

[0003] Electronic devices are constantly evolving and developing, providing users with flexibility and adaptability. As electronic devices become increasingly adaptable, users carry their devices with them every day, and as the world becomes an increasingly global community, users who may speak different languages ​​may interact frequently.

[0004] Typically, communication between users or people who speak different languages ​​may involve a translator using a device (via a voice translation system) to translate one language into the user's language, or text being input into a device (via a voice translation system) to perform the translation. Each of these methods can disrupt the fluency or sense of connection in the conversation between the people attempting to communicate. For example, a voice translation system that translates a first person's speech into a second person's speech may allow the second person to understand the first, but in many cases, the second user may now have to speak or type their language into the voice translation system to receive a translation into the first person's language. Depending on the response, this process can take time, and in some examples, the translation may be inefficient. One example where this is particularly evident can be artificial reality devices.

[0005] Artificial reality is a form of reality that has been adjusted in some way before being presented to a user. Artificial reality can include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality (HR), or some combination and / or derivative thereof. Artificial reality content can include entirely computer-generated content or computer-generated content combined with captured (e.g., real-world) content. Artificial reality content can include video, audio, haptic feedback, or some combination thereof, any of which can be presented in single-channel or multi-channel (e.g., stereoscopic video that provides a three-dimensional (3D) effect to the viewer). Furthermore, in some embodiments, artificial reality can also be associated with applications, products, accessories, services, or some combination thereof, which are used, for example, to create content in artificial reality and / or otherwise use in artificial reality (e.g., to perform actions in artificial reality). Head-mounted displays (HMDs), which include one or more near-eye displays, are typically used to present visual content to users for use in artificial reality applications.

[0006] Given the aforementioned drawbacks, it may be beneficial to provide an efficient and reliable speech translation method that facilitates more fluent dialogue. Summary of the Invention

[0007] Methods and systems for voice translation between users via an HMD or other electronic device (e.g., a smartphone, tablet, smartwatch, or any electronic device capable of communicating with an HMD) are described.

[0008] According to one aspect of the present invention, a method is provided, the method comprising: detecting one or more audio signals associated with speech data of at least one first user, and determining that the speech data is associated with a first language; translating one or more words of the speech data associated with the first language into one or more other words associated with a second language different from the first language; presenting the one or more other words translated into the second language as one or more text items to a display of a head-mounted device (HMD) of at least one second user, or outputting the one or more other words of the second language as audio content to at least one second user; and generating, in response to the speech data, a first set of one or more responses in the first language and a second set of corresponding one or more other responses in the second language presented via the display, wherein the one or more other responses include translations of the one or more responses in the first language into the second language.

[0009] Optionally, the method further includes: in response to detecting the selection of a first response among one or more other responses, outputting audio data of one or more words associated with the selected first response, so that at least one second user can speak the one or more words to at least one first user.

[0010] Optionally, when one or more other words translated into a second language are input as one or more text items and displayed on the HMD's screen, one or more content items of the real-world environment can be viewed via the screen.

[0011] Optionally, while one or more content items in the real-world environment can be viewed on the HMD's display, a first set of one or more replies and a second set of one or more corresponding other replies can also be viewed on the HMD's display.

[0012] Optionally, the method further includes generating one or more pronunciations of one or more other words translated into a second language.

[0013] Optionally, the method further includes: outputting audio data of one or more pronunciations from the HMD so that at least one second user can speak one or more pronunciations.

[0014] Optionally, the method further includes: in response to detecting one or more audio signals and based on at least one setting associated with the device, determining that the second language is at least one language spoken by at least one second user, so as to facilitate the translation of one or more words in the speech data.

[0015] Optionally, presentation includes presenting one or more other words translated into a second language as one or more text items to the display, overlaying visual content items of the real-world environment captured by the display.

[0016] Alternatively, HMD may include smart glasses, augmented reality devices, or virtual reality devices.

[0017] Optionally, the HMD is configured to present one or more augmented reality content items, one or more virtual reality content items, one or more mixed reality content items, or a combination thereof, via a display.

[0018] Optionally, the method further includes: determining that the voice data of at least one first user is associated with an active or current dialogue with at least one second user; and providing a display with translations of a first set of one or more other words or one or more replies and a second set of corresponding one or more other replies during the dialogue, so that the first user can use the first set of one or more other words or one or more replies or the second set of one or more replies to respond to the voice data of at least one first user in a first language.

[0019] According to another aspect of the present invention, an apparatus is provided, comprising: one or more processors; and at least one memory storing a plurality of instructions, which, when executed by the one or more processors, cause the apparatus to: detect one or more audio signals associated with voice data of at least one first user and determine that the voice data is associated with a first language; translate one or more words of the voice data associated with the first language into one or more other words associated with a second language different from the first language; present the one or more other words translated into the second language as one or more text items to a display of the apparatus of the at least one second user, or output the one or more other words of the second language as audio content to the at least one second user; and, in response to the voice data, generate a first set of one or more responses in the first language presented via the display and a second set of corresponding one or more other responses in the second language, wherein the one or more other responses include translations of the one or more responses in the first language into the second language.

[0020] Alternatively, the device may include a head-mounted display, smart glasses, an augmented reality device, or a virtual reality device.

[0021] Optionally, when one or more processors further execute the plurality of instructions, the device is configured to: in response to detecting the selection of a first response among one or more other responses, output audio data of one or more words associated with the selected first response, so that at least one second user can speak the one or more words to at least one first user.

[0022] Optionally, when one or more other words translated into a second language are input as one or more text items and presented on the display, one or more content items of the real-world environment can be viewed via the display.

[0023] Optionally, while one or more content items in the real-world environment can be viewed on the display, a first set of one or more replies and a second set of one or more corresponding other replies can also be viewed on the display.

[0024] Optionally, when one or more processors further execute the plurality of instructions, the device is configured to generate one or more pronunciations of one or more other words translated into a second language.

[0025] Optionally, when one or more processors further execute the plurality of instructions, the device is configured to: in response to detecting one or more audio signals and based on at least one setting associated with the device, determine that the second language is at least one language spoken by at least one second user, so as to translate one or more words of the speech data.

[0026] According to another aspect of the invention, a non-transitory computer-readable medium is provided, the non-transitory computer-readable medium storing a plurality of instructions, which, when executed, cause: detecting one or more audio signals associated with speech data of at least one first user, and determining that the speech data is associated with a first language; translating one or more words of the speech data associated with the first language into one or more other words associated with a second language different from the first language; presenting the one or more other words translated into the second language as one or more text items to a display of a head-mounted device of at least one second user, or outputting the one or more other words of the second language as audio content to at least one second user; and, in response to the speech data, generating a first set of one or more responses in the first language presented via the display and a second set of corresponding one or more other responses in the second language, wherein the one or more other responses include translations of the one or more responses in the first language into the second language.

[0027] Optionally, when executed, the plurality of instructions also cause: in response to determining that a selection of a first response among one or more other responses has been detected, audio data of one or more words associated with the selected first response is output so that at least one second user can speak one or more words to at least one first user.

[0028] In various examples, the system and / or method may receive an audio signal associated with speech spoken by a person / user in their first language via a device associated with the user. One or more text items associated with the audio signal may be identified. The text may be displayed in a second language via one or more graphical user interfaces. One or more machine learning models may develop a list of responses associated with responses to one or more text items in the first language. One or more machine learning models may utilize neural networks to establish associations between one or more text items, previous responses to one or more similar text items, and / or one or more contexts associated with one or more text items. One or more machine learning models and / or artificial intelligence (AI) models may generate the list of responses. The list of responses may be provided to the user in both the first and second languages ​​via one or more graphical user interfaces. Each of the multiple responses may include one or more pronunciations associated with each or more words in the multiple responses. The one or more pronunciations may be associated with the first language to assist / help the user in pronouncing one or more translated phrases or one or more words, etc. For example, the user may select a response, and the selected response may be output as audio via the HMD's audio device in a staccato manner to help the user respond by speaking fluently to a person who speaks the first language, for example, during a real-time conversation.

[0029] One or more machine learning models can be trained based on statistical models to analyze relationships between large amounts of data, learning patterns, and / or the following: words; phrases; natural language patterns; and / or previously selected responses associated with one or more users. In various examples, one or more machine learning models can leverage one or more neural networks to develop associations between received audio signals and associated text, natural language patterns, and / or previously selected responses. As described above, the list of responses for the user can be generated and provided by one or more graphical user interfaces of a device (e.g., one or more HMDs, smartphones, tablets, smartwatches, computing devices, or communication devices, etc.). The list of responses can be in the form of customizable text, one or more images, one or more videos, or any combination thereof.

[0030] In one example aspect of this disclosure, a method is provided. The method may include: detecting one or more audio signals associated with speech data of at least one first user, and determining that the speech data is associated with a first language. The method may include: translating one or more words of the speech data associated with the first language into one or more other words associated with a second language different from the first language. The method may include: presenting the one or more other words translated into the second language as one or more text items to a display of a head-mounted device of at least one second user, or outputting the one or more other words in the second language as audio content to at least one second user. The method may include: in response to the speech data, generating a first set of one or more responses in the first language presented via the display, and a second set of corresponding one or more other responses in the second language. The one or more other responses include translations of the one or more responses in the first language into the second language.

[0031] In another example aspect of this disclosure, an apparatus is provided. The apparatus may include one or more processors and a memory including computer program code instructions. The memory and computer program code instructions are configured to use at least one of the one or more processors to cause the apparatus to perform at least the following operations: detecting one or more audio signals associated with speech data of at least one first user, and determining that the speech data is associated with a first language. The memory and computer program code are further configured to use the one or more processors to cause the apparatus to translate one or more words of the speech data associated with the first language into one or more other words associated with a second language different from the first language. The memory and computer program code are further configured to use the one or more processors to cause the apparatus to present the one or more other words translated into the second language as one or more text items to a display of the apparatus of at least one second user, or to output one or more other words in the second language as audio content to at least one second user. The memory and computer program code are further configured to use the one or more processors to cause the apparatus, in response to the speech data, to generate a first set of one or more responses in the first language and a second set of corresponding one or more other responses in the second language presented via the display. The one or more other responses include translations of the one or more responses in the first language into the second language.

[0032] In another example aspect of this disclosure, a computer program product is provided. The computer program product may include at least one non-transitory computer-readable medium including computer-executable program code instructions stored therein. The computer-executable program code instructions may include program code instructions configured to detect one or more audio signals associated with speech data of at least one first user and determine that the speech data is associated with a first language. The computer program product may also include program code instructions configured to translate one or more words of the speech data associated with the first language into one or more other words associated with a second language different from the first language. The computer program product may further include program code instructions configured to present the one or more other words translated into the second language as one or more text items to a display of a head-mounted device of at least one second user, or to output one or more other words in the second language as audio content to at least one second user. The computer program product may further include program code instructions configured to, in response to the speech data, generate a first set of one or more responses in the first language and a second set of corresponding one or more other responses in the second language presented via the display. The one or more other responses include translations of one or more responses in the first language into a second language.

[0033] Additional advantages will be set forth in part in the description that follows, or may be learned through practice. These advantages will be realized and obtained through the elements and combinations particularly pointed out in the appended claims. It should be understood that, as claimed, the foregoing general description and the following detailed description are merely exemplary and illustrative, and not restrictive. Attached Figure Description

[0034] The invention and the following detailed description can be further understood when read in conjunction with the accompanying drawings. Examples of the disclosed subject matter are shown in the drawings to illustrate the subject matter; however, the disclosed subject matter is not limited to the specific methods, compositions, and apparatuses disclosed. Furthermore, the drawings are not necessarily drawn to scale. In the drawings: Figure 1 A sample HMD is shown according to the examples in this disclosure.

[0035] Figure 2 An example HMD environment for speech translation is shown according to the examples in this disclosure.

[0036] Figure 3 An example method for speech translation based on the examples in this disclosure is shown.

[0037] Figure 4A flowchart illustrating the generation of a response list according to an example of this disclosure is shown.

[0038] Figure 5 An example of a machine learning framework based on one or more examples of this disclosure is shown.

[0039] Figure 6A A speech translation system according to an example of this disclosure is shown.

[0040] Figure 6B A speech translation system according to an example of this disclosure is shown.

[0041] Figure 7 An example block diagram of an HMD device according to the present disclosure is shown.

[0042] Figure 8 An example processing system based on the examples in this disclosure is shown.

[0043] Figure 9 The example shown in this disclosure illustrates the operation of translating voice data associated with one or more conversations or sessions of a user.

[0044] The accompanying drawings depict various examples for illustrative purposes only. Those skilled in the art will readily recognize from the following discussion that alternative examples of the structures and methods shown herein can be employed without departing from the principles described herein. Detailed Implementation

[0045] Some embodiments of the invention will now be described more fully below with reference to the accompanying drawings, which illustrate some, but not all, of the embodiments of this disclosure. In fact, various embodiments of this disclosure may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Similar reference numerals throughout refer to similar elements.

[0046] As used herein, the terms “data,” “content,” “information,” and similar terms are used interchangeably to refer to data that can be sent, received, and / or stored, according to embodiments of this disclosure. Furthermore, as used herein, the term “exemplary” is not intended to convey any qualitative assessment but merely to convey illustrative examples. Therefore, any use of such terms should not be construed as limiting the scope of embodiments of this disclosure.

[0047] As defined herein, “computer-readable storage medium” means a non-transitory physical or tangible storage medium (e.g., volatile or non-volatile storage device), and “computer-readable storage medium” can be distinguished from “computer-readable transmission medium” which refers to electromagnetic signals.

[0048] As described herein, an "application" can refer to a computer software package that performs specific functions for a user, and / or, in some cases, performs specific functions for one or more other applications. One or more applications may be run using an operating system (OS) and other supporting programs. In some examples, one or more applications may request one or more services from other entities and communicate with other entities via an application programming interface (API).

[0049] As used herein, “artificial reality” can refer to a form of reality that has been modified in some way before being presented to a user. Artificial reality can include, for example, virtual reality, augmented reality, mixed reality, metaverse reality, or some combination and / or derivative thereof. Artificial reality content can include entirely computer-generated content or computer-generated content combined with captured (e.g., real-world) content. In some instances, artificial reality can also be associated with applications, products, accessories, services, or some combination thereof, which are used, for example, to create content in artificial reality or otherwise used in artificial reality (e.g., to perform actions in artificial reality).

[0050] As referred to in this article, “artificial reality content” can include content such as video, audio, haptic feedback, or some combination thereof, any of which can be presented to the user in a single or multiple channel (e.g., stereoscopic video that gives the viewer a three-dimensional effect).

[0051] As used herein, virtual reality can refer to an immersive virtual reality world / immersive augmented reality world, in which multiple augmented reality devices can be used in a network (e.g., a virtual reality network) where users in the network may have, but do not necessarily have, one or more social connections. A virtual reality network can be associated with a three-dimensional virtual world, online games (e.g., video games), one or more content items such as images, videos, non-fungible tokens (NFTs), etc., where the content items may be purchased, for example, using digital currencies (e.g., cryptocurrencies) and / or other suitable currencies.

[0052] As referred to in this article, "staccato" or "staccato style" can refer to or indicate that one or more users, one or more people, or one or more individuals speak individual words one after another, with clear pauses between words in multiple words.

[0053] It will be understood that the methods and systems described herein are not limited to specific methods, specific components, or particular implementations. It will also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.

[0054] Exemplary Artificial Reality System This disclosure generally relates to systems and methods for speech translation via audio and / or text using speakers and / or microphones associated with electronic devices (e.g., HMDs). Figure 1 An example HMD 100 associated with artificial reality content according to aspects of this disclosure is illustrated. The HMD 100 can be configured to present augmented reality content items, virtual reality content items, mixed reality content items, or combinations thereof via a display (e.g., one or more displays 108). The HMD 100 may include a frame 102 (e.g., eyeglasses frame), one or more cameras 104, one or more displays 108, one or more voice translation systems 114, and one or more audio devices 110 (e.g., speakers / microphones). In some examples, the one or more voice translation systems 114 may be referred to herein as one or more voice translation components 112. The one or more displays 108 can be configured to direct images and / or video onto a surface 106 (e.g., a user's eyes or another structure). In some examples, the HMD 100 may be implemented in the form of augmented reality glasses (e.g., smart glasses). Thus, the one or more displays 108 may be at least partially transparent to visible light to allow a user to view a real-world environment through the one or more displays 108. One or more audio devices 110 (e.g., speakers / microphones) can provide the user with audio associated with augmented reality content and capture audio signals.

[0055] The tracking surface 106 can be beneficial for graphics rendering or user peripheral input. In many systems, the HMD 100 design may include one or more cameras 104 (e.g., remote). Figure 2 One or more front-facing cameras of the main user 120 or one or more rear-facing cameras facing the main user 120. One or more cameras 104 can track monocular or binocular eye movements (e.g., gaze) or gazes associated with the main user 120. In some examples, a person (e.g., Figure 2The HMD 100 may include an eye-tracking system for tracking the vergence movement of the main user 120. One or more cameras 104 may be eye-tracking systems. Depending on the orientation and viewing angle of the one or more cameras 104, the one or more cameras 104 may capture images and / or video of one or more areas, and / or capture video and / or images associated with surface 106 (e.g., the eyes of the main user 120, or other areas of the face (e.g., the face of user 120)). In an example where the one or more cameras 104 are facing backward toward the main user 120, the one or more cameras 104 may capture images and / or video associated with surface 106. In an example where the one or more cameras 104 are facing forward away from the main user 120, the one or more cameras 104 may capture images and / or video of one or more areas (e.g., one or more areas in a real-world environment). The HMD 100 may be designed to have both front-facing and rear-facing cameras (e.g., one or more cameras 104). Multiple cameras 104 may be available for detecting reflections or other movement (e.g., flash images and / or any other suitable features) on surface 106. One or more cameras 104 may be positioned at different locations on frame 102. One or more cameras 104 may be positioned along a portion of the width of frame 102. In some other examples, one or more cameras 104 may be arranged on one side of frame 102 (e.g., the side of frame 102 closest to the eyes). Alternatively, in some examples, one or more cameras 104 may be positioned on one or more displays 108. In some examples, one or more cameras 104 may be sensors, or a combination of cameras and sensors, for tracking a user's single or binocular eyes (e.g., surface 106).

[0056] One or more audio devices 110 may be positioned at different locations on the frame 102 or in any other configuration, such as, but not limited to, one or more headphones communicatively connected to the HMD 100 or peripheral devices. One or more audio devices 110 may be positioned along a portion of the width of the frame 102. In some other examples, one or more audio devices 110 may be arranged on the side of the frame 102 (e.g., the side of the frame 102 closest to the ears). In some examples, one or more audio devices 110 may be sensors, or combinations of speakers, microphones, and / or sensors, for capturing and generating sound associated with the user.

[0057] One or more speech translation components 114 may be configured to receive one or more audio signals associated with speech data (e.g., sound data) in a language (e.g., one or more first languages) of one or more users, and may translate the received audio signals in the first language into one or more other languages, as described more fully below. One or more speech translation components 114 may provide the audio of the translated language to one or more displays 108 to present the audio of the translated language (e.g., one or more second languages) to one or more users via the one or more displays 108, the translated language being displayed as text input in a real-world environment also captured by one or more cameras, and output via the one or more displays 108. The text input associated with the audio of the translated language may be overlaid (e.g., superimposed) on the content of the real-world environment captured in the field of view of one or more cameras 104 and output via the one or more displays 108. One or more speech translation components 114 may also provide the audio of the translated language to one or more audio devices 110 to enable the one or more audio devices 110 to output the audio of the translated language. In some examples, the audio of the translated language can be output in real time (e.g., simultaneously) via audio device 110 and the text input of the translated language can be presented on one or more displays 108 of HMD 100.

[0058] Exemplary System Architecture Figure 2 An exemplary environment for speech translation is illustrated. This exemplary environment may include a system 101 for facilitating speech translation. A primary user 120 may be associated with an HMD 100, a mobile device 111, or a smartwatch 112. Users can associate with these devices based on links to user profiles. A base station 121 may be used for wide area network access (e.g., cellular systems) or local area network access (e.g., Wi-Fi). The HMD 100, mobile device 111, and / or smartwatch 112 may communicate directly with each other or through the base station 121 (e.g., via Bluetooth, near field communication, ultra-wideband, or any other suitable communication connection). A line-of-sight (LOS) region 125 may be determined based on the gaze of one or more primary users 120 and / or the video (or image) capture area of ​​one or more cameras 104 of the HMD 100, wherein a person 130 may be located within the LOS region 125.

[0059] In one example, an inertial measurement unit (IMU) (e.g., Figure 7The IMU (Implantable Memory Unit 56) can be used to determine whether inertial movement of the smartwatch 112, mobile device 111, and / or HMD 100 can be used as a trigger for selection of a response associated with a response list. The IMU can be calibrated for one or more different muscle and / or body movements. Artificial intelligence can be used to determine the trigger for response selection. In some other examples, inertial movement of the smartwatch 112, mobile device 111, and / or HMD 100 can be determined by eye gaze tracking (EGT), electromyography (EMG), or any other suitable method to determine one or more inertial movements and / or body movements. In some examples, the trigger can be associated with one or more inputs, which are related to user speech captured by one or more audio devices 110.

[0060] Exemplary system operation Some example aspects of this disclosure may provide one or more machine learning models to help provide real-time language translation (e.g., during a conversation between users). In this respect, examples of this disclosure may address a problem associated with a situation where one or more first users who are fluent in one or more first languages ​​are able to respond to one or more other users who are fluent in one or more second languages ​​(e.g., one or more second users) even if the one or more first users may not understand / speak one or more second languages.

[0061] To further complicate matters, one or more second users may not have or may not wear smart glasses (e.g., HMDs), and in some instances, the speakers included in the smart glasses worn by one or more first users may be sensitive, such that only the one or more first users wearing the smart glasses can hear the audio output from the speakers of the smart glasses. Furthermore, one or more first users may not have other hardware (e.g., external displays, etc.) that they could use to read the translated audio to one or more second users.

[0062] In this respect, the examples disclosed herein can address the problem associated with a mode in which one or more first users fluent in one or more first languages ​​are able to respond to another one or more users fluent in one or more second languages ​​(e.g., one or more second users) in situations where the one or more first users may not understand or speak one or more second languages.

[0063] In this regard, exemplary aspects of this disclosure may receive and / or detect audio signals (e.g., speech data) from a person speaking one or more first languages, and may translate the speech data into one or more different second languages ​​that the user wearing the smart glasses may prefer / understand or speak fluently, based on the received audio signals. Exemplary aspects may (e.g., automatically) generate one or more responses (e.g., based on captured video of the conversation) for a dialogue between the person and the user wearing the smart glasses. In some examples, the responses may be presented on a graphical user interface (e.g., a panel on one side (e.g., the left side) of the graphical user interface). One or more machine learning models of the exemplary aspects may automatically translate possible responses to one or more words in response to detected translated utterances from a person speaking one or more first languages. In instances where the user selects a response, the response may be output, for example, read aloud one word at a time in a manner such as “repeat after me,” allowing the user wearing the smart glasses to speak each of the one or more words aloud to the person in one or more first languages ​​that the person speaks and understands.

[0064] In an AR environment, the user / wearer of the smart glasses may not like, speak, or understand another language, but repeating the same speech as me can help with one or more real-time translations by displaying a live stream translation of the other language. For example, the AR environment could be displayed on the smart glasses showing a conversation between the users, where the real-time translation could be displayed on the smart glasses' display (e.g., as captions accompanying the conversation video).

[0065] Applications / implementations of one or more machine learning models capable of generating responses can accelerate conversational interactions between users, especially when no keyboard or other text input device is available for the user to use during the conversation.

[0066] Figure 3An example method 300 for speech translation according to an example of this disclosure is shown. Method 300 may be performed by one or more of a processor (e.g., processor 91), a memory (e.g., ROM 93 or RAM 82), one or more microphones (e.g., speaker / microphone 81), and / or a memory controller (e.g., memory controller 92). In some example aspects, method 300 (and steps 310, 320, 330, 340, 350, 360 of method 300) may be performed by an HMD (e.g., HMD 100), a communication device (e.g., processing system 800), or a UE (e.g., UE 30). In step 310, an audio signal from a person (e.g., person 130) or the environment may be captured by a microphone (e.g., speaker / microphone 81) without user input, and the audio signal may be temporarily stored in temporary memory (e.g., RAM 82). The captured audio signals and stored in a memory such as RAM 82 can be overwritten or replaced at the end of each / predetermined time period between received audio signals. The predetermined time period can be a predetermined number of seconds. The time period during which RAM 82 stores data (e.g., audio signals) before it is replaced can be determined by the user of the device (e.g., HMD 100, processing system 800, UE 30, etc.), device settings stored in a storage device (e.g., storage device 97), or when the predetermined time period between received audio signals approaches or reaches its end (e.g., the predetermined time period expires).

[0067] For illustrative and not limiting purposes, for example, in an instance where a user utilizes HMD 100, one or more audio devices 110 may capture a first audio signal associated with person 130 in LOS 125 for up to 30 seconds. In this example, there may be 25 seconds before one or more audio devices 110 receive a second audio signal associated with that person. Thus, device settings and / or user settings may be configured to store or retain data (e.g., the first audio signal) captured from one or more audio devices 110 until 20 seconds have elapsed since one or more inputs associated with one or more first audio signals were captured. Therefore, after 20 seconds (which could be a period of time where the received audio signal may not be associated with speech, or a period of time where there is a small amount of non-speech audio signal, or a period of time where no audio signal is received), one or more first audio signals or data captured by (e.g., associated with the audio device) speaker / microphone 81 may be overwritten by subsequently captured audio content. The audio content can be, for example, a second audio signal, which can be associated with the speech or voice data of a user (e.g., user 120, person 130) for the next time period (e.g., a few seconds (e.g., a 30-second period following the 20-second period)). In some examples, machine learning models (e.g., one or more machine learning models 510) can be used to determine when / at what moment the audio signal capture can be rewritten based on its content; for example, the one or more machine learning models can determine that the person associated with the audio signal has finished speaking.

[0068] In an alternative example, the captured audio signal stored in a memory such as RAM 82 can be overwritten or replaced every predetermined time period (e.g., a predetermined number of seconds). The time period during which RAM 82 can store the data (e.g., the audio signal) before it is replaced can be determined by the user of the HMD (e.g., HMD 100) and / or based on device settings stored in a storage device (e.g., storage device 97). For example, in an instance where the user is capturing a first audio signal associated with a person 130 in LOS 125 for 30 seconds using HMD 100 and / or one or more audio devices 110, there may be 25 seconds before a second audio signal associated with that person is received. Therefore, as an example, device settings and / or user settings can be configured (e.g., specified) to facilitate the storage or retention of data captured by one or more audio devices for a predetermined time period, such as 45 seconds. Thus, every 45 seconds during the capture of an audio signal or data captured by a speaker / microphone 81 associated with an audio device can be overwritten by the next 45 seconds of audio capture (e.g., other audio signals).

[0069] In step 320, the processor (e.g., processor 91) may refer to memory (e.g., ROM 93 or RAM 82) to translate the audio signal received in step 310 and provide text associated with the received audio signal. The received audio signal may be in a first language, and processor 91 may translate the received audio signal in the first language into a different second language associated with the user (e.g., user 120). Furthermore, processor 91 may provide / generate text associated with the first language in both the first and second languages. For example, the user of HMD 100 may be having a conversation with person 130, and HMD 100 may receive / capture the audio of the conversation as an audio signal in the first language associated with person 130. Processor 91 may translate the received audio signal in the first language (e.g., Spanish) into a different second language associated with the user (e.g., user 120). This translation may be provided to the user (e.g., user 120) in the second language associated with user 120 (e.g., English) as text (e.g., presented via a display device) and / or as an audio signal via one or more audio devices 110.

[0070] In step 330, a translation may be provided to user 120. This translation may include visual (e.g., text items) components and / or audio (e.g., audio signals) components associated with the second language, based on received audio signals corresponding to the first language. The translation may be provided via a display (e.g., one or more displays 108) and / or audio device (e.g., one or more audio devices 110, speakers / microphones 81) of the HMD (e.g., HMD 100) or other communication device (e.g., processing system 800, UE 30). In one example, the translation may also be provided via a peripheral device associated with HMD 100 (e.g., mobile device 111 or smartwatch 112, etc.).

[0071] In step 340, a response list may be determined by referring to memory (e.g., ROM 93 or RAM 82). This memory may be referenced by one or more machine learning models (e.g., one or more machine learning models 510) that can determine the response list. The response list may be associated with the received audio signal to facilitate communication (e.g., dialogue) between user 120 and person 130. The response list may be associated with one or more contexts related to the received audio signal in step 310. As described in more detail herein, the response list may be determined based on an evaluation of one or more contexts of the audio signal, a historical database, or machine learning operations (also referred to herein as one or more machine learning models), such as a large language model (LLM) user. The historical database may include books, newspapers, articles, conversations, television (TV) programs, movies, or video clips, and / or any combination thereof. In step 350, the response list may be provided to user 120. The response list may include text components and / or audio components. The response list may also include each response in a first language (e.g., Spanish or Italian) and a second language (e.g., English or German) for user 120 to understand each response. The response list may be provided via a display (e.g., one or more displays 108) of an HMD (e.g., HMD 100) associated with user 120. In one example, the translation may also be provided via a peripheral device associated with HMD 100 (e.g., mobile device 111 or smartwatch 112). In one example, the audio components of / associated with the response list may not be used until the user (e.g., user 120) triggers a selection of the response list. In some examples, each of the multiple responses in the response list may include pronunciations associated with that response (e.g., one or more words of that response), configured so that user 120 can reply to that person (e.g., person 130) in their first language (e.g., Spanish or Italian).

[0072] Steps 310, 320, 330, and 340 are expected to be performed at the same time or simultaneously. In some examples, steps 310, 320, 330, 340, and 350 may be performed at the same time or simultaneously.

[0073] In step 360, a user (e.g., user 120) may trigger a selection from a list of replies / responses. The trigger can be any audio, visual, motion, eye-tracking, or touch captured by a device (e.g., HMD 100, processing system 800, UE 30). In an example of audio triggering, the user may begin speaking a word or string of words associated with a reply captured by a microphone of HMD 100 (e.g., one or more audio devices 110). Audio can also be determined by any other suitable method, function, or sensor. In an example of touch triggering, the user may touch a button or area of ​​HMD 100 or a peripheral device (e.g., mobile device 111 or smartwatch 112, etc.) to select a reply from a list of replies. The touch can be determined by satisfying a pressure sensor threshold, which can be set by device settings and stored in a storage device (e.g., storage device 97). Touch can also be determined by any other suitable method, function, or sensor. In a vision-triggered example, a camera (e.g., one or more cameras 104, 54) can be configured to detect or determine hand movements and / or gestures to initiate a selection of a response from a response list. For example, the camera (e.g., one or more cameras 104, 54) can be configured to determine instances where the user's hand is raised in the line of sight (e.g., LOS 125) to a predetermined level or height associated with each response in the response list and the user's line of sight. Thus, in instances where the user's hand can be raised to a predetermined level or height, the selection can be determined by the HMD 100. In an eye-tracking-triggered example, the user can view responses in the response list associated with the display of the HMD 100. Eye movements can be determined based on the achievement / satisfaction of a predetermined saccade threshold associated with the eye-tracking system, wherein the predetermined saccade threshold can be set by device settings and can be stored in a storage device (e.g., storage device 97). Eye movements can also be determined by any other suitable method, function, or sensor. In a motion-triggered example, the user can perform a movement or motion toward a response in the response list associated with the display of the HMD 100 (e.g., one or more displays 108). Motion can be determined using a motion sensor or an IMU (e.g., IMU 56). The IMU can be an electronic device that uses a combination of the following to measure and report specific forces, angular velocities, and orientation of a device (e.g., HMD 100): an accelerometer; a gyroscope; and in some cases, a magnetometer. The IMU can determine whether a trigger has exceeded or reached a predetermined threshold of the HMD 100. In another example of motion triggering, a user can perform a movement or motion toward a selected response, or can move a hand or finger in a predetermined manner that indicates selection of a response from a list of responses associated with the HMD 100's display (e.g., one or more displays 108).Movement can be determined by achieving / meeting predetermined EMG thresholds, where EMG can be captured by EMG sensors associated with HMD 100 and / or peripheral devices (e.g., mobile device 111 or smartwatch 112, etc.). EMG thresholds can be set via device settings and can be stored in a storage device (e.g., storage device 97). EMG sensors can measure one or more muscle responses and / or electrical activity in response to neural stimulation of one or more muscles of the user that respond to movement or motion. Movement can also be determined by any other suitable method, function, or sensor.

[0074] In some examples, where an option has been selected, the device (e.g., HMD 100) may provide pronunciation associated with the selected response to help user 120 respond in the first language associated with a person (e.g., person 130). In some examples, audio device 110 may provide audio corresponding to the response. In some examples, audio device 110 may provide audio corresponding to the response in a "sing-along" manner (e.g., one word at a time), allowing the user to hear how the words of the selected response are pronounced and repeat those words to respond to person 130.

[0075] Figure 4A flowchart for generating a list of responses according to an example of this disclosure is shown. At box 402, a device (e.g., HMD 100) may refer to a storage device (e.g., storage device 97) to evaluate and utilize a profile associated with a user (e.g., user 120). The profile may include any number of data types, including but not limited to one or more styles associated with the user, one or more languages ​​(e.g., one or more designated / spoken languages ​​of the user), or any other suitable data. One or more styles may be based on previous responses selected by the user, wherein the previously selected responses and one or more contexts in which those responses were selected may include the one or more styles. At box 404, the device (e.g., HMD 100) may receive an audio signal (e.g., data) associated with speech / sound data, wherein the speech / sound data may be in a language different from the language associated with the user (e.g., a language the user does not speak). At box 410, the device (e.g., HMD 100) may receive and compare the data associated with the profile and the audio signal, and may determine, based on the received data, whether a translation can be generated by a speech translation system (e.g., one or more speech translation components 114). In this respect, a voice translation system can be configured to determine that translation may be necessary by comparing data in a profile (e.g., a user profile) with received audio signals (e.g., audio signals of a person's (e.g., person 130) speech). In one example, the profile may instruct the user to speak a language such as English (or other languages ​​such as German, etc.), while the audio signal may be determined to be a language such as Spanish (or other languages ​​such as Portuguese, etc.), thus indicating that the received audio signal (e.g., one or more audio signals of person 130's speech) may need to be translated into the language associated with the profile (e.g., English) (e.g., the instruction in the user profile instructing the user to speak English). In other examples, the language instructing the user to speak in the profile (e.g., the user profile) can be any other suitable language (e.g., besides English), and the language of the audio signal can be any other suitable language (e.g., besides Spanish), such as, but not limited to, French, Italian, German, Portuguese, Dutch, Russian, Hebrew, Arabic, Mandarin, Hindi, Japanese, and any other suitable language.

[0076] In box 412, a device (e.g., HMD 100) can determine a temporal range of speech, where audio signal data can be stored for a temporary time period according to device (e.g., HMD 100) settings. This time range can include predetermined time periods corresponding to one or more sentences or portions of a dialogue (e.g., a conversation between users), where the time increment can be any suitable time increment (e.g., milliseconds, seconds, minutes, hours, etc.) associated with the received audio signal associated with (e.g., person 130) speech / sound data. In box 420, one or more machine learning models (e.g., one or more machine learning models 510) can be applied to analyze the audio signal of the speech data of a person (e.g., person 130) associated with a profile (e.g., a user profile) based on the audio signal captured within the time range. In some examples, the device (e.g., HMD 100) can execute (e.g., associated with box 410) one or more machine learning models. One or more machine learning models (e.g., one or more machine learning models 510) can analyze the audio signal to determine one or more contexts of the dialogue and can generate a list of responses to the received audio signal associated with the person (e.g., person 130). In various examples, at box 422, one or more machine learning models can use training data (e.g., training data 520) to develop and predict associations between at least one of profile data, audio signal data, and / or one or more contexts of the audio signal data. The device (e.g., HMD 100) can implement one or more machine learning models. The training data (e.g., training 520) can be historical data or associated with profile data, audio signal data (e.g., audio signal data in various languages), context data, and / or pronunciation. Context data can be determined based on one or more defined conversational contexts (e.g., morning greetings, evening greetings, dinner topics, breakfast topics, sports topics, etc.). In other examples, the training data can be associations between profile data (e.g., one or more styles) and / or audio signal data.

[0077] In box 424, the device (e.g., HMD 100) may utilize / implement a neural network to leverage training data (e.g., training data 520) or use one or more machine learning model techniques to help analyze profile data and audio signal data that may be associated with historical data (e.g., historical data of training data 520). For example, the neural network may be trained on historical data of training data 520, which indicates communication (e.g., conversations) between two or more people with various contextual backgrounds. In many instances, historical data may include, but is not limited to, books, movies, news articles, magazines, television programs, and / or previous conversations of other users. Previous conversations of other users, books, movies, news articles, magazines, television programs, etc., may be in a variety of different languages, including but not limited to English, Spanish, French, Italian, German, Portuguese, Dutch, Russian, Hebrew, Arabic, Mandarin, Hindi, Japanese, and any other suitable language. The neural network may be implemented by one or more machine learning models (e.g., one or more machine learning models 510). In some examples, one or more machine learning models (e.g., one or more machine learning models 510) may be any large language model.

[0078] In box 430, a device (e.g., HMD 100) can generate a response list based on the correlation between a user profile, received audio signals, and / or historical data. The generated response list can (e.g., directly) utilize data from one or more machine learning models. In the example, the generated response list can be compiled on the display of the device (e.g., HMD 100) (e.g., one or more displays 108) or through one or more graphical user interfaces (e.g., displays / touchpads / one or more interfaces 42) associated with peripheral devices (e.g., mobile devices 111, smartwatches 112, etc.).

[0079] In box 440, a trigger for selecting a response can be determined by a device (e.g., HMD 100). This trigger can be any audio, visual, motion, eye-tracking, and / or touch captured by the device. In an example of audio triggering, a user can begin speaking a word or string of words associated with a response captured by a microphone of HMD 100 (e.g., one or more audio devices 110). Audio can also be determined by any other suitable method, function, or sensor. In an example of touch triggering, a user can touch a button or area of ​​HMD 100 or a peripheral device (e.g., mobile device 111 or smartwatch 112, etc.) to select a response from a list of responses. Touch can be determined by a pressure sensor threshold being implemented / met, which can be set by device settings and stored in a storage device (e.g., storage device 97). Touch can also be determined by any other suitable method, function, or sensor. In an example of visual triggering, one or more cameras (e.g., one or more cameras 104, camera 54) can be configured to detect or determine hand movements and / or gestures to initiate a selection of a response from a list of responses.

[0080] For example, one or more cameras (e.g., one or more cameras 104, camera 54) can be configured to determine instances where a user's hand is raised in their line of sight (e.g., LOS 125) to a predetermined level or height associated with each reply in the reply list and the user's line of sight. Thus, in instances where the user's hand can be raised to the predetermined level or height, the HMD 100 can determine which reply to select from the reply list. In an eye-tracking triggered example, the user can view replies in the reply list associated with the display of the HMD 100 (e.g., one or more displays 108). Eye movements can be determined by achieving / meeting a predetermined saccade threshold associated with an eye-tracking system (e.g., a rear camera (e.g., one or more rear cameras 104)), wherein the predetermined saccade threshold can be set by device settings and can be stored in a storage device (e.g., storage device 97). Eye movements can also be determined by any other suitable method, function, or sensor. As an example of motion triggering, motion can be captured by the motion sensor of the HMD 100 (e.g., an IMU (e.g., IMU 56)), where the motion sensor can be configured to determine minute changes in motion to determine the selection of a response from a list of responses. Motion can also be determined by any other suitable method, function, or sensor. In some examples, in instances where a response has been selected from a list of responses by triggering, one or more pronunciations associated with the selected response can be generated and provided to the user (e.g., user 120) to respond in a first language (e.g., Spanish, Portuguese, etc.) associated with the language of the person (e.g., person 130). At this point, the one or more pronunciations can be presented, for example, as text input to a display (e.g., one or more displays 108) and / or a graphical user interface (e.g., a display / touchpad / one or more interfaces 42), and in some examples, as audio output of the pronunciation can be presented by the audio device of the HMD 100 (e.g., one or more audio devices 110, speakers / microphones 81).

[0081] In box 450, the selected response may be provided to the user via a device (e.g., HMD 100). In box 452, the device (e.g., HMD 100) may present / provide the selected response as an alert, notification, image, explanatory text, email, message, etc., where the response may include pronunciation associated with the words of the response. In some examples, the response may be presented via a peripheral device (e.g., mobile device 111, smartwatch 112). In some examples, the presentation of the response may include at least one of the following: amplification of the selected response; or addition of pronunciation to the response. Pronunciation may be applied to the first language of a person (e.g., person 130) (e.g., Spanish or Portuguese, etc.) and a second language corresponding to the language of user 120 (e.g., English or German, etc.). In box 454, audio (e.g., in audio form) may similarly be presented by the device (e.g., HMD 100) via an audio device (e.g., one or more audio devices 110), where the words of the selected response may be presented as audio to the ears of the user (e.g., user 120). In one example, words can be presented to the user one word at a time or staccato, allowing the user to repeat (in their own voice) each word presented to one or both ears of the user via one or more audio devices 110. For example, staccato could include instances where the user (e.g., user 120) selects a response, such that one or more words associated with the selected response can be output (e.g., read aloud) as audio by one or more audio devices 110 in a “repeat after me” manner, allowing the user (e.g., the wearer of HMD 100) to loudly repeat (as heard by one or both ears of the user) one word at a time to the person (e.g., person 130) in that person's language (e.g., Spanish). This enhances the user's conversation with the person because the user can speak a language different from that person's language (e.g., a first language (e.g., Spanish or Portuguese, etc.)) (e.g., a second language (e.g., English or German, etc.)).

[0082] Figure 5 A machine learning and training model according to an example of this disclosure is shown. A frame 500 associated with one or more machine learning models 510 can be remotely hosted. Alternatively, the frame 500 can reside within a device, such as... Figure 1 The HMD 100 shown Figure 7 UE 30 in Figure 8 The processing system 800, or by Figure 1One or more other devices (e.g., mobile device 111, smartwatch 112) shown are used for processing. One or more machine learning models 510 can be operatively coupled to training data 520 stored in memory such as training database 550 or another database (e.g., ROM 93, RAM 82). In some examples, one or more machine learning models 510 can be used with Figure 3 , Figure 4 and / or Figure 9 The operations are associated with these components. In some other examples, one or more machine learning models 510 may be associated with other operations. In some examples, one or more machine learning models 510 may be examples of one or more speech translation components 114. One or more machine learning models 510 may be implemented by one or more machine learning components and / or another device (e.g., a server and / or a computing system). For example, in some example aspects, one or more machine learning models 510 may be implemented by HMD 100, UE 30, processing system 800, mobile device 111, or smartwatch 112, etc.

[0083] The training data 520 used by one or more machine learning models 510 can be pre-trained, fixed, or periodically updated. Alternatively, the training data 520 can be updated in real time based on evaluations performed by one or more machine learning models 510 in non-training mode. This can be illustrated by a bidirectional arrow connecting one or more machine learning models 510 to stored training data 520, which can be stored in a training database 550. Some other examples of training data 520 may include, but are not limited to, multiple items identified as being associated with networks (e.g., the Internet, social networks, etc.), platforms, or systems (e.g., system 101).

[0084] In some examples, training data 520 may be obtained based on user data of users associated with the system (e.g., system 101). User data may be data or information relating to past / historical interactions / conversations and / or current (e.g., real-time) interactions / conversations of users associated with the system. In some examples, these users may choose to join the system so that the use of user data can be used as training data 520. The users' past / historical interactions / conversations and / or the users' current interactions / conversations may include translating audio content spoken in one or more languages ​​by at least a subset of multiple users into one or more other languages ​​spoken by one or more other subsets of those users. In this way, one or more machine learning models 510 may utilize training data 520 to translate content (e.g., users' audio content (e.g., speech / voice data)) from one or more languages ​​into another one or more different languages. Some training data 520 of one or more machine learning models 510 may include user data associated with contextual information / variables related to user interactions / conversations, and one or more machine learning models 510 may generate one or more responses (e.g., automatic responses) with the user data in response to translated content (e.g., translated speech / sound data) determined by one or more machine learning models 510.

[0085] Figure 6A A speech translation system according to an example of this disclosure is shown. Figure 6AThe view seen by a user (e.g., user 120) when viewed / watched via a display (e.g., one or more displays 108) of an HMD (e.g., HMD 100) can be displayed. In this respect, in addition to text input 602 and / or text input 604 described below, the user can also view / watch real-world objects / content in the user's environment via the display (e.g., one or more displays 108). User 120 wearing HMD 100 can receive audio signals via audio devices of HMD 100 (e.g., one or more audio devices 110). In some examples, one or more audio signals can also be received via a communication-connected device or peripheral device (e.g., mobile device 111, smartwatch 112, or any other suitable device). When one or more audio devices 110 receive audio signals (e.g., voice data, sound data, etc.) associated with a person (e.g., person 130), one or more speech translation components 114 can begin providing translation to a second language associated with the first language of the received audio signal. The translation can be provided via text input 602 in a first language (e.g., Spanish) and text input 604 in a second language (e.g., English) corresponding to the original language (and / or associated audio output) or the received audio signal. For example, when person 130 speaks the term “cómo” in the first language (e.g., Spanish), one or more speech translation components 114 can translate the spoken term (e.g., “cómo”) into “how” in the second language (e.g., English) in real time (e.g., simultaneously), and the device (e.g., HMD 100) can display the translated “how” in the second language (e.g., English) text input 604 via one or more displays 108. In some examples, the device (e.g., HMD 100) can display the translated “how” in the second language (e.g., English) in real time (e.g., simultaneously with the person speaking “cómo”). Continuing the previous example, as person 130 continues to say / pronounce “estás” in their first language (e.g., Spanish), one or more speech translation units 114 can translate the spoken term “estás” into “are you” in a second language (e.g., English) text input 604 in real time, and HMD 100 can display the remaining / other part of the translation, “are you,” via one or more displays 108. Thus, translating the speech associated with person 130 saying “cómo estás” produces a second-language translation, displayed as “how are you” via one or more displays 108 of HMD 100.In some examples, audio (e.g., speech data of person 130) translated from a first language to a second language can be provided to user 120 as text input via one or more displays 108 and (e.g., simultaneously) as audio output of the translated text (e.g., audio output of the English translation "how are you" based on the Spanish translation "cómo estás") via one or more audio devices 110 of HMD 100. For example, as a result of receiving an audio signal corresponding to the phrase "cómo estás" in the first language (e.g., Spanish), the phrase "how are you" in the second language can be provided to user 120 as audio output via one or more audio devices 110 associated with HMD 100.

[0086] Figure 6B A speech translation system based on an example of this disclosure is further illustrated. Figure 6B This can show the view seen by a user (e.g., user 120) when viewed / watched via a display (e.g., one or more displays 108) of an HMD (e.g., HMD 100). In this respect, in addition to the reply list 610 (e.g., replies 601, 603, 605, 612, 614, 616) and / or audio buttons (e.g., audio buttons 606, 607, 608) described below, the user can also see / watch real-world objects / content in the user's environment via the display (e.g., one or more displays 108). Next... Figure 6A For example, in response to receiving an audio signal associated with the speech of a person (e.g., person 130), such as in Figure 6A When a Chinese person speaks "cómoestás," a response list 610 can be generated (e.g., automatically generated) by one or more speech translation systems 114, and can be provided via a display associated with HMD 100 (e.g., one or more displays 108). The response list 610 can be generated by one or more speech translation systems 114 based on... Figure 3 and Figure 4 The method used to determine or provide this information is for illustrative and not limiting purposes. Figure 6B In the example, the response list 610 for the voice "cómo estás" may include corresponding translations of three responses 601, 603, and 605 in a first language (e.g., Spanish) and three responses 612, 614, and 616 in a second language (e.g., English). The response list may also include or be associated with one or more audio buttons 606, 607, and 608, which may be... Figure 3 and Figure 4Any of the various triggers described herein may be initiated. The reply list (e.g., replies 601, 603, 605, 612, 614, 616) and / or audio buttons (e.g., audio buttons 606, 607, 608) may be overlaid (e.g., superimposed) on content items in the real-world environment captured / visible in the field of view of one or more displays 108.

[0087] Triggering one or more of the audio buttons 606, 607, and 608 can initiate audio presented to the user 120 via one or more audio devices 110 associated with the HMD 100. This audio may be presented in a first language (e.g., Spanish) to help the user 120 facilitate conversation with the person 130. The received audio signal, which may be captured by the user's HMD 100's audio devices (e.g., one or more audio devices 110, speakers / microphones 81) and associated with the person 130's voice data, may be in the first language and... Figure 6B In the example, user 120 can try to communicate with first person 130 in person 130’s first language by using HMD 100 (e.g., even if user 120’s primary language is a second language, such as English).

[0088] Thus, for example, in response to detecting user 120's selection of response 601 and / or triggering (e.g., selection) of audio button 606, audio button 606 can cause an audio device (e.g., one or more audio devices 110, speaker / microphone 81) to output the text "haciendolo bien!" associated with audio button 606 in a first language (e.g., Spanish) that person 130 can understand and / or speak fluently. In this way, user 120 can read the corresponding response (e.g., response 612 "Doing great!") associated with user 120's selection in the user's language (e.g., a second language (e.g., English)) and can have the translation associated with the selected response (e.g., response 601 "haciendolo bien!") spoken to person 130 as audio content. This can be beneficial because Figure 6BIn this example, person 130 may not be using smart glasses (e.g., may not be wearing HMD 100), but may communicate with user 120 using person 130's own voice (e.g., voice content). HMD 100 can detect the audio content of person 130 speaking and can translate the audio content of person 130 speaking in one language into a different language preferred by user 120 (e.g., English) to display the translated audio content to user 120 via one or more displays 108, and can automatically present one or more responses to user 120 so that user can speak with person 130 (e.g., respond to the person speaking to user). In this respect, HMD can accelerate and enhance the conversation between user 120 and person 130, and enable user 120 to converse with person 130 in a language that user 120 may not be fluent in and may not understand, but person 130 can speak and understand that language fluently. By utilizing exemplary aspects of this disclosure, when user 120 converses with person 130, user 120 does not need to use a keyboard or one or more other devices to type a translation of audio content to show to person 130 (e.g., manually). For example, user 120 does not need to display audio / text content translated into person 130's language in order to communicate (e.g., converse) with person 130.

[0089] In some alternative examples, in response to detecting user 120's selection of a response (response 601) and / or triggering (e.g., selection) of an audio button (e.g., audio button 606), the audio button may cause an audio device (e.g., one or more audio devices 110, speakers / microphones 81) to output audio content (e.g., “haciendolo bien!”) associated with the audio button in a first language (e.g., Spanish) that is understandable to and / or fluently spoken by person 130. In this alternative example, user 120 may utilize HMD 100 to output audio associated with the selected response (e.g., response 601 “haciendolo bien!”) and associated with the audio button (e.g., audio button 606) in a first language (e.g., Spanish) (e.g., via one or more audio devices 110, speakers / microphones 81), causing HMD 100 (e.g., via speakers / microphones 81) to directly output audio associated with the selected response (e.g., response 601 “haciendolo bien!”). In this alternative example, user 120 may not need to speak the audio of the chosen reply (e.g., reply 601 “haciendolo bien!”) to person 130 in order to talk to person 130.

[0090] In some examples, the audio associated with audio buttons 606, 607, and 608 can be presented in a “sing-along” or staccato manner to help user 120 pronounce words in their first language (e.g., Spanish), where user 120 may speak a second language (e.g., English). For example, when person 130 asks a question in their first language (e.g., Spanish) “cómo estás”, a list of responses can be generated and can be provided (e.g., via one or more speech translation components 114) in a second language (e.g., English) corresponding to user 120. For example, one such response could be “pretty good” associated with response 614. Accompanying response 614 could be a translation of response 603 in the first language (e.g., Spanish) generated by one or more speech translation systems 114, such as “bastante bien” corresponding to “pretty good” in the second language (e.g., English). As another example, accompanying response 616 could be a translation of response 605 in a first language (e.g., Spanish) generated by one or more speech translation systems 114, such as "nada mal" corresponding to "Notbad" in a second language (e.g., English). In an alternative example, one or more pronunciations associated with responses 601, 603, 605 in the first language (e.g., Spanish) could be provided to the user as audio to help user 120 speak / recite the response to person 130.

[0091] For example, in some instances, a user can select settings associated with audio buttons 606, 607, and 608 to output audio to the speakers of the HMD 100 (e.g., speaker / microphone 81) associated with corresponding responses (e.g., responses 601, 603, and 605), allowing one or both ears of user 120 to hear one or more pronunciations of the words in the corresponding response (e.g., “haciendolobien!”, “bastante bien”, “nada mal”). In this example, user 120 can utilize one or more pronunciations of individual words from a plurality of words output to one or both ears of user 120 by the speakers (e.g., speaker / microphone 81) to enable user 120 to speak words directly to person 130 based on those one or more pronunciations (e.g., when conversing with person 130).

[0092] Exemplary communication device Figure 7 A block diagram of an example hardware / software architecture for user equipment (UE) 30 is shown. In some examples, UE 30 may be an example of a mobile device 111, a smartwatch 112, or an HMD 100. Figure 7As shown, UE 30 (also referred to herein as node 30) may include: a processor 32; non-removable memory 44; removable memory 46; a speaker / microphone 38; a keypad 40; a display, touchpad, and / or one or more interfaces 42; a power supply 48; a global positioning system (GPS) chipset 50; and other peripheral devices 52. UE 30 may also include a camera 54. In one example, camera 54 may be a smart camera configured to sense / capture images appearing within one or more bounding boxes and capable of capturing one or more videos. IMU 56 may be an electronic device that uses a combination of the following to measure and report specific forces, angular velocities, and orientation of the device (e.g., UE 30): an accelerometer; a gyroscope; and in some instances, a magnetometer. IMU 56 may also determine the inertial movement of the device (e.g., UE 30). Additionally, IMU 56 may be a sensor (e.g., a motion sensor) configured to determine changes in motion of the device (e.g., UE 30). UE 30 may also include communication circuitry, such as transceiver 34 and transmit / receive element 36. It will be appreciated that UE 30 may include any sub-combination of the foregoing elements while remaining consistent with the example.

[0093] Processor 32 can be a dedicated processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and a state machine. Generally, processor 32 can execute computer-executable instructions stored in the memory of node 30 (e.g., memory 44 and / or memory 46) to perform various required functions of the node. For example, processor 32 can perform signal encoding, data processing, power control, input / output processing, and / or any other function that enables node 30 to operate in a wireless or wired environment. Processor 32 can run application layer programs (e.g., a browser) and / or radio access-layer (RAN) programs and / or other communication programs. Processor 32 can also perform security operations, such as authentication, security key protocols, and / or encryption operations (e.g., at the access layer and / or application layer).

[0094] Processor 32 is coupled to its communication circuitry (e.g., transceiver 34 and transmit / receive element 36). Processor 32 can control the communication circuitry by executing computer-executable instructions to enable node 30 to communicate with other nodes via the network to which it is connected.

[0095] Transmitting / receiving element 36 can be configured to transmit signals to or receive signals from other nodes or networked devices. For example, in one example, transmitting / receiving element 36 can be an antenna configured to transmit and / or receive radio frequency (RF) signals. Transmitting / receiving element 36 can support various networks and air interfaces, such as wireless local area networks (WLANs), wireless personal area networks (WPANs), and cellular networks. In yet another example, transmitting / receiving element 36 can be configured to transmit and receive both RF signals and optical signals. It will be appreciated that transmitting / receiving element 36 can be configured to transmit and / or receive any combination of wireless or wired signals.

[0096] Transceiver 34 can be configured to modulate signals to be transmitted by transmitting / receiving element 36 and demodulate signals received by transmitting / receiving element 36. As described above, node 30 can have multi-mode capabilities. Therefore, transceiver 34 can include multiple transceivers enabling node 30 to communicate via various radio access technologies (RATs), such as Universal Terrestrial Radio Access (UTRA) and Institute of Electrical and Electronics Engineers (IEEE 802.11).

[0097] Processor 32 can access information from any type of suitable memory and store data in any type of suitable memory, such as non-removable memory 44 and / or removable memory 46. For example, as described above, processor 32 can store session context in its memory. Non-removable memory 44 may include random access memory (RAM), read-only memory (ROM), hard disk, or any other type of memory storage device. Removable memory 46 may include a subscriber identity module (SIM) card, memory stick, and secure digital (SD) memory card, etc. In other examples, processor 32 can access information from memory that is not physically located on node 30 (e.g., located on a server or home computer) and store data in that memory.

[0098] The processor 32 can receive power from the power source 48 and can be configured to distribute this power to other components in the node 30 and / or control the power supply to other components in the node 30. The power source 48 can be any suitable device for powering the node 30. For example, the power source 48 can include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, and fuel cells, etc.

[0099] The processor 32 may also be coupled to a GPS chipset 50, which can be configured to provide location information (e.g., longitude and latitude) about the current location of node 30. It will be appreciated that node 30 can acquire location information using any suitable location determination method, while remaining consistent with the example.

[0100] Exemplary computing system Figure 8 An example schematic diagram of an example processing system 800 is shown, which can be a component of a system or may be... Figure 7 The processing system 800 is part of UE 30. In some other examples, the processing system 800 may be one or more components of HMD 100. The processing system 800 is merely one example of a suitable processing system 800 within a device (e.g., a mobile phone, laptop, tablet, or any device with messaging capabilities) and is not intended to impose any limitation on the scope or functionality of the examples of the methods described herein. The processing system 800 may include a computer or server and may be primarily controlled by computer-readable instructions, which may be in the form of software, regardless of where or how such software is stored or accessed. Such computer-readable instructions may be executed within a processor (e.g., processor 91) to cause the processing system 800 to operate. In operation, processor 91 fetches, decodes, and executes instructions, and transmits information to and from other resources via the main data transmission path (bus 80) of the processing system 800. Bus 80 connects multiple components within the processing system 800 and defines the medium for data exchange. Bus 80 typically includes data lines for transmitting data, address lines for transmitting addresses, and control lines for transmitting interrupts and operating bus 80.

[0101] In a particular example, bus 80 includes hardware, software, or both of which couple multiple components of system 800 to each other. By way of example and not limitation, bus 80 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth interconnect, a low-pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus, or a combination of two or more of these buses. Where appropriate, bus 80 may include one or more buses. Although this disclosure describes and illustrates a particular bus, this disclosure considers any suitable bus or interconnect.

[0102] The memories coupled to bus 80 include RAM 82 and ROM 93. These memories may include circuitry that allows for the storage and retrieval of information. ROM 93 typically contains stored data that is not easily modified. Data stored in RAM 82 can be read or changed by processor 91 or other hardware devices. In some examples, access to RAM 82 and / or ROM 93 may be controlled by a memory controller. The memory controller may provide address translation functionality, which translates virtual addresses into physical addresses when instructions are executed. The memory controller may also provide memory protection functionality, which isolates processes within the system and separates system processes from user processes. Therefore, a program running in first mode can only access memory mapped by its own process virtual address space; the program cannot access memory within the virtual address space of another process unless memory sharing has been established between the processes.

[0103] In some examples, I / O interface 86 includes hardware, software, or both, providing one or more interfaces for processing communication between system 800 and one or more I / O devices. Where appropriate, processing system 800 may include one or more of these I / O devices. These one or more I / O devices enable communication between an individual and processing system 800. By way of example and not limitation, I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet computer, touchscreen, video camera, another suitable I / O device, or a combination of two or more of these I / O devices. I / O devices may include one or more sensors. This disclosure contemplates any suitable I / O device and any suitable I / O interface for such I / O devices. Where appropriate, I / O interface 86 may include one or more devices or software drivers that enable processor 91 to drive one or more of these I / O devices. Where appropriate, I / O interface 86 may include one or more I / O interfaces. Although this disclosure describes and illustrates specific I / O interfaces, this disclosure considers any suitable I / O interface.

[0104] In some examples, storage device 97 includes a mass storage device for data or instructions. By way of example and not limitation, storage device 97 may include a hard disk drive (HDD), flash memory, random access memory (RAM), read-only memory (ROM), non-volatile read-only memory (NVROM), or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, storage device 97 may include removable or non-removable (or fixed) media. Where appropriate, storage device 97 may be located internally or externally to processing system 800. In some examples, storage device 97 is non-volatile solid-state memory. In a particular example, storage device 97 includes read-only memory (ROM). This disclosure contemplates mass storage devices in any suitable physical form. Where appropriate, storage device 97 may include one or more storage control units facilitating communication between processor 91 and storage device 97. Where appropriate, storage device 97 may include one or more storage devices 97. Although this disclosure describes and illustrates specific storage devices, it contemplates any suitable storage device.

[0105] In some examples, communication interface 84 includes hardware, software, or both that provides one or more interfaces for communication (e.g., packet-based communication) between processing system 800 and one or more other processing systems 800 or one or more networks. By way of example, and not limitation, communication interface 84 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wire-based networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks such as Wi-Fi networks. This disclosure considers any suitable network and any suitable communication interface for that network. By way of example, and not limitation, processing system 800 may communicate with one or more portions of an ad hoc network, personal area network (PAN), local area network (LAN), wide area network (WAN), metropolitan area network (MAN), or the Internet, or a combination of two or more of these networks. One or more portions of one or more of these networks may be wired or wireless. As an example, the processing system 800 may communicate with the following networks: wireless PAN (WPAN) (e.g., Bluetooth WPAN), Wi-Fi networks, Wi-MAX networks, cellular telephone networks (e.g., Global System for Mobile Communications (GSM) networks), or other suitable wireless networks, or combinations of two or more of these networks. Where appropriate, the processing system 800 may include any suitable communication interface 84 for any of these networks. Where appropriate, the communication interface 84 may include one or more communication interfaces 84. Although specific communication interfaces are described and shown in this disclosure, any suitable communication interface is contemplated in this disclosure.

[0106] The components of the processing system 800 may include a processor 91, RAM 82, ROM 93, memory controller 92, storage device 97, input / output (I / O) interface 86, communication interface 84, and bus 80. Although this disclosure describes and illustrates a particular processing system having a particular number of components arranged in a particular manner, this disclosure contemplates any suitable processing system having any suitable number of components arranged in any suitable manner.

[0107] In some examples, ROM 93 includes main memory for storing instructions to be executed by processor 91 or data to be operated by processor 91. RAM 82 may include temporary memory for transfers to main memory (e.g., ROM 93) that may be determined by processor 91. By way of example and not limitation, processing system 800 may load instructions from storage device 97 or another source (e.g., another processing system 800) into ROM 93. Processor 91 may then load these instructions from ROM 93 into internal registers or internal cache memory. To execute these instructions, processor 91 may retrieve and decode them from internal registers or internal cache memory. During or after the execution of these instructions, processor 91 may write one or more results (which may be intermediate or final results) into internal registers or internal cache memory. Processor 91 may then write one or more of these results into ROM 93 or RAM 82. In a particular example, processor 91 executes only instructions in one or more internal registers or internal cache memory or ROM 93 or RAM 82 (instead of storage device 97 or other locations), and operates only on data in one or more internal registers or internal cache memory or ROM 93 or RAM 82 (instead of storage device 97 or other locations).

[0108] Figure 9 An example flowchart illustrating operations for translating voice data associated with one or more conversations or sessions related to a user, according to an example of this disclosure, is shown. In operation 900, a device (e.g., HMD 100) can detect one or more audio signals associated with voice data of at least one first user (e.g., person 130) and can determine that the voice data is associated with a first language (e.g., Spanish, Portuguese, etc.). In operation 902, the device (e.g., HMD 100) can translate one or more words of the voice data associated with the first language into one or more other words associated with a second language different from the first language (e.g., English, German, etc.).

[0109] In operation 904, the device (e.g., HMD 100) may present one or more other words translated into a second language as one or more text items (e.g., text item 602, text item 604) to the display device of at least one second user's device (e.g., one or more displays 108, displays / touchpads / one or more interfaces 42), or output one or more other words in the second language as audio content to the at least one second user. The audio content may be output by the device's (e.g., HMD 100's) audio devices (e.g., one or more audio devices 110, speakers / microphones 81). In some examples, the audio content (e.g., the audio content of "haciendolo bien", "bastante bien", "nada mal") may be output in response to the selection / triggering of one or more corresponding audio buttons (e.g., audio buttons 606, 607, 608).

[0110] In operation 906, a device (e.g., HMD 100) may, in response to voice data, generate a first set of one or more responses in a first language (e.g., responses 601, 603, 605) presented via a display device, and a second set of one or more corresponding other responses in a second language (e.g., responses 612, 614, 616). The one or more other responses may include translations of the one or more responses in the first language into the second language. In some examples, one or more other devices (e.g., UE 30, processing system 800, etc.) may implement flowchart operations (e.g., operations 900, 902, 904, 906).

[0111] Alternative embodiments In this document, where appropriate, one or more computer-readable non-transitory storage media may include one or more semiconductor-based integrated circuits (ICs) or other ICs (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), hard disk drives (HDDs), hybrid hard disk drives (HHDs), optical disks, optical disc drives (ODDs), magneto-optical disks, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, security digital cards or security digital drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these storage media. Where appropriate, computer-readable non-transitory storage media may be volatile, non-volatile, or a combination of volatile and non-volatile.

[0112] In some examples, the processing system 800 may incorporate a speaker / microphone 81 for capturing audio signals or providing audio associated with the AR system to support augmented reality functionality. In such examples, the processing system 800 may further include, for example, one or more speakers or audio sensors. Multiple speakers or audio sensors may be coupled to the processor 91 via a bus 80 and used to manage the transmission of control signaling data between the processor 91 and the audio sensors 81.

[0113] It should be recognized that, in application, the examples of the various methods and apparatuses described herein are not limited to the details of the construction and arrangement of the various components set forth in the following description or shown in the accompanying drawings. These methods and apparatuses can be implemented in other examples and can be practiced or performed in various ways. Examples of specific implementations are provided herein for illustrative purposes and are not intended to limit the scope of these specific implementations. In particular, the various actions, elements, and features described in conjunction with any one or more examples are not intended to exclude similar effects in any other examples. These methods are contemplated for application to users or groups. For example, an alarm associated with an motivator can be determined by the mood or other health information of a group. The motivator can be a group-related activity rather than an individual-related activity.

[0114] In this document, unless otherwise expressly indicated or indicated by the context, "or" is inclusive rather than exclusive. Therefore, in this document, unless otherwise expressly indicated or indicated by the context, "A or B" means "A, B, or both." Furthermore, unless otherwise expressly indicated or indicated by the context, "and" is both common and separate. Therefore, in this document, unless otherwise expressly indicated or indicated by the context, "A and B" means "A and B, commonly or separately."

[0115] The foregoing descriptions of the examples have been presented for illustrative purposes; such descriptions are not intended to be exhaustive or to limit the patent right to the precise form disclosed. Those skilled in the art will recognize that many modifications and variations are possible in light of this disclosure.

[0116] The scope of this disclosure covers all changes, substitutions, variations, alterations, and modifications to the exemplary examples described or shown herein, which will be understood by those skilled in the art. The scope of this disclosure is not limited to the exemplary examples described or shown herein. Furthermore, although this disclosure describes and illustrates various examples herein as including specific components, elements, features, functions, operations, or steps, those skilled in the art will understand that any example of these examples may include any combination or arrangement of any component, element, feature, function, operation, or step described or shown anywhere herein. Moreover, references in the appended claims to apparatus or systems, or components in apparatus or systems (that are adapted, arranged, capable, configured, implemented, operable, or usable to perform a particular function) cover that apparatus, system, or component (whether or not the apparatus, system, component, or the particular function is activated, turned on, or unlocked), provided that the apparatus, system, or component is so adapted, arranged, capable, configured, implemented, operable, or usable. Furthermore, although this disclosure describes or illustrates specific examples to provide specific advantages, specific examples may not provide these advantages, or may provide some or all of these advantages.

[0117] Finally, the language used in this specification has been chosen primarily for readability and instruction purposes, and may not have been selected to define or limit the subject matter of the invention. Therefore, the scope of the patent right is not limited by this specific embodiment, but rather by any of the claims published in this application. Thus, the disclosure of each example is intended to illustrate, not limit, the scope of the patent right, as set forth in the following claims.

Claims

1. A method comprising: Detect one or more audio signals associated with voice data of at least one first user, and determine that the voice data is associated with a first language; Translate one or more words of the speech data associated with the first language into one or more other words associated with a second language different from the first language; The device presents the one or more other words translated into the second language as one or more text items to the display of a head-mounted device (HMD) of at least one second user, or outputs the one or more other words in the second language as audio content to the at least one second user. as well as In response to the voice data, a first set of one or more responses in the first language and a second set of one or more corresponding other responses in the second language are generated and presented via the display, wherein the one or more other responses include translations of the one or more responses in the first language into the second language.

2. The method according to claim 1, further comprising: In response to detecting the selection of a first response among the one or more other responses, audio data of one or more words associated with the selected first response is output so that the at least one second user can speak the one or more words to the at least one first user.

3. The method according to claim 1 or 2, wherein, When the one or more other words translated into the second language are input as one or more text items and presented on the display of the HMD, one or more content items of the real-world environment can be viewed on the display.

4. The method according to any one of the preceding claims, wherein, While one or more content items in a real-world environment can be viewed on the display of the HMD, a first set of one or more replies and a second set of corresponding one or more other replies can also be viewed on the display of the HMD.

5. The method according to any one of the preceding claims, further comprising: Generating one or more pronunciations of the one or more other words translated into the second language, wherein, optionally, the method further includes: The HMD outputs audio data of the one or more pronunciations so that the at least one second user can pronounce the one or more pronunciations.

6. The method according to any one of the preceding claims, further comprising: In response to detecting the one or more audio signals and based on at least one setting associated with the device, it is determined that the second language is at least one language spoken by the at least one second user, so as to translate the one or more words of the speech data.

7. The method according to any one of the preceding claims, wherein: The presentation includes presenting the one or more other words translated into the second language as the one or more text items to the display, overlaying visual content items of the real-world environment captured by the display.

8. The method according to any one of the preceding claims, wherein, The HMD includes smart glasses, augmented reality devices, or virtual reality devices.

9. The method according to any one of the preceding claims, wherein, The HMD is configured to present one or more augmented reality content items, one or more virtual reality content items, one or more mixed reality content items, or a combination thereof, via the display.

10. The method according to any one of the preceding claims, further comprising: Determine that the voice data of the at least one first user is associated with the active or current conversation of the at least one second user; as well as During the dialogue, the display is provided with translations of a first set of the one or more other words or the one or more replies and a second set of the corresponding one or more other replies, so that the first user can use the one or more other words or the first set of the one or more replies or the second set of the one or more replies to reply to the voice data of the at least one first user in the first language.

11. An apparatus comprising: One or more processors; as well as At least one memory storing instructions that, when executed by the one or more processors, cause the device to: Detect one or more audio signals associated with voice data of at least one first user, and determine that the voice data is associated with a first language; Translate one or more words of the speech data associated with the first language into one or more other words associated with a second language different from the first language; The device displays the one or more other words translated into the second language as one or more text items to at least one second user, or outputs the one or more other words in the second language as audio content to the at least one second user; as well as In response to the voice data, a first set of one or more responses in the first language and a second set of one or more corresponding other responses in the second language are generated and presented via the display, wherein the one or more other responses include translations of the one or more responses in the first language into the second language.

12. The apparatus of claim 11, and one or more of the following: a) Wherein, the device includes a head-mounted display, smart glasses, augmented reality devices, or virtual reality devices; or b) Among them, When the one or more processors further execute the instructions, the device is configured to: In response to detecting the selection of a first response among the one or more other responses, audio data of one or more words associated with the selected first response is output so that the at least one second user can speak the one or more words to the at least one first user; c) Wherein, when the one or more other words translated into the second language are input as one or more text items and presented via the display, one or more content items of the real-world environment can be viewed through the display; or d) Wherein, while one or more content items in a real-world environment can be viewed on the display, a first set of one or more replies and a second set of corresponding one or more other replies can also be viewed on the display.

13. The apparatus according to claim 11 or 12, and one or more of the following: a) wherein, when the one or more processors further execute the instructions, the apparatus is configured to: Generate one or more pronunciations of the one or more other words translated into the second language; or b) Among them, When the one or more processors further execute the instructions, the device is configured to: In response to detecting the one or more audio signals and based on at least one setting associated with the device, it is determined that the second language is at least one language spoken by the at least one second user, so as to translate the one or more words of the speech data.

14. A non-transitory computer-readable medium storing instructions that, when executed, cause: Detect one or more audio signals associated with voice data of at least one first user, and determine that the voice data is associated with a first language; Translate one or more words of the speech data associated with the first language into one or more other words associated with a second language different from the first language; The translation of the one or more other words into the second language is presented as one or more text items to the display of a head-mounted device of at least one second user, or the translation of the one or more other words into the second language is output as audio content to the at least one second user; as well as In response to the voice data, a first set of one or more responses in the first language and a second set of one or more corresponding other responses in the second language are generated and presented via the display, wherein the one or more other responses include translations of the one or more responses in the first language into the second language.

15. The computer-readable medium according to claim 14, wherein, When the instruction is executed, it also causes: In response to determining the selection of a first response among the one or more other responses, audio data of one or more words associated with the selected first response is output so that the at least one second user can speak the one or more words to the at least one first user.