Methods, apparatuses and computer program products for providing large language models to facilitate live translations

The HMD-based speech translation system addresses inefficiencies in conventional systems by offering real-time translation and response capabilities, enabling seamless language conversion and user interaction in artificial reality environments.

WO2025151486A9PCT designated stage expired Publication Date: 2025-08-28META PLATFORMS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/010692
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-09
Filing Date
2025-01-08
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Conventional speech translation systems disrupt the flow of conversation and are inefficient, particularly in artificial reality environments, as they require users to type or speak in their language for translation, which can be time-consuming and disruptive.

Method used

A method and apparatus for real-time speech translation using a head-mounted device (HMD) that detects audio signals, translates speech from a first language to a second language, and presents the translation as text or audio to the user, allowing simultaneous viewing of the real-world environment and enabling users to respond in the translated language through generated replies.

Benefits of technology

Facilitates fluent conversational interactions by providing immediate and non-disruptive language translation, enhancing user experience in artificial reality environments by allowing users to respond in the translated language without the need for additional input devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025010692_28082025_PF_FP_ABST
    Figure US2025010692_28082025_PF_FP_ABST
Patent Text Reader

Abstract

A system and method for translating speech are provided. The system may detect audio signals associated with speech data of a first user(s) and determine the speech data is associated with a first language. The system may translate words of the speech data associated with the first language to other words associated with a second language. The system may present the other words translated in the second language as text items to a display of a device of a second user(s) or output the other words in the second language as audio to the second user(s). The system may generate, in response to the speech data, a first set of replies in the first language and a corresponding second set of other replies in the second language presented via the display. The other replies may include a translation in the second language of the replies in the first language.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS, APPARATUSES AND COMPUTER PROGRAM PRODUCTS FOR PROVIDING LARGE LANGUAGE MODELS TO FACILITATE LIVE TRANSLATIONSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application No. 63 / 619,136, filed January 9, 2024, entitled “Large Language Models in Live Translation,"’.TECHNOLOGICAL FIELD

[0002] Examples of the present disclosure relate generally to methods, apparatuses, and computer program products for a speech translation system.BACKGROUND

[0003] Electronic devices are constantly changing and evolving to provide a user with flexibility and adaptability. With increasing adaptability in electronic devices users are taking and keep their devices on their person daily for activities and as the world is becoming a global community users who may speak different languages may frequently interact.

[0004] Conventionally , communication between users or persons speaking differing languages may involve an interpreter using a device to translate a language to a user’s language (via a speech translation system), or entering text into a device that performs a translation (via a speech translation system). Each of these methods may disrupt the flow of conversation or feeling of connectedness between the persons attempting to communicate. For example, speech translation systems translating the speech of a first person to that of a second person may allow the second person to understand the first person but, in many examples, the second user may now either speak or type in their language on the speech translation system to receive a translation to the first person’s language. This process may take time depending on the reply and in some examples the translation may be inefficient. One such case where this is particularly apparent may be with artificial reality devices.

[0005] Artificial reality is a form of reality that has been adjusted in some manner before presentation to a user, which may include, for example, a virtual reality' (VR), an augmented reality (AR), a mixed reality (MR), a hybrid reality (HR), or some combination or derivative thereof. Artificial reality content may include completely computer-generated content or computer-generated content combined with captured (e.g., real-world) content. The artificial reality content may include video, audio, haptic feedback, or some combination thereof, any of which may be presented in a single channel or in multiple channels (such as stereo video that produces a three-dimensional (3D) effect to the viewer). Additionally, in some instances, artificial reality may be associated with applications, products, accessories, services, or some combination thereof, that may be used, for example, to create content in an artificial reality' orare otherwise used in (e.g.. to perform activities in) an artificial reality. Head-mounted displays (HMDs) including one or more near-eye displays may often be used to present visual content to a user for use in artificial reality' applications.

[0006] In view' of the foregoing drawbacks, it may be beneficial to provide an efficient and reliable method for speech translation that may foster a more fluent conversational pattern.BRIEF SUMMARY

[0007] Methods and systems are described for speech translation between a user and a person via HMDs or other electronic devices (e.g., smartphones, tablets, smartwatches, or any electronic device capable of communicating w ith a HMD).

[0008] According to an aspect of the present invention there is provided a method comprising: detecting one or more audio signals associated with speech data of at least one first user and determining that the speech data is associated with a first language; translating one or more words of the speech data associated w ith the first language to one or more other words associated with a second language different from the first language; presenting the one or more other words translated in the second language as one or more items of text to a display of a head mounted device (HMD) of at least one second user or outputting the one or more other words in the second language as audio content to the at least one second user; and generating, in response to the speech data, a first set of one or more replies in the first language and a corresponding second set of one or more other replies in the second language presented via the display, wherein the one or more other replies comprises a translation in the second language of the one or more replies in the first language.

[0009] Optionally, the method further comprises: outputting, in response to detection of a selection of a first reply of the one or more other replies, audio data of one or more words associated with the selected first reply to enable the at least one second user to speak the one or more words to the at least one first user.

[0010] Optionally, one or more content items of a real-world environment are view able by the display of the HMD while the one or more other words translated in the second language as the one or more items of text input are simultaneously being presented via the display.

[0011] Optionally the first set of the one or more replies and the corresponding second set of the one or more other replies are view able by the display of the HMD simultaneously with one or more content items of a real-world environment being viewable by the display of the HMD.

[0012] Optionally, the method further comprises: generating one or more pronunciations of the one or more other words translated in the second language.

[0013] Optionally, the method further comprises: outputting, by the HMD, audio data of the one or more pronunciations to enable the at least one second user to speak the one or more pronunciations.

[0014] Optionally, the method further comprises: determining, in response to the detecting the one or more audio signals and based on at least one setting associated with the apparatus, that the second language is at least one language spoken by the at least one second user to facilitate the translating of the one or more words of the speech data.

[0015] Optionally, the presenting comprises presenting the one or more other words translated in the second language as the one or more items of text to the display overlaid on viewable content items of a real-world environment captured by the display.

[0016] Optionally, the HMD comprises smart glasses, an augmented reality device, or a virtual reality device.

[0017] Optionally, the HMD is configured to present one or more augmented reality content items, virtual reality content items, hybrid reality content items, or a combination thereof, via the display.

[0018] Optionally, the method further comprises: determining that the speech data of the at least one first user is associated with an active or current conversation with the at least one second user; and providing the translating of the one or more other words or the first set of the one or more replies and the corresponding second set of the one or more other replies to the display during the conversation to enable the first user to utilize the one or more other words, or the first set of the one or more replies or the second set of the one or more replies to reply, in the first language, to the speech data of the at least one first user.

[0019] According to a further aspect of the present invention there is provided an apparatus comprising: one or more processors; and at least one memory storing instructions, that when executed by the one or more processors, cause the apparatus to: detect one or more audio signals associated with speech data of at least one first user and determine that the speech data is associated with a first language; translate one or more words of the speech data associated with the first language to one or more other words associated with a second language different from the first language; present the one or more other words translated in the second language as one or more items of text to a display of the apparatus of at least one second user or output the one or more other words in the second language as audio content to the at least one second user; and generate, in response to the speech data, a first set of one ormore replies in the first language and a corresponding second set of one or more other replies in the second language presented via the display, wherein the one or more other replies comprises a translation in the second language of the one or more replies in the first language.

[0020] Optionally, the apparatus comprises a head mounted display, smart glasses, an augmented reality device, or a virtual reality device.

[0021] Optionally, when the one or more processors further execute the instructions, the apparatus is configured to: output, in response to detection of a selection of a first reply of the one or more other replies, audio data of one or more words associated with the selected first reply to enable the at least one second user to speak the one or more words to the at least one first user.

[0022] Optionally, one or more content items of a real-world environment are viewable by the display while the one or more other words translated in the second language as the one or more items of text input are simultaneously being presented via the display.

[0023] Optionally, the first set of the one or more replies and the corresponding second set of the one or more other replies are viewable by the display simultaneously with one or more content items of a real-world environment being viewable by the display.

[0024] Optionally, when the one or more processors further execute the instructions, the apparatus is configured to: generate one or more pronunciations of the one or more other words translated in the second language.

[0025] Optionally, when the one or more processors further execute the instructions, the apparatus is configured to: determine, in response to the detect the one or more audio signals and based on at least one setting associated with the apparatus, that the second language is at least one language spoken by the at least one second user to facilitate the translate of the one or more words of the speech data.

[0026] According to a further aspect of the present invention there is provided a non- transitory computer-readable medium storing instructions that, when executed, cause: detecting one or more audio signals associated with speech data of at least one first user and determining that the speech data is associated with a first language; translating one or more w ords of the speech data associated with the first language to one or more other words associated with a second language different from the first language; presenting the one or more other words translated in the second language as one or more items of text to a display of a head mounted device of at least one second user or outputting the one or more other words in the second language as audio content to the at least one second user; and generating,in response to the speech data, a first set of one or more replies in the first language and a corresponding second set of one or more other replies in the second language presented via the display, wherein the one or more other replies comprises a translation in the second language of the one or more replies in the first language.

[0027] Optionally, the instructions, when executed, further cause: outputting, in response to determination of a selection of a first reply of the one or more other replies, audio data of one or more words associated with the selected first reply to enable the at least one second user to speak the one or more words to the at least one first user.

[0028] In various examples, systems and / or methods may receive, via a device associated with a user, an audio signal associated with speech spoken by a person / user in a first language. An item(s) of text associated with the audio signal may be determined. The text may be displayed in a second language via a graphical user interface(s). A machine learning model(s) may develop a list of replies associated with a reply to the item(s) of text in the first language. The machine learning model(s) may utilize a neural network to develop an association between the item(s) of text, previous replies to one or more similar texts and / or a context(s) associated with the item(s) of text. The machine learning model(s) and / or an artificial intelligence (Al) model may generate the list of replies. The list of replies may be provided to the user via a graphical user interface(s) in a first language and a second language. Each of the replies may comprise a pronunciation(s) associated with each word(s) of the replies. The pronunciation(s) may be associated with the first language to aid / help a user in pronouncing a translated phrase(s), word(s), or the like. For instance, a user may select a reply, and the selected reply may be output to the user as audio via an audio device of an HMD in a staccato fashion to help the user reply by speaking to a person that is fluent in speaking the first language (e.g.. during a conversation in real-time).

[0029] The machine learning model(s) may be trained based on statistical models to analyze vast amounts of data, learning patterns and / or connections between words, phrases, natural language patterns, and / or previously selected replies associated with a user(s). In various examples, the machine learning model(s) may utilize one or more neural networks to develop associations between received audio signals and the associated text, natural language patterns, and / or previously selected replies. As described above, a list of replies to a user may be generated and provided by a graphical user interface(s) of a device (e.g., a HMD(s), smartphones, tablets, smartwatches, computing devices, communication devices, or the like). The list of replies may be in the form of customizable text, an image(s), a video(s). or any combination thereof.

[0030] In one example aspect of the present disclosure, a method is provided. The method may include detecting one or more audio signals associated with speech data of at least one first user and determining that the speech data is associated with a first language. The method may include translating one or more words of the speech data associated with the first language to one or more other words associated with a second language different from the first language. The method may include presenting the one or more other words translated in the second language as one or more items of text to a display of a head mounted device of at least one second user or outputting the one or more other words in the second language as audio content to the at least one second user. The method may include generating, in response to the speech data, a first set of one or more replies in the first language and a corresponding second set of one or more other replies in the second language presented via the display. The one or more other replies includes a translation in the second language of the one or more replies in the first language.

[0031] In another example aspect of the present disclosure, an apparatus is provided. The apparatus may include one or more processors and a memory including computer program code instructions. The memory and computer program code instructions are configured to, with at least one of the processors, cause the apparatus to at least perform operations including detecting one or more audio signals associated with speech data of at least one first user and determining that the speech data is associated with a first language. The memory and computer program code are also configured to, with the processor(s). cause the apparatus to translate one or more words of the speech data associated with the first language to one or more other words associated with a second language different from the first language. The memory' and computer program code are also configured to, with the processor(s), cause the apparatus to present the one or more other words translated in the second language as one or more items of text to a display of the apparatus of at least one second user or output the one or more other words in the second language as audio content to the at least one second user. The memory and computer program code are also configured to, with the processor(s), cause the apparatus to generate, in response to the speech data, a first set of one or more replies in the first language and a corresponding second set of one or more other replies in the second language presented via the display. The one or more other replies includes a translation in the second language of the one or more replies in the first language.

[0032] In yet another example aspect of the present disclosure, a computer program product is provided. The computer program product may include at least one non-transitory computer-readable medium including computer-executable program code instructions storedtherein. The computer-executable program code instructions may include program code instructions configured to detect one or more audio signals associated with speech data of at least one first user and determine that the speech data is associated with a first language. The computer program product may further include program code instructions configured to translate one or more words of the speech data associated with the first language to one or more other words associated with a second language different from the first language. The computer program product may further include program code instructions configured to present the one or more other words translated in the second language as one or more items of text to a display of a head mounted device of at least one second user or output the one or more other words in the second language as audio content to the at least one second user. The computer program product may further include program code instructions configured to generate, in response to the speech data, a first set of one or more replies in the first language and a corresponding second set of one or more other replies in the second language presented via the display. The one or more other replies includes a translation in the second language of the one or more replies in the first language.

[0033] Additional advantages will be set forth in part in the description which follows or may be learned by practice. The advantages will be realized and attainted by means of the elements and combinations particularly pointed out in the appended claims. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive, as claimed.BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The summary’, as well as the following detailed description, is further understood when read in conjunction with the appended drawings. For the purpose of illustrating the disclosed subject matter, there are shown in the drawings examples of the disclosed subject matter; however, the disclosed subject matter is not limited to the specific methods, compositions, and devices disclosed. In addition, the drawings are not necessarily drawn to scale. In the drawings:

[0035] FIG. 1 illustrates an example HMD in accordance with an example of the present disclosure.

[0036] FIG. 2 illustrates an example HMD environment for speech translation in accordance with an example of the present disclosure.

[0037] FIG. 3 illustrates an example method of speech translation, in accordance with an example of the present disclosure.

[0038] FIG. 4 illustrates a flow chart for generating a list of replies, in accordance with anexample of the present disclosure.

[0039] FIG. 5 illustrates an example of a machine learning framework in accordance with one or more examples of the present disclosure.

[0040] FIG. 6A illustrates a speech translation system in accordance with an example of the present disclosure.

[0041] FIG. 6B illustrates a speech translation system in accordance with an example of the present disclosure.

[0042] FIG. 7 illustrates an example block diagram of an HMD device, in accordance with an example of the present disclosure.

[0043] FIG. 8 illustrates an example processing system, in accordance with an example of the present disclosure.

[0044] FIG. 9 illustrates an example flowchart illustrating operations for translating speech data associated with a conversation(s), dialogue(s), or the like associated with users in accordance with an example of the present disclosure.

[0045] The figures depict various examples for purposes of illustration only. One skilled in the art will readily recognize from the following discussion that alternative examples of the structures and methods illustrated herein may be employed without departing from the principles described herein.DETAILED DESCRIPTION

[0046] Some embodiments of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all embodiments of the disclosure are show n. Indeed, various embodiments of the disclosure may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Like reference numerals refer to like elements throughout.

[0047] As used herein, the terms “data,” “content,” “information” and similar terms may be used interchangeably to refer to data capable of being transmitted, received and / or stored in accordance w ith embodiments of the disclosure. Moreover, the term “exemplary ”, as used herein, is not provided to convey any qualitative assessment, but instead merely to convey an illustration of an example. Thus, use of any such terms should not be taken to limit the scope of embodiments of the disclosure.

[0048] As defined herein a “computer-readable storage medium,” which refers to a non- transitory, physical or tangible storage medium (e g., volatile or non-volatile memory' device), may be differentiated from a “computer-readable transmission medium,” which refers to an electromagnetic signal.

[0049] As referred to herein, an “application” may refer to a computer software package that may perform specific functions for users and / or, in some cases, for another application(s). An application(s) may utilize an operating system (OS) and other supporting programs to function. In some examples, an application(s) may request one or more services from, and communicate with, other entities via an application programming interface (API).

[0050] As referred to herein, “artificial reality” may refer to a form of immersive reality that has been adjusted in some manner before presentation to a user, which may include, for example, a virtual reality, an augmented reality, a mixed reality, a hybrid reality, Metaverse reality or some combination or derivative thereof. Artificial reality content may include completely computer-generated content and / or computer-generated content combined with captured (e.g., real -world) content. In some instances, artificial reality may be associated with applications, products, accessories, services, or some combination thereof, that may be used, for example, to create content in an artificial reality or are otherwise used in (e.g., to perform activities in) an artificial reality.

[0051] As referred to herein, “artificial reality content” may refer to content such as video, audio, haptic feedback, or some combination thereof, any of which may be presented in a single channel or in multiple channels (such as stereo video that produces a three- dimensional effect to the viewer) to a user.

[0052] As referred to herein, a Metaverse may denote an immersive virtual / augmented reality world in which augmented reality’ devices may be utilized in a network (e.g., a Metaverse network) in which there may, but need not, be one or more social connections among users in the network. The Metaverse network may be associated with three- dimensional virtual worlds, online games (e.g., video games), one or more content items such as, for example, non-fungible tokens (NFTs) and in which the content items may, for example, be purchased with digital currencies (e.g., cryptocurrencies) and other suitable currencies.

[0053] As referred to herein, “staccato” or a “staccato manner” may refer to or denote a user(s). a person(s), an individual(s) speaking separate words one by one with distinct pauses between each of the words.

[0054] It is to be understood that the methods and systems described herein are not limited to specific methods, specific components, or to particular implementations. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.Exemplary Artificial Reality System

[0055] The present disclosure is generally directed to systems and methods of speech translation via audio and / or text utilizing speakers and / or microphones associated with an electronic device, such as a HMD. FIG. 1 illustrates an example HMD 100 associated with artificial reality content in accordance with example aspects of the present disclosure. The HMD 100 may be configured to present augmented reality content items, virtual reality’ content items, hybrid reality content items, or a combination thereof, via a display (e.g., display(s) 108). HMD 100 may include frame 102 (e.g., an eyeglasses frame), a camera(s) 104, a display(s) 108, a speech translation system(s) 114, and an audio device(s) 110 (e.g., speakers / microphone). In some examples, the speech translation system(s) 114 may be referred to herein as a speech translation component(s) 112. Display(s) 108 may be configured to direct images and / or videos to a surface 106 (e.g., a user’s eye or another structure). In some examples, the HMD 100 may be implemented in the form of augmented- reality glasses (e.g., smart glasses). Accordingly, display (s) 108 may be at least partially transparent to visible light to allow the user to view a real-world environment through the display(s) 108. The audio device(s) 110 (e.g., speakers / microphones) may provide audio associated with augmented-reality content to users and capture audio signals.

[0056] Tracking of surface 106 may be beneficial for graphics rendering or user peripheral input. In many systems, HMD 100 design may include one or more cameras 104 (e.g.. a front facing camera(s) away from a primary user 120 of FIG. 2 or a rear facing camera(s) towards a primary user 120). Camera(s) 104 may track movement (e.g., gaze) of an eye(s) of primary user 120 or line of sight associated with primary' user 120. In some examples, a person (e.g., person 130 of FIG. 2) may be in the line of sight of user 120. HMD 100 may include an eye tracking system to track the vergence movement of primary user 120. One or more cameras 104 may be the eye tracking system. Camera(s) 104 may capture images and / or videos of an area(s), and / or capture video and / or images associated with surface 106 (e.g., eyes of primary' user 120 or other areas of a face (e.g., a face of user 120) depending on the directionality’ and view of camera(s) 104. In examples in which camera(s) 104 is rear facing towards a primary user 120. camera(s) 104 may capture images and / or videos associated with surface 106. In examples in which camera(s) 104 is front facing away from primary' user 120, camera(s) 104 may capture images and / or videos of an area(s) (e.g., an area(s) of a real-world environment). HMD 100 may be designed to have both front facing and rear facing cameras (e.g., camera(s) 104). There may be multiple cameras 104 that may be used to detect the reflection off of surface 106 or other movements (e.g., glint imagesand / or any other suitable characteristics). Camera(s) 104 may be located on frame 102 in different positions. Camera(s) 104 may be located along a width of a section of frame 102. In some other examples, the camera(s) 104 may be arranged on one side of frame 102 (e.g., a side of frame 102 nearest to the eye). Alternatively, in some examples, the camera(s) 104 may be located on display(s) 108. In some examples, camera(s) 104 may be sensors or a combination of cameras and sensors to track an eye(s) (e.g., surface 106) of a user.

[0057] Audio device(s) 110 may be located on frame 102 in different positions or any other configuration such as, but not limited to, headphone(s) communicatively connected to HMD 100, a peripheral device, or the like. Audio device(s) 110 may be located along a width of a section of frame 102. In some other examples, the audio device(s) 110 may be arranged on sides of frame 102 (e.g., a side of frame 102 nearest to an ear). In some examples, audio device(s) 110 may be sensors or a combination of speakers, microphones, and / or sensors to capture and produce sound associated with a user.

[0058] The speech translation component(s) 114 may be configured to receive one or more audio signals associated with, for example, speech data (e.g., voice data) of a user(s) in a language (e.g.. a first language(s)) and may translate audio of the audio signals in the received first language to one or more other languages, as described more fully below. The speech translation component(s) 114 may provide the audio in the translated language to the display (s) 108 to present to a user(s) by the display (s) 108 the audio in the translated language (e.g., a second language(s)) as text input shown in a real-world environment also captured in a field of view of a camera(s) and output by the display(s) 108. The text input associated with the audio in the translated language may be overlaid (e.g., superimposed) on the content of the real -world environment captured in the field of view of a camera(s) 104 and output by the display(s) 108. The speech translation component(s) 114 may also provide the audio in the translated language to the audio device(s) 110 to enable the audio device(s) 110 to output the audio in the translated language. In some examples, the audio in the translated language may be output by the audio device 110 in real-time (e.g., simultaneously) with the presentation of the text input in the translated language by the display(s) 108 of the HMD 100.Exemplary System Architecture

[0059] FIG. 2 illustrates an exemplary' environment for speech translation. The exemplary environment may include a system 101 to facilitate the speech translation. Primary user 120 may be associated with HMD 100, mobile device 111, or smartwatch 112. Users may be associated with devices based on being linked to user profiles. Base station 121 may be forwide area network access (e.g., a cellular system) or local area network access (e.g., Wi-Fi). HMD 100, mobile device 111. and / or smartwatch 112, may be communicatively connected with each other directly (e.g., via Bluetooth, near field communication, ultra-wideband, or any other suitable form of communicative connection) or by base station 121. Line-of-sight (LOS) area 125 may be based on a determination of the gaze(s) of primary user 120 and / or video (or an image) capture area of camera(s) 104 of HMD 100, in which a person 130 may be positioned in the LOS area 125.

[0060] In an example, an inertial measurement unit (IMU) (e.g., IMU 56 of FIG. 7) may be used to determine whether the inertial movement of smartwatch 112, mobile device 111, and / or HMD 100 may be used as a trigger for a selection of a reply associated with a list of replies. The IMU may be calibrated for different muscle(s) and / or body movements. Artificial intelligence may be used to determine the triggers for reply selection. In some other examples, inertial movement of smartwatch 112, mobile device 111, and / or HMD 100 may be determined via eye gaze tracking (EGT), electromyography (EMG), or any other suitable way to determine inertial and / or body movement(s). In some examples, the trigger may be associated with an input(s) associated with the user speaking being captured via audio device(s) 110.Exemplary System Operation

[0061] Some example aspects of the present disclosure may provide a machine learning model(s) to assist in providing live translations of languages in real time (e.g., during a conversation among users). In this regard, the examples of the present disclosure may address problems associated with a manner in which a first user(s) that may be fluent in speaking a first language(s) may be able to reply to another user(s) (e.g., a second user(s)) that may be fluent in speaking a second different language(s) in an instance in which the first user(s) may not understand / speak the second different language(s).

[0062] Further complicating matters, the second user(s) may not have or may not be wearing smart glasses (e.g., an HMD) and in some instances the speakers included in the smart glasses worn by the first user(s) may be sensitive such that only the first user(s) wearing the smart glasses may hear audio output from the speakers of the smart glasses. Additionally, the first user(s) may not have other hardware (e.g., an external display, etc.) that the first user(s) may use to read from to read translated audio to the second user(s).

[0063] In this regard, the examples of the present disclosure may address problems associated with a manner in which a first user(s) fluent in speaking a first language(s) is able to reply to another user(s) (e.g., a second user(s)) fluent in speaking a second differentlanguage(s) in an instance in which the first user(s) may not understand / speak the second different language(s).

[0064] In this regard, the exemplary aspects of the present disclosure may receive and / or detect audio signals (e.g., speech data) of a person speaking in a first language(s) and based on the received audio signals may translate the speech data to a second different language(s) that the user wearing the smart glasses may prefer / understand or be fluent in speaking. The exemplary aspects may generate (e.g.. automatically) one or more replies (e.g.. based on a captured video of a conversation) for conversation between the person and the user wearing the smart glasses. In some examples, the replies may be presented on a graphical user interface (e.g., a panel on a side (e.g., the left side) of the graphical user interface). A machine learning model(s) of the exemplary aspects may automatically translate possible replies to one or more words responsive to the detected translated utterance from the person that speaks the first language(s). In an instance in which the reply is selected by the user, the reply may be output, for example read off in a ‘‘repeat after me” sty le one word at a time so that the user wearing the smart glasses may speak each word(s) out loud to the person in the first language(s) spoken and understood by the person.

[0065] The repeat after me style of speech may help with the live translation(s) by displaying live streaming translations in another language, that may not be preferred, spoken, or understood by a user / wearer of smart glasses in an AR environment. For instance, the AR environment may be shown on the smart glasses illustrating a conversation among users in which a live translation may be shown on the display of the smart glasses (e g., as a caption along with the video of the conversation).

[0066] The application / implementation of the machine learning model(s) that may generate the replies may speed up the interactions of conversation among users, especially when no keyboard or other text input device may be present for the users to utilize during conversation.

[0067] FIG. 3 illustrates an example method 300 for speech translation, in accordance with an example of the present disclosure. The method 300 may be performed by one or more of processors (e.g.. processor 91), memories (e.g.. ROM 93 or RAM 82). microphone(s) (e.g.. speaker / mi crophone 81), and / or a memory controller (e g., memory controller 92). In some example aspects, the method 300 (and the steps 310, 320, 330, 340, 350, 360 of method 300) may be performed by an HMD (e.g.. HMD 100), a communication device (e.g., processing system 800), or a UE (e.g.. UE 30). At step 310, audio signals from a person (e.g., person 130) or environment may be captured via a microphone (e.g., speaker / mi crophone 81) andmay be stored temporarily in temporary memory' (e.g., RAM 82), without an input from a user. Audio signals captured and stored in a memory such as RAM 82 may be overwritten or replaced every / each lapse in a predetermined time period between received audio signals. The predetermined time period may be a predetermined number of seconds. The time period in which RAM 82 may store data (e.g., audio signals) before the data is replaced may be determined by the user of a device (e.g., HMD 100, processing system 800, UE 30, etc.), device settings stored in a storage device (e.g., storage 97), or when the lapse in predetermined time period between the received audio signals is approached or reached (e.g., the predetermined time period expires).

[0068] For purposes of illustration and not of limitation, for example, in an instance in which a user utilizes a HMD 100, the audio device(s) 110 may be capturing a first audio signal associated with a person 130 in the LOS 125 for 30 seconds. In this example, there may be another 25 seconds before a second audio signal associated with the person is received by the audio device(s) 110. As such, the device settings and / or user settings may be configured to store or retain the data captured (e.g., the first audio signal) from the audio device(s) 110 until 20 seconds has lapsed after capture of the input(s) associated with the first audio signal(s). Thus, after 20 seconds (which may be a time period in which audio signals received may not be associated with speech, or in which there are small amounts of nonspeech audio signals, or no audio signals are received) the first audio signal(s) or data captured via speaker / microphone 81 (e.g.. associated with an audio device) may be overwritten by a subsequent captured audio content. The audio content may, for example, be the second audio signal which may be associated with speech, voice data of a user (e.g., user 120, person 130), or the like of a next time period (e.g., seconds (e.g., a subsequent 30 second time period after the expiration of the 20 seconds, etc.)). In some examples, a machine learning models (e.g., machine learning model(s) 510) may be utilized to determine when / an instance in which the capture of an audio signal may be overwritten based on the content of the audio signal, for example, the machine learning model(s) may determine that the person associated with the audio signal has finished talking.

[0069] In alternative examples, audio signals captured and stored in a memory such as RAM 82 may be overwritten or replaced every predetermined time period (e.g., a predetermined number of seconds). The time period in which RAM 82 may store data (e.g., audio signals) before the data is replaced may be determined by the user of the HMD (e.g.. HMD 100), and / or may be determined based on device setings stored in a storage device (e.g., storage 97). For example, in an instance in which a user utilizes a HMD 100, and / or theaudio device(s) 110 may be capturing a first audio signal associated with a person 130 in the LOS 125 for 30 seconds. There may be 25 seconds before a second audio signal associated with the person is received, and as such, as an example, the device settings and / or user settings may be configured (e.g., designated) to facilitate storage or retaining the data captured by the audio device(s) for a predetermined time period such as for example 45 seconds. Thus, every 45 seconds during the capture of audio signals or data captured via a speaker / mi crophone 81, associated with an audio device, may be overwritten by the next 45 seconds of audio capture (e g., other audio signals).

[0070] At step 320, a processor (e.g., processor 91) may reference a memory (e.g., ROM 93 or RAM 82) to translate the audio signal received in step 310 and provide text associated with the audio signal received. The audio signal received may be in a first language, and processor 91 may translate the received audio signal in the first language to a second different language associated with a user (e.g., user 120). Further, processor 91 may provide / generate text associated with the first language in both the first language and the second language. For example, a user of HMD 100 may be having a conversation w ith a person 130, and the HMD 100 may receive / capture audio of the conversation as an audio signal in a first language associated with the person 130. The processor 91 may perform a translation of the received audio signal in the first language (e.g., Spanish) to a second different language associated with the user (e.g., user 120). The translation may be provided in text (e.g., presented via a display device) and / or an audio signal to the user (e.g., user 120) via audio device(s) 110 in the second language (e.g., English) associated with the user 120.

[0071] At step 330, the translation may be provided to the user 120. The translation may comprise a visual (e.g., a text item) and / or an audio (e.g., audio signals) component associated with the second language based on the received audio signals corresponding to the first language. The translation may be provided via a display (e.g., display(s) 108) and / or audio device (e.g., audio device(s) 110, speaker / microphone 81) of a HMD (e.g., HMD 100) or other communication device (e.g., processing system 800, UE 30). In an example, the translation may also be provided via a peripheral device (e.g.. mobile device 111, smartwatch 112, or the like) associated with HMD 100.

[0072] At step 340, a memory (e.g., ROM 93 or RAM 82) may be referenced to determine a list of replies. The memory may be referenced by a machine learning model(s) (e.g., machine learning model(s) 510), which may determine the list of replies. The list of replies may be associated with a received audio signal to foster communication (e.g.. conversation) betw een the user 120 and the person 130. The list of replies may be associatedwith a context(s) associated with the received audio signal of step 310. As described in more detail herein, the list of replies may be determined based on an assessment of the context(s) of the audio signals, a historical database, or the user of machine learning operations (also referred to as the machine learning model(s) herein), such as large language models (LLMs). The historical database may comprise books, newspapers, articles, conversations, television (TV) shows, movies, video clips or the like, and / or any combination thereof. At step 350, the list of replies may be provided to the user 120. The list of replies may comprise a text and / or an audio component. The list of replies may further comprise each reply in both the first language (e.g., Spanish, or Italian, etc.) and the second language (e.g., English, or German, etc.) for the user 120 to understand each reply. The list of replies may be provided via a display (e.g.. display(s) 108) of a HMD (e.g., HMD 100) associated with the user 120. In an example, the translation may also be provided via a peripheral device (e.g., mobile device 111, smartwatch 112, or the like) associated with HMD 100. In an example, the audio component of / associated with the list of replies may not be used until the user (e.g., user 120) triggers a selection of the list of replies. In some examples, each of the replies of the list of replies may comprise pronunciations associated with the reply (e.g., the word(s) of the reply) configured such that the user 120 may respond to the person (e.g., person 130) in the first language (e.g., Spanish, or Italian, etc.).

[0073] It is contemplated that step 310, step 320, step 330, and step 340 may occur at the same time or simultaneously. In some examples, step 310, step 320, step 330, step 340, and step 350 may occur at the same time or simultaneously.

[0074] At step 360, a user (e.g., user 120) may trigger a selection of the list of replies / responses. A trigger may be any audio, visual, motion, eye movement, or touch captured by a device (e.g., HMD 100, processing system 800. UE 30). In the example of an audio trigger, a user may begin to say a word or string of words associated with a reply- captured via a microphone (e.g., audio device(s) 110) of the HMD 100. Audio may also be determined by any other suitable method, function, or sensor. In the example of a touch trigger, a user may touch a button or an area of HMD 100 or peripheral device (e.g., mobile device 111, smartwatch 112. or the like) to select a reply of / from the hst of replies. The touch may be determined via a pressure sensor threshold being met, in which the pressure sensor threshold may be set via device settings and may be stored in a storage device (e.g., storage 97). The touch may also be determined by any other suitable method, function, or sensor. In the example of a visual trigger, a camera (e.g., camera(s) 104, camera 54) may be configured to discern or determine hand movements and / or gestures to initiate a selection of a reply ofthe list of replies. For example, the camera (e.g., camera(s) 104, camera 54) may be configured to determine an instance in which a user's hand is raised in the line of sight (e.g., LOS 125) to a predetermined level or height associated with each reply of the list of replies and the line of sight of a user. As such, in an instance in which the user’s hand may be raised to the predetermined level or height, a selection may be determined by the HMD 100. In the example of an eye movement trigger, a user may look at a reply of the list of replies associated with the display of HMD 100. The eye movement may be determined based on a predetermined glance threshold associated with an eye tracking system being met / satisfied, in which the predetermined glance threshold may be set via device settings and may be stored in a storage device (e.g., storage 97). Eye movement may also be determined by any other suitable method, function, or sensor. In the example of a motion trigger, a user may perform a motion or movement toward a reply of the list of replies associated with a display (e.g., display(s) 108) of HMD 100. The motion may be determined via a motion sensor or an IMU (e.g., IMU 56). The IMU may be an electronic device that measures and reports specific force, angular rate, orientation of a device (e.g., HMD 100) using a combination of accelerometers, gyroscopes, and in some cases magnetometers. The IMU may determine whether a trigger has surpassed or reached a predetermined threshold of the HMD 100. In another example of a motion trigger, a user may perform a motion or movement toward a reply for selection or may move a hand, finger, or the like in a predetermined manner that indicates a selection of a reply of the list of replies associated with the display (e.g., display(s) 108) of HMD 100. The motion may be determined via a predetermined EMG threshold being met / satisfied, in which the EMG may be captured via EMG sensors associated with HMD 100 and / or a peripheral device (e.g., mobile device 111, smartwatch 112, or the like). The EMG threshold may be set via device settings and may be stored in a storage device (e.g., storage 97). EMG sensors may measure muscle(s) response and / or electrical activity in response to a nerve’s stimulation of the muscle(s) of a user in response to movement or motion. Motion may also be determined by any other suitable method, function, or sensor.

[0075] In some examples, in an instance in which a selection has been chosen a pronunciation associated with the reply for selection may be provided to a device (e g., HMD 100) to aid the user 120 to respond in the first language associated with the person (e.g., person 130). In some examples, the audio device 110 may provide audio corresponding to the reply. In some examples, the audio device 110 may provide audio corresponding to the reply in a “sing-along” (e.g., one word at a time) fashion such that the user may hear how topronounce the words of a selected reply and repeat the words to respond to the person 130.

[0076] FIG. 4 illustrates a flow chart for generating a list of replies, in accordance with an example of the present disclosure. At block 402, a device (e.g., HMD 100) may reference a storage device (e.g., storage 97) to assess and utilize a profile associated with a user (e.g., user 120). The profile may comprise any number of ty pes of data including, but not limited to, one or more of a style(s) associated with the user, one or more languages (e.g., one or more designated / spoken languages of the user), or any other suitable data. The style(s) may be based on previous replies chosen by the user, in which the previously selected replies and the context(s) in which the replies were chosen may comprise the style(s). At block 404, a device (e g., HMD 100) may receive audio signals (e.g., data) associated with speech / voice data, in which the speech / voice data may be of a different language from a language associated with the user (e.g., a language not spoken by the user). At block 410, the device (e.g., HMD 100) may receive and compare data associated with the profile and the audio signal, and based on the data received a determination may be made on whether a translation may be generated via a speech translation system (e.g., speech translation component(s) 114). In this regard, a speech translation system may be configured to determine that a translation may be needed due to a comparison between the data in the profile (e.g., a user profile) and the received audio signals (e.g., audio signals of the speech of a person (e.g., person 130)). In an example, the profile may indicate that the user speaks a language such as, for example, English (or other language (e.g.. German, etc.)) and the audio signals may be determined to be in a language such as Spanish (or other language (e.g., Portuguese, etc.)), indicating a potential need to translate the audio signals received (e.g., speech audio signal(s) of person 130) to the language (e.g., English) associated with the profile (e.g., an indication in the user profile indicating that the user speaks English). In other examples, the language a user speaks indicated in a profile (e.g., user profile) may be any other suitable language (e.g., other than English) and the language of the audio signals may be any other suitable language (e.g., other than Spanish) such as for example languages including, but not limited to, French, Italian, German, Portuguese. Dutch, Russian, Hebrew, Arabic, Mandarin, Hindi, Japanese and any other suitable languages.

[0077] At block 412, the device (e.g., HMD 100) may determine a time range of speech, in which audio signal data may be stored for a temporary period of time dependent on the settings of the device (e.g.. HMD 100). The time range may comprise a predetermined time period corresponding to a sentence(s) or portion of conversation (e.g., conversation among users), in which the increment of time may be any suitable increment of time (e.g.,milliseconds, seconds, minutes, hours, etc.) associated with a received audio signal associated with speech / voice data (e.g.. of person 130). At block 420, a machine learning model(s) (e.g., machine learning model(s) 510) may be applied to analyze the audio signal of the speech data of the person (e.g., person 130) in association with the profile (e.g., user profile) based on the audio signal captured within the time range. In some examples, a device (e.g., HMD 100) may execute the machine learning model(s) (e.g., associated with block 410). One or more machine learning models (e.g., machine learning model(s) 510) may analyze audio signals to determine context(s) of the conversation and may generate a list of replies to the received audio signal associated with a person (e.g., person 130). In various examples, at block 422 training data (e.g., training data 520) may be utilized by the machine learning model(s) to develop and predict associations between at least one of profile data, audio signal data. context(s) of the audio signal data, and / or the like. The device (e.g., HMD 100) may implement the machine learning model(s). The training data (e.g., training 520) may be historical data or associated with profile data, audio signal data (e.g., audio signal data in various languages), contextual data, pronunciations, and / or the like. The contextual data may be determined based on one or more determined contexts of conversations (e.g., morning greetings, nighttime greetings, dinner topics, breakfast topics, sports topics, etc.) In other examples, the training data may be associations between profile data (e.g., style(s)), and / or audio signal data.

[0078] At block 424. a neural network may be utilized / implemented by a device (e.g., HMD 100) to assist in utilizing the training data (e.g., training data 520) or assisting with the machine learning model(s) techniques to analyze the profile data and audio signal data that may be associated with historical data (e.g., historical data of training data 520). For example, the neural network may be trained on historical data, of the training data 520, indicating communication (e.g., conversations) between two or more persons with various contextual backgrounds. In many instances, historical data may comprise, but is not limited to, books, movies, news articles, magazines, TV shows, previous conversations of other users, and / or the like. The previous conversations of other users, the books, movies, news articles, magazines. TV shows, etc. may be in various different languages including, but not limited to, English, Spanish, French, Italian, German, Portuguese, Dutch, Russian, Hebrew Arabic, Mandarin, Hindi. Japanese and any other suitable languages. The neural network may be implemented by a machine learning model(s) (e.g., machine learning model(s) 510). In some examples, the machine learning model(s) (e.g., machine learning model(s) 510) may be any large language model.

[0079] At block 430. a device (e.g., HMD 100) may generate a list of replies based on the association between the user profile, received audio signals, and / or historical data. The generated list of replies may utilize data (e.g., directly) from the machine learning model(s). In examples, the generated list of replies may be compiled on a display (e.g., display(s) 108) of the device (e.g., HMD 100) or via a graphical user interface(s) (e.g., display / touchpad / interface(s) 42) associated with a peripheral device (e.g., mobile device 111, smartwatch 112, etc.).

[0080] At block 440, a trigger indicating a selection of a reply may be determined via a device (e.g., HMD 100). The trigger may be any audio, visual, motion, eye movement, and / or touch captured by a device. In the example of an audio trigger, a user may begin to say a word or string of words associated with a reply captured via a microphone (e.g., audio device(s) 110) of the HMD 100. Audio may also be determined by any other suitable method, function, or sensor. In the example of touch, a user may touch a button or an area of HMD 100 or a peripheral device (e.g., mobile device 111, smartwatch 112, or the like) to select a reply of the list of replies. The touch may be determined via a pressure sensor threshold being met / satisfied, in which the pressure sensor threshold may be set via device settings and may be stored in the storage device (e.g., storage 97). Touch may also be determined by any other suitable method, function, or sensor. In the example of a visual trigger, a camera(s) (e.g., camera(s) 104, camera 54) may be configured to discern or determine hand movements and / or gestures to initiate a selection of a reply of the list of replies.

[0081] For example, the camera(s) (e.g., camera(s) 104, camera 54) may be configured to determine an instance in which a user’s hand is raised in the line of sight (e.g., LOS 125) to a predetermined level or height associated with each reply of the list of replies and the line of sight of a user. As such, in an instance in which the user’s hand may be raised to the predetermined level or height, a selection of a reply from the list of replies may be determined by HMD 100. In the example of an eye movement trigger, a user may look at a reply of the list of replies associated with the display (e.g., display(s) 108) of HMD 100. The eye movement may be determined via a predetermined glance threshold, associated with an eye tracking system (e.g., a rear facing camera (e.g., a rear facing camera(s) 104), being met / satisfied, in which the predetermined glance threshold may be set via device settings and may be stored in the storage device (e.g., storage 97). Eye movement may also be determined by any other suitable method, function, or sensor. As an example of a motion trigger, motion may be captured via a motion sensor (e.g.. an IMU (e.g., IMU 56)) of the HMD 100. in which the motion sensor may be configured to determine slight changes in motion to determine aselection of a reply from the list of replies. Motion may also be determined by any other suitable method, function, or sensor. In some examples, in an instance in which a selection of a reply from the list of replies has been chosen via a trigger, a pronunciation(s) associated with the selected reply may be generated and may be provided to a user to aid the user (e.g., user 120) to respond in a first language (e.g., Spanish, Portuguese, etc.) associated with a language of a person (e.g., person 130). In this regard, the pronunciation(s) may be presented to a display (e.g.. display(s) 108) and / or a graphical user interface (e.g.. display / touchpad / interface(s) 42), for example as text input and in some examples may be presented as an audio output of the pronunciation output via an audio device (e.g., audio device(s) 110, speaker / microphone 81) of the HMD 100.

[0082] At block 450. the selected reply may be provided to the user via a device (e.g., HMD 100). At block 452, a device (e.g., HMD 100) may pres ent / pro vide the selected reply in the form of an alert, a notification, image, instructional text, email, message, among others, in which the reply may comprise pronunciations associated with the words of the reply. In some examples, the reply may be presented via a peripheral device (e.g., mobile device 111, smartwatch 112). In some examples, the presentation of the reply may include at least one of, an amplification of the reply selected, or pronunciations being added to the reply. The pronunciations may apply to the first language (e.g., Spanish or Portuguese, etc.) of a person (e.g., person 130) and the second language (e.g., English or German, etc.) which may correspond to a language of the user 120. At block 454. an audio (e.g., audio form) may similarly be presented by a device (e.g., HMD 100) via an audio device (e.g., audio device(s) 110), in which the w ords of the selected reply may be presented as the audio to the ears of a user (e.g., user 120). In an example, the words may be presented to the user in a word by word or staccato manner to allow the user to repeat (in the user’s own voice) after each word presented via the audio device(s) 110 to the ear(s) of the user. For example, the staccato manner may entail an instance in which a reply is selected by the user (e.g., user 120), such that one or more words associated with the selected reply may be output (e g., read off) as audio by the audio device(s) 110 in a “repeat after me” style one word at a time such that the user (e.g., the wearer of the HMD 100) may speak out loud repeating the output word(s) one word at a time (that an ear(s) of the user hears) to the person (e.g., person 130) in the language (e.g., Spanish) of the person. In this manner, the conversation of the user and the person may be enhanced since the user may speak a language (e.g., a second language (e.g., English or German, etc.)) different than a language (e.g., a first language (e.g., Spanish or Portuguese, etc.) of the person.

[0083] FIG. 5 illustrates a machine learning and training model, in accordance with an example of the present disclosure. The framework 500 associated with the machine learning model(s) 510 may be hosted remotely. Alternatively, the framework 500 may reside within a device such as, for example, the HMD 100 shown in FIG. 1, the UE 30 of FIG. 7, the processing system 800 of FIG. 8 or be processed by another device(s) (e.g., mobile device 111, smartwatch 112) shown in FIG. 1. The machine learning model(s) 510 may be operably coupled to the stored training data 520 in a memory such as training database 550 or another database (e.g., ROM 93, RAM 82). In some examples, the machine learning model(s) 510 may be associated with operations of FIG. 3, FIG. 4 and / or FIG. 9. In some other examples, the machine learning model(s) 510 may be associated with other operations. In some examples, the machine learning model(s) 510 may be an example of speech translation component(s) 114. The machine learning model(s) 510 may be implemented by one or more machine learning components and / or another device (e.g., a server and / or a computing system). For example, in some example aspects, the machine learning model(s) 510 may be implemented by the HMD 100, the UE 30, the processing system 800, the mobile device 111, or the smartwatch 112. etc.

[0084] The training data 520 employed by the machine learning model(s) 510 may be pre-trained, fixed or updated periodically. Alternatively, the training data 520 may be updated in real-time based upon the evaluations performed by the machine learning model(s) 510 in a non-training mode. This may be illustrated by the double-sided arrow connecting the machine learning model(s) 510 and stored training data 520 which may be stored in the training database 550. Some other examples of the training data 520 may include, but are not limited to, items of content determined as being associated with a network (e g., the Internet, a social network, etc ), a platform, system (e.g., system 101) or the like.

[0085] In some example aspects, the training data 520 may be obtained based on user data of users associated with a system (e.g., system 101). The user data may be data, information or the like involving past / historical interactions / conversations and / or current (e.g., real time) interactions / conversations of the users associated with the system. In some example aspects, these users may opt in with the system to enable usage of user data to be utilized as training data 520. The past / historical interactions / conversations of users and / or current interactions / conversations of the users may include translations of audio content in one or more languages spoken by at least a subset of the users to one or more other different languages spoken by one or more other subsets of the users. In this manner, the machine learning model(s) 510 may utilize the training data 520 to translate content (e.g., audiocontent (e.g., speech / voice data) of a user) from one language(s) to another different language(s). Some of the training data 520 of the machine learning model(s) 510 may include user data associated with contextual information / variables associated with the interactions / conversations of users that may be utilized by the machine learning model(s) 510 to generate one or more replies (e.g., automated replies) in response to translated content (e.g., translated speech / voice data) determined by the machine learning model(s) 510.

[0086] FIG. 6A illustrates a speech translation system in accordance with an example of the present disclosure. FIG. 6A may illustrate the view a user (e.g., user 120) sees when looking / viewing through / viathe display (e.g., display(s) 108) of a HMD (e.g., HMD 100). In this regard, the user may see / view real-world objects / content in an environment of the user via the display (e.g., display(s) 108) in addition to the text input 602 and / or the text input 604 described below. The user 120 wearing HMD 100 may receive audio signals via an audio device (e.g., audio device(s) 110) of HMD 100. In some examples, the audio signal(s) may also be received via a communicatively connected device or a peripheral device (e.g., mobile device 111, smartwatch 112. or any other suitable device). As the audio device(s) 110 receives audio signals (e.g., speech data, voice data, etc.) associated with a person (e.g., person 130), the speech translation component(s) 114 may begin to provide a translation associated with a first language of the received audio signal to a second language. The translation may be provided via text (and / or associated audio output) corresponding to the original language or text input 602 in a first language (e.g., Spanish) of the received audio signal and the text input 604 in second language (e g., English). For example, as the person 130 says the term, in a first language (e.g., Spanish), “como,” the speech translation component(s) 114 may translate the spoken term (e.g., “como”) in real-time (e.g., simultaneously) to “how” in a second language such as English and the device (e.g., HMD 100) may display via display(s) 108, the translation “how” in the text input 604 in the second language (e.g., English). In some examples, the device (e.g., HMD 100) may display the translation “how” in the second language (e.g., English) in real-time (e.g., simultaneously with the person speaking “como”). Continuing with the previous example, as the person 130 continues to say / speak “estas” in the first language (e.g., Spanish), the speech translation component(s) 114 may translate the spoken term “estas” in real-time to “are you” in the text input 604 of second language such as English and the HMD 100 may display via display(s) 108 the rest / remainder of the translation “are you”. Thus, translating the speech associated with the person 130, who said, “como estas”, resulting in a translation to the second language, “how are you” displayed via display (s) 108 of HMD 100. In some examples, thecorresponding translation of the audio (e.g., speech data of person 130) from the first language to the second language may be provided to the user 120 via the display (s) 108 as the text input and (e.g., simultaneously) via audio device(s) 1 10 of HMD 100 as audio output of the translated text (e.g., an audio output of “how are you” in English based on the translated text “como estas” in Spanish). For example, the phrase “how are you,” in the second language, may be provided as audio output to the user 120 via audio device(s) 110 associated with HMD 100 as a result of receiving audio signals corresponding to the phrase “como estas” in a first language (e.g., Spanish).

[0087] FIG. 6B further illustrates a speech translation system in accordance with an example of the present disclosure. FIG. 6B may illustrate the view a user (e.g., user 120) sees when looking / viewing through / via the display (e.g.. display(s) 108) of a HMD (e.g., HMD 100). In this regard, the user may see / view real-world objects / content in an environment of the user via the display (e.g., display(s) 108) in addition to the list of replies 610 (e.g., replies 601, 603, 605, 612, 614, 616) and / or the audio buttons (e.g., audio buttons 606, 607, 608) described below. Following the example of FIG. 6A, in reply to receiving the audio signal associated with speech of a person (e.g.. person 130). for example the person speaking “como estas” as in FIG. 6A, a list of replies 610 may be generated (e.g., automatically generated) by the speech translation system(s) 114 and may be provided via a display (e.g., display(s) 108) associated with the HMD 100. The list of replies 610 may be determined or provided by the speech translation system(s) 114 based on the methods of FIG. 3 and FIG. 4. For purposes of illustration and not of limitation, in the example of FIG. 6B, the list of replies 610 to the speech “como estas” may comprise three replies 601, 603 and 605 in the first language (e.g., Spanish) and the corresponding translation of the three replies 612, 614, 616 in the second language (e.g., English). The list of replies may also comprise, or be associated with, one or more audio buttons 606, 607, 608 which may be initiated by any of the triggers described in FIG. 3 and FIG. 4. The list of relies (e.g., replies 601, 603, 605, 612, 614, 616), and / or the audio buttons (e.g., audio buttons 606, 607, 608) may be overlaid (e.g., superimposed) on content items in the real-world environment captured / viewable in the field of view of the display (s) 108.

[0088] Triggering one or more of the audio buttons 606, 607, 608 may initiate audio being presented to the user 120 via audio device(s) 110 associated with HMD 100. The audio may be presented in the first language (e.g., Spanish) to aid the user 120 in facilitating conversation with the person 130. The received audio signals associated with speech data of the person 130, which may be captured by an audio device (e.g., audio device(s) 110,speaker / microphone 81) of the HMD 100 of the user, may be in the first language and in the example of FIG. 6B, the user 120 may attempt to converse with the first person 130 in the first language of the person 130 by utilizing the HMD 100 (e.g., even though the primary language of the user 120 may be the second language e.g., English).

[0089] As such for example, in response to detection of a selection of the reply 601 byuser 120, and / or a trigger (e.g., selection) of audio button 606, the audio button 606 may cause the audio device (e.g.. audio device(s) 110, speaker / microphone 81) to output audio content of text '‘haciendolo bien!” associated with audio button 606 in the first language (e.g., Spanish), which the person 130 may understand and / or in which the person 130 may be fluent speaking. In this manner, the user 120 may be able to read the corresponding reply (e.g., reply 612 '’Doing great!’7) associated with the user’s 120 choice for selection in the user’s language (e.g., the second language (e.g., English)) and may speak the translation associated with the selected reply (e.g., reply 601 “haciendolo bien!”) as the audio content to the person 130. This may be beneficial since the person 130 in this example of FIG. 6B may not be utilizing smart glasses (e.g., may not be wearing an HMD 100) and instead may be communicating with the user 120 with the person 130’s own voice (e.g., speech content). The HMD 100 may detect the audio content of person 130 speaking and may translate the audio content of the person 130 speaking in a language to a preferred different language (e.g., English) of the user 120 for display via display (s) 108 of the translated audio content to the user 120 and may automatically present one or more replies to the user 120 to enable the user to speak to the person 130 (e.g., in reply to the person speaking to the user). In this regard, the HMD is able speed up and enhance the conversation between the user 120 and the person 130 and enables the user 120 to speak with the person 130 in a language in which the user 120 may not be fluent and may not understand, but in which the person 130 is fluent in speaking and understands the language. By utilizing the exemplary aspects of the present disclosure, the user 120 need not utilize a keyboard or other device(s) to type a translation of audio content to show the person 130 (e.g., manually show) while the user 120 and the person 130 converse. For example, the user 120 need not show the person 130 a display of translated audio / text content in the person’s 130 language in order to communicate (e.g., conversate) with the user 120.

[0090] In some alternative example aspects, in response to detection of a selection of a reply (the reply 601) by user 120, and / or a trigger (e.g., selection) of an audio button (e.g., audio button 606), the audio button may cause the audio device (e.g.. audio device(s) 110, speaker / microphone 81) to output audio content of text (e.g., “haciendolo bien!”) associatedwith audio button in the first language (e.g., Spanish), which the person 130 may understand and / or in which the person 130 may be fluent speaking. In this alternate example, the user 120 may utilize the HMD 100 to output audio (e.g., by audio device(s) 110, speaker / mi crophone 81) associated with the selected reply (e.g., reply 601 “haciendolo bien!”) and associated with the audio button (e.g., audio button 606) in the first language (e.g., Spanish) such that the HMD 100 directly outputs (e.g., via speaker / mi crophone 81) the audio associated with the selected reply (e.g., reply 601 "‘haciendolo bien!”) to the person 130. In this alternative example, the user 120 may not need to speak the audio of the selected reply (e.g., reply 601 “haciendolo bien!”) to the person 130 in order to conversate with the person 130.

[0091] In some examples, the audio associated with the audio buttons 606, 607, 608 maybe presented in a “sing-along” or staccato manner to aid the user 120 in pronunciation of the words in a first language (e.g., Spanish), in which the user 120 may speak a second language (e.g., English). For example, when asked by a person 130, in a first language (e.g., Spanish), “como estas” a list of replies may be generated and may be provided (e.g., by speech translation component(s) 1 14) in a second language (e.g.. English) corresponding to the user 120. For instance, one such reply may be “pretty good” associated with reply 614. Accompanied with the reply 614 may be a translation of the reply 603, generated by the speech translation system(s) 114, in the first language (e.g., Spanish), such as for example “bastante bien” corresponding to “pretty good” in the second language (e.g., English). As another example, accompanied with the reply 616 may be a translation of the reply 605, generated by the speech translation system(s) 114, in the first language (e.g., Spanish), such as for example “nada mal” corresponding to “Not bad” in the second language (e.g., English). In an alternate example, one or more pronunciations associated with the replies 601, 603, 605 in the first language (e.g., Spanish) may be provided as audio to the user to aid the user 120 in saying / speaking the replies to the person 130.

[0092] For instance, in some example aspects, the user may select a setting associated with the audio buttons 606, 607, 608 to output the audio associated with the corresponding replies (e.g., replies 601. 603, 605) to a speaker (e.g., speaker / microphone 81) of the HMD 100 such that an ear(s) of the user 120 may hear the pronunciation(s) of the words (e.g., “haciendolo bien!”, “bastante bien”, “nada mal”) of the associated replies (e.g., replies 601, 603, 605). In this example, the user 120 may utilize the pronunciation(s) of each of the words output by the speaker (e.g., speaker / microphone 81) to the ear(s) of the user 120 to enable the user 120 to speak the words, based on the pronunciation(s), directly to the person 130 (e.g..while conversating with the person 130).Exemplary Communication Device

[0093] FIG. 7 illustrates a block diagram of an example hardware / software architecture of user equipment (UE) 30. In some example aspects, the UE 30 may be examples of the mobile device 111, the smartwatch 112 or the HMD 100. As shown in FIG. 7, the UE 30 (also referred to herein as node 30) may include a processor 32, non-removable memory 44, removable memory 46, a speaker / microphone 38. a keypad 40, a display, touchpad, and / or interface(s) 42, a power source 48, a global positioning system (GPS) chipset 50, an IMU 56 and other peripherals 52. The UE 30 may also include a camera 54. In an example, the camera 54 may be a smart camera configured to sense / capture images appearing within one or more bounding boxes and may capture video(s). The IMU 56 may be an electronic device that measures and reports specific force, angular rate, orientation of a device (e.g., UE 30) using a combination of accelerometers, gy roscopes, and in some instances magnetometers. The IMU 56 may also determine inertial movement of a device (e.g., UE 30). Additionally, the IMU 56 may be a sensor (e.g., a motion sensor) configured to determine changes in motion of a device (e.g.. UE 30). The UE 30 may also include communication circuitry, such as a transceiver 34 and a transmit / receive element 36. It will be appreciated that the UE 30 may include any sub-combination of the foregoing elements while remaining consistent with an example.

[0094] The processor 32 may be a special purpose processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors in association with a DSP core, a controller, a microcontroller, Application Specific Integrated Circuits (ASICs), Field Programmable Gate Array (FPGAs) circuits, any other type of integrated circuit (IC), a state machine, and the like. In general, the processor 32 may execute computer-executable instructions stored in the memory (e.g., memory 44 and / or memory 46) of the node 30 in order to perform the various required functions of the node. For example, the processor 32 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the node 30 to operate in a wireless or wired environment. The processor 32 may run application-layer programs (e.g.. browsers) and / or radio access-layer (RAN) programs and / or other communications programs. The processor 32 may also perform security operations such as authentication, security key agreement, and / or cry ptographic operations, such as at the access-layer and / or application layer for example.

[0095] The processor 32 is coupled to its communication circuitry (e.g., transceiver 34and transmit / receive element 36). The processor 32, through the execution of computer executable instructions, may control the communication circuitry in order to cause the node 30 to communicate with other nodes via the network to which it is connected.

[0096] The transmit / receive element 36 may be configured to transmit signals to, or receive signals from, other nodes or networking equipment. For example, in an example, the transmit / receive element 36 may be an antenna configured to transmit and / or receive radio frequency (RF) signals. The transmit / receive element 36 may support various networks and air interfaces, such as wireless local area network (WLAN), wireless personal area network (WPAN), cellular, and the like. In yet another example, the transmit / receive element 36 may be configured to transmit and receive both RF and light signals. It will be appreciated that the transmit / receive element 36 may be configured to transmit and / or receive any combination of wireless or wired signals.

[0097] The transceiver 34 may be configured to modulate the signals that are to be transmitted by the transmit / receive element 36 and to demodulate the signals that are received by the transmit / receive element 36. As noted above, the node 30 may have multi-mode capabilities. Thus, the transceiver 34 may include multiple transceivers for enabling the node 30 to communicate via multiple radio access technologies (RATs), such as universal terrestrial radio access (UTRA) and Institute of Electrical and Electronics Engineers (IEEE 802.11), for example.

[0098] The processor 32 may access information from, and store data in, any type of suitable memory, such as the non-removable memory 44 and / or the removable memory 46. For example, the processor 32 may store session context in its memory, as described above. The non-removable memory 44 may include RAM, ROM, a hard disk, or any other type of memory storage device. The removable memory 46 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like. In other examples, the processor 32 may access information from, and store data in, memory that is not physically located on the node 30, such as on a server or a home computer.

[0099] The processor 32 may receive power from the power source 48 and may be configured to distribute and / or control the power to the other components in the node 30. The power source 48 may be any suitable device for powering the node 30. For example, the powder source 48 may include one or more dry' cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, and the like.

[0100] The processor 32 may also be coupled to the GPS chipset 50, which may beconfigured to provide location information (e.g., longitude and latitude) regarding the current location of the node 30. It will be appreciated that the node 30 may acquire location information by way of any suitable location-determination method while remaining consistent with an example.Exemplary Computing System

[0101] FIG. 8 illustrates an example schematic of an example processing system 800 that may implement components of a system or may be part of the UE 30 of FIG. 7. In some other example aspects, the processing system 800 may be a component(s) of the HMD 100. The processing system 800 is only one example of a suitable processing system 800 within a device (e g., mobile phone, laptop, tablet, or any device with messaging capabilities) and is not intended to suggest any limitation as to the scope of use or functionality of examples of the methodology described herein. The processing system 800 may comprise a computer or server and may be controlled primarily by computer readable instructions, which may be in the form of software, wherever, or by whatever means such software is stored or accessed. Such computer readable instructions may be executed within a processor, e.g., processor 91, to cause processing system 800 to operate. In operation, processor 91 fetches, decodes, and executes instructions, and transfers information to and from other resources via the processing systems 800 main data-transfer path, bus 80. Bus 80 connects the components in processing system 800 and defines the medium for data exchange. Bus 80 typically includes data lines for sending data, address lines for sending addresses, and control lines for sending interrupts and for operating the bus 80.

[0102] In particular examples, bus 80 includes hardware, software, or both coupling components of processing system 800 to each other. As an example and not by way of limitation, bus 80 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), a HYPERTRANSPORT (HT) interconnect, an Industry' Standard Architecture (ISA) bus, an INFINIBAND interconnect, a low-pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus. a serial advanced technology attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus or a combination of two or more of these. Bus 80 may include one or more buses, where appropriate. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnection.

[0103] Memories coupled to bus 80 include RAM 82 and ROM 93. Such memories mayinclude circuitry that allows information to be stored and retrieved. ROMs 93 generally contain stored data that cannot easily be modified. Data stored in RAM 82 may be read or changed by processor 91 or other hardware devices. In some examples, access to RAM 82 and / or ROM 93 may be controlled by memory controller. A Memory controller may provide an address translation function that translates virtual addresses into physical addresses as instructions are executed. Memory controllers may also provide a memory protection function that isolates processes within the system and isolates system processes from user processes. Thus, a program running in a first mode may access only memory mapped by its own process virtual address space; it cannot access memory' within another process’s virtual address space unless memory' sharing between the processes has been set up.

[0104] In some examples, I / O interface 86 includes hardware, software, or both, providing one or more interfaces for communication between processing system 800 and one or more I / O devices. Processing system 800 may include one or more of these I / O devices, where appropriate. One or more of these I / O devices may enable communication between a person and processing system 800. As an example, and not by way of limitation, a I / O device may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, sty lus, tablet, touch screen, video camera, another suitable I / O device, or a combination of tw o or more of these. An I / O device may include one or more sensors. This disclosure contemplates any suitable I / O devices and any suitable I / O interfaces for them. Where appropriate, I / O interface 86 may include one or more device or software drivers enabling processor 91 to drive one or more of these I / O devices. I / O interface 86 may include one or more I / O interfaces, w here appropriate. Although this disclosure describes and illustrates a particular I / O interface, this disclosure contemplates any suitable I / O interface.

[0105] In some examples, storage 97 includes mass storage for data or instructions. As an example, and not by way of limitation, storage 97 may include a hard disk drive (HDD), flash memory, random access memory (RAM), read only7memory (ROM), non-volatile read only memory' (NVROM) or a Universal Serial Bus (USB) drive or a combination of tw o or more of these. Storage 97 may include removable or non-removable (or fixed) media, where appropriate. Storage 97 may be internal or external to processing system 800, where appropriate. In some examples, storage 97 is non-volatile, solid-state memory. In particular examples, storage 97 includes read-only memory7(ROM). This disclosure contemplates mass storage taking any suitable physical form. Storage 97 may include one or more storage control units facilitating communication between processor 91 and storage 97, where appropriate. Where appropriate, storage 97 may include one or more storages 97. Althoughthis disclosure describes and illustrates particular storage, this disclosure contemplates any suitable storage.

[0106] In some examples, communication interface 84 includes hardware, software, or both providing one or more interfaces for communication (such as, for example, packet-based communication) between processing system 800 and one or more other processing systems 800 or one or more networks. As an example, and not by way of limitation, communication interface 84 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI network. This disclosure contemplates any suitable network and any suitable communication interface for it. As an example, and not by way of limitation, processing system 800 may communicate with an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area netw ork (MAN), or one or more portions of the Internet or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. As an example, processing system 800 may communicate with a wireless PAN (WPAN) (such as, for example, a BLUETOOTH WPAN), a WI-FI network, a WI-MAX netw ork, a cellular telephone network (such as, for example, a Global System for Mobile Communications (GSM) network), or other suitable wireless netw ork or a combination of two or more of these. Processing system 800 may include any suitable communication interface 84 for any of these networks, where appropriate. Communication interface 84 may include one or more communication interfaces 84, where appropriate. Although this disclosure describes and illustrates a particular communication interface, this disclosure contemplates any suitable communication interface.

[0107] The components of processing system 800 may include processor 91, RAM 82. ROM 93, memory controller 92, storage 97, input / output (I / O) interface 86, communication interface 84, and bus 80. Although the present disclosure describes and illustrates a particular processing system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable processing system having any suitable number of any suitable components in any suitable arrangement.

[0108] In some examples, ROM 93 includes main memory for storing instructions for processor 91 to execute or data for processor 91 to operate on. Whereas RAM 82 may include temporary memory for possible transfer to main memory (e.g., ROM 93) when determined by the processor 91. As an example, and not by way of limitation, processing system 800 mayload instructions from storage 97 or another source (such as, for example, another processingsystem 800) to ROM 93. Processor 91 may then load the instructions from ROM 93 to an internal register or internal cache. To execute the instructions, processor 91 may retrieve the instructions from the internal register or internal cache and decode them. During or after execution of the instructions, processor 91 may write one or more results (which may be intermediate or final results) to the internal register or internal cache. Processor 91 may then write one or more of those results to ROM 93 or RAM 82. In particular examples, processor 91 executes only instructions in one or more internal registers or internal caches or in ROM 93 or RAM 82 (as opposed to storage 97 or elsewhere) and operates only on data in one or more internal registers or internal caches or in ROM 93 or RAM 82 (as opposed to storage 97 or elsewhere).

[0109] FIG. 9 illustrates an example flowchart illustrating operations for translating speech data associated with a conversation(s), dialogue(s), or the like associated with users according to an example of the present disclosure. At operation 900, a device (e.g., HMD 100) may detect one or more audio signals associated with speech data of at least one first user (e.g., person 130) and may determine that the speech data is associated with a first language (e.g.. Spanish, Portuguese, etc.). At operation 902, a device (e.g.. HMD 100) may translate one or more words of the speech data associated with the first language to one or more other words associated with a second language (e.g., English, German, etc.) different from the first language.

[0110] At operation 904. a device (e.g., HMD 100) may present the one or more other w ords translated in the second language as one or more items of text (e.g., text item 602, text item 604) to a display device (e.g., display(s) 108, display / touchpad / interface(s) 42) of the device of at least one second user or output the one or more other w ords in the second language as audio content to the at least one second user. The audio content may be output by an audio device (e.g., audio device(s) 110, speaker / microphone 81) of the device (e.g., HMD 100). In some examples, the audio content (e.g., audio content of “haciendolo bien”, “bastante bien”, “nada mal”) may be output in response to a selection / trigger of a corresponding audio button(s) (e.g., audio buttons 606, 607, 608).

[0111] At operation 906. a device (e.g., HMD 100) may generate, in response to the speech data, a first set of one or more replies (e.g., replies 601, 603, 605) in the first language and a corresponding second set of one or more other replies (e.g., replies 612, 614, 616) in the second language presented via the display device. The one or more other replies may comprise a translation in the second language of the one or more replies in the first language. In some examples, one or more other devices (e.g., UE 30, processing system 800, etc.) mayimplement the flowchart operations (e.g.., operations 900, 902, 904, 906) of FIG. 9. Alternative Embodiments

[0112] Herein, a computer-readable non-transitory storage medium or media may include one or more semiconductor-based or other integrated circuits (ICs) (such, as for example, field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical discs, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM-drives, SECURE DIGITAL cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these, where appropriate. A computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile, where appropriate.

[0113] In some examples, the processing system 800 may incorporate a speaker / microphone 81 for purposes of capturing audio signals or providing audio associated with AR systems to support augmented reality functionality. In such examples, the processing system 800 may further include, for example, one or more speakers or audio sensors. A Plurality of speaker or audio sensors may be coupled, via bus 80, with processor 91, and operates to manage transfer of control signaling data between a processor 91 and the audio sensor 81.

[0114] It is to be appreciated that examples of the methods and apparatuses described herein are not limited in application to the details of construction and the arrangement of components set forth in the following description or illustrated in the accompanying drawings. The methods and apparatuses are capable of implementation in other examples and of being practiced or of being carried out in various ways. Examples of specific implementations are provided herein for illustrative purposes only and are not intended to be limiting. In particular, acts, elements and features described in connection with any one or more examples are not intended to be excluded from a similar role in any other examples. It is contemplated that methods may apply to the user or to the group. For example, energizer related alerts may be determined by the groups mood or other wellness information. Energizers may be group related activity rather than an individual related activity.

[0115] Herein, “or” is inclusive and not exclusive, unless expressly indicated otherwise or indicated otherwise by context. Therefore, herein, “A or B” means “A, B, or both,” unless expressly indicated otherwise or indicated otherwise by context. Moreover, “and” is both joint and several, unless expressly indicated otherwise or indicated otherwise by context.Therefore, herein, “A and B"’ means “A and B, jointly or severally,'’ unless expressly indicated otherwise or indicated otherwise by context.

[0116] The foregoing description of the examples has been presented for the purpose of illustration; it is not intended to be exhaustive or to limit the patent rights to the precise forms disclosed. Persons skilled in the relevant art can appreciate that many modifications and variations are possible in light of the disclosure.

[0117] The scope of this disclosure encompasses all changes, substitutions, variations, alterations, and modifications to the example examples described or illustrated herein that a person having ordinary' skill in the art would comprehend. The scope of this disclosure is not limited to the example examples described or illustrated herein. Moreover, although this disclosure describes and illustrates respective examples herein as including particular components, elements, feature, functions, operations, or steps, any of these examples may include any combination or permutation of any of the components, elements, features, functions, operations, or steps described or illustrated anywhere herein that a person having ordinary skill in the art would comprehend. Furthermore, reference in the appended claims to an apparatus or system or a component of an apparatus or system being adapted to. arranged to, capable of, configured to, enabled to, operable to, or operative to perform a particular function encompasses that apparatus, system, component, whether or not it or that particular function is activated, turned on, or unlocked, as long as that apparatus, system, or component is so adapted, arranged, capable, configured, enabled, operable, or operative. Additionally, although this disclosure describes or illustrates particular examples as providing particular advantages, particular examples may provide none, some, or all of these advantages.

[0118] Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the patent rights be limited not by this detailed description, but rather by any claims that issue on an application based hereon. Accordingly, the disclosure of the examples is intended to be illustrative, but not limiting, of the scope of the patent rights, which is set forth in the following claims.

Claims

WHAT IS CLAIMED:

1. A method comprising: detecting one or more audio signals associated with speech data of at least one first user and determining that the speech data is associated with a first language; translating one or more words of the speech data associated with the first language to one or more other words associated with a second language different from the first language; presenting the one or more other words translated in the second language as one or more items of text to a display of a head mounted device (HMD) of at least one second user or outputting the one or more other words in the second language as audio content to the at least one second user; and generating, in response to the speech data, a first set of one or more replies in the first language and a corresponding second set of one or more other replies in the second language presented via the display, wherein the one or more other replies comprises a translation in the second language of the one or more replies in the first language.

2. The method of claim 1. further comprising: outputting, in response to detection of a selection of a first reply of the one or more other replies, audio data of one or more words associated with the selected first reply to enable the at least one second user to speak the one or more words to the at least one first user.

3. The method of claim 1 or 2, wherein one or more content items of a real-world environment are viewable by the display of the HMD while the one or more other words translated in the second language as the one or more items of text input are simultaneously being presented via the display.

4. The method of any preceding claim, wherein the first set of the one or more replies and the corresponding second set of the one or more other replies are viewable by the display of the HMD simultaneously with one or more content items of a real-world environment being viewable by the display of the HMD.

5. The method of any preceding claim, further comprising: generating one or more pronunciations of the one or more other words translated in the second language, in which case optionally further comprising: outputting, by the HMD, audio data of the one or more pronunciations to enable the at least one second user to speak the one or more pronunciations.

6. The method of any preceding claim, further comprising: determining, in response to the detecting the one or more audio signals and based onat least one setting associated with the apparatus, that the second language is at least one language spoken by the at least one second user to facilitate the translating of the one or more words of the speech data.

7. The method of any preceding claim, wherein: the presenting comprises presenting the one or more other words translated in the second language as the one or more items of text to the display overlaid on viewable content items of a real-world environment captured by the display.

8. The method of any preceding claim, wherein the HMD comprises smart glasses, an augmented reality device, or a virtual reality' device.

9. The method of any preceding claim, wherein the HMD is configured to present one or more augmented reality content items, virtual reality content items, hybrid reality content items, or a combination thereof, via the display.

10. The method of any preceding claim, further comprising: determining that the speech data of the at least one first user is associated with an active or cunent conversation with the at least one second user; and providing the translating of the one or more other words or the first set of the one or more replies and the corresponding second set of the one or more other replies to the display during the conversation to enable the first user to utilize the one or more other words, or the first set of the one or more replies or the second set of the one or more replies to reply, in the first language, to the speech data of the at least one first user.1 1. An apparatus comprising: one or more processors; and at least one memory' storing instructions, that when executed by the one or more processors, cause the apparatus to: detect one or more audio signals associated with speech data of at least one first user and determine that the speech data is associated with a first language; translate one or more words of the speech data associated with the first language to one or more other words associated with a second language different from the first language; present the one or more other words translated in the second language as one or more items of text to a display of the apparatus of at least one second user or output the one or more other words in the second language as audio content to the at least one second user; and generate, in response to the speech data, a first set of one or more replies in thefirst language and a corresponding second set of one or more other replies in the second language presented via the display, wherein the one or more other replies comprises a translation in the second language of the one or more replies in the first language.

12. The apparatus of claim 11, and any one or more of: a) wherein the apparatus comprises a head mounted display, smart glasses, an augmented reality device, or a virtual reality device; or b) wherein when the one or more processors further execute the instructions, the apparatus is configured to: output, in response to detection of a selection of a first reply of the one or more other replies, audio data of one or more words associated with the selected first reply to enable the at least one second user to speak the one or more words to the at least one first user; or c) wherein one or more content items of a real-world environment are viewable by the display while the one or more other words translated in the second language as the one or more items of text input are simultaneously being presented via the display; or d) wherein the first set of the one or more replies and the corresponding second set of the one or more other replies are viewable by the display simultaneously with one or more content items of a real- world environment being viewable by the display.

13. The apparatus of claim 11 or 12, and any one or more of: a) wherein when the one or more processors further execute the instructions, the apparatus is configured to: generate one or more pronunciations of the one or more other words translated in the second language; or b) wherein when the one or more processors further execute the instructions, the apparatus is configured to: determine, in response to the detect the one or more audio signals and based on at least one setting associated with the apparatus, that the second language is at least one language spoken by the at least one second user to facilitate the translate of the one or more words of the speech data.

14. A non-transitory computer-readable medium storing instructions that, when executed, cause: detecting one or more audio signals associated with speech data of at least one first user and determining that the speech data is associated with a first language; translating one or more words of the speech data associated with the first language toone or more other words associated with a second language different from the first language; presenting the one or more other words translated in the second language as one or more items of text to a display of a head mounted device of at least one second user or outputting the one or more other words in the second language as audio content to the at least one second user; and generating, in response to the speech data, a first set of one or more replies in the first language and a corresponding second set of one or more other replies in the second language presented via the display, wherein the one or more other replies comprises a translation in the second language of the one or more replies in the first language.

15. The computer-readable medium of claim 14, wherein the instructions, when executed, further cause: outputting, in response to determination of a selection of a first reply of the one or more other replies, audio data of one or more words associated with the selected first reply to enable the at least one second user to speak the one or more words to the at least one first user.