Generating user interface in augmented reality environment

By using the AR system to analyze audio data and display related content items in an augmented reality environment, the problem of difficulty for users to perform operations and view teaching content at the same time is solved, and the operation efficiency and synchronization of content and actions are improved.

CN119998782APending Publication Date: 2025-05-13SNAP INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070406.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-04
Filing Date
2023-10-03
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to view teaching content while users perform teaching or operational steps, resulting in users frequently pausing or rewinding the content, affecting efficiency.

Method used

By using the AR system in an augmented reality environment, the audio data is analyzed to identify commands and display relevant content items in the user interface, allowing the user to view the content while performing an action.

Benefits of technology

It enables users to view content while performing the actions of the teaching process without frequent pauses or rewinds, improving the efficiency and synchronization of the workflow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998782A_ABST
    Figure CN119998782A_ABST
Patent Text Reader

Abstract

An augmented reality (AR) content system is provided. The AR content system may analyze audio input obtained from a user to generate a search request. The AR content system may obtain search results in response to the search request and determine a layout that displays the search results. The search results may be displayed in a user interface within the AR environment according to the layout. The AR content system may also analyze the audio input to detect commands executed regarding the content displayed in the user interface.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority declaration

[0002] This patent application claims the benefit of priority to U.S. patent application Serial No. 17 / 959,985, filed on October 4, 2022, which is incorporated herein by reference in its entirety. Technical Field

[0003] The present disclosure generally relates to generating user interfaces in an augmented reality environment. Background Art

[0004] A head-mounted device may be implemented with a transparent or translucent display through which a user of the head-mounted device can view the surrounding environment. Such a device enables a user to view the surrounding environment through the transparent or translucent display, and also to see objects generated for display that appear as part of the surrounding environment and / or superimposed on the surrounding environment (e.g., virtual objects such as renderings of 2D or 3D graphics models, images, videos, text, etc.). This is commonly referred to as "augmented reality" or "AR". A head-mounted device may also completely block the user's field of view and display a virtual environment through which the user can move or be moved. This is commonly referred to as "virtual reality" or "VR". As used herein, unless the context indicates otherwise, the term "AR" refers to either or both of augmented reality and virtual reality as traditionally understood.

[0005] Users of head mounted devices can access and use computer software applications to perform various tasks or participate in entertainment activities. In order to use the computer software application, the user interacts with a user interface provided by the head mounted device. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] To easily identify the discussion of any particular element or action, the highest-order digit or digits in a reference number refer to the figure number in which the element is first introduced.

[0007] Figure 1 is a perspective view of a head mounted device according to one or more examples.

[0008] Figure 2 Based on one or more examples Figure 1 Additional views of the head-mounted device.

[0009] Figure 3 is a diagrammatic representation of a machine in the form of a computing device according to one or more examples within which a set of instructions may be executed to cause the machine to perform any one or more of the methodologies discussed herein.

[0010] Figure 4is a diagram of an environment including one or more systems for determining a layout of content items and displaying information of the content items in an augmented reality environment according to one or more examples.

[0011] Figure 5 is a diagram including an architecture of a system for determining a content template for arranging information displayed in a user interface of an augmented reality environment, according to one or more examples.

[0012] Figure 6 is a diagram illustrating a user interface generated by a client device and displayed in an augmented reality environment according to one or more examples.

[0013] Figure 7 is a flow diagram of a process for determining an arrangement of information in a user interface displayed in an augmented reality environment according to one or more examples.

[0014] Figure 8 is a user interface including results of a search request displayed in an augmented reality environment according to one or more examples.

[0015] Fig. 9 is a user interface that includes information of content items and a menu of commands that can be executed with respect to display of the information according to one or more examples.

[0016] Fig.10 is a block diagram illustrating a software architecture within which the present disclosure may be implemented, according to some examples.

[0017] Fig.11 is shown according to some examples Figure 1 A block diagram of the details of the head mounted device.

[0018] Fig.12 is a diagrammatic representation of a networking environment in which the present disclosure may be deployed, according to some examples. DETAILED DESCRIPTION

[0019] In many augmented reality systems, users can interact with virtual objects displayed in their environment. Input modalities that can be used with AR systems are hand tracking and direct manipulation of virtual objects (DMVO), in which the user is provided with a user interface displayed to the user in an AR overlay with a two-dimensional (2D) or three-dimensional (3D) rendering. The rendering is a 2D or 3D graphical model in which the virtual objects located in the model correspond to interactive elements of the user interface. In this way, the user perceives the virtual objects as objects within the overlay in the user's field of view of the real world scene when wearing the AR system, or perceives the virtual objects as objects within the virtual world as seen by the user when wearing the AR system. In order to allow the user to manipulate the virtual objects, the AR system detects the user's hands and tracks the movement, position, and / or location of the hands to determine the user's interaction with the virtual objects. In addition, the AR system can respond to commands provided by the user to determine the user's interaction with the virtual objects.

[0020] In existing systems, which are generally not AR systems, a user is generally unable to view and interact with objects in their environment while accessing content via a computing device. For example, in a situation where a user is performing steps of a recipe, the user views the instructions for the recipe on a computing device, and then shifts their attention away from the instructions displayed on the computing device to follow the steps of the recipe, such as by interacting with the recipe ingredients and kitchen tools. The user is unable to simultaneously view the instructions for the recipe and the ingredients and kitchen tools used to perform the instructions for the recipe within their field of view. The same scenario exists in many types of instructional content, where the user shifts their attention away from the instructions to perform the steps included in the instructions.

[0021] In addition, the instruction content is usually accessed via at least one of a continuous video or text content page. In these instances, the various steps of the teaching process are continuously presented to the user. Typically, the user watches the video or reads the content, and then stops the video or otherwise stops watching the teaching content to perform actions related to the teaching process. If the user does not stop the video, the teaching process will continue regardless of whether the user has performed the actions of the previously presented teaching process. Therefore, the actions being performed by the user become out of sync with the instructions being displayed, and the user frequently pauses or rewinds, and then plays or replays the content.

[0022] The implementation of the augmented reality system described herein can enable a user to view content while performing actions of a teaching process without frequently pausing or rewinding the content. In one or more examples, the AR system can analyze audio data to determine one or more commands included in the audio data. One or more commands can be related to a search request to obtain content related to one or more keywords included in the audio data. In response to the search request, search results including content items can be returned, wherein the content items can include video content, image content, text content, augmented reality content, audio content, or one or more combinations thereof.

[0023] Audio input can also be used to access content included in the content item. In various examples, the content of the content item included in the search results can be presented in a user interface displayed in an augmented reality environment. In at least some examples, a head-mounted computing device is used to display the user interface in an augmented reality environment. In one or more illustrative examples, a user interface can be displayed so that a user can view the user interface and objects included in a real-world scene. In this way, a user can view the instructional content presented in one or more user interfaces while performing actions on the instructional content of an item included in a real-world scene. Therefore, compared to existing systems, users can switch from viewing instructional content to performing actions related to the instructional content with minimal interruption.

[0024] In addition, the AR content system can analyze audio data obtained from the user to browse the teaching content. The teaching content can be arranged according to the discrete steps of the teaching process. For example, multiple user interfaces are used to present the teaching content, wherein each user interface provides content corresponding to the discrete steps of the teaching process. When the user completes the steps of the teaching process, the user can provide an audio input including a command to navigate to the next step of the teaching process. Therefore, the user is able to complete the current step of the teaching process while accessing the content of the current step, without having to frequently pause the playback of the teaching content as in the existing system to prevent the teaching content from moving to the next step before the user is ready.

[0025] Other technical features may be readily apparent to those skilled in the art from the following drawings, descriptions and appended claims.

[0026] Figure 1 According to some examples, a head mounted device (e.g., Figure 1100). The glasses 100 may include a frame 102 made of any suitable material, such as plastic or metal, including any suitable shape memory alloy. In one or more examples, the frame 102 includes a first optical element holder or a left optical element holder 104 (e.g., a display or lens holder) and a second optical element holder or a right optical element holder 106 connected by a bridge 112. A first optical element or a left optical element 108 and a second optical element or a right optical element 110 may be disposed in the respective left optical element holder 104 and the right optical element holder 106. The right optical element 110 and the left optical element 108 may be lenses, displays, display components, or a combination of the foregoing. Any suitable display component may be disposed in the glasses 100.

[0027] The frame 102 additionally includes a left arm or temple portion 122 and a right arm or temple portion 124. In some examples, the frame 102 can be formed from a single piece of material to have a unitary or integrated construction.

[0028] The glasses 100 may include a computing device such as a computer 120, which may be of any suitable type to be carried by the frame 102, and in one or more examples, the computing device may be of a suitable size and shape to be at least partially disposed in one of the temple components 122 or temple components 124. The computer 120 may include one or more processors with memory, wireless communication circuitry, and a power source. As discussed below, the computer 120 includes a low-power circuitry, a high-speed circuitry, and a display processor. Various other examples may include these elements in different configurations or integrated together in different ways. Additional details of various aspects of the computer 120 may be implemented as shown in the data processor 1002 discussed below.

[0029] The computer 120 additionally includes a battery 118 or other suitable portable power supply. In some examples, the battery 118 is disposed in the left temple portion 122 and is electrically coupled to the computer 120 disposed in the right temple portion 124. The glasses 100 may include a connector or port (not shown) suitable for charging the battery 118, a wireless receiver, transmitter or transceiver (not shown), or a combination of such devices.

[0030] The glasses 100 include a first or left camera 114 and a second or right camera 116. Although two cameras are depicted, other examples contemplate the use of a single or additional (i.e., more than two) cameras. In one or more examples, the glasses 100 include any number of input sensors or other input / output devices in addition to the left camera 114 and the right camera 116. Such sensors or input / output devices may additionally include biometric sensors, positioning sensors, motion sensors, etc.

[0031] In some examples, left camera 114 and right camera 116 provide video frame data for use by glasses 100 to extract 3D information from a real-world scene.

[0032] The glasses 100 may also include a touchpad 126 mounted to one or both of the left temple component 122 and the right temple component 124, or integrated with one or both of the left temple component 120 and the right temple component 122. The touchpad 126 is generally arranged vertically, approximately parallel to the temple of the user in some examples. As used herein, approximately vertical alignment means that the touchpad is more vertical than horizontal, although it may be more vertical than the vertical. Additional user input can be provided by one or more buttons 128, in the example shown, one or more buttons 228 are set on the outer upper edges of the left optical element holder 104 and the right optical element holder 106. One or more touchpads 126 and buttons 128 provide the following means, by which the glasses 100 can receive input from the user of the glasses 100.

[0033] Figure 2 The glasses 100 are shown from the user's perspective. Figure 1 Many of the components shown in FIG. have been omitted. Figure 1 As described, Figure 2 The illustrated eyeglasses 100 include left and right optical elements 108, 110 secured within left and right optical element holders 104, 106, respectively.

[0034] The glasses 100 include a front optical assembly 202 including a right projection portion 204 and a right near-eye display 206 , and a front optical assembly 210 including a left projection portion 212 and a left near-eye display 216 .

[0035] In some examples, the near-eye display is a waveguide. The waveguide includes a reflective structure or a diffractive structure (e.g., a grating and / or an optical element such as a mirror, a lens, or a prism). The light 208 emitted by the projection portion 204 encounters the diffractive structure of the waveguide of the near-eye display 206, which directs the light toward the right eye of the user to provide an image on or in the right optical element 110 that overlays the view of the real-world scene seen by the user. Similarly, the light 214 emitted by the projection portion 212 encounters the diffractive structure of the waveguide of the near-eye display 216, which directs the light toward the left eye of the user to provide an image on or in the left optical element 108 that overlays the view of the real-world scene seen by the user. The combination of the GPU, the forward optical assembly 202, the left optical element 108, and the right optical element 110 provides an optical engine of the glasses 100. The glasses 100 use the optical engine to generate an overlay of the user's view of the real-world scene, including displaying a user interface to the user of the glasses 100.

[0036] However, it should be understood that other display technologies or configurations may be utilized within the optical engine to display images to the user in the user's field of view. For example, instead of the projection portion 204 and waveguide, an LCD, LED or other display panel or surface may be provided.

[0037] In use, information, content, and various user interfaces will be presented to the user of the glasses 100 on the near-eye display. As described in more detail herein, the user can then use the touch pad 126 and / or buttons 128, associated devices (e.g., Fig. 9 The user may interact with the glasses 100 through voice input or touch input on a client device 1026 shown in FIG. 1 , and / or hand movements, position, and orientation detected by the glasses 100 .

[0038] Figure 3 is a diagrammatic representation of a computing device 300 within which instructions 310 (e.g., software, programs, applications, applet, applications, or other executable code) may be executed that cause the computing device 300 to perform any one or more of the methodologies discussed herein. The computing device 300 may be used as Figure 1The computer 120 of the glasses 100. For example, the instructions 310 can cause the computing device 300 to perform any one or more of the methods described herein. The instructions 310 transform a general, unprogrammed computing device 300 into a specific computing device 300 that is programmed to perform the functions described and shown in the described manner. The computing device 300 can operate as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the computing device 300 can operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The computing device 300 may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a PDA, an entertainment media system, a cellular phone, a smart phone, a mobile device, a head-mounted device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web device, a network router, a network switch, a network bridge, or any machine capable of sequentially or otherwise executing instructions 310 specifying actions to be taken by the computing device 300. In addition, while a single computing device 300 is shown, the term "machine" may also be understood to include a collection of machines that individually or collectively execute instructions 310 to perform any one or more of the methodologies discussed herein.

[0039] Computing device 300 may include processor 302, memory 304, and I / O components 306 that may be configured to communicate with each other via bus 344. In some examples, processor 302 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an ASIC, a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, processor 308 and processor 312 that execute instructions 310. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") that can execute instructions simultaneously. Although Figure 3 Multiple processors 302 are shown, but computing device 300 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.

[0040] The memory 304 includes a main memory 314, a static memory 316, and a storage unit 318, which are all accessible by the processor 302 via a bus 344. The main memory 304, the static memory 316, and the storage unit 318 store instructions 310 that implement any one or more of the methods or functions described herein. During execution of the instructions 310 by the computing device 300, the instructions 310 may also reside, in whole or in part, within the main memory 314, within the static memory 316, within the machine-readable medium 320 within the storage unit 318, within one or more of the processors 302 (e.g., within a cache memory of a processor), or within any suitable combination thereof.

[0041] I / O components 306 may include various components that receive input, provide output, generate output, send information, exchange information, capture measurements, etc. The specific I / O components 306 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine will be less likely to include such a touch input device. It is to be understood that I / O components 306 may include Figure 3 306. In various examples, the I / O component 306 may include an output component 328 and an input component 332. The output component 328 may include a visual component (e.g., a display such as a plasma display panel (PDP), a light emitting diode (LED) display, a liquid crystal display (LCD), a projection unit, or a cathode ray tube (CRT)), an acoustic component (e.g., a speaker), a tactile component (e.g., a vibration motor, a resistance mechanism), other signal generators, etc. The input component 332 may include an alphanumeric input component (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input component), a point-based input component (e.g., a mouse, a touch pad, a trackball, a joystick, a motion sensor, or other pointing instrument), a tactile input component (e.g., a physical button, a touch screen that provides the location and / or force of a touch or touch gesture, or other tactile input component), an audio input component (e.g., a microphone), etc.

[0042] In some examples, I / O component 306 may include: biometric component 334, motion component 336, environment component 338, and positioning component 340, as well as various other components. For example, biometric component 334 includes components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition), etc. Motion component 336 may include an inertial measurement unit, an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope), etc. Environmental components 338 include, for example, lighting sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers that detect ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), acoustic sensor components (e.g., one or more microphones that detect background noise), proximity sensor components (e.g., infrared sensors that detect nearby objects), gas sensors (e.g., gas detection sensors that detect concentrations of hazardous gases for safety or measure pollutants in the atmosphere), or other components that can provide indications, measurements, or signals associated with the surrounding physical environment. Positioning components 340 may include position sensor components (e.g., GPS receiver components), altitude sensor components (e.g., an altimeter or barometer that detects air pressure from which altitude can be derived), orientation sensor components (e.g., an inertial measurement unit (IMU)), etc.

[0043] Various technologies may be used to implement communications. I / O component 306 also includes a communication component 342 that is operable to couple computing device 300 to network 322 or device 324 via coupling 330 and coupling 326, respectively. For example, communication component 342 may include a network interface component or other suitable device to connect to network 322. In other examples, communication component 342 may include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, Parts (e.g. Low energy consumption), Wi- Device 324 may be another machine or any of a variety of peripherals (eg, a peripheral coupled via USB).

[0044] In addition, the communication component 342 can detect the identifier or include a component operable to detect the identifier. For example, the communication component 342 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting the following: a one-dimensional barcode, such as a universal product code (UPC) barcode; a multi-dimensional barcode, such as a Quick Response (QR) code, Aztec code, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcode, and other optical codes) or an acoustic detection component (e.g., a microphone for an audio signal to identify the tag). In addition, various information can be derived via the communication component 342, such as location via Internet Protocol (IP) geolocation, location via Wi-Fi, etc. Signal triangulation to derive location, location via detection of NFC beacon signals that can indicate a specific location, etc.

[0045] Various memories (e.g., memory 304, main memory 314, static memory 316, and / or memory of processor 302) and / or storage unit 318 may store one or more sets of instructions and data structures (e.g., software) that implement or are used by any one or more of the methods or functions described herein. These instructions (e.g., instructions 310) when executed by processor 302 cause various operations to implement the disclosed examples.

[0046] Instructions 310 may be sent or received via a network interface device (e.g., a network interface component included in communication component 342) using a transmission medium and using any of a number of well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)) over network 322. Similarly, instructions 310 may be sent or received to device 324 via coupling 326 (e.g., a peer-to-peer coupling) using a transmission medium.

[0047] Figure 44 is a diagram of an environment 400 including one or more systems according to one or more examples, wherein the one or more systems are used to determine the layout of content items and display information of the content items in an augmented reality environment. The environment 400 includes an augmented reality (AR) content system 402. The AR content system 402 can analyze input received from one or more users to obtain content in response to a request from one or more users. In addition, the AR content system 402 can determine the layout of content within a user interface based on one or more features of the content. In addition, the AR content system 402 can determine a location within a real-world scene for displaying content based on the layout. The AR content system 402 can also enable one or more users to interact with content by providing multiple actions that one or more users can take with respect to the content.

[0048] The AR content system 402 can include an audio input processing system 404 to receive and analyze audio input 406 generated by a user 408. In one or more examples, the audio input 406 can be captured by one or more sensors of the client device 410. For example, the client device 410 includes one or more microphones to capture the audio input 406 generated by the user 408. In at least some examples, the client device 410 begins capturing the audio input 406 in response to one or more activation commands provided by the user 408 corresponding to activating the client device 410 to capture audio data.

[0049] Client device 410 can execute an instance of client application 412. The processing resources and memory resources of client device 410 can execute multiple applications such as client application 412. In one or more examples, client application 412 can include messaging functionality that enables a user of client application 412 to send messages to and receive messages from other users of client application 412. In one or more additional examples, client application 412 can include social networking functionality that enables a user of client application 412 to share content with other users of client application 412 and / or access content created by other users of client application 412. In one or more illustrative examples, client application 412 can include information about Fig.11 At least one of the messaging client 1102 or the application 1104 is described in more detail. In various examples, the audio input 406 can be captured during execution of an instance of the client application 412 by the client device 410.

[0050] Additionally, in one or more examples, at least a portion of the operations described with respect to the AR content system 402 can be performed by the client device 410. In one or more further examples, at least a portion of the operations described with respect to the AR content system 402 can be performed by one or more computing devices different from the client device 410. For illustration, at least a portion of the operations described with respect to the AR content system 402 can be performed by a distributed computing system such as a cloud computing system. In at least some additional examples, the operations described with respect to the AR content system 402 can be performed by a combination of a computing device including the client device 410 and one or more computing devices of the distributed computing system.

[0051] In at least some examples, the AR content system 402 can cause augmented reality content to be displayed within a real-world scene. The augmented reality content item may include a program code that can be executed to perform one or more functions. In various examples, the augmented reality content item can be executed within the client application 412. For example, an instance of the client application 412 can be activated by the client device 410, and one or more user interfaces of the client application 412 can be displayed via the client device 410. The augmented reality content item can be selected while viewing one or more user interfaces of the client application 412, and the augmented reality content item is executed to activate one or more functions corresponding to the selected augmented reality content item. In at least some examples, the augmented reality content item can change the appearance of at least one of one or more objects or one or more locations within the real-world scene.

[0052] In various examples, the audio input 406 may be captured in response to content provided via the client application 412. For illustration, the audio input 406 may be captured in response to at least one of audio content, video content, image content, text content, or augmented reality content generated by the client application 412. Additionally, the client device 410 may include one or more cameras, such as a camera 414, to capture at least one of video or image content of a real-world scene in which the user 408 and the client device 410 are located. The video content captured by the camera 414 may include at least one of a series of images or image streams captured during a period of time. In various examples, the camera 414 may capture video of the real-world scene in response to input from the user 408. The image captured by the camera 414 may be within the field of view of the camera 414. The field of view of the camera 414 may correspond to a portion of the environment that may be imaged by the camera 414 at a given time, and may be based on the focal length of the lens of the camera 414 and the size of the sensor of the camera 414.

[0053] In one or more examples, client device 410 may include multiple computing devices having processing resources and memory resources. For example, client device 410 may include at least one of a head mounted device, a wearable device, or a mobile computing device such as a smart phone. In various examples, client device 410 may include multiple computing devices operating in conjunction with each other. For illustration, a head mounted device may operate in conjunction with at least one of a wearable device or a mobile computing device, or a wearable device may operate in conjunction with a mobile computing device. In one or more illustrative examples, client device 410 may include Figure 1 100 for glasses.

[0054] The audio input processing system 404 may include an audio-to-text system 416 that analyzes the audio input 406 to generate text data 418 corresponding to the audio input 406. In one or more examples, the audio input processing system 404 may generate an audio file using the audio input 406 and provide the audio file to the audio-to-text system 416. The audio file may have one or more formats, such as a Moving Picture Experts Group (MPEG) Audio Layer 3 (MP3) format, an M4A format, a Free Lossless Audio Codec (FLAC) format, a Waveform Audio File (WAV) format, a Windows Media Audio (WMA) format, an Advanced Audio Coding (AAC) format, or one or more combinations thereof. In various examples, the audio input processing system 404 may generate one or more audio files based on the audio input 406 by performing one or more analog-to-digital conversion techniques. In at least some examples, the audio input processing system 404 may provide a modified version of the audio input 406 to the audio-to-text system 416 after performing one or more pre-processing operations on the audio input 406. For example, the audio input processing system 404 may perform one or more signal processing techniques to reduce background noise present in the audio input 406 .

[0055] The audio-to-text system 416 can perform one or more feature extraction operations to generate text data 418 based on the audio input 406. In one or more examples, the audio-to-text system 416 can perform one or more automatic speech recognition (ASR) techniques to generate the text data 418 based on the audio input 406. In at least some examples, the audio-to-text system 416 can perform one or more natural language processing techniques to generate the text data 418 using the audio input 406. In various examples, the audio-to-text system 416 can perform one or more machine learning techniques to generate the text data 418 based on the audio input 406. In one or more illustrative examples, the audio-to-text system 416 can perform one or more hidden Markov models to generate the text data 418 based on the audio input 406. In one or more additional illustrative examples, the audio-to-text system 416 can perform one or more neural networks to generate the text data 418 based on the audio input 406. In one or more additional illustrative examples, audio-to-text system 416 may execute one or more deep feed-forward neural networks to generate text data 418 corresponding to audio input 406 .

[0056] The audio input processing system 404 may also include a text analysis system 420 that analyzes the text data 418. The text analysis system 420 may analyze the text data 418 to identify one or more keywords included in the text data 418. In one or more examples, the text analysis system 420 may determine a similarity measure between at least one of the words or phrases included in the text data 418 with respect to the one or more keywords. The text analysis system 420 may determine the similarity measure based on at least one of the number of letters or the order of letters of the one or more words in the text data 418 with respect to the arrangement of letters of the one or more keywords. In a scenario where the similarity measure between the one or more words included in the text data 418 and the one or more keywords is at least a threshold similarity measure, the text analysis system 420 determines that the one or more keywords are included in the audio input 406.

[0057] In one or more additional examples, the keywords identified by the AR content system 402 may be associated with one or more additional words or phrases having a meaning similar to the meaning of the keywords. In these cases, the text analysis system 420 determines a similarity measure between the meaning of one or more words included in the text data 418 relative to the meaning of the one or more keywords. For example, the text analysis system 420 analyzes the one or more words included in the text data 418 relative to one or more keywords and a synonym group corresponding to the one or more keywords. In one or more additional examples, the text analysis system 420 may perform one or more machine learning techniques to determine whether the meaning of one or more words included in the text data 418 corresponds to one or more keywords identified by the AR content system 402. For illustration, the text analysis system 420 may perform one or more natural language processing techniques to determine that at least a portion of the text data 418 includes one or more keywords or includes at least one of one or more words corresponding to the meaning of one or more keywords. In one or more illustrative examples, text analysis system 420 may execute one or more neural networks to determine that at least a portion of text data 418 includes at least one of one or more keywords or one or more words corresponding to the meaning of the one or more keywords.

[0058] In various examples, the AR content system 402 may identify keywords that cause the AR content system 402 to perform a plurality of different actions. In one or more examples, the keywords identified by the AR content system 402 may be related to retrieving content from one or more sources, where the content may be accessed using the client application 412. In one or more additional examples, the keywords identified by the AR content system 402 may be related to rendering and displaying content in one or more user interfaces displayed via the client application 412. In one or more illustrative examples, the keywords identified by the AR content system 402 may be related to retrieving content and rendering and displaying content in an augmented reality environment using a user interface generated by the client application 412. For example, the AR content system 402 identifies a plurality of keywords corresponding to the retrieval of content displayed within a real-world scene. The AR content system 402 may identify one or more first keywords corresponding to a command to retrieve content from one or more data sources and one or more second keywords corresponding to a command related to the display of augmented reality content in a real-world scene. In addition to commands, the text analysis system 420 may also determine one or more additional keywords included in the audio input 406. For illustration, where the text analysis system 420 determines that the audio input 406 includes one or more commands for retrieving content from one or more content sources, the text analysis system 420 determines one or more additional keywords included in the audio input 406 that correspond to characteristics of the content to be retrieved. For example, the one or more additional keywords may correspond to search terms related to the content that the user 408 desires to retrieve.

[0059] In response to determining that the audio input 406 includes one or more keywords related to the retrieval of content, the audio input processing system 404 can generate a search request 422. The search request 422 can include one or more search terms included in the audio input 406. The audio input processing system 404 can send the search request 422 to one or more content database servers 424. The one or more content database servers 424 can be physically or logically coupled to at least one of the one or more content databases 426. The one or more content databases 426 can store content that can be displayed via one or more user interfaces generated in conjunction with the client application 412. The one or more content databases 426 can store at least one of text content, image content, video content, audio content, or augmented reality content that can be accessed using the client application 412.

[0060] The one or more content database servers 424 can at least one of manage, control or maintain storage and retrieval of content from the one or more content databases 426. In one or more examples, the one or more content databases 426 can be at least one of controlled, maintained or managed by one or more content providers. In one or more illustrative examples, the one or more content databases 426 can include a first content database that is at least one of controlled, maintained or managed by a first content provider and a second content database that is at least one of controlled, maintained or managed by a second content provider. In various examples, the one or more content databases 426 can be at least one of controlled, maintained or managed by one or more search engines.

[0061] One or more content database servers 424 may analyze the search request 422 and generate search results 428 in response to the search request 422. The search results 428 may indicate one or more content items 430 that meet one or more criteria included in the search request 422. For example, the one or more content items 430 included in the search results 428 may be related to one or more search terms included in the search request 422. In various examples, the search results 428 may include an ordered list of the one or more content items 430. In one or more illustrative examples, the search request 422 may include a phrase such as "how to ride a bike?" In this case, the one or more content items 430 included in the search results 428 may include at least one of a web page, a video, a message content, a social media post, or other content related to learning to ride a bike.

[0062] The AR content system 402 may include a content presentation system 432 that determines one or more arrangements of content included in one or more content items 430 within a user interface displayed by the client device 410 in a real-world scene. The content presentation system 432 may include a content item identification system 434 for determining one or more characteristics of the one or more content items 430. In one or more examples, the content item identification system 434 may determine one or more content formats associated with the one or more content items 430. The one or more content formats may correspond to at least one of one or more file types of the one or more content items 430 or one or more technologies for accessing content of the one or more content items 430. The technology for accessing content of the one or more content items 430 may correspond to one or more software technologies executed to access content of the one or more content items 430, one or more hardware technologies executed to access content of the one or more content items 430, or one or more combinations thereof. In one or more examples, content item identification system 434 may determine that one or more content items 430 include at least one of text content, audio content, image content, video content, or augmented reality content.

[0063] The content presentation system 432 may also include a content item display system 436 that determines one or more layouts of content included in the one or more content items 430 based on the characteristics of the one or more content items 430 determined by the content item recognition system 434. For example, the content item display system 436 determines that the one or more content items 430 including text content can be arranged according to one or more first layouts. In addition, the content item display system 436 may determine that the one or more content items 430 including a combination of text content and image content can be arranged according to one or more second layouts. In addition, the content item display system 436 may determine that the one or more content items 430 including a combination of text content and video content can be arranged according to one or more third layouts. In another example, the content item display system 436 may determine that the one or more content items 430 including augmented reality content can be arranged according to one or more fourth layouts. The content item display system 436 may also determine that the one or more content items 430 including at least one of text content, video content, or image content and augmented reality content can be arranged according to one or more fifth layouts.

[0064] In various examples, the content item display system 436 can determine the layout of the user interface for the content included in the one or more content items 430 based on the corresponding sources of the one or more content items 430. In one or more examples, the respective sources of the content items can generate content items with one or more characteristics, such as generating content items including content with one or more formats and / or content with one or more arrangements. The content item display system 436 can determine the layout of the content items 430 with one or more characteristics, wherein the layout includes parts for one or more types of content in the user interface. For example, the content item display system 436 identifies one or more layouts of the one or more content items 430, the one or more layouts having one or more parts of the user interface for text content, one or more parts of the user interface for video content, one or more parts of the user interface for image content, one or more parts of the user interface for augmented reality content, or one or more combinations thereof. In one or more illustrative examples, a content source can provide a content item with video content and text content. In these scenarios, content item display system 436 determines a layout of the content item that includes a first portion within the user interface for displaying video content and a second portion within the user interface for displaying text content.

[0065] In one or more illustrative examples, one or more content items 430 may include teaching content. In these scenarios, the content included in one or more content items 430 may include multiple actions to be performed by user 408. For example, one or more content items 430 include teaching content related to one or more recipes, teaching content related to performing vehicle maintenance, teaching content related to repairing objects that cannot operate normally, teaching content related to building objects, other operation content, etc. In various examples, one or more content items 430 may be arranged so that at least one of video content, text content, audio content, image content, or augmented reality content is presented in discrete steps that are ordered in a manner to achieve a desired result. For illustration, content item 430 may be related to a recipe for baking bread that includes four steps. In one or more examples, content item 430 may include a first text content and a first video content related to a first step, a second text content and a second video content related to a second step, a third text content and a third video content related to a third step, and a fourth text content and a fourth video content related to a fourth step. In these scenarios, content item display system 436 makes content related to the various steps accessible in a sequential order via a corresponding user interface of client application 412 .

[0066] In at least some examples, at least a portion of the one or more content databases 426 may include one or more curated databases storing instructional content that has been arranged so that the various steps of a process can be accessed in discrete portions. The one or more content databases 426 storing curated content may be generated by one or more third-party content sources. In addition, a service provider that at least one of controls, maintains, manages, or creates the client application 412 may obtain instructional content from one or more content sources and modify the content obtained from the one or more content sources so that the content is arranged and stored in the one or more content databases 426 according to the various steps that can be accessed in discrete portions.

[0067] In one or more additional examples, one or more content items 430 may include teaching content without being divided into multiple steps arranged in discrete parts. In these cases, the content item display system 436 modifies one or more content items 430 so that the modified version of one or more content items 430 includes multiple parts, wherein each part corresponds to the discrete steps of the teaching content. For example, the content item display system 436 analyzes one or more content items 430 and determines that the content item 430 includes teaching content. For illustration, the content item display system 436 may analyze at least one of the words, phrases, or images included in at least one of the text content, image content, or video content to determine that the content item 430 includes teaching content. In one or more illustrative examples, the content item display system 436 may determine a similarity measure between at least one of the words, phrases, or images of the content item 430 relative to at least one of the words, phrases, or images of the content item 430 that has been previously identified as having teaching content. In various examples, the content item display system 436 may perform one or more machine learning techniques to generate one or more models based on training data including content items that have been previously identified as having teaching content. One or more models may be executed to determine the similarity metric. Where the similarity metric is at least a threshold similarity metric for content item 430, content item display system 436 determines that content item 430 includes instructional content.

[0068] In response to determining that the content item 430 includes teaching content, the content item display system 436 can determine the portion of the content item 430 corresponding to the discrete steps of the teaching process. In one or more examples, the content item display system 436 can determine the portion of the text content included in the content item 430 corresponding to the various steps of the teaching process. The content item display system 436 can also determine the image corresponding to the various steps of the teaching process included in the content item 430. In addition, the content item display system 436 can also determine one or more portions of the video content corresponding to the various steps of the teaching process. For example, the content item display system 436 determines the start timestamp and the end timestamp of the portion of the video content included in the content item 430 corresponding to the various steps in the teaching process. Additionally, the content item display system 436 can determine the augmented reality content items corresponding to the various steps of the teaching process included in the content item 430.

[0069] In one or more illustrative examples, the content item display system 436 may perform one or more machine learning techniques to determine discrete portions of the content item 430 corresponding to individual instruction steps. One or more machine learning techniques may be performed by the content item display system 436 to generate one or more models based on training data. The training data may include at least one of a content item or a portion of a content item, wherein discrete portions of at least one of a content item or a portion of a content item correspond to individual steps of a teaching process. In at least some examples, the content item display system 436 may use one or more machine learning techniques to generate multiple computational models corresponding to different types of instructional content. For illustration, the content item display system 436 may generate a first computational model and a second computational model, the first computational model corresponding to identifying discrete portions of a content item corresponding to individual steps of a recipe, and the second computational model corresponding to identifying discrete portions of a content item corresponding to individual steps of a teaching process for assembling furniture.

[0070] In response to determining the portions of content item 430 corresponding to the steps of the teaching process, content item display system 436 may arrange the discrete portions of content item 430 so that the discrete portions of content item 430 may be accessed via client application 412 according to a sequence corresponding to the teaching process. For example, content item display system 436 generates first user interface data including a first portion of content item 430 corresponding to a first step of the teaching process and second user interface data including a second portion of content item 430 corresponding to a second step of the teaching process. Content item display system 436 may also generate metadata indicating the order in which the portions of content item 430 are to be displayed. For illustration, content item display system 436 may generate metadata indicating that the second user interface is displayed after the first user interface.

[0071] In a number of additional implementations, the content item display system 436 can determine a location within the real-world scene at which to display the content of the content item 430. In one or more examples, the content item display system 436 can cause one or more user interfaces to be displayed relative to the one or more locations within the real-world scene, wherein the one or more user interfaces include the content included in the content item 430. In various examples, the content item display system 436 can determine that the location at which the content of the content item 430 is displayed corresponds to the gaze of the user 408. In one or more illustrative examples, the AR content system 402 can include a gaze tracking system 438 that determines the location of the field of view of the gaze of the user 408. In at least some examples, the gaze tracking system 438 can analyze camera data obtained from the client device 410 to determine the location of the gaze of the user 408. Additionally, the gaze tracking system 438 can analyze data obtained from one or more inertial measurement unit (IMU) sensors to determine the location of the gaze of the user 408. Furthermore, the gaze tracking system 438 can analyze camera data obtained from one or more cameras external to the client device 410 to determine the location of the gaze of the user 408. In one or more illustrative examples, gaze tracking system 438 can determine at least one of the field of view of user 408 or the center of the field of view of user 408, and provide gaze tracking information to content item display system 436. Content item display system 436 can then cause the content of content item 430 to be displayed within the field of view of user 408, e.g., at a central location of the field of view of user 408.

[0072] In one or more examples, when the gaze of user 408 changes, gaze tracking system 438 can determine the new position of the field of view of the gaze of user 408, and provide the new position of the gaze of user 408 to content item display system 436. Then, content item display system 436 can move the position of the content of content item 430 displayed within the real world scene. In one or more additional examples, the position of the content of content item 430 displayed within the real world scene can be a fixed position. The fixed position can be determined by content item display system 436 based on input from user 408. Additionally, the fixed position can correspond to an object located in the real world scene. In these scenes, content item display system 436 performs one or more object recognition techniques to identify one or more objects located in the real world scene. For example, content item display system 436 analyzes information captured by one or more cameras 414 of client device 410 to determine objects that may be suitable for display of content. To illustrate, the content item display system 436 can analyze information captured by one or more cameras 414 of the client device 410 to identify a television, wall, table, screen, appliance, or other surface on which content can be displayed. The object identified by the content item display system 436 on which the content is displayed can be related to the subject matter included in the content. In one or more illustrative examples, the content item display system 436 can determine that content related to recipes is to be displayed on a refrigerator or other appliance, or determine that media content such as a movie or television show is to be displayed on a television.

[0073] In various examples, the audio input processing system 404 can provide one or more content commands 440 to the content presentation system 432. The one or more content commands 440 can be related to one or more actions that can be performed by the client application 412 with respect to the content included in the one or more content items 430. For example, the one or more content commands 440 can be related to selecting one or more user interface elements included in one or more user interfaces displayed using the client application 412. To illustrate, the client application 412 can display a user interface in an augmented reality environment that includes search results 428 generated in response to the search request 422. The user interface can include a user interface element, such as an icon, corresponding to a single content item 430 included in the search results 428, where the user interface element is selectable to cause the client application 412 to display a user interface including the content of the selected content item 430 in the augmented reality environment.

[0074] One or more content commands 440 may also be associated with display of content included in one or more user interfaces generated with respect to client application 412 and displayed in an augmented reality environment. In one or more examples, one or more content commands 440 may be associated with actions that modify display characteristics of content included in one or more user interfaces displayed in conjunction with client application 412, such as one or more content magnification operations that increase or decrease at least one of the appearance of content displayed in one or more user interfaces. Additionally, one or more content commands 440 may be associated with a location within a real-world scene for displaying content. For example, one or more content commands 440 are associated with causing content to be displayed at a fixed location or to move with the gaze of user 408. In addition, one or more content commands 440 may be associated with browsing content included in content item 430. In various examples, one or more content commands 440 may be associated with selecting one or more options from one or more menus corresponding to options for browsing content included in one or more content items 430. In one or more illustrative examples, content item 430 may include instructional content, and one or more content commands 440 may be associated with accessing at least one of one or more subsequent steps or one or more previous steps in the instructional process with respect to a current step of the instructional process.

[0075] In one or more examples, a set of commands can be selected based on the content displayed in the user interface. For example, the AR content system 402 determines that a first user interface including search results 428 is to be displayed. The AR content system 402 can then identify a first set of commands corresponding to interacting with the search results 428, such as selecting one or more of the content items 430 in the search results 428. In these scenarios, the AR content system 402 can cause a first user interface including the search results 428 and at least a portion of the first set of commands to be displayed. In at least some examples, the AR content system 402 can identify the first set of commands during the time period in which the first user interface is displayed. In response to navigating to the second user interface, the AR content system 402 can then identify a second set of commands. For illustration, after selecting a content item 430 included in the search results 428, the content of the content item 430 can be displayed in the second user interface. The AR content system 402 can determine a second set of commands corresponding to the second user interface, and display at least a portion of the second set of commands in the second user interface together with the content of the selected content item. In one or more illustrative examples, the second set of commands may correspond to at least one of: setting one or more locations in the real-world scene where the additional user interface is displayed, modifying one or more display characteristics of the content item 430 in the additional user interface, or browsing the instructional content of the content item 430. In these cases, the second set of commands may be recognized by the AR content system 402 during the time period when the second user interface is displayed. In various examples, the AR content system 402 may not recognize the second set of commands during the time period when the first user interface is displayed, and may not recognize the first set of commands during the time period when the second user interface is displayed. In this way, the processing and memory resources of the AR content system 402 may be minimized.

[0076] In at least some examples, one or more content commands 440 can be determined by text analysis system 420. For example, in addition to analyzing text data 418 to identify terms of search request 422, text analysis system 420 analyzes text data 418 to identify at least one of the words or phrases corresponding to one or more content commands 440. For example, text analysis system 420 analyzes text data 418 relative to one or more keywords corresponding to the one or more content commands by determining a similarity measure between at least one of the words or phrases included in the text data and at least one of the words or phrases of the one or more content commands 440. In one or more additional examples, text analysis system 420 can determine a similarity measure between the meaning of one or more words included in text data 418 relative to the meaning of one or more keywords associated with one or more content commands 440.

[0077] In one or more illustrative examples, the AR content system 402 may analyze the audio input 406 obtained from the user 408, and determine the content to be provided to the user 408 in response to the audio input 406. The AR content system 402 may also determine the arrangement of the content within one or more user interfaces. In various examples, the AR content system 402 may generate content item data 442 corresponding to one or more content items 430 identified by the AR content system 402 based on the audio input 406. Additionally, the AR content system 402 may generate content arrangement data 444 corresponding to the layout of the content included in one or more content items 430 included in the content item data 442. The content item data 442 and the content arrangement data 444 may be used to generate one or more content user interfaces 446. The one or more content user interfaces 446 may include the content included in the content item data 442 displayed according to the layout corresponding to the content arrangement data 444. In one or more examples, the content user interface 446 may be displayed within a camera view 448 of the client device 410. To illustrate, content user interface 446 may be displayed within an augmented reality environment such that user 408 may view content included in content user interface 446 in addition to objects included in the real-world scene in which user 408 is located.

[0078] In various examples, content item data 442 may include instructional content divided into a plurality of discrete portions corresponding to steps of an instructional process. Instructional content may be arranged in content user interface 446 based on one or more formats of the instructional content included in content item data 442. For example, content arrangement data 444 indicates at least one of: a portion of content user interface 446 for displaying text content included in content item data 442, a portion of content user interface 446 for displaying image content included in content item data 442, a portion of content user interface 446 for displaying video content included in content item data 442, or a portion of content user interface 446 for displaying augmented reality content included in content item data 442. In one or more examples, when a user browses instructional content included in content item data 442, the content displayed within content user interface 446 may be modified. For illustration, when the user 408 navigates from a first step of the instructional content to a second step of the instructional content, the content user interface 446 may be modified from displaying at least one of the text content, video content, image content, or augmented reality content of the first step of the instructional content to displaying at least one of the text content, video content, image content, or augmented reality content of the second step of the instructional content.

[0079] In addition, in one or more scenes, the position of content user interface 446 within the real-world scene can be fixed. In one or more examples, content user interface 446 can be displayed relative to the position of an object, such as a wall or a television, in real-world space. In one or more additional examples, the position of content user interface 446 within the real-world scene can correspond to a fixed position indicated by user 408. For example, user 408 provides a command to fix the position of content user interface 446 within the real-world scene. Additionally, the position of content user interface 446 within the real-world scene can be modified. To illustrate, when the gaze of user 408 changes, the position of content user interface 446 within the real-world scene can move to track the position of the gaze of user 408.

[0080] Figure 5 is a diagram of an architecture 500 including a system for determining a content template for arranging information displayed in a user interface of an augmented reality environment according to one or more examples. The architecture 500 may include a content item display system 436. In one or more examples, the content item display system 436 may analyze the search results 428 to generate content arrangement data 444 indicating the location of one or more portions of the content item to be displayed within the content user interface. For example, each content item 430 included in the search results 428 has one or more content item features 502. In one or more examples, the content item display system 436 may generate the one or more content item features 502 by analyzing the content items 430. In one or more additional examples, the one or more content item features 502 may be indicated in metadata provided to the AR content system 402 along with the search results 428.

[0081] One or more content item features 502 may indicate one or more formats of content included in content item 430. For illustration, content item features 502 may indicate that content item 430 includes at least one of text content, image content, video content, or augmented reality content. One or more content item features 502 may also indicate a source of content item 430. The source of content item 430 may indicate a content provider that generates search results 428 and provides content item 430 included in search results 428. In one or more illustrative examples, the source of content item 430 may include an e-commerce service provider, a media content provider, a social network content provider, a search engine, one or more combinations thereof, and the like.

[0082] The content item display system 436 may analyze one or more content item features 502 to generate content arrangement data 444. In one or more examples, the content item display system 436 may analyze one or more content item features 502 with respect to features of a plurality of content templates 504. Based on one or more features of the content displayed via the user interface, the plurality of content templates 504 may include different layouts of the content within the user interface. In various examples, the content templates 504 may indicate the location of content in different formats. For example, the first content template 506 corresponds to the first content template feature set 508 and has a first content layout 510. In addition, the second content template 512 may correspond to the second content template feature set 514 and have a second content layout 516. In various examples, the first content layout 510 may include a first arrangement of a portion of a content user interface for displaying at least one of text content, video content, image content, or augmented reality content, and the second content layout 516 may include a second arrangement of a portion of a content user interface for displaying at least one of text content, video content, image content, or augmented reality content.

[0083] In one or more illustrative examples, the first content template 506 may correspond to a first content source, such as an e-commerce source, that can provide content items including image content and text content. For example, the content item 430 obtained from the first source includes one or more images related to a product and text content related to the product, such as a product description, a product review, etc. In these scenarios, the first content layout 510 indicates a first portion of a content user interface for displaying text content and a second portion of a content user interface for displaying image content. In one or more additional illustrative examples, the second content template 512 may correspond to a second content source, such as a video content provider, that can provide content items including video content and text content. For illustration, the content item 430 obtained from the second source may include one or more videos and text content related to one or more videos, such as a summary of a video, a comment related to a video, etc. In these cases, the second content layout 516 indicates a first portion of a content user interface for displaying video content and a second portion of a content user interface for displaying text content.

[0084] Despite Figure 5, but the content template 504 may include at least a third content template. In one or more examples, the third content template may correspond to a third content source, such as a social media content provider, that provides content items including text content and at least one of video content, image content, or augmented reality content. In various examples, the content item 430 obtained from the social media content provider may include text content corresponding to a social media post or a social media message, such as a description associated with a social media post, a comment associated with a social media post, and the like, and at least one of one or more images, one or more videos, or one or more augmented reality content items associated with a social media post or a social media message. In these cases, the third layout may include a portion that displays text content in the content user interface and at least one additional portion that displays at least one of video content, image content, or augmented reality content in the content user interface.

[0085] The content item display system 436 may analyze the one or more content item features 502 relative to the first content template feature set 508 and the second content template feature set 514 to determine a content layout to be applied to the content item 430. In one or more examples, the content item display system 436 may determine a first similarity metric between the one or more content item features 502 and the first content template feature set 508 and a second similarity metric between the one or more content item features 502 and the second content template feature set 514. The content item display system 436 may analyze the first similarity metric and the second similarity metric relative to a threshold value to determine whether to display the content included in the content item 430 according to the first content template 506 or the second content template 512. In one or more additional examples, the content item display system 436 may determine a ranking based on the first value of the first similarity metric and the second value of the second similarity metric to determine whether to display the content of the content item 430 according to the first content template 506 or the second content template 512. After determining whether to display content of content item 430 based on first content template 506 or second content template 512, content item display system 436 can generate content arrangement data 444. For example, content item display system 436 can generate content arrangement data 444 including first content template 506 or second content template 512 based on the first similarity metric or the second similarity metric.

[0086] Figure 66 is a diagram illustrating a user interface generated by a client device 410 and displayed in an augmented reality environment 600 according to one or more examples. The augmented reality environment 600 may include a real-world scene in which multiple objects are located. For example, the augmented reality environment 600 may include a first object 602, a second object 604, and a third object 606. Additionally, the client device 410 and the user 408 may be located in the augmented reality environment 600.

[0087] Client device 410 may cause one or more user interfaces to be displayed at one or more locations in augmented reality environment 600. Figure 6 In the illustrative example of FIG. 4 , user interface 608 is displayed at first location 610 within augmented reality environment 600. First location 610 can be characterized according to first real-world coordinates. User interface 608 can display content corresponding to content item 612. In one or more examples, user interface 608 can be displayed within first field of view 614 of user 408. When the gaze of user 408 shifts to second field of view 616, user interface 608 can be displayed at second location 618 within augmented reality environment 600. Second location 618 can be characterized according to second real-world coordinates.

[0088] In at least some examples, the position of user interface 608 within augmented reality environment 600 can be fixed in response to one or more commands from user 408. In various examples, the position of user interface 608 within augmented reality environment 600 can be fixed for a period of time. In one or more additional examples, the position of user interface 608 within augmented reality environment 600 can be fixed until a command is received indicating that the position of user interface 608 can be moved relative to the gaze of user 408. In one or more further examples, the position of user interface 608 within augmented reality environment 600 can correspond to the position of an object located in augmented reality environment 600. For example, user interface 608 is displayed at the position of second object 604. In one or more illustrative examples, in response to one or more commands from user 408, the position of user interface 608 can be moved from second position 618 to the position of second object 604. In one or more additional illustrative examples, user interface 608 can be displayed at the position of second object 604 based on one or more features of content item 612. For illustration, user interface 608 can be displayed at the position of second object 604 based on content item 612 including video content.

[0089] Figure 7A flowchart of an example process 700 for determining the arrangement of information in a user interface displayed in an augmented reality environment according to one or more examples is shown. Implementations of process 700 may be embodied in computer-readable instructions for execution by one or more processors, such that operations of the process may be performed in part or in whole by functional components of at least one of one or more client devices or one or more server systems. Thus, in some cases, the processes described below are examples for reference. However, in other implementations, the description of the process 700 may be performed in part or in whole by functional components of at least one of one or more client devices or one or more server systems. Figure 7 At least some of the operations of the described example processes may be deployed on various other hardware configurations. Figure 7 The example processes described are not intended to be limited to being performed by one or more server systems or one or more client devices described herein, and may be implemented in whole or in part by one or more additional components. Although the described flowcharts may show the operations as sequential processes, many of the operations may be performed in parallel or simultaneously. In addition, the order of the operations may be rearranged. The process terminates when its operations are completed. The process may correspond to a method, a program, an algorithm, etc. The operations of the method may be performed in whole or in part, may be performed in combination with some or all of the operations in other methods, and may be performed by any number of different systems (e.g., the systems described herein) or any part thereof (e.g., a processor included in any one of the systems).

[0090] Process 700 may include obtaining audio data captured by one or more microphones at operation 702. In one or more examples, the audio data may be generated by an individual who is a user of a client application. For example, the individual has an account with a service provider that controls, maintains, or creates at least one of the client applications. In various examples, a client device that executes an instance of a client application may be operated by an individual. One or more microphones may be located in a client device operated by an individual. In one or more illustrative examples, the client device may include a head-mounted device, such as glasses 100. One or more microphones may be located in the head-mounted device. In one or more additional examples, at least a portion of the one or more microphones may be placed in an external position relative to the head-mounted device within an augmented reality environment. In addition, the head-mounted device may include one or more cameras to capture at least one of image content or video content of a real-world scene.

[0091] In addition, process 700 may include analyzing audio data at operation 704 to generate text data corresponding to at least one of the one or more words or one or more phrases included in the audio data. Process 700 may include generating a search request including one or more keywords extracted from the text data at operation 706. In one or more examples, the search request may be sent to a content source to process the search request and generate search results based on one or more keywords included in the search request. In various examples, the content source may be specified in the search request. For example, one or more keywords indicate the content source of the search request. In one or more additional examples, one or more keywords may be analyzed to determine the content source to send the search request. For illustration, a search request for video content may be sent to a source of video content. Additionally, a search request for at least one of a product or service may be sent to an e-commerce content source. In addition, a search request related to social media content may be sent to a social media content source. In various examples, the search request may be sent to multiple content sources.

[0092] Process 700 may also include obtaining the search results of one or more content items corresponding to one or more keywords of the search request at operation 708. Each content item included in the search results may include at least one of text content, video content, image content, audio content or augmented reality content. In addition, process 700 may include determining one or more features of the content items in one or more content items at operation 710. One or more features may include at least one of the source of the content item or the format of the content item. In one or more examples, the content included in one or more content items may be analyzed to determine one or more formats of the content included in one or more content items. For example, one or more content items may be analyzed to determine whether one or more content items include at least one of text content, video content, image content or augmented reality content.

[0093] In various examples, the corresponding source of one or more content items can be indicated in metadata obtained about one or more content items. In these scenarios, the source of the content item is extracted from the metadata associated with the content item. Additionally, the source of the content item can also be determined based on an analysis of one or more features of the content item. For illustration, a content item obtained from an e-commerce source can include a first feature set, a content item obtained from a media content provider can include a second feature set, and a content item obtained from a social network content provider can include a third feature set.

[0094] Process 700 may also include determining the layout of the content included in the content item based on one or more features of the content item at operation 712. In one or more examples, one or more features of the content item may be analyzed relative to multiple feature sets corresponding to multiple content templates to determine a similarity measure between one or more features of the content item and one or more feature sets in multiple feature sets of the content template. Each content template in the multiple content templates may indicate the corresponding arrangement of the content in one or more user interfaces. For example, each content template indicates at least one of: one or more first parts of the user interface for displaying text content, one or more second parts of the user interface for displaying video content, one or more third parts of the user interface for displaying image content, or one or more fourth parts of the user interface for displaying augmented reality content. In various examples, a content template may be selected from multiple content templates based on a similarity measure. In one or more illustrative examples, multiple similarity measures may be determined based on one or more features of the content item relative to features associated with each content template.

[0095] In one or more illustrative examples, a content template corresponding to the highest similarity measure can be selected. In one or more additional illustrative examples, a template corresponding to a similarity measure that is at least a threshold similarity measure can be selected. In various examples, the selected content template can be based on the source of the content item. In addition, the selected content template can be based on one or more content formats included in the content item. For example, a first template is selected to display the content of a content item including text content, and a second template can be selected to display the content of a content item including at least one of text content and video content or image content.

[0096] Additionally, process 700 may include causing a user interface including content items presented according to a layout to be displayed in an augmented reality environment at operation 714. The augmented reality environment may include a real-world scene, and the user interface is displayed relative to a position in the real-world scene. In one or more examples, the field of view of the gaze of an individual may be determined. For example, the field of view of the gaze of a user of a client application is determined. In various examples, a position within the real-world scene for displaying the user interface may be determined, the position corresponding to the field of view of the gaze of the individual. In one or more illustrative examples, the field of view of the gaze of an individual may be determined based on camera data from one or more cameras included in the augmented reality environment. In one or more additional illustrative examples, the field of view of the gaze of an individual may be determined based on sensor data from one or more inertial measurement unit sensors included in the augmented reality environment. In at least some examples, at least one of the camera data or the sensor data may be captured by a head-mounted device worn by an individual. In one or more examples, the position of the user interface may change based on a change in the field of view of the individual. For example, the field of view of the gaze of the individual changes from a first position to a second position. In these scenarios, the position of the user interface within the real-world scene moves from a first position to a second position. In one or more additional examples, the user interface can be displayed relative to the location of an object included in the augmented reality environment.

[0097] In one or more examples, one or more commands related to the data displayed in the user interface can be obtained. One or more commands can be audible commands. One or more commands can also correspond to one or more gestures made by an individual. In at least some examples, one or more commands related to one or more gestures and one or more audible words can be conveyed. In the case where one or more commands correspond to audible words or phrases, additional audio data captured by one or more microphones are analyzed to generate additional text data corresponding to at least one of the one or more additional words or one or more additional phrases included in the additional audio data. The additional words or additional phrases corresponding to the command included in the additional text data can be different from the words or phrases corresponding to the generated search request included in the text data. The additional text data can then be analyzed to identify one or more commands included in the text data. In various examples, at least one of the one or more additional words or one or more additional phrases can be analyzed relative to at least one of the one or more words or one or more phrases of at least one command to determine a similarity measure. In this way, a command can be determined based on the value of the similarity measure.

[0098] In one or more illustrative examples, one or more commands may correspond to fixing the position of the user interface at a position within the real-world scene. In these scenes, when the field of view of the individual's gaze changes from a first position to a second position, the position of the user interface within the real-world scene remains the same. In one or more additional illustrative examples, one or more commands may correspond to modifying the display characteristics of the content item. For example, the command causes the appearance of the content item in the user interface to be modified. For illustration, the magnification level of at least one of the text content, image content, video content, or augmented reality content of the content item can be modified based on a command to modify the display characteristics of the content item.

[0099] In various examples, one or more commands may be associated with a selection of one or more user interface elements included in a user interface. In one or more illustrative examples, at least a portion of the one or more user interface elements may correspond to an option included in a menu displayed in the user interface. In one or more additional illustrative examples, at least a portion of the one or more user interface elements may correspond to a content item included in the search results. In one or more additional illustrative examples, at least a portion of the one or more user interface elements may correspond to at least a portion of the content of the content item.

[0100] In one or more examples, a user interface element may be selected based on the field of view of an individual's gaze. For example, the field of view of an individual's gaze corresponds to a given user interface element. In various examples, the appearance of a user interface element may change in response to the user interface element being within the field of view of an individual's gaze. In one or more illustrative examples, in response to determining that at least a threshold amount of user interface elements are within the central portion of the field of view of an individual's gaze, the appearance of the user interface element may be modified. In the case where the appearance of a user interface element is modified due to the user interface element being within the field of view of an individual's gaze, one or more commands about the user interface element may be obtained. For illustration, when the user interface element is within at least a threshold amount of the center of the field of view of an individual's gaze, the user interface element may be selected based on one or more commands. In at least some examples, one or more actions may be performed in response to the selection of a user interface element. For example, a content item is selected from a list of content items in response to one or more commands, and content corresponding to the content item may be displayed in a user interface.

[0101] In one or more examples, the augmented reality content item can be executed in response to launching the augmented reality content item being executed within the client application. Figure 7The augmented reality content item may include computer readable code executed within the client application. In one or more examples, after launching the augmented reality content item, audio data obtained with respect to operation 702 may be captured, and operations 704, 706, 708, 710, 712, and 714 and with respect to the augmented reality content item may be performed when the augmented reality content item is executed within the client application. Figure 7 Other operations described.

[0102] In one or more additional examples, regarding Figure 7 At least a portion of the described operations may be performed in response to one or more activation actions performed by an individual, such as a user of the head mounted device. For example, audio data obtained with respect to operation 702 may be captured, and operations 704, 706, 708, 710, 712, and 714 and with respect to operation 706 may be performed in response to one or more activation words or one or more activation phrases spoken by the user. Figure 7 In one or more additional examples, audio data obtained with respect to operation 702 can be captured, and operations 704, 706, 708, 710, 712, and 714 and with respect to operation 702 can be performed in response to one or more activation gestures or one or more other activation inputs provided by a user. Figure 7 Other operations described.

[0103] Figure 8 FIG. 8 is a user interface 800 including results of a search request displayed in an augmented reality environment according to one or more examples. Figure 8 In the illustrative example of , user interface 800 includes first search result 802, second search result 804, and third search result 806. Search results 802, 804, 806 may be provided in response to a search request having one or more criteria. First search result 806 may include a first thumbnail image 808 corresponding to the content of first search result 802. Additionally, second search result 804 may include a second thumbnail image 810 corresponding to the content of second search result 804. Furthermore, third search result 806 may include a third thumbnail image 812 corresponding to the content of third search result 806.

[0104] The user interface 800 may also include command text 814. The command text 814 may include at least one of a word or phrase that may be selected by a user to perform one or more actions with respect to a feature of the user interface 800. In at least some examples, the command text 814 may correspond to one or more commands currently recognized by the AR content system 402. Figure 8In the illustrative example of , command text 814 can correspond to a selection of one or more of search results 802, 804, 806. Additionally, user interface 800 can include audio input text 816. Audio input text 816 can correspond to audio input provided by the user. For illustration, when the user provides audio input, text generated by AR content system 402 based on the audio input can be displayed as audio input text 816. In this way, the user can see how AR content system 402 interprets the audio input obtained from the user. Figure 8 In the illustrative example of , the audio input text 816 corresponds to the user's selection of the search results 802 , 804 , 806 .

[0105] In addition, the user interface 800 may indicate a search result that is a target selection of the user by displaying a target selection having different visual characteristics from search results of non-target selections. Figure 8 In the illustrative example of , the second search result 804 can be a target selection and is displayed larger than the first search result 802 and the third search result 806. The target selection can be selected in response to a command from the user. In one or more examples, the target selection can be determined based on the user's gaze. In various examples, when the user's gaze moves, the target selection can also change to correspond to the change of the user's gaze. In this way, when the user's gaze shifts from one search result to another search result, the search result corresponding to the target selection can change, and the display characteristics of the search result can also change. In at least some examples, the user can identify multiple target selections. In these scenarios, a command indicating that multiple selections are to be made can be provided, and the user can move his gaze to select multiple search results.

[0106] Fig. 9 9 is a user interface 900 including a content item 902 and a menu 904 including a plurality of commands executable with respect to the content item 902, according to one or more examples. Fig. 9 In the illustrative example of , content item 902 includes video content 906 and text content 908. The content of content item 902 may be arranged according to a layout determined by AR content system 402. For example, video content 906 is displayed in a first portion of user interface 900 dedicated to video, and text content 908 may be displayed in a second portion of user interface 900 dedicated to text.

[0107] The menu 904 may include a first command text 910 corresponding to a first command that may be provided by a user, a second command text 912 corresponding to a second command that may be provided by a user, and a third command text 914 corresponding to a third command that may be provided by a user. In one or more examples, the menu 904 may indicate one or more commands corresponding to display characteristics of the content item 902. For example, the menu 904 indicates at least one of the following: one or more commands to increase the size of one or more parts of the content item 902, or one or more commands to reduce the size of one or more parts of the content item 902. In a scenario where the content item 902 includes instructional content, the menu 904 may indicate one or more commands to browse one or more steps of the instructional content. For illustration, the menu 904 may indicate at least one of a command to move to the next step of the instructional content or a command to move to the previous step of the instructional content.

[0108] The menu 904 may also indicate one or more commands related to the display of the user interface 900 within the augmented reality environment. In one or more examples, the menu 904 may indicate one or more commands to fix the position of the user interface 900 in the real world scene. In one or more additional examples, the menu 904 may indicate one or more commands to fix the position of the user interface 900 relative to objects included in the real world scene. In one or more further examples, the menu 904 may indicate one or more commands to move the position of the user interface 900 relative to the position of the user. For example, the menu 904 includes one or more commands to move the user interface 900 relative to the user's gaze.

[0109] Additionally, the user interface 900 may include audio input text 916. The audio input text 916 may correspond to the audio input provided by the user. For illustration, when the user provides audio input, the text generated by the AR content system 402 based on the audio input may be displayed as the audio input text 916. Fig. 9 In the illustrative example of , the audio input text 916 corresponds to multiple steps of navigating the instructional content.

[0110] Fig.101 is a block diagram showing a networked system 1000 including details of the glasses 100 according to some examples. The networked system 1000 includes the glasses 100, a client device 1026, and a server system 1032. The client device 1026 may be a smart phone, a tablet computer, a tablet phone, a laptop computer, an access point, or any other such device capable of connecting to the glasses 100 using a low-power wireless connection 1036 and / or a high-speed wireless connection 1034. The client device 1026 is connected to the server system 1032 via a network 1030. The network 1030 may include any combination of wired and wireless connections. The server system 1032 may be one or more computing devices that are part of a service or network computing system. The client device 1026 and any elements of the server system 1032 and the network 1030 may be connected using a wireless network, such as a wireless network or a wireless network. Fig.12 and Figure 3 The details of the software architecture 1204 or computing device 300 described in the embodiment of the present invention are implemented.

[0111] The glasses 100 include a data processor 1002, a display 1010, one or more cameras 1008, and additional input / output elements 1016. The input / output elements 1016 may include microphones, audio speakers, biometric sensors, additional sensors, or additional display elements integrated with the data processor 1002. Examples of input / output elements 1016 are as follows: Fig.12 and Figure 3 For example, the input / output element 1016 may include any I / O component 306, including the output component 328, the motion component 336, etc. Figure 2 An example of display 1010 is described in . In the specific examples described herein, display 1010 includes displays for the left and right eyes of a user.

[0112] Data processor 1002 includes image processor 1006 (eg, video processor), GPU and display driver 1038, tracking module 1040, interface 1012, low power circuitry 1004, and high speed circuitry 1020. Components of data processor 1002 are interconnected by bus 1042.

[0113] The interface 1012 refers to any source of user commands provided to the data processor 1002. In one or more examples, the interface 1012 is a physical button that sends a user input signal from the interface 1012 to the low-power processor 1014 when pressed. The low-power processor 1014 can process pressing such a button and then immediately releasing it as a request to capture a single image, and vice versa. The low-power processor 1014 can process pressing such a button for a first period of time as a request to capture video data when the button is pressed and stop video capture when the button is released, wherein the video captured when the button is pressed is stored as a single video file. Alternatively, pressing the button for a long period of time can capture a still image. In some examples, the interface 1012 can be any mechanical switch or physical interface capable of accepting user input associated with a data request from the camera 1008. In other examples, the interface 1012 can have a software component, or can be associated with a command received wirelessly from another source, such as from the client device 1026.

[0114] Image processor 1006 includes circuitry for receiving signals from camera 1008 and processing those signals from camera 1008 into a format suitable for storage in memory 1024 or for transmission to client device 1026. In one or more examples, image processor 1006 (e.g., a video processor) includes a microprocessor integrated circuit (IC) customized for processing sensor data from camera 1008, and volatile memory used by the microprocessor in operation.

[0115] The low power circuit system 1004 includes a low power processor 1014 and a low power wireless circuit system 1018. These elements of the low power circuit system 1004 can be implemented as separate elements or can be implemented on a single IC as part of a single system on a chip. The low power processor 1014 includes logic for managing other elements of the glasses 100. As described above, for example, the low power processor 1014 can accept user input signals from the interface 1012. The low power processor 1014 can also be configured to receive input signals or command communications from the client device 1026 via the low power wireless connection 1036. The low power wireless circuit system 1018 includes circuit elements for implementing a low power wireless communication system. Bluetooth TM Smart, also known as Bluetooth TM Low power consumption is a standard implementation of a low power wireless communication system that may be used to implement the low power wireless circuitry 1018. In other examples, other low power communication systems may be used.

[0116] High-speed circuit system 1020 includes a high-speed processor 1022, a memory 1024, and a high-speed wireless circuit system 1028. High-speed processor 1022 can be any processor capable of managing high-speed communications and operations for any general-purpose computing system for data processor 1002. High-speed processor 1022 includes processing resources for managing high-speed data transmission over high-speed wireless connection 1034 using high-speed wireless circuit system 1028. In some examples, high-speed processor 1022 executes an operating system such as a LINUX operating system or a program such as a UNIX operating system. Fig.12 The high-speed processor 1022 uses the software architecture of the data processor 1002 to manage data transmission with the high-speed wireless circuit system 1028, in addition to any other responsibilities. In some examples, the high-speed wireless circuit system 1028 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE)

[0117] 802.11 communication standard, which is also referred to herein as Wi-Fi. In other examples, high-speed wireless circuitry 1028 can implement other high-speed communication standards.

[0118] The memory 1024 includes any storage device capable of storing camera data generated by the camera 1008 and the image processor 1006. Although the memory 1024 is shown as being integrated with the high-speed circuitry 1020, in other examples, the memory 1024 may be a separate, independent element of the data processor 1002. In some such examples, electrical wiring may provide a connection from the image processor 1006 or the low-power processor 1014 to the memory 1024 through a chip including the high-speed processor 1022. In other examples, the high-speed processor 1022 may manage the addressing of the memory 1024 so that the low-power processor 1014 will initiate the high-speed processor 1022 any time a read or write operation involving the memory 1024 is needed.

[0119] The tracking module 1040 estimates the pose of the glasses 100. For example, the tracking module 1040 uses image data and associated inertial data from the camera 1008 and the positioning component 340, as well as GPS data, to track the position and determine the pose of the glasses 100 relative to a reference frame (e.g., a real-world scene environment). The tracking module 1040 continuously collects and uses updated sensor data describing the movement of the glasses 100 to determine an updated three-dimensional pose of the glasses 100, which indicates changes in relative position and orientation relative to physical objects in the real-world scene environment. The tracking module 1040 allows the glasses 100 to visually place virtual objects relative to physical objects within the user's field of view via the display 1010.

[0120] The GPU and display driver 1038 can use the pose of the glasses 100 to generate frames of virtual content or other content to be presented on the display 1010 when the glasses 100 are operating in a conventional augmented reality mode. In this mode, the GPU and display driver 1038 generate updated frames of virtual content based on the updated three-dimensional pose of the glasses 100, which reflects changes in the user's position and orientation relative to physical objects in the user's real scene environment.

[0121] One or more functions or operations described herein may also be performed in an application resident on the glasses 100 or on the client device 1026 or on a remote server. For example, one or more functions or operations described herein may be performed by one of the applications 1206, such as the messaging application 1246.

[0122] Fig.11 1 is a block diagram illustrating an example messaging system 1100 for exchanging data (e.g., messages and associated content) over a network. The messaging system 1100 includes multiple instances of client devices 1026 that host multiple applications including a messaging client 1102 and other applications 1104. The messaging client 1102 is communicatively coupled to other instances of the messaging client 1102 (e.g., hosted on respective other client devices 1026), a messaging server system 1106, and a third-party server 1108 via a network 1030 (e.g., the Internet). The messaging client 1102 may also communicate with the locally hosted application 1104 using an application program interface (API).

[0123] The messaging clients 1102 are able to communicate and exchange data with other messaging clients 1102 and with a messaging server system 1106 via the network 1030. The data exchanged between the messaging clients 1102 and between the messaging clients 1102 and the messaging server system 1106 includes functions (e.g., commands for activating functions) and payload data (e.g., text, audio, video, or other multimedia data).

[0124] The messaging server system 1106 provides server-side functionality to a particular messaging client 1102 via the network 1030. Although some functions of the messaging system 1100 are described herein as being performed by the messaging client 1102 or by the messaging server system 1106, it may be a design choice whether some functions are located within the messaging client 1102 or within the messaging server system 1106. For example, it may be technically preferable to initially deploy some technologies and functions within the messaging server system 1106, but then migrate the technologies and functions to the messaging client 1102 where the client device 1026 has sufficient processing power.

[0125] The messaging server system 1106 supports various services and operations provided to the messaging clients 1102. Such operations include sending data to the messaging clients 1102, receiving data from the messaging clients 1102, and processing data generated by the messaging clients 1102. As examples, the data may include message content, client device information, geographic location information, media enhancements and overlays, message content persistence conditions, social network information, and live event information. The data exchange within the messaging system 1100 is activated and controlled by functions available via the user interface (UI) of the messaging client 1102.

[0126] Turning now specifically to the messaging server system 1106, an application program interface (API) server 1110 is coupled to and provides a programming interface to an application server 1114. The application server 1114 is communicatively coupled to a database server 1116, which facilitates access to a database 1120 that stores data associated with messages processed by the application server 1114. Similarly, a web server 1124 is coupled to and provides a web-based interface to the application server 1114. To this end, the web server 1124 handles incoming network requests via the Hypertext Transfer Protocol (HTTP) and several other related protocols.

[0127] The application program interface (API) server 1110 receives and sends message data (e.g., commands and message payloads) between the client device 1026 and the application server 1114. In particular, the application program interface (API) server 1110 provides a set of interfaces (e.g., routines and protocols) that can be called or queried by the messaging client 1102 to activate the functionality of the application server 1114. The application program interface (API) server 1110 exposes various functions supported by the application server 1114, including: account registration; login functionality; sending messages from a particular messaging client 1102 to another messaging client 1102 via the application server 1114; sending media files (e.g., images or videos) from the messaging client 1102 to the messaging server 1112 and for possible access by another messaging client 1102; setting of media data collections (e.g., stories); retrieving a friend list of a user of the client device 1026; retrieving such collections; retrieving messages and content; adding and removing entities (e.g., friends) from an entity graph (e.g., a social graph); locating friends within a social graph; and opening application events (e.g., related to the messaging client 1102).

[0128] The application server 1114 hosts multiple server applications and subsystems, including, for example, a messaging server 1112, an image processing server 1118, and a social network server 1122. The messaging server 1112 implements multiple message processing technologies and functions, particularly those related to the aggregation and other processing of content (e.g., text and multimedia content) included in messages received from multiple instances of the messaging client 1102. As will be described in more detail, text and media content from multiple sources can be aggregated into collections of content (e.g., referred to as stories or galleries). These collections are then made available to the messaging client 1102. In view of the hardware requirements for other processor- and memory-intensive data processing, such processing can also be performed on the server side by the messaging server 1112.

[0129] The application server 1114 also includes an image processing server 1118 that is dedicated to performing various image processing operations, typically with respect to images or videos within the payload of messages sent from or received at the messaging server 1112.

[0130] The social network server 1122 supports various social networking functions and services and makes these functions and services available to the messaging server 1112. To this end, the social network server 1122 maintains and accesses an entity graph within the database 1120. Examples of functions and services supported by the social network server 1122 include identifying other users in the messaging system 1100 with whom a particular user has a relationship or who the particular user is "following", and also includes identifying interests and other entities of a particular user.

[0131] The messaging client 1102 may notify the user of the client device 1026 or other users associated with such a user (e.g., "friends") of activities occurring in a shared or shareable session. For example, the messaging client 1102 may provide notifications related to current or recent use of a game by one or more members of a user group to participants in a conversation (e.g., a chat session) in the messaging client 1102. One or more users may be invited to join an active session or initiate a new session. In some examples, a shared session may provide a shared augmented reality experience in which multiple people may collaborate or participate.

[0132] Fig.12 1200 is a block diagram illustrating a software architecture 1204 that may be installed on any one or more of the devices described herein. The software architecture 1204 is supported by hardware, such as a machine 1202 including a processor 1220, a memory 1226, and an I / O component 1238. In this example, the software architecture 1204 may be conceptualized as a stack of layers, where each layer provides specific functionality. The software architecture 1204 includes layers such as an operating system 1212, a library 1208, a framework 1210, and an application 1206. In operation, the application 1206 invokes an API call 1250 through the software stack and receives a message 1252 in response to the API call 1250.

[0133] The operating system 1212 manages hardware resources and provides public services. The operating system 1212 includes, for example, a kernel 1214, services 1216, and drivers 1222. The kernel 1214 serves as an abstraction layer between the hardware layer and other software layers. For example, the kernel 1214 provides functions such as memory management, processor management (e.g., scheduling), component management, networking, and security settings. Services 1216 can provide other public services to other software layers. Drivers 1222 are responsible for controlling or interfacing with the underlying hardware. For example, drivers 1222 may include display drivers, camera drivers, or Low-power drivers, flash drives, serial communication drivers (e.g., Universal Serial Bus (USB) drivers), WI- Drivers, audio drivers, power management drivers, etc.

[0134] The libraries 1208 provide low-level common infrastructure used by the applications 1206. The libraries 1208 may include system libraries 1218 (eg, C standard libraries) that provide functions such as memory allocation functions, string manipulation functions, math functions, and the like. In addition, the library 1208 may include an API library 1224, such as a media library (e.g., a library for supporting presentation and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), a graphics library (e.g., an OpenGL framework for rendering graphics content in two dimensions (2D) and three dimensions (3D) on a display, GLMotif for implementing a user interface), an image feature extraction library (e.g., OpenIMAJ), a database library (e.g., SQLite providing various relational database functions), a web library (e.g., WebKit providing web browsing functions), etc. The library 1208 may also include various other libraries 1228 to provide many other APIs to the application 1206.

[0135] The framework 1210 provides a high-level common infrastructure used by the applications 1206. For example, the framework 1210 provides various graphical user interface (GUI) functions, advanced resource management, and advanced location services. The framework 1210 can provide a wide range of other APIs that can be used by the applications 1206, some of which may be specific to a particular operating system or platform.

[0136] In an example, applications 1206 may include a home application 1236, a contacts application 1230, a browser application 1232, a book reader application 1234, a location application 1242, a media application 1244, a messaging application 1246, a game application 1248, and a variety of other applications such as third-party applications 1240. Applications 1206 are programs that execute functions defined in the program. Various programming languages ​​may be used to create one or more of the applications 1206 constructed in various ways, such as an object-oriented programming language (e.g., Objective-C, Java, or C++) or a procedural programming language (e.g., C or assembly language). In a specific example, third-party applications 1240 (e.g., applications created by an entity other than the vendor of a particular platform using ANDROID TM or IOS TM Software Development Kit (SDK) can be used to develop applications on platforms such as IOSTM ANDROID TM , Mobile software running on the mobile operating system of the phone or another mobile operating system. In this example, the third-party application 1240 can activate the API call 1250 provided by the operating system 1212 to facilitate the functions described herein.

[0137] "Carrier signal" refers to any intangible medium that can store, encode or carry instructions for execution by a machine and includes a digital or analog communications signal or other intangible medium that facilitates communication of such instructions. Instructions may be sent or received over a network using a transmission medium via a network interface device.

[0138] "Client Device" refers to any machine that interfaces with a communications network to obtain resources from one or more server systems or other client devices. A client device may be, but is not limited to, a mobile phone, a desktop computer, a laptop computer, a portable digital assistant (PDA), a smart phone, a tablet computer, an ultrabook, a netbook, a laptop computer, a multiprocessor system, a microprocessor-based or programmable consumer electronics product, a game console, a set-top box, or any other communications device that a user may use to access a network.

[0139] “Communications Network” means one or more parts of a network, which may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the Public Switched Telephone Network (PSTN), a Plain Old Telephone Service (POTS) network, a cellular telephone network, a wireless network, a Wi-Fi network, a Wi-Fi hotspot ... The coupling may be a network, another type of network, or a combination of two or more such networks. For example, the network or a portion of the network may include a wireless network or a cellular network, and the coupling may be a code division multiple access (CDMA) connection, a global system for mobile communications (GSM) connection, or other types of cellular or wireless couplings. In this example, the coupling may implement any of various types of data transmission technologies, such as single carrier radio transmission technology (1xRTT), evolution data optimized (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rate for GSM evolution (EDGE) technology, including the third generation partnership project (3GPP) of 3G, fourth generation wireless (4G) network, universal mobile telecommunications system (UMTS), high speed packet access (HSPA), world wide interoperability for microwave access (WiMAX), long term evolution (LTE) standards, other data transmission technologies defined by various standard setting organizations, other long distance protocols, or other data transmission technologies.

[0140] "Component" refers to a device, physical entity or logic with the following boundaries, which is defined by function or subroutine calls, branch points, APIs or other technologies that provide partitioning or modularization for specific processing or control functions. Components can be combined with other components via their interfaces to perform machine processing. Components can be packaged functional hardware units designed for use with other components, and part of a program that generally performs a specific function of related functions. Components can constitute software components (e.g., codes implemented on machine-readable media) or hardware components. "Hardware components" are tangible units that can perform some operations and can be configured or arranged in a specific physical manner. In various examples, one or more computer systems (e.g., independent computer systems, client computer systems, or server computer systems) or one or more hardware components (e.g., processors or processor groups) of a computer system can be configured by software (e.g., applications or application parts) to operate to perform hardware components of some operations described herein. Hardware components can also be implemented mechanically, electronically, or any suitable combination thereof. For example, a hardware component may include a dedicated circuit system or logic that is permanently configured to perform some operations. The hardware component may be a dedicated processor, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC). The hardware component may also include a programmable logic or circuit system that is temporarily configured to perform some operations by software. For example, the hardware component may include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a specific machine (or a specific component of a machine) that is customized to perform the configured function, and is no longer a general-purpose processor. It should be understood that the decision to mechanically implement the hardware component in a dedicated and permanently configured circuit system or in a temporarily configured (e.g., configured by software) circuit system can be driven for cost and time considerations. Accordingly, the phrase "hardware component" (or "hardware-implemented component") should be understood to include a tangible entity, that is, an entity that is physically constructed, permanently configured (e.g., hardwired) or temporarily configured (e.g., programmed) to operate in a specific manner or perform some operations described herein. Considering an example in which a hardware component is temporarily configured (e.g., programmed), the hardware component may not be configured or instantiated at any time. For example, where a hardware component includes a general purpose processor that is configured by software to become a special purpose processor, the general purpose processor may be configured as different special purpose processors (e.g., including different hardware components) at different times. For example, software configures one or more specific processors accordingly to constitute a specific hardware component at one time and to constitute different hardware components at different times. Hardware components may provide information to other hardware components and may receive information from other hardware components. Thus, the described hardware components may be considered to be communicatively coupled.In the case of multiple hardware components being present at the same time, communication can be achieved by signal transmission between or among two or more hardware components in the hardware components (e.g., through appropriate circuits and buses). In examples where multiple hardware components are configured or instantiated at different times, communication between such hardware components can be achieved, for example, by storing information in a memory structure accessible to multiple hardware components and retrieving information in the memory structure. For example, a hardware component can perform an operation and store the output of the operation in a memory device coupled to it in communication. Then, other hardware components can access the memory device at a subsequent time to retrieve the stored output and process it. The hardware component can also initiate communication with an input or output device, and can operate on resources (e.g., a collection of information). The various operations of the example methods described herein can be performed by a temporary configuration (e.g., by software) or one or more processors that are permanently configured to perform related operations. Whether it is a temporary configuration or a permanent configuration, such a processor can constitute a processor-implemented component that operates to perform one or more operations or functions described herein. As used herein, a "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the method described herein can be implemented in part by a processor, wherein specific one or more processors are examples of hardware. For example, some operations in the operation of the method can be performed by one or more processors or the parts implemented by the processor. In addition, one or more processors can also operate to support the execution of related operations in the "cloud computing" environment or operate as "software as a service" (SaaS). For example, some operations in the operation can be performed by a group of computers (as an example of a machine including a processor), wherein these operations can be accessed via a network (for example, the Internet) and via one or more appropriate interfaces (for example, API). The execution of some operations in the operation can be distributed between processors, reside in a single machine, and deployed across multiple machines. In some examples, a processor or a part implemented by a processor can be located in a single geographical location (for example, in a home environment, an office environment, or a server cluster). In other examples, a processor or a part implemented by a processor can be distributed across multiple geographical locations.

[0141] "Computer-readable media" refers to both machine storage media and transmission media. Therefore, these terms include both storage devices / media and carrier / modulated data signals. The terms "machine-readable medium", "computer-readable medium" and "device-readable medium" mean the same thing and can be used interchangeably in this disclosure.

[0142] "Machine storage medium" refers to a single or multiple storage devices and / or media (e.g., centralized or distributed databases, and / or associated caches and servers) that store executable instructions, routines, and / or data. The term includes, but is not limited to, solid-state memory and optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage media, computer storage media, and / or device storage media include: non-volatile memory, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGA, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage medium", "device storage medium", "computer storage medium" mean the same thing and may be used interchangeably in this disclosure. The terms "machine storage medium", "computer storage medium", and "device storage medium" expressly exclude carrier waves, modulated data signals, and other such media, some of which are covered by the term "signal media".

[0143] A "processor" refers to any circuit or virtual circuit (a physical circuit simulated by logic executed on an actual processor) that manipulates data values ​​according to control signals (e.g., "commands," "opcodes," "machine codes," etc.) and produces associated output signals that are applied to operate a machine. A processor may be, for example, a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), or any combination thereof. A processor may also be a multi-core processor having two or more independent processors (sometimes referred to as "cores") that can execute instructions simultaneously.

[0144] "Signal medium" refers to any intangible medium that is capable of storing, encoding or carrying instructions executed by a machine, and includes digital or analog communication signals or other intangible media that facilitate the communication of software or data. The term "signal medium" may be considered to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure.

[0145] Changes and modifications may be made to the disclosed examples without departing from the scope of the present disclosure. These and other changes or modifications are intended to be included within the scope of the present disclosure as expressed in the appended claims.

[0146] In view of the above-mentioned implementation methods of the subject matter, the present application discloses the following list of examples, wherein one feature of a separate example or more than one feature of an example adopted in combination and optionally combined with one or more features of one or more additional examples are additional examples that also fall within the disclosure of the present application.

[0147] Example 1 is a computer-implemented method, comprising: obtaining, by a computing system comprising one or more processors and a memory, audio data captured by one or more microphones; analyzing, by the computing system, the audio data to generate text data corresponding to at least one of one or more words or one or more phrases included in the audio data; generating, by the computing system, a search request comprising one or more keywords extracted from the text data; obtaining, by the computing system, search results indicating one or more content items corresponding to the one or more keywords in the search request; determining, by the computing system, one or more features of a content item among the one or more content items, the one or more features comprising at least one of a source of the content item or a format of the content item; determining, by the computing system, a layout of content included in the content item based on the one or more features of the content item; and causing, by the computing system, a user interface to be displayed in an augmented reality environment, the user interface comprising content of the content item presented according to the layout.

[0148] In Example 2, the subject matter of Example 1 includes: determining, by a computing system, a field of view of a gaze of an individual; and determining, by the computing system, a location within a real-world scene for displaying a user interface, the location corresponding to at least a portion of the field of view of a gaze of the individual.

[0149] In Example 3, the subject matter of Example 2 includes: obtaining, by a computing system, camera data from one or more cameras included in an augmented reality environment; obtaining, by a computing system, sensor data from one or more inertial measurement unit sensors included in the augmented reality environment; and analyzing, by the computing system, at least one of the camera data or the sensor data to determine a field of view of an individual's gaze.

[0150] In Example 4, the subject matter of Example 2 or Example 3 includes: determining, by a computing system, that an individual's gaze field of view has changed from a first position to a second position; and moving, by the computing system, a position within a real-world scene for displaying a user interface from the first position to the second position.

[0151] In Example 5, the subject matter of Examples 2 to 4 includes: obtaining, by a computing system, additional audio data captured by one or more microphones; and analyzing, by the computing system, the additional audio data to generate additional text data corresponding to at least one of one or more additional words or one or more additional phrases included in the additional audio data.

[0152] In Example 6, the subject matter of Example 5 includes: determining, by a computing system, that the additional text data includes a command to fix the position of a user interface within a real-world scene; determining, by the computing system, that the individual's gaze field of view has changed from a first position to a second position; and causing the computing system to keep the user interface displayed at the said position.

[0153] In Example 7, the subject matter of Example 5 includes: determining, by the computing system, that the additional text data includes a command to modify a display characteristic of the content item; and modifying, by the computing system, an appearance of the content item within the user interface based on the command to modify the display characteristic.

[0154] In Example 8, the subject matter of Examples 1 to 7 includes: obtaining, by a computing system, camera data from one or more cameras included in an augmented reality environment; analyzing, by the computing system, the camera data to determine one or more objects located in the augmented reality environment; and causing, by the computing system, a user interface to display a location of an object among the one or more objects.

[0155] In Example 9, the subject matter of Examples 1 to 8 includes: causing an additional user interface to be displayed by a computing system in an augmented reality environment, the additional user interface including a first user interface element corresponding to a first content item among one or more content items and a second user interface element corresponding to a second content item among the one or more content items; obtaining, by the computing system, camera data from one or more cameras included in the augmented reality environment; obtaining, by the computing system, sensor data from one or more inertial measurement unit sensors included in the augmented reality environment; analyzing, by the computing system, at least one of the camera data or the sensor data to determine a field of view of an individual's gaze; determining, by the computing system, that the field of view of the individual's gaze corresponds to a position of the first user interface element; and modifying, by the computing system, one or more display characteristics of the first user interface element.

[0156] In Example 10, the subject matter of Example 9 includes: obtaining, by a computing system, additional audio data captured by one or more microphones; analyzing, by the computing system, the additional audio data to generate additional text data corresponding to at least one of one or more additional words or one or more additional phrases included in the additional audio data; determining, by the computing system, that the additional text data includes a command to access content of at least one content item; determining, by the computing system, that a first user interface element has been selected based on a location of a field of view of an individual's gaze and based on the command; and causing, by the computing system, to display another user interface in an augmented reality environment, the other user interface including content of the first content item.

[0157] Example 11 is a computing device comprising: one or more processors; and a memory storing instructions, which, when executed by the one or more processors, cause the computing device to perform operations including the following: obtaining audio data captured by one or more microphones; analyzing the audio data to generate text data corresponding to at least one of one or more words or one or more phrases included in the audio data; generating a search request including one or more keywords extracted from the text data; obtaining search results indicating one or more content items corresponding to the one or more keywords in the search request; determining one or more features of a content item among the one or more content items, the one or more features including at least one of a source of the content item or a format of the content item; determining a layout of content included in the content item based on the one or more features of the content item; and causing a user interface to be displayed in an augmented reality environment, the user interface including content of the content item presented according to the layout.

[0158] In Example 12, the subject matter of Example 11 includes a memory storing additional instructions that, when executed by one or more processors, cause the computing device to perform additional operations including: analyzing one or more features of a content item relative to multiple feature sets corresponding to multiple content templates to determine a similarity measure between the one or more features and one or more feature sets in the multiple feature sets, each of the multiple content templates indicating a corresponding arrangement of content within one or more user interfaces; and determining, based on the similarity measure, the following content template among the multiple content templates: the content of the content item is displayed through the content template.

[0159] In Example 13, the subject matter of Example 11 or Example 12 includes a memory storing additional instructions that, when executed by one or more processors, cause the computing device to perform additional operations including: analyzing one or more features of a content item to determine a source of the content item; and determining a content template from a plurality of content templates based on the source of the content item, wherein the content template indicates a corresponding arrangement of content of the content item within a user interface.

[0160] In Example 14, the subject matter of Examples 11 to 13 includes a content template indicating a first portion of the user interface for displaying text content and a second portion of the user interface for displaying at least one of image content or video content.

[0161] In Example 15, the subject matter of Examples 11 to 14 includes: a content item including instructional content, the instructional content indicating multiple steps of a teaching process; and a memory storing additional instructions, which, when executed by one or more processors, causes the computing device to perform additional operations including: analyzing the content item to determine multiple discrete parts of the content item corresponding to respective steps of the multiple steps of the teaching process; and generating multiple user interfaces, so that each of the multiple user interfaces includes content corresponding to each step of the teaching process.

[0162] In Example 16, the subject matter of Example 15 includes: a user interface comprising first content of a content item corresponding to a first step of a teaching process; and a memory storing additional instructions, which, when executed by one or more processors, cause the computing device to perform additional operations comprising: obtaining additional audio data captured by one or more microphones; analyzing the additional audio data to generate additional text data corresponding to at least one of one or more additional words or one or more additional phrases included in the additional audio data; determining that the additional text data includes a command to navigate to a second step of the teaching process; and causing an additional user interface to be displayed in an augmented reality environment, the additional user interface comprising the additional content of the content item corresponding to the second step of the teaching process.

[0163] In Example 17, the subject matter of Examples 11 to 16 includes a memory storing additional instructions that, when executed by one or more processors, cause the computing device to perform additional operations including: obtaining additional audio data captured by one or more microphones; analyzing the additional audio data to generate additional text data corresponding to at least one of one or more additional words or one or more additional phrases included in the additional audio data; analyzing the additional text data to determine multiple similarity metrics between at least one of the one or more additional words or one or more additional phrases and one or more words in a plurality of commands; determining that the additional text data corresponds to a command in a plurality of commands based on a similarity metric in the plurality of similarity metrics; and causing an action corresponding to the command to be executed.

[0164] Example 18 is one or more non-transitory computer-readable storage media, comprising computer-readable instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations including: obtaining audio data captured by one or more microphones; analyzing the audio data to generate text data corresponding to at least one of one or more words or one or more phrases included in the audio data; generating a search request including one or more keywords extracted from the text data; obtaining search results indicating one or more content items corresponding to the one or more keywords in the search request; determining one or more features of a content item in the one or more content items, the one or more features including at least one of a source of the content item or a format of the content item; determining a layout of content included in the content item based on the one or more features of the content item; and causing a user interface to be displayed in an augmented reality environment, the user interface including content of the content item presented according to the layout.

[0165] In Example 19, the subject matter of Example 18 includes: a user interface including a command menu, the command menu including multiple commands that can be selected to perform one or more actions with respect to a content item; and one or more non-transitory computer-readable storage media, the one or more non-transitory computer-readable storage media including additional computer-readable instructions, which, when executed by one or more hardware processors, cause the one or more hardware processors to perform additional operations including: analyzing additional audio data to generate additional text data corresponding to at least one of one or more additional words or one or more additional phrases included in the additional audio data; determining that the additional text data corresponds to a command among multiple commands; and causing an action corresponding to the command to be performed with respect to the content of the content item.

[0166] In Example 20, the subject matter of Example 19 includes additional computer-readable instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform additional operations including causing a command menu to be displayed in a first portion of the user interface, causing one or more user interface elements corresponding to one or more content items to be displayed in a second portion of the user interface, and causing one or more words of text data to be displayed in a third portion of the user interface.

Claims

1. A computer-implemented method comprising: obtaining, by a computing system including one or more processors and memory, audio data captured by one or more microphones; analyzing, by the computing system, the audio data to generate text data corresponding to at least one of the one or more words or the one or more phrases included in the audio data; generating, by the computing system, a search request including one or more keywords extracted from the text data; obtaining, by the computing system, search results indicating one or more content items corresponding to the one or more keywords in the search request; determining, by the computing system, one or more characteristics of a content item of the one or more content items, the one or more characteristics comprising at least one of a source of the content item or a format of the content item; determining, by the computing system, a layout of content included in the content item based on the one or more characteristics of the content item; as well as A user interface is caused to be displayed, by the computing system, in an augmented reality environment, the user interface including content of the content item presented according to the layout.

2. The computer-implemented method of claim 1 , comprising: determining, by the computing system, a field of view of gaze of the individual; as well as A location within a real-world scene for displaying the user interface is determined by the computing system, the location corresponding to at least a portion of a field of view of a gaze of the individual.

3. The computer-implemented method of claim 2, comprising: obtaining, by the computing system, camera data from one or more cameras included in the augmented reality environment; obtaining, by the computing system, sensor data from one or more inertial measurement unit sensors included in the augmented reality environment; as well as At least one of the camera data or the sensor data is analyzed by the computing system to determine a field of view of gaze of the individual.

4. The computer-implemented method of claim 2, comprising: determining, by the computing system, that a field of view of the individual's gaze has changed from a first location to a second location; as well as A location within the real-world scene for displaying the user interface is moved, by the computing system, from the first location to the second location.

5. The computer-implemented method of claim 2, comprising: obtaining, by the computing system, additional audio data captured by the one or more microphones; as well as The additional audio data is analyzed by the computing system to generate additional text data corresponding to at least one of the one or more additional words or the one or more additional phrases included in the additional audio data.

6. The computer-implemented method of claim 5, comprising: determining, by the computing system, that the additional text data includes a command to fix a position of the user interface within the real-world scene; determining, by the computing system, that a field of view of the individual's gaze has changed from a first location to a second location; as well as The user interface is caused to remain displayed at the location by the computing system.

7. The computer-implemented method of claim 5, comprising: determining, by the computing system, that the additional text data includes a command to modify a display characteristic of the content item; as well as The appearance of the content item within the user interface is modified by the computing system based on the command to modify the display characteristic.

8. The computer-implemented method of claim 1 , comprising: obtaining, by the computing system, camera data from one or more cameras included in the augmented reality environment; analyzing, by the computing system, the camera data to determine one or more objects located in the augmented reality environment; as well as The computing system causes the user interface to display a location of an object among the one or more objects.

9. The computer-implemented method of claim 1 , comprising: causing, by the computing system, to display an additional user interface in the augmented reality environment, the additional user interface comprising a first user interface element corresponding to a first content item of the one or more content items and a second user interface element corresponding to a second content item of the one or more content items; obtaining, by the computing system, camera data from one or more cameras included in the augmented reality environment; obtaining, by the computing system, sensor data from one or more inertial measurement unit sensors included in the augmented reality environment; analyzing, by the computing system, at least one of the camera data or the sensor data to determine a field of view of gaze of an individual; determining, by the computing system, that a field of view of gaze of the individual corresponds to a location of the first user interface element; as well as One or more display characteristics of the first user interface element are modified by the computing system.

10. The computer-implemented method of claim 9, comprising: obtaining, by the computing system, additional audio data captured by the one or more microphones; analyzing, by the computing system, the additional audio data to generate additional text data corresponding to at least one of the one or more additional words or the one or more additional phrases included in the additional audio data; determining, by the computing system, that the additional textual data includes a command to access content of at least one content item; determining, by the computing system, that the first user interface element has been selected based on a location of a field of view of a gaze of the individual and based on the command; as well as Another user interface is caused to be displayed, by the computing system, in the augmented reality environment, the other user interface including content of the first content item.

11. A computing device comprising: one or more processors; as well as a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform operations including: obtaining audio data captured by one or more microphones; analyzing the audio data to generate text data corresponding to at least one of the one or more words or the one or more phrases included in the audio data; generating a search request including one or more keywords extracted from the text data; obtaining search results indicating one or more content items corresponding to the one or more keywords in the search request; determining one or more characteristics of a content item of the one or more content items, the one or more characteristics comprising at least one of a source of the content item or a format of the content item; determining a layout of content included in the content item based on the one or more characteristics of the content item; as well as A user interface is caused to be displayed in an augmented reality environment, the user interface including content of the content item presented according to the layout.

12. The computing device according to claim 11, wherein: The memory stores additional instructions that, when executed by the one or more processors, cause the computing device to perform additional operations including: analyzing the one or more features of the content item relative to a plurality of feature sets corresponding to a plurality of content templates to determine a similarity measure between the one or more features and one or more of the plurality of feature sets, respective content templates of the plurality of content templates indicating a respective arrangement of content within one or more user interfaces; as well as Based on the similarity measure, a content template among the plurality of content templates is determined through which content of the content item is displayed.

13. The computing device of claim 11, the memory storing additional instructions that, when executed by the one or more processors, cause the computing device to perform additional operations including: analyzing the one or more characteristics of the content item to determine a source of the content item; and determining a content template from a plurality of content templates based on a source of the content item, wherein: The content template indicates a corresponding arrangement of content of the content item within the user interface.

14. The computing device of claim 11, wherein: The content template indicates a first portion of the user interface for displaying text content and a second portion of the user interface for displaying at least one of image content or video content.

15. The computing device of claim 11, wherein: The content item includes instruction content, the instruction content indicating a plurality of steps of a teaching process; and The memory stores additional instructions that, when executed by the one or more processors, cause the computing device to perform additional operations including: analyzing the content item to determine a plurality of discrete portions of the content item corresponding to respective ones of a plurality of steps of the instructional process; as well as A plurality of user interfaces are generated, so that each of the plurality of user interfaces includes content corresponding to each step of the teaching process.

16. The computing device of claim 15, wherein: The user interface includes first content of the content item corresponding to a first step of the teaching process; and The memory stores additional instructions that, when executed by the one or more processors, cause the computing device to perform additional operations including: obtaining additional audio data captured by the one or more microphones; analyzing the additional audio data to generate additional text data corresponding to at least one of the one or more additional words or the one or more additional phrases included in the additional audio data; determining that the additional text data includes a command to navigate to a second step of the teaching process; as well as An additional user interface is caused to be displayed in the augmented reality environment, the additional user interface including additional content of the content item corresponding to the second step of the teaching process.

17. The computing device of claim 11, wherein: The memory stores additional instructions that, when executed by the one or more processors, cause the computing device to perform additional operations including: obtaining additional audio data captured by the one or more microphones; analyzing the additional audio data to generate additional text data corresponding to at least one of the one or more additional words or the one or more additional phrases included in the additional audio data; analyzing the additional text data to determine a plurality of similarity measures between at least one of the one or more additional words or the one or more additional phrases and one or more words in a plurality of commands; determining, based on a similarity measure in the plurality of similarity measures, that the additional text data corresponds to a command in the plurality of commands; as well as Causes the action corresponding to the command to be executed.

18. One or more non-transitory computer-readable storage media comprising computer-readable instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations comprising: obtaining audio data captured by one or more microphones; analyzing the audio data to generate text data corresponding to at least one of the one or more words or the one or more phrases included in the audio data; generating a search request including one or more keywords extracted from the text data; obtaining search results indicating one or more content items corresponding to the one or more keywords in the search request; determining one or more characteristics of a content item of the one or more content items, the one or more characteristics comprising at least one of a source of the content item or a format of the content item; determining a layout of content included in the content item based on the one or more characteristics of the content item; as well as A user interface is caused to be displayed in an augmented reality environment, the user interface including content of the content item presented according to the layout.

19. The one or more non-transitory computer-readable storage media of claim 18, wherein: The user interface includes a command menu including a plurality of commands selectable to perform one or more actions with respect to the content item; and The one or more non-transitory computer-readable storage media include additional computer-readable instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform additional operations including: analyzing the additional audio data to generate additional text data corresponding to at least one of the one or more additional words or the one or more additional phrases included in the additional audio data; determining that the additional text data corresponds to a command among the plurality of commands; as well as An action corresponding to the command is caused to be performed with respect to the content of the content item.

20. The one or more non-transitory computer-readable storage media of claim 19, comprising additional computer-readable instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform additional operations comprising: The command menu is displayed in a first portion of the user interface, one or more user interface elements corresponding to the one or more content items are displayed in a second portion of the user interface, and one or more words of the text data are displayed in a third portion of the user interface.