Generating augmented reality content including translations

By implementing an augmented reality system, the function of automatically translating identifiers when viewing objects was achieved, which solved the problems of low translation efficiency and high error rate in the existing system and improved the user experience.

CN120418801APending Publication Date: 2025-08-01SNAP INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380086135.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-16
Filing Date
2023-12-15
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing systems cannot efficiently translate object identifiers from the first language to the second language when users are viewing objects in the environment, which requires users to switch their attention and increases the possibility of translation errors.

Method used

With the augmented reality system, users can mark objects by gestures, audio, or touch input while viewing them. The system automatically recognizes and translates the object's identifier and displays the translation results, including text and audio content, in real time on the user interface.

Benefits of technology

It implements the function of automatically translating identifiers when users view objects in the environment, which improves translation efficiency and reduces the possibility of user distraction and translation errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120418801A_ABST
    Figure CN120418801A_ABST
Patent Text Reader

Abstract

An augmented reality (AR) translation system is provided. The AR translation system may analyze the camera data to determine an object included in a field of view of a camera of the user device. Augmented reality content may be provided that includes a visual translation of an object included in a field of view from a primary language to an additional language of a user. The translated auditory version may also be provided as part of augmented reality content. The user may also add an object in the field of view to a list of translated objects associated with the user based on at least one of a touch input, an audio input, or a gesture input.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority Claim

[0002] This patent application claims the benefit of priority of U.S. Patent Application Serial No. 18 / 082,969, filed on December 16, 2022, which is hereby incorporated by reference in its entirety. Technical Field

[0003] The present disclosure generally relates to generating augmented reality content including translations. Background Art

[0004] A head-mounted device can be implemented with a transparent or translucent display through which a user of the head-mounted device can view the surrounding environment. Such a device enables the user to view the surrounding environment through the transparent or translucent display and also enables the user to see objects (e.g., virtual objects such as renderings of 2D or 3D graphical models, images, videos, text, etc.) generated for display as part of and / or superimposed on the surrounding environment. This is commonly referred to as "augmented reality" or "AR". The head-mounted device can also completely occlude the user's field of view and display a virtual environment through which the user can move or be moved. This is commonly referred to as "virtual reality" or "VR". As used herein, unless the context otherwise indicates, the term "AR" refers to either or both of augmented reality and virtual reality as conventionally understood.

[0005] A user of a head-mounted device can access and use computer software applications to perform various tasks or engage in entertainment activities. To use a computer software application, the user interacts with a user interface provided by the head-mounted device. Brief Description of the Drawings

[0006] To facilitate identification of discussion of any particular element or act, one or more of the most significant digits in the reference numerals refer to the figure number in which the element is first introduced.

[0007] Figure 1 is a perspective view of a head-mounted device according to one or more examples.

[0008] Figure 2 is according to one or more examples of Figure 1 additional views of the head-mounted device.

[0009] Figure 3 is a graphical representation of a machine in the form of a computing device within which a set of instructions can be executed to cause the machine to perform any one or more of the methods discussed herein.

[0010] Figure 4is a diagram of a computing architecture including one or more systems for generating augmented reality content, the augmented reality content including translations of identifiers of objects located in an environment.

[0011] Figure 5 is a diagram of a computing architecture including one or more systems for tracking objects stored in an object translation list of a user for translating augmented reality content items.

[0012] Figure 6 is a diagram of a computing architecture for generating augmented reality content, the augmented reality content including text content and audio content of identifiers of objects that have been translated from a first language to a second language.

[0013] Figure 7 is a flowchart of a process for adding an object to an object translation list of a user of a user application and generating augmented reality content, the augmented reality content including a translation corresponding to an identifier of the object.

[0014] Figure 8 is a flowchart of a process for generating augmented reality content, the augmented reality content including translations of words of objects included in an object translation list of a user of a user application.

[0015] Figure 9 is a diagram including a user interface, the user interface including a view of a real-world scene and including augmented reality content corresponding to translations of objects included in the real-world scene.

[0016] Figure 10 is a block diagram showing a software architecture in which the present disclosure may be implemented according to one or more examples.

[0017] Figure 11 is a diagram showing according to one or more examples Figure 1 [[ID=३०]]details of a head-mounted device.

[0018] Figure 12 is a graphical representation of a networked environment in which the present disclosure may be deployed according to one or more examples. Detailed Description

[0019] In many augmented reality (AR) systems, users can interact with virtual objects displayed in their environment. An input modality that can be used with an AR system is hand tracking combined with direct manipulation of virtual objects (DMVO), where the user is provided with a user interface that is displayed to the user as an AR overlay having two-dimensional (2D) or three-dimensional (3D) rendering. The rendering is a graphical model in 2D or 3D, where the virtual objects located within the model correspond to the interaction elements of the user interface. In this way, the user perceives the virtual objects as objects within an overlay in the user's field of view of the real-world scene when wearing the AR system, or perceives the virtual objects as objects within a virtual world as seen by the user when wearing the AR system. To allow the user to manipulate the virtual objects, the AR system detects the user's hand and tracks the movement, position, and / or orientation of the hand to determine the user's interaction with the virtual objects. Additionally, the AR system can determine the user's interaction with the virtual objects in response to commands provided by the user.

[0020] There are various ways in which an individual attempts to learn a language. In some cases, the individual takes a course in person, takes a course online, or does both in order to learn the language. In other cases, the individual utilizes a user application that is executed on a computing device such as a mobile phone, a laptop computing device, or a tablet computing device, and the user application is customized to teach the vocabulary and grammar of one or more languages. In existing systems that rely on a computing device and a user application for teaching a language, when users are viewing their environment, they typically cannot identify the identifiers of the objects within their environment. That is, when an individual moves from one place to another within a location, the individual encounters a number of objects for which the individual can recall the identifier of the object in a first language but not in a second language. Using the existing systems in these instances, the individual can access the user application executed on the computing device to perform a translation operation and determine the identifier of the object in the second language. This is generally an inefficient process because the individual spends time taking their eyes off the object being viewed in their environment in order to focus their attention on entering the identifier of the object in the first language and requesting a translation into the second language. Thus, existing systems lack the ability to translate the identifier of an object from a first language into a second language when the individual is viewing the object in their environment.

[0021] The implementation of the augmented reality system described herein enables a user to view a translation of an object identifier while viewing an object in the user's environment. In various examples, a user can tag an object in their environment while viewing the object. The tag can indicate that the object identifier is to be translated from a first language to a second language. In one or more examples, a user can tag an object by placing the object within the field of view of a camera device of a mobile computing device and providing an object tagging input. The object tagging input can include audio input, gesture input, or a combination thereof. The object tagging input can also include a touch input that corresponds to a user touching a portion of the mobile computing device when the object is within the field of view of the camera device of the mobile computing device. In response to tagging the object, the object identifier can be stored in a data store that includes a user's data structure for translation augmented reality (AR) content items to be executed within a user application. The data structure can indicate the user's default language and additional languages selected by the user related to translation. Additionally, the data structure can indicate the objects and / or object identifiers for which the user has requested translation. Further, the data structure can indicate the location of the objects for which the user has requested translation.

[0022] After the user has tagged an object, as the user moves in their environment, the translation of the object identifier can be displayed in the user interface of the user device. For example, the location of the user device can be determined, and the object tagged by the user for translation can be identified, where the object corresponds to the location and is included within the field of view of the camera device of the user device. The translations of these objects can then be displayed as augmented reality content overlaid on a live view of the scene including the objects. Additionally, audio content of the pronunciation of the translation of the object in the center of the field of view of the camera device can be played. In this way, when the user views an object in their environment, the translation of the object identifier can be directly displayed for the user to view via the user interface of the user device (such as a smart phone or a head-mounted device), thereby showing a live view of the user's environment. Thus, the user can more easily associate the translated identifier with the object because the translated identifier can be automatically presented to the user when the user views the object, and the user does not have to shift their attention from viewing the object to viewing the translation of the object identifier displayed on a computing device. Further, the user can avoid the inefficiencies caused by existing systems in which the user views an object in their environment, shifts their attention away from the environment to provide input to a translation application executing on the user device, and then obtains a translation of the object identifier. In these cases, the likelihood of translation errors in existing systems increases due to errors in the user input provided to obtain the translation of the object identifier.

[0023] Other technical features will be readily apparent to those skilled in the art from the following drawings, description, and claims.

[0024] Figure 1 is a perspective view of an AR system in the form of a head-mounted device (e.g., Figure 1 glasses 100). The glasses 100 may include a frame 102 made of any suitable material such as plastic or metal, any suitable material including any suitable shape memory alloy. In one or more examples, the frame 102 includes a first optical element holder or left optical element holder 104 (e.g., a display or lens holder) and a second optical element holder or right optical element holder 106 connected by a bridge portion 112. A first optical element or left optical element 108 and a second optical element or right optical element 110 may be disposed within the left optical element holder 104 and the right optical element holder 106, respectively. The right optical element 110 and the left optical element 108 may be lenses, displays, display components, or combinations of the foregoing. Any suitable display component may be provided in the glasses 100.

[0025] The frame 102 additionally includes a left arm piece or left temple piece 122 and a right arm piece or right temple piece 124. In some examples, the frame 102 may be formed from a single piece of material to have a unified or integral construction.

[0026] The glasses 100 may include a computing device such as a computer 120, which may be of any suitable type for being carried by the frame 102, and in one or more examples, the computing device may have a suitable size and shape to be partially disposed within one of the temple pieces 122 or temple piece 124. The computer 120 may include one or more processors and a memory, wireless communication circuitry, and a power source. As discussed below, the computer 120 includes low-power circuitry, high-speed circuitry, and a display processor. Various other examples may include these elements configured differently or integrated together in different ways.

[0027] The computer 120 additionally includes a battery 118 or other suitable portable power supply. In some examples, the battery 118 is disposed within the left temple piece 122 and is electrically coupled to the computer 120 disposed within the right temple piece 124. The glasses 100 may include a connector or port (not shown) suitable for charging the battery 118, a wireless receiver, a transmitter, or a transceiver (not shown), or a combination of such devices.

[0028] Glasses 100 include a first or left camera device 114 and a second or right camera device 116. Although two camera devices are depicted, other examples contemplate the use of a single or additional (i.e., more than two) camera devices. In one or more examples, in addition to the left camera device 114 and the right camera device 116, the glasses 100 further include any number of input sensors or other input / output devices. Such sensors or input / output devices may additionally include biometric sensors, position sensors, motion sensors, and the like.

[0029] In some examples, the left camera device 114 and the right camera device 116 provide video frame data for the glasses 100 to use to extract 3D information from the real scene.

[0030] The glasses 100 may also include a touchpad 126 that is mounted to or integrated with one or both of the left temple piece 122 and the right temple piece 124. The touchpad 126 is typically arranged vertically, and in some examples, the touchpad 124 is approximately parallel to the user's temple. As used herein, being typically vertically aligned means that the touchpad is more vertical than horizontal, although potentially more vertical than said vertical. Additional user input may be provided by one or more buttons 128, which in the illustrated example are disposed on the outer upper edges of the left optical element holder 104 and the right optical element holder 106. The one or more touchpads 126 and buttons 128 provide a means by which the glasses 100 can receive input from a user of the glasses 100.

[0031] Figure 2 Glasses 100 are shown from the user's perspective. For clarity, Figure 1 several elements shown in Figure 1 are omitted. As Figure 2 described,

[0032] The glasses 100 include a forward optical assembly 202 that includes a right projector 204 and a right near-eye display 206; and a forward optical assembly 210 that includes a left projector 212 and a left near-eye display 216.

[0033] In some examples, the near-eye display is a waveguide. The waveguide includes a reflective or diffractive structure (e.g., a grating and / or optical elements such as mirrors, lenses, or prisms). The light 208 emitted by the projector 204 encounters the diffractive structure of the waveguide of the near-eye display 206, which directs the light towards the user's right eye to provide an image that superimposes a view of the real-world scene seen by the user on or within the right optical element 110. Similarly, the light 214 emitted by the projector 212 encounters the diffractive structure of the waveguide of the near-eye display 216, which directs the light towards the user's left eye to provide an image that superimposes a view of the real-world scene seen by the user on or within the left optical element 108. The combination of the GPU, the forward optical assembly 202, the left optical element 108, and the right optical element 110 provides the optical engine of the glasses 100. The glasses 100 use the optical engine to generate a superimposition of the user's view of the real-world scene, including displaying a user interface to the user of the glasses 100.

[0034] However, it should be understood that other display technologies or configurations can be utilized within the optical engine to display images to the user within the user's field of view. For example, instead of providing the projector 204 and the waveguide, an LCD, an LED, or other display panel or surface can be provided.

[0035] In use, the user of the glasses 100 will be presented with information, content, and various user interfaces on the near-eye display. As described in more detail herein, the user can then interact with the glasses 100 using the touchpad 126 and / or the buttons 128, voice input or touch input on an associated device (e.g., Figure 10 the mobile device 1050 shown in), and / or hand movements, positions, and orientations detected by the glasses 100.

[0036] Figure 3 is a graphical representation of a computing device 300 within which instructions 310 (e.g., software, programs, applications, applets, apps, other executable code) can be executed to cause the computing device 300 to perform any one or more of the methods discussed herein. The computing device 300 can be used as Figure 1a computer 120 of the glasses 100. For example, the instruction 310 may cause the computing device 300 to execute any one or more of the methods described herein. The instruction 310 transforms the general, unprogrammed computing device 300 into a particular computing device 300 programmed to perform the described and illustrated functions in the described manner. The computing device 300 may operate as a stand-alone device or may be coupled (e.g., networked) to other machines. In a networked deployment, the computing device 300 may operate in a server-client network environment with the capabilities of a server machine or a client machine, or as a peer machine in a peer-to-peer (or distributed) network environment. The computing device 300 may include, but is not limited to: server computers, client computers, personal computers (PCs), tablet computers, laptop computers, notebooks, set-top boxes (STBs), PDAs, entertainment media systems, cellular telephones, smart phones, mobile devices, head-mounted devices (e.g., smart watches), smart home devices (e.g., smart appliances), other smart devices, web devices, network routers, network switches, network bridges, or any machine capable of sequentially or otherwise executing the instruction 310 specifying the actions to be taken by the computing device 300. Additionally, although a single computing device 300 is shown, the term "machine" may also be understood to include a collection of machines that individually or jointly execute the instruction 310 to perform any one or more of the methods discussed herein.

[0037] The computing device 300 may include a processor 302, a memory 304, and an I / O component 306, which may be configured to communicate with each other via a bus 344. In some examples, the processor 302 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an ASIC, a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, a processor 308 and a processor 312 that execute the instruction 310. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") that may execute instructions simultaneously. Although Figure 3 multiple processors 302 are shown, the computing device 300 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.

[0038] Memory 304 includes main memory 314, static memory 316, and storage unit 318, all of which are accessible by processor 302 via bus 344. Main memory 304, static memory 316, and storage unit 318 store instructions 310 for implementing any one or more of the methods or functions described herein. During execution of instructions 310 by computing device 300, instructions 310 may also reside, in whole or in part, within main memory 314, within static memory 316, within machine-readable medium 320, within storage unit 318, within one or more of processors 302 (e.g., within a cache memory of the processor), or in any suitable combination thereof.

[0039] I / O component 306 may include various components for receiving input, providing output, generating output, transmitting information, exchanging information, capturing measurements, etc. The specific I / O components 306 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine will likely not include such a touch input device. It is to be understood that I / O component 306 may include Figure 3 many other components not shown. In various examples, I / O component 306 may include output component 328 and input component 332. Output component 328 may include visual components (e.g., a display such as a plasma display panel (PDP), a light emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., a vibration motor, a resistance mechanism), other signal generators, etc. Input component 332 may include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input components), pointing-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instrument), haptic input components (e.g., a physical button, a touch screen that provides the location and / or force or touch gesture of a touch, or other haptic input components), audio input components (e.g., a microphone), etc.

[0040] In some examples, the I / O component 306 may include: a biometric component 334, a motion component 336, an environmental component 338, a positioning component 340, and various other components. For example, the biometric component 334 includes components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biometric signals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying people (e.g., voice recognition, retina recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), etc. The motion component 336 may include an inertial measurement unit, an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotational sensor component (e.g., a gyroscope), etc. The environmental component 338 includes, for example, a lighting sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers for detecting ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones for detecting background noise), a proximity sensor component (e.g., an infrared sensor for detecting nearby objects), a gas sensor (e.g., a gas detection sensor for detecting the concentration of hazardous gases for safety or measuring pollutants in the atmosphere), or other components that can provide indications, measurement results, or signals associated with the surrounding physical environment. The positioning component 340 may include a position sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer for detecting the air pressure from which altitude can be obtained), an orientation sensor component (e.g., an inertial measurement unit (IMU)), etc.

[0041] A variety of techniques can be used to implement communication. The I / O component 306 also includes a communication component 342, and the communication component 342 is operable to couple the computing device 300 to the network 322 or the device 324 via a coupling 330 and a coupling 326, respectively. For example, the communication component 342 may include a network interface component or another suitable device that interfaces with the network 322. In additional examples, the communication component 342 may include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, components (e.g., low power consumption), components, and other communication components that provide communication via other modalities. The device 324 may be another machine or any of a variety of peripheral devices (e.g., a peripheral device coupled via USB).

[0042] In addition, communication component 342 can detect an identifier or include components operable to detect an identifier. For example, communication component 342 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., for detecting optical sensors described below: one-dimensional barcodes, such as Universal Product Code (UPC) barcodes; multi-dimensional barcodes, such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying an audio signal of a marker). Additionally, various information can be obtained via communication component 342, such as a location via Internet Protocol (IP) geolocation, a location via signal triangulation, a location via detecting an NFC beacon signal that can indicate a specific location, etc.

[0043] Various memories (e.g., memory 304, main memory 314, static memory 316, and / or the memory of processor 302) and / or storage unit 318 can store one or more sets of instructions and data structures (e.g., software) implemented or used by any one or more of the methods or functions described herein. These instructions (e.g., instruction 310), when executed by processor 302, cause various operations to implement the disclosed examples.

[0044] Instructions 310 can be transmitted or received over network 322 via a network interface device (e.g., the network interface component included in communication component 342), using a transmission medium and using any one of several well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, instructions 310 can be transmitted or received to device 324 via coupling 326 (e.g., a peer-to-peer coupling) using a transmission medium.

[0045] Figure 4is a diagram of a computing architecture 400 that includes one or more systems for generating augmented reality content according to one or more examples, the augmented reality content including translations of identifiers of objects located in an environment. The computing architecture 400 can include one or more user devices 402. The one or more user devices 402 can be operated by a user 404. The one or more user devices 402 can include a number of computing devices having processing resources and storage resources. For example, the one or more user devices 402 can include at least one of a head-mounted device, a wearable device, or a mobile computing device (such as a smart phone, a tablet computing device, a laptop computing device, a portable gaming device, etc.). In one or more illustrative examples, a wearable device can include a computing device worn on a part of a user's body, such as jewelry, a wrist-worn device, contact lenses, a hearing aid, or one or more combinations thereof. In various examples, the one or more user devices 402 can include multiple computing devices that operate in combination with each other. For illustration, the one or more user devices 402 can include a head-mounted device that operates in combination with at least one of a wearable device or a mobile computing device. In one or more additional examples, the one or more user devices 402 can include a wearable device that operates in combination with a mobile computing device. In one or more illustrative examples, the one or more user devices 402 include Figure 1 glasses 100 of

[0046] The processing resources and storage resources of the one or more user devices 402 can execute a number of applications, such as a user application 406. In one or more examples, the user application 406 can include a messaging function that enables the user 404 to send messages to other users of the user application 406 and receive messages from other users of the user application 406. In one or more additional examples, the user application 406 can include a social networking function that enables the user 404 to share content with other users of the user application 406 and / or access content created by other users of the user application 406. In one or more illustrative examples, the user application 406 includes at least one of an interaction client 1104 or an application 1106 described in more detail Figure 11 below.

[0047] The one or more user devices 402 can also execute a translated AR content item 408. The translated AR content item 408 can include software code that can be executed within the user application 406. For example, the translated AR content item 408 can include computer-readable instructions that can be executed to perform a number of functions within the user application 406. In various examples, the translated AR content item 408 can be executed after instantiating an instance of the user application 406. In Figure 4In the illustrative example, the translation AR content item 408 can be executed to display augmented reality content corresponding to the translation of the identifier of an object in the real-world scene. In one or more examples, the user 404 can be associated with multiple AR content items with respect to the account of the user application 406. Each AR content item can be executable to perform a different set of functions to generate augmented reality content within the user application 406. In at least some examples, a set of AR content items associated with the account of the user 404 corresponding to the user application 406 can be displayed as an array of user interface elements, and the user interface elements can be individually selected to execute the corresponding AR content item. In one or more illustrative examples, an array of user interface elements corresponding to a number of augmented reality content items is displayed in a carousel arrangement.

[0048] One or more user devices 402 may also include one or more imaging devices, such as the imaging device 410. The imaging device 410 can capture images of the real-world scene in the environment where one or more user devices 402 are located. In one or more examples, the imaging device 410 can capture video content of the real-world scene in the environment where one or more user devices 402 are located. The video content can include at least one of a series of images or an image stream captured over a period of time. In various examples, the imaging device 410 can capture video of the real-world scene in response to an input from the user 404. The images captured by the imaging device 410 can be within the field of view of the imaging device 410. The field of view can correspond to a portion of the real-world scene of the environment that can be imaged by the imaging device 410 at a given time and can be based on the focal length of the lens of the imaging device 410 and the size of the sensor of the imaging device 410. In at least some examples, the imaging device 410 can capture a real-time / current view of the real-world scene within the field of view. Although not shown in Figure 4 the illustrative example, one or more user devices 402 may also include a number of audio capture devices. By way of illustration, one or more user devices 402 can include a number of microphones to capture audio content generated in the environment where one or more user devices 402 are located. In one or more illustrative examples, one or more user devices 402 include one or more microphones to capture audio content in combination with the video content captured by the imaging device 410. One or more user devices 402 may also include one or more speakers to play audio content corresponding to the augmented reality content displayed within the translation AR content item 408.

[0049] The computing architecture 400 also includes an Augmented Reality (AR) translation system 412. The AR translation system 412 can analyze data generated by one or more camera devices 410 of one or more user devices 402 to identify an object located in a real-world scene. The AR translation system 412 can determine a translation of an identifier of the object and cause the translation to be displayed in a user interface that includes a view of the real-world scene in which the object is located. In at least some examples, the view of the real-world scene can include a live view of the real-world scene. In one or more illustrative examples, the identifier of the object includes at least one of one or more words, one or more symbols, one or more characters, or one or more phrases for identifying the object in at least one language. In one or more additional illustrative examples, the identifier of the object includes at least one of a noun or an adjective corresponding to the object. In various examples, the AR translation system 412 can generate augmented reality content that indicates the identifier of the object in a first language and the translation of the identifier of the object in a second language.

[0050] The AR translation system 412 can include an object detection system 416 that analyzes camera device data 414 captured by the camera device 410 to identify one or more objects located in a real-world scene. In various examples, the object detection system 416 can implement one or more machine learning algorithms to identify the objects indicated by the camera device data 414. In one or more examples, the object detection system 416 can implement one or more artificial neural networks to analyze the camera device data 414 to identify objects within the real-world scene. By way of illustration, the object detection system 416 implements one or more convolutional neural networks to identify the objects indicated by the camera device data 414. Additionally, the object detection system 416 can implement one or more residual neural networks to determine that the objects are indicated by the camera device data 414. In one or more additional examples, the object detection system 416 can implement at least one of a k-nearest neighbor artificial neural network, a support vector machine algorithm, or a random forest algorithm to identify the objects indicated by the camera device data 414. In at least some examples, one or more machine learning algorithms implemented by the object detection system 416 can be trained according to a training data set that includes at least one of image content or video content of a number of different objects. Additionally, the object detection system 416 can implement one or more classification machine learning techniques to analyze the camera device data 414 to identify one or more objects in the real-world scene. By way of illustration, the object detection system 416 can implement one or more support vector machines for the camera device data 414 to identify one or more objects in the real-world scene.

[0051] In one or more illustrative examples, object detection system 416 analyzes camera device data 414 to determine a number of at least one of a contour, an edge, a color, a chromaticity, a texture, or a shape, which may be used to determine one or more candidate regions that may include one or more objects of interest. Object detection system 416 may also implement a convolutional neural network for extracting features from one or more candidate regions. Additionally, object detection system 416 may implement one or more support vector machines to classify one or more objects included in one or more candidate regions based on features extracted by the convolutional neural network from the one or more candidate regions. In various examples, one or more machine learning techniques implemented by object detection system 416 may be trained using previously captured images that include one or more of the objects of interest and are labeled as including one or more objects of interest.

[0052] In one or more additional illustrative examples, object detection system 416 implements one or more gaze tracking techniques to determine a position of a field of view of a gaze of user 404. In at least some examples, object detection system 416 may analyze camera device data 414 to determine the position of the gaze of user 404. Additionally, object detection system 416 may analyze data obtained from one or more inertial measurement unit (IMU) sensors of one or more user devices 402 to determine the position of the gaze of user 404. Further, object detection system 416 may analyze additional camera device data obtained from one or more camera devices external to one or more user devices 402 to determine the position of the gaze of user 404. In one or more illustrative examples, object detection system 416 determines at least one of a field of view of user 404 or a center of a field of view of user 404. In response to determining at least one of a field of view of user 404 or a center of a field of view of user 404, object detection system 416 may identify one or more objects included in the field of view of user 404 and / or one or more objects included in the center of the field of view of user 404. In one or more examples, when the gaze of user 404 changes, object detection system 416 may determine a new position of a field of view of the gaze of user 404. Object detection system 416 may then determine one or more additional objects included in at least one of the new fields of view of user 404, or one or more additional objects included in the center of the new fields of view of user 404.

[0053] The AR translation system 412 may also include a location recognition system 418. The location recognition system 418 may determine the location of one or more user devices 402. In one or more examples, the location recognition system 418 may determine the location of one or more user devices 402 based on Global Positioning System (GPS) data obtained from one or more user devices 402. In one or more additional examples, the location recognition system 418 may determine the location of one or more user devices 402 based on the Internet Protocol address of one or more user devices 402. In one or more further examples, the location recognition system 418 may determine the location of one or more user devices 402 based on triangulation data obtained from one or more wide area wireless communication networks corresponding to one or more user devices 402.

[0054] Additionally, the location recognition system 418 may determine the location of one or more user devices 402 by analyzing camera device data 414. In various examples, the location recognition system 418 may determine the location of one or more user devices 402 based on the arrangement of objects in the real-world scene. For example, the object detection system 416 may determine a number of objects in the real-world scene and the spatial relationships between the number of objects. In at least some examples, the object detection system 416 may determine the real-world coordinates of each of the number of objects. In one or more additional examples, the object detection system 416 may determine the distances between the objects included in the number of objects in the real-world scene. In one or more further examples, the object detection system 416 may determine a direction indicator indicating the direction or heading between the objects included in the number of objects in the real-world scene.

[0055] In one or more illustrative examples, the location recognition system 418 classifies the arrangement of objects in the real-world scene as part of the corresponding location. In various examples, the corresponding location of the arrangement of objects may correspond to an identifier. In one or more examples, the identifier of the location of the arrangement of objects may be provided by the user 404. In one or more additional examples, the identifier of the location of the arrangement of objects may be provided by the location recognition system 418. In at least some examples, the location recognition system 418 may store the arrangement of objects in the real-world scene in combination with the location based on an input from the user 404 indicating that the arrangement of objects is to be stored in combination with the location and that the translation of the identifier of at least one object in the arrangement of objects has been requested or previously determined.

[0056] In addition, the AR translation system 412 may include a translation object data management system 420. The translation object data management system 420 may cause information related to objects included in a real-world scene to be stored in one or more databases 422. The one or more databases 422 may be at least one database physically or logically connected to the AR translation system 412. The one or more databases 422 may be at least one database located locally or remotely with respect to one or more computing devices implementing the AR translation system 412. In at least some examples, the one or more databases 422 and the AR translation system 412 may be implemented as part of a cloud-based computing architecture.

[0057] The one or more databases 422 may store an object translation list 424. The object translation list 424 may indicate objects corresponding to a user of the user application 406 that have identifiers that have been translated from a first language to a second language. For each object, the object translation list 424 may indicate at least one of the identifier of the object in the first language and the identifier of the object in the second language. In one or more additional examples, the object translation list 424 may indicate the default language of the user of the user application 406 and one or more additional languages used to generate the translation of the identifier of the object. The object translation list 424 may also indicate the location of the object having an identifier that has been translated from a first language to a second language.

[0058] In one or more illustrative examples, the translation object data management system 420 causes information related to objects having identifiers translated from at least one language to another language to be stored in and retrieved from one or more object translation lists 424 of a user (e.g., user 404) of the user application 406. For example, as user 404 moves through the environment, the object detection system 416 may analyze the camera device data 414 to identify one or more objects in the environment. The object detection system 416 may operate in conjunction with the translation object data management system 420 to determine whether an object is included in the object translation list 424 of user 404. In various examples, the object detection system 416 may generate at least one of a classification or label identifying an object detected in the real-world scene and provide the classification and / or label to the translation object data management system 420. The translation object data management system 420 may then use the classification and / or label of the object and query the object translation list 424 of user 404 for at least one of the classification or label of the object.

[0059] In the case where an object does not exist in the object translation list 424 of user 404, the AR translation system 412 provides one or more options to user 404 to add the object to the object translation list 424 of user 404. For example, in response to determining that the object identified by the object detection system 416 does not exist in the object translation list 424 of user 404, the AR translation system 412 can cause one or more user interface elements to be displayed that can be selected to add the object to the object translation list 424 of user 404. In one or more illustrative examples, the one or more user interface elements are displayed as augmented reality content in a user interface that displays a view of a real-world scene including the object. In various examples, in response to the selection of a user interface element, one or more user devices 402 can provide user input 426 to the AR translation system 412 to add the object identified by the object detection system 416 to the object translation list 424 of user 404. In at least some examples, the one or more options for adding an object to the object translation list 424 of a user can correspond to options for translating the identifier of the object from a first language to a second language. By way of illustration, in response to the selection of a user interface element to obtain a translation of the identifier of an object, the AR translation system 412 causes the object to be added to the object translation list 424 of user 404.

[0060] In the case where an object is stored in the object translation list 424 of user 404, the translated object data management system 420 can operate in combination with at least one of the object detection system 416 or the text content translation system 428 to generate a translation of the identifier of the object from a first language to a second language. In one or more examples, in response to the object detection system 416 identifying the presence of an object in a real-world scene captured by the imaging device 410, the translated object data management system 420 can obtain a translation of the identifier of the object stored in one or more databases 422. In various examples, the translation of the identifier of the object can be stored in the object translation list 424 of user 404 or in combination with the object translation list 424 of user 404.

[0061] In one or more additional examples, in response to the object detection system 416 identifying the presence of an object in the real-world scene captured by the imaging device 410, the text content translation system 428 may generate a translation of the identifier of the object. In various examples, the text content translation system 428 may determine a first identifier of the object in a first language. In at least some examples, the first identifier of the object may be a default language. The default language may be selected by the user 404. Additionally, the default language may be determined based on the location of the user 404. Further, the default language may be determined by an entity that performs at least one of maintaining, controlling, or managing the AR translation system 412. The text content translation system 428 may translate the first identifier into one or more second identifiers in one or more second languages. In one or more examples, at least one of the one or more second languages may be selected by the user 404.

[0062] The text content translation system 428 may implement one or more computational algorithms to generate a translation of the identifier of an object located in a real-world scene. In various examples, the text content translation system 428 may implement one or more machine learning techniques to generate a translation of the identifier of the object. For example, the text content translation system 428 may implement one or more artificial neural networks to generate a translation of the identifier of the object. In one or more illustrative examples, the text content translation system 428 implements one or more recurrent neural networks to generate a translation of the identifier of the object. In one or more additional illustrative examples, the text content translation system 428 implements one or more convolutional neural networks to generate a translation of the identifier of the object.

[0063] The text content translation system 428 may also obtain a translation to generate a translation of the identifier of the object by making calls to one or more application programming interfaces (APIs) 430 from one or more third-party systems 432. In one or more examples, the text content translation system 428 may determine the identifier of the object detected by the object detection system 416 in a first language and generate one or more calls to one or more APIs 430 to request a translation of the identifier in one or more second languages. The text content translation system 428 may obtain a translation of the identifier of the object in one or more second languages from one or more third-party systems 432. In at least some examples, the one or more third-party systems may include one or more translation services that are controlled, maintained, or managed by an entity different from the entity that performs at least one of controlling, maintaining, or managing the AR translation system 412.

[0064] In one or more examples, the text content translation system 428 can generate text content that includes translations of the identifiers of the objects detected by the object detection system 416. In various examples, the text content translation system 428 can generate augmented reality content to be displayed in a user interface that includes a view of a real-world scene that includes one or more objects corresponding to the text after translation. In at least some examples, the augmented reality content generated by the text content translation system 428 can be presented by a translation AR content item 408 executed within the user application 406. In one or more illustrative examples, the text content generated by the text content translation system 428 includes the identifier of the object in a first language and one or more additional translations of the identifier in one or more second languages. The text content translation system 428 can cause the identifier of the object in the first language and the translation of the identifier in one or more second languages to be displayed near the object. In at least some examples, the user interface element corresponding to the translation of the identifier of the object can include a virtual object that can be manipulated and / or controlled by the user 404.

[0065] The AR translation system 412 can also include an audio content system 434. The audio content system 434 can generate audio content corresponding to the translations of the identifiers of the objects detected by the object detection system 416. In one or more examples, the audio content system 434 can generate audio content corresponding to one or more pronunciations of the identifier of the object in a second language, where the identifier of the object in the second language corresponds to the translation of the identifier of the object in the first language. In one or more additional examples, the audio content system 434 can generate audio content corresponding to a first pronunciation of a first identifier of the object in the first language and a second pronunciation of a second identifier of the object in the second language. In various examples, the audio content generated by the audio content system 434 for an object can be played in combination with the display of the text content generated by the text content translation system 428 for the object.

[0066] In one or more examples, the audio content system 434 can generate an audio file of audio content having pronunciations with identifiers corresponding to objects, based on the text content generated by the text content translation system 428. The audio content system 434 can implement one or more computing techniques to convert the text generated by the text content translation system 428 into audio content. For example, the audio content system 434 can implement one or more hidden Markov models to generate audio content based on the translated text content generated by the text content translation system 428. In one or more additional examples, the audio content system 434 can implement at least one of one or more convolutional neural networks, one or more recurrent neural networks, or at least one long short-term memory, to generate audio content based on the translated text content generated by the text content translation system 428. In one or more further examples, the audio content system 434 can implement at least one of one or more encoders, one or more decoders, or one or more vocoders, to generate audio content based on the text content generated by the text content translation system 428. In yet some additional examples, the audio content system 434 can implement one or more transformer-based machine learning techniques to generate audio content based on the text content generated by the text content translation system 428.

[0067] The AR translation system 412 can generate translation output data 436 that is accessible to one or more user devices 402. The translation output data 436 can include at least one of the text content generated by the text content translation system 428 or the audio content generated by the audio content system 434. For example, the translation output data 436 can include one or more translated text contents corresponding to identifiers of objects located in a view of a real-world scene captured by the imaging device 410, and audio content corresponding to one or more translated audible pronunciations. The translation output data 436 can include augmented reality content that includes the text content generated by the text content translation system 428 and the audio content generated by the audio content system 434. In various examples, the AR translation system 412 can generate at least one of image content, video content, or animation content, based on at least one of the text content generated by the text content translation system 428 or the audio content generated by the audio content system 434. By way of illustration, the AR translation system 412 can generate one or more animations that are displayed in relation to one or more objects included in a real-world scene, where the one or more animations include at least one of the text content generated by the text content translation system 428 or the audio content generated by the audio content system 434.

[0068] The augmented reality content included in the translation output data 436 can be accessed via one or more output devices of one or more user devices 402. By way of illustration, the audio content included in the translation output data 436 can be accessed via the speakers of one or more user devices 402, and at least one of the text content, image content, video content, or animation content included in the translation output data 436 can be accessed via one or more display devices of one or more user devices 402. In one or more illustrative examples, at least one of the text content, image content, video content, or animation content included in the translation output data 436 is displayed in one or more user interfaces generated by the translation AR content item 408. In at least some examples, one or more user interfaces generated by the translation AR content item 408 based on the translation output data 436 can include a view of the real-world scene captured by the camera device 410, such as a live view, and the view of the real-world scene includes one or more objects having identifiers translated by the AR translation system 412.

[0069] Although several operations are described as being performed by the AR translation system 412, at least a portion of the operations described as being performed by the AR translation system 412 can be performed by one or more user devices 402. For example, one or more operations described as being performed by the object detection system 416, one or more operations described as being performed by the location recognition system 418, one or more operations described as being performed by the translated object data management system 420, one or more operations described as being performed by the text content translation system 428, one or more operations described as being performed by the audio content system 434, or one or more combinations thereof can be performed by one or more user devices 402 associated with the user application 406.

[0070] Figure 5 FIG. 500 is a diagram of a computing architecture of one or more systems including an object translation list for tracking objects stored by a user of a translation augmented reality content item according to one or more examples. The computing architecture 500 includes one or more user devices 402 operated by a user 404. The computing architecture 500 also includes an AR translation system 412 and one or more databases 422. One or more user devices 402 can include one or more camera devices, such as the camera device 410. The camera device 410 can have a field of view 502. The field of view 502 can be based on the positioning of one or more user devices 402. In one or more examples, one or more user devices 402 can include a head-mounted device, and the field of view 502 can be based on the positioning of the head of the user 404. In various examples, the field of view �02 can correspond to the gaze of the user 404.

[0071] The field of view 502 may correspond to the imaging device data 414 generated by the imaging device 410, and the imaging device data 414 is provided to the AR translation system 412. In Figure 5 the illustrative example, the imaging device 410 captures at least one of image content or video content of the real-world scene 504 included in the field of view 502. The real-world scene 504 may include a first object 506, a second object 508, and a third object 510. The objects 506, 508, 510 may be positioned in an arrangement. In at least some examples, the arrangement of the first object 506, the second object 508, and the third object 510 relative to each other may indicate the location of the real-world scene.

[0072] In Figure 5 the illustrative example, one or more databases 422 store the object translation list 512 of the user 404. The object translation list 512 may correspond to the user identifier 514 corresponding to the user 404. In one or more examples, the user identifier 514 may correspond to the identifier of the user 404 within the user application 406, such as a username, alias, login identifier, one or more combinations thereof, etc. In these scenarios, the user identifier 514 is selected by the user 404. The user identifier 514 may also correspond to at least one of one or more symbols or one or more characters assigned to the user 404 by an entity that performs at least one of maintaining, managing, or controlling the AR translation system 412 to identify the user 404 within the AR translation system 412 and within other systems operating in combination with the user application 406. In at least some examples, the user identifier 514 is used to store the data of the user 404 and retrieve the data of the user 404 relative to one or more databases 422.

[0073] The user identifier 514 may correspond to the default language 516. In one or more examples, the default language 516 may be selected by the user 404. In one or more additional examples, the default language 516 may be determined based on the location of the user 404, such as the language commonly used for communication in the location of the user 404. In one or more further examples, the default language 516 may be selected by an entity that performs at least one of controlling, managing, or maintaining the user application 406. In at least some examples, the default language 516 may be modified. For example, the default language 516 may be modified by the user 404 and / or based on a change in the location of the user 404. In various examples, the default language 516 corresponds to the language that the user 404 typically uses to communicate with other users of the user application 406 and / or the language in which the user 404 is fluent in at least one of oral communication or written communication.

[0074] The user identifier 518 can be associated with a first location 518 corresponding to first location data 520. In one or more examples, the first location data 520 can indicate the first location 518 based on GPS location data. The first location data 520 can also indicate the first location 518 based on a first arrangement of objects included in the first location 518. For example, the first location data 520 can indicate the real-world coordinates of the arrangement of objects corresponding to the first location 518. In one or more illustrative examples, the first location data 520 indicates the first real-world coordinates of a first object 506, the second real-world coordinates of a second object 508, and the third real-world coordinates of a third object 510 corresponding to the location of the user 404.

[0075] One or more components of the AR translation system 412, such as the location recognition system 418 described with respect to Figure 4 can determine the location of the user 404 based on the first location data 520. For example, the location recognition system 418 can analyze current location data obtained from the user device 402 relative to the first location data 520 to determine whether the user 404 is located in the first location 518. In one or more examples, the AR translation system 412 can obtain current GPS coordinates from one or more user devices 402 and analyze the current GPS coordinates relative to the GPS coordinates included in the first location data 520. In the case where the current GPS coordinates correspond to the GPS coordinates included in the first location data 520, the AR translation system 412 can determine that the user 404 is located in the first location 518. In at least some examples, the AR translation system 412 can determine a quantitative measurement indicative of the amount of similarity between the current GPS coordinates obtained from one or more user devices 402 and the GPS coordinates included in the first location data 520. In various examples, the AR translation system 412 can determine that the user 404 is located in the first location 518 in response to the quantitative measurement of similarity being at least a threshold. By way of illustration, the AR translation system 412 determines that the user 404 is located in the first location 518 in response to determining that the current GPS coordinates obtained from one or more user devices 402 are within at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% of the GPS coordinates included in the first location data 520.

[0076] Additionally, the AR translation system 412 can analyze the camera device data 414 to determine the arrangement of objects in the current location of the user 404 and analyze the arrangement of the objects relative to the first arrangement of the objects included in the first location data 520. In various examples, the AR translation system 412 can analyze several objects in the current location of the user 404, the positioning of the objects in the current location of the user 404 relative to each other, or at least one of the distances between several first objects included in the first location 518 of the user 404, the positioning of the first objects included in the first location data 520 relative to each other, and / or the distances between the first objects included in the first location data 520. In a scenario where the first object 506, the second object 508, and the third object 510 are included in the first location data 520, the AR translation system 412 analyzes the arrangement of the objects included in the camera device data 414 to determine a quantitative measurement of the amount indicating the similarity between the arrangement of the objects included in the camera device data 514 and the arrangement of the first object 506, the second object 508, and the third object 510 relative to each other.

[0077] In one or more illustrative examples, the AR translation system 412 analyzes the real-world coordinates of the objects indicated by the camera device data 414 relative to the first real-world coordinates of the first object 506, the second real-world coordinates of the second object 508, and the third real-world coordinates of the third object 510. For example, in a case where the current location of the user 404 includes three objects, the AR translation system 412 can analyze the real-world coordinates of each object indicated by the camera device data 414 relative to the real-world coordinates of the first object 506, the real-world coordinates of the second object 508, and the real-world coordinates of the third object 510 to determine whether the corresponding sets of real-world coordinates correspond to each other. In a case where the camera device data 414 includes an object having real-world coordinates with at least a threshold amount of similarity to the first real-world coordinates of the first object 506, another object having real-world coordinates with at least a threshold amount of similarity to the second real-world coordinates of the second object 508, and an additional object having real-world coordinates with at least a threshold amount of similarity to the third real-world coordinates of the third object 510, the AR translation system 412 determines that the current location of the user 404 corresponds to the first location 518. In at least some examples, when the real-world coordinates of each object included in the camera device data 414 are within at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% of the real-world coordinates of the corresponding objects included in the first location data 520, the real-world coordinates of the objects included in the camera device data 414 can have a threshold amount of similarity to the objects included in the first location data 520.

[0078] In one or more examples, the AR translation system 412 can analyze a combination of GPS coordinates and camera device data 414 to determine the location of the user 404. For example, the AR translation system 412 can determine a first quantitative measurement indicative of the amount of similarity between the current GPS coordinates indicating one or more user devices 402 and the GPS coordinates included in the first location data 520, and a second quantitative measurement indicative of the amount of similarity between the arrangement of objects indicated by the camera device data 414 and an additional arrangement of objects included in the first location data 520. In the case where the first quantitative measurement corresponds to a first threshold and the second quantitative measurement corresponds to a second threshold, the AR translation system 412 determines that the user 404 is located in the first location 518.

[0079] The first position 518 may also correspond to one or more first translated objects 522. The one or more first translated objects 522 may include one or more objects having identifiers that represent translations made by the user 404. In at least some examples, the one or more first translated objects 522 may be associated with the first position 518 in the one or more databases 422 according to a first tag corresponding to the first position 518. In one or more examples, the AR translation system 412 may obtain an object tagging input 524 from one or more user devices 402, the object tagging input 524 indicating a request by the user 404 to translate the identifier of an object from a default language 516 into one or more first translation languages 526. By way of illustration, when the second object 508 moves into the field of view 502 of the imaging device 410, the AR translation system 412 determines that the user 404 is located in the first position 518 and that the second object 508 is not present in the one or more first translated objects 522. The AR translation system 412 may then cause one or more user interface elements to be displayed by the user application 406, the one or more user interface elements being selectable by the user 404 to generate a translation of the identifier of the second object 508. In response to the selection of a user interface element to generate a translation of the identifier of the second object 508, the object tagging input 524 is provided to the AR translation system 412. In response to the object tagging input 524, the AR translation system 412 causes the identifier of the second object 508 to be translated from the default language 516 into one or more additional languages. The AR translation system 412 may then add the second object 508 to the one or more first translated objects 522. In the case where one or more of the additional languages are included in the one or more first translation languages, the AR translation system 412 may indicate at least a portion of the one or more first translation languages 526 corresponding to the second object 508. In the case where one or more of the additional languages are not present in the first translation language 526, the AR translation system 412 may add the one or more additional languages to the first translation language 526 and indicate that the second object 508 corresponds to the one or more additional languages added to the first translation language 526.

[0080] In various examples, one or more translations for one or more first translated objects 522 may be stored in one or more databases 422. In these cases, the AR translation system 412 retrieves a translation of the identifier of the first translated object 522 from the one or more databases 422 in response to detecting that the user 404 is located in the first location 518 and the one or more first translated objects 522 are in the field of view 502 of the imaging device 410. In one or more additional examples, one or more translations for the first translated object 522 may be retrieved from one or more translation services, such as one or more third-party translation services. In these scenarios, the AR translation system 412 may generate one or more requests, such as one or more API calls, to obtain a translation of the first translated object 522 from the default language 516 to one or more first translation languages 526 in response to the AR translation system 412 determining that the user 404 is located in the first location 518 and the first translated object 522 is in the field of view 502 of the imaging device 410.

[0081] In Figure 5 an illustrative example of, the user identifier 514 corresponds to a second location 528 that is different from the first location 518. The second location 528 may correspond to second location data 530. The second location data 530 may include at least one of GPS coordinates corresponding to the second location 528 or an arrangement of objects. Additionally, the second location 528 may correspond to a second translated object 532 that corresponds to one or more objects having an identifier for which the user 404 has requested a translation. Further, the second translated object 532 may correspond to one or more second translation languages 534. The one or more second translation languages 534 may be different from the one or more first translated objects 522. In one or more additional examples, the one or more second translation languages 534 may be the same as the one or more first translation languages 526.

[0082] Although Figure 5 the illustrative example of indicates that the user identifier 514 corresponds to the first location 518 and the second location 528, in one or more additional examples, the user identifier 514 may correspond to fewer locations. In one or more additional examples, the user identifier 514 may correspond to a greater number of locations than the first location 518 or the second location 528.

[0083] Figure 6FIG. 600 is a diagram of a computing architecture 600 for generating augmented reality content according to one or more examples, the augmented reality content including text content and audio content of identifiers of objects in one or more languages. The computing architecture 600 may include one or more user devices 402 operated by a user 404. The one or more user devices 402 may include a camera device 410 that generates camera device data 414 accessible by an AR translation system 412. The camera device data 414 may include at least one of image content or video content of a real-world scene. In at least some examples, the camera device data 414 may include a data stream captured by the camera device 410. In various examples, the camera device data 414 may correspond to a real-time view of a real-world scene. The one or more user devices 402 may also execute an instance of a user application 406, where a translated AR content item 408 is executed within the instance of the user application 406.

[0084] An object detection system 416 may analyze the camera device data 414 and determine object data 602. The object data 602 may indicate one or more objects located within the field of view of the camera device 410. In one or more examples, the object detection system 416 may determine a central region of the field of view of the camera device 410, and at least a portion of the object data 602 may indicate one or more objects included in the center of the field of view of the camera device 410. The object detection system 416 may also analyze the camera device data 414 to generate object layout data 604. The object layout data 604 may indicate the arrangement of the objects included in the camera device data 414. For example, the object layout data 604 may indicate the spatial arrangement of the objects included in the real-world scene corresponding to the camera device data 414. In one or more examples, the object layout data 604 may indicate the real-world coordinates of one or more objects located in the real-world scene. In one or more additional examples, the object layout data 604 may indicate the distances between the objects included in the real-world scene, and angles and / or vectors indicating the spatial arrangement of the objects in the real-world scene relative to each other.

[0085] In one or more examples, the location recognition system 418 can analyze at least one of object data 602, object layout data 604, or GPS coordinates obtained from one or more user devices 402 to determine a location identifier 606. The location identifier 606 can indicate the location of the user 404. The location recognition system 418 can operate in combination with the translated object data management system 420 to access location data 608. In various examples, the translated object data management system 420 can retrieve the location data 608 from the translated object list 424 of the user 404 based on the user identifier 610. In at least some examples, the translated object data management system 420 can use the user identifier 610 to retrieve at least one of GPS coordinates or object arrangement data from the translated object list 424 of the user 404.

[0086] In at least some examples, the location recognition system 418 can analyze the location data 608 accessed via one or more databases 422 relative to the current location data accessed via one or more user devices 402 to determine the location of the user 404. In various examples, the location recognition system 418 can analyze the current GPS coordinates accessed via one or more user devices 402 relative to a set or more sets of GPS coordinates corresponding to one or more locations and included in the location data 608 to determine the location of the user 404. In one or more additional examples, the location recognition system 418 can analyze the arrangement of objects in the real-world scenario indicated by the object layout data 604 relative to the one or more arrangements of objects included in the location data 608. In response to determining at least a threshold level of similarity between at least a portion of the location data 608 and at least one of the object layout data 604 or additional location data accessed via one or more user devices 402, the location recognition system 418 determines the location identifier 606 corresponding to the location of the user 404.

[0087] The translated object data management system 420 can use the location identifier 606 to generate translated object data 612. The translated object data 612 can indicate one or more translated objects 614 corresponding to the location of the user 404. The one or more translated objects 614 can correspond to one or more objects in the translation object list 424 of the user 404 for which an object identifier translation has been requested. In one or more examples, the translated object data 612 can include one or more identifiers of the one or more translated objects 614. The one or more identifiers of the one or more translated objects 614 can correspond to at least one of one or more symbols or one or more characters that identify the one or more translated objects 614. In one or more illustrative examples, the one or more identifiers of the one or more translated objects 614 uniquely identify each of the one or more translated objects 614. In one or more additional illustrative examples, the one or more identifiers of the one or more translated objects 614 correspond to identifiers in a given language that correspond to each of the one or more translated objects 614. For example, the translated object data 612 can indicate the respective identifiers of each of the one or more translated objects 614 in a default language. In various examples, the default language can be selected by the user 404.

[0088] The translated object data 612 can be provided to the text content translation system 428. The text content translation system 428 can determine the translation of the identifiers of the one or more translated objects 614. For example, the text content translation system 428 can generate text content corresponding to the identifiers of the translated objects 614 in an additional language different from the default language. The text content translation system 428 can implement one or more computational algorithms to generate the translation of the identifiers of the one or more translated objects 614. In one or more additional examples, the text content translation system 428 can obtain the translation of the identifiers of the one or more translated objects 614 from one or more translation services. In various examples, one or more of the translation services can be controlled, managed, or maintained by one or more entities different from one or more of the entities that control, manage, or maintain the AR translation system 412.

[0089] The text content translation system 428 can generate text translation data 616. The text translation data 616 can indicate at least one of characters or symbols of one or more identifiers of one or more translated objects 614 in an additional language different from the default language. In one or more examples, the additional language can be selected by the user 404 or at least one of entities that control, maintain, or manage the AR translation system 412. In at least some examples, the text translation data 616 can include at least one of a noun or an adjective corresponding to an identifier of one or more translated objects 614. In one or more illustrative examples, the text translation data 616 includes at least one of one or more first characters or first symbols of a first identifier of a translated object 614 in the default language, and at least one of one or more second characters or second symbols of a second identifier of the translated object 614 in the additional language.

[0090] The text content translation system 428 can provide the text translation data 616 to the audio content system 434. The audio content system 434 can generate audio translation data 618 based on the text translation data 616. For example, the audio content system 434 can determine the pronunciation of one or more identifiers of one or more translated objects 614 included in the text translation data 616. By way of illustration, the audio content system 434 uses the text content of the identifiers included in the text translation data 616 to generate the audio translation data 618 including the audible pronunciation of the identifiers. In one or more illustrative examples, the audio content system 434 generates the audio translation data 618 including a first audible pronunciation of a first identifier of a translated object 614 in a first language based on the first text content of the first identifier included in the text translation data 616. Additionally, the audio content system 434 can generate the audio translation data 618 including a second audible pronunciation of a second identifier of a translated object 614 in a second language based on the second text content of the second identifier included in the text translation data 616. In one or more illustrative examples, the audio translation data 618 includes the first audible pronunciation and the second audible pronunciation. In at least some examples, the first language corresponding to the first pronunciation can be the default language, and the second language corresponding to the second pronunciation can be the additional language for which the user 404 has requested a translation of the identifier of the object.

[0091] In one or more examples, the text content translation system 428 may provide text translation data 616 to one or more user devices 402, and the audio content system 434 may provide audio translation data 618 to one or more user devices 402. The text translation data 616 and the audio translation data 618 may be accessible via the translated AR content item 408. In one or more examples, the text translation data 616 may include augmented reality content to be displayed in combination with the translated AR content item 408, and the audio translation data 618 may correspond to the augmented reality content to be played in combination with the display of the text translation data 616. In one or more illustrative examples, when a translation of an identifier of an object is displayed by one or more display devices of one or more user devices 402 in combination with the translated AR content item 408, the pronunciation of the translation is played at least once via one or more speakers of one or more user devices 402. In one or more additional illustrative examples, the translation of an identifier of an object is displayed by one or more display devices of one or more user devices 402 in combination with the translated AR content item 408, and the pronunciation of the translation may be played via one or more speakers of one or more user devices 402 in response to additional user input. In one or more further illustrative examples, when text content including a translation of an identifier of an object is not displayed, the translation of the identifier of the object is played via one or more speakers in response to user input. In this way, in one or more scenarios, the text content corresponding to the translation of the identifier and the audio content corresponding to the pronunciation of the translation may be independently accessed and consumed by the user 404.

[0092] In at least some examples, at least one of the AR translation system 412 or the translated AR content item 408 can track the gaze of the user 404 to determine one or more objects for which to provide at least one of the text translation data 616 or the audio translation data 618. In various examples, the gaze of the user 404 can correspond to the center of the field of view of one or more camera devices 410 of one or more user devices 402. In one or more examples, at least one of the text translation data 616 or the audio translation data 618 can be responsive to determining that an object in the center of the field of view of the camera device 410 is accessed by the translated AR content item 408. For example, in one or more examples, one or more objects can be within the field of view of the camera device 410. The AR translation system 412 or at least one of the one or more user devices 402 can analyze the camera device data 414 to determine the center of the field of view of the camera device 410 and at least one object within the center of the field of view of the camera device 410. In one or more illustrative examples, at least one of the text translation data 616 or the audio translation data 618 is generated by the AR translation system 412 for an object included in the center of the field of view of the camera device 410. In one or more additional illustrative examples, in a case where multiple objects are included in the field of view of the camera device 410, the AR translation system 412 determines at least one of the text translation data 616 or the audio translation data 618 for each of the multiple objects captured in the field of view of the camera device 410. In these scenarios, in one or more examples, the translated AR content item 408 causes at least one of the text translation data 616 or the audio translation data 618 to be displayed and / or played for the object among the multiple objects in the center of the field of view of the camera device 410. In these cases, in one or more additional examples, the translated AR content item 408 can cause at least one of the text translation data 616 or the audio translation data 618 to be displayed and / or played for each of the multiple objects in the field of view of the camera device 410.

[0093] Figure 7 and Figure 8 FIGS. 700 and 800 are flowcharts illustrating example processes for generating augmented reality content related to the translation of identifiers of objects captured in the field of view of one or more camera devices, according to one or more examples. The implementation of processes 700 and 800 can be embodied in computer-readable instructions that are executed by one or more processors such that the operations of the processes can be performed, in part or in whole, by functional components of at least one of one or more client devices or one or more server systems. Thus, in some cases, the processes described below are presented by way of example with reference to their examples. However, in other implementations, with respect to Figure 7 and Figure 8At least some of the operations in the described example processing may be deployed on various other hardware configurations. Thus, with respect to Figure 7 and Figure 8 the described example processing is not intended to be limited to being performed by one or more server systems or one or more client devices described herein, but may be implemented in whole or in part by one or more additional components. Although the described flowcharts may show operations as sequential processes, many of the operations may be performed in parallel or simultaneously. Additionally, the order of the operations may be rearranged. When the operations of the processing are complete, the processing terminates. The processing may correspond to a method, program, algorithm, etc. The operations of the method may be performed in whole or in part, may be performed in combination with some or all of the operations in other methods, and may be performed by any number of different systems (e.g., the systems described herein) or any part thereof (e.g., a processor included in any one of the systems).

[0094] Figure 7 is a flowchart of a processing 700 for adding an object to a user's object translation list of a user application and generating augmented reality content, the augmented reality content including a translation corresponding to an identifier of the object. At 702, the processing 700 may include obtaining camera device data including at least one of image content or video content captured in a field of view of a camera device of a user device. In one or more examples, the user device may include a head-mounted device. In one or more illustrative examples, the head-mounted device includes glasses. In one or more additional examples, the user device may include a wearable device. The wearable device may include an activity tracker and / or a watch and may also include other wearable devices such as contact lenses, jewelry, or other items worn on at least one of the user's ears, eyes, or other parts of the body. In at least some examples, the user device may include multiple camera devices. In these scenarios, the field of view includes a combined field of view of the multiple camera devices.

[0095] Additionally, at 704, process 700 may include analyzing camera device data to identify one or more objects indicated by the camera device data. In one or more examples, one or more object detection machine learning techniques may be used to analyze the camera device data to identify one or more objects indicated by the camera device data. In various examples, the camera device data may also be analyzed to determine the arrangement of the objects included in the real-world scene captured by the camera device. The arrangement of the objects may indicate the spatial relationships between the objects included in the real-world scene. In at least some examples, the camera device data may be analyzed to determine the real-world coordinates of one or more objects included in the real-world scene. In one or more additional examples, the camera device data may be analyzed to track the user's gaze. Additionally, the camera device data may be analyzed to identify one or more objects within the user's gaze. In one or more illustrative examples, the user's gaze is determined by identifying the center of the field of view of one or more cameras of the user device.

[0096] At 706, process 700 may further include identifying user input indicating a translation request to include an identifier of an object corresponding to at least one of the words or phrases from the one or more objects in a first language into an additional language. In one or more examples, the first language may correspond to a default language. In various examples, the default language and the additional language may be selected by the user. In one or more additional examples, at least one of the default language or the additional language may be based on one or more locations of the user. In one or more illustrative examples, the default language corresponds to a language in which the user is at least proficient or fluent, and the additional language may correspond to a language that the user is attempting to learn. In at least some examples, the user input indicating the translation request may correspond to at least one of audio input or one or more gestures captured by the camera device. In one or more additional examples, user interface elements may be displayed in the user interface that may be selected to generate the translation request. The user interface elements may be displayed in response to determining that the object does not exist in the user's object translation list.

[0097] Additionally, at 708, process 700 can include storing an identifier of an object in a data store in combination with an identifier of a user of the user device and a user's object translation list. The user's object translation list can indicate one or more objects for which the user has requested that the identifier be translated from a default language into an additional language. In at least some examples, the object translation list can indicate one or more locations of the one or more objects. Each location in the one or more locations can correspond to a group of one or more of the following objects having at least one of one or more words, one or more characters, one or more symbols, or one or more phrases translated from a default language into at least one additional language.

[0098] Process 700 can include: at 710, determining a translation of at least one of one or more words, one or more characters, one or more symbols, or one or more phrases corresponding to an identifier of an object in an additional language. In one or more examples, one or more machine learning translation algorithms can be used to determine the translation. In one or more additional examples, one or more API calls to at least one translation service can be used to obtain the translation. In one or more illustrative examples, one or more API calls to a third-party translation service are used to obtain the translation.

[0099] At 712, process 700 can include causing a translated identifier including at least one of the translated words, phrases, characters, or symbols to be displayed in augmented reality content in a user interface including a view of the object. In one or more examples, the augmented reality content can include at least one of text content, image content, video content, or animation content displayed in the user interface. In at least some examples, the augmented reality content can be displayed near the object within the user interface, where the user interface includes a view of a real-world scene containing the object. In one or more illustrative examples, the user interface includes a live view of a real-world scene. In one or more additional examples, the augmented reality content can include audio content. By way of illustration, an audio file is generated that includes an audible pronunciation of at least one of the words or phrases of the identifier of the object in at least one of the default language or the additional language. The audio file can be sent to the user device for playback in combination with the translated text, video, or image augmented reality content being displayed.

[0100] In one or more examples, in response to determining that a user is located at a location corresponding to an object, a translation of an identifier of the object can be displayed. The location can be indicated in the user's object translation list. In one or more examples, the arrangement of one or more objects in a real-world scene can be used to determine the location of a user of a user device. In one or more additional examples, the location of the user can be determined based on GPS coordinates accessed via the user device.

[0101] In various examples, when the field of view of a camera device changes, additional translations of additional objects in the field of view of the camera device are determined. For example, the field of view of the camera device can change from a first field of view to a second field of view. In these scenarios, additional camera device data is analyzed to determine the additional objects in the second field of view. In one or more examples, in response to an input from a user, an additional object can be added to the user's translation list to obtain a translation of an identifier of the additional object. Additionally, augmented reality content including at least one of text content, video content, image content, animation content, or audio content corresponding to the translation of the identifier of the additional object can be generated.

[0102] Figure 8 is a flowchart of a process 800 for generating augmented reality content according to one or more examples, the augmented reality content including translations corresponding to identifiers of objects included in a user's object translation list. At 802, process 800 can include obtaining camera device data including at least one of image content or video content captured in a field of view of a camera device of a user device. In one or more examples, the user device can include a head-mounted device. In one or more illustrative examples, the head-mounted device includes glasses. In one or more additional examples, the user device can include a wearable device. The wearable device can include an activity tracker and / or a watch and can also include other wearable devices such as contact lenses, jewelry, or other items worn on at least one of the user's ears, eyes, or other parts of the body. In at least some examples, the user device can include multiple camera devices. In these scenarios, the field of view includes a combined field of view of the multiple camera devices.

[0103] The processing 800 may further include: at 804, analyzing the camera device data to determine one or more objects indicated by the camera device data. In various examples, the camera device data may be analyzed in response to a translated AR content item being executed within an instance of a user application being executed by the user device to identify one or more objects. For example, the AR translation content item may be activated in response to an input from the user. When the translated AR content item is activated and executed within an instance of the user application, the camera device data captured by the camera device is analyzed to determine the objects indicated by the camera device data. Additionally, during activation of the translated AR content item, operations for translating identifiers of the objects indicated by the camera device data and generating augmented reality content related to the translation may also be activated.

[0104] Additionally, at 806, the processing 800 may include determining that the object is included in the user's object translation list. For example, the camera device data may be analyzed to determine characteristics of the object, such as one or more contours, one or more edges, one or more colors, one or more textures, or one or more chromaticities. The characteristics of the object may be used to determine the identifier corresponding to the object. By way of illustration, the characteristics of the object are analyzed relative to the characteristics of other objects to determine the amount of similarity between the characteristics of the object and the characteristics of one or more additional objects. If the amount of similarity is at least a threshold level, an identifier is assigned to the object. In these scenarios, the identifier of the object may be analyzed relative to the identifiers of one or more additional objects included in the object translation list to determine whether the object is included in the object translation list. In one or more illustrative examples, the object may be identified as a lamp, and the object translation list is parsed to determine whether the translation for the word "lamp" is associated with the object translation list.

[0105] Furthermore, the processing 800 may include: at 808, determining a translation of at least one of one or more words, one or more symbols, one or more phrases, or one or more characters of the object in an additional language. At 810, the processing 800 may include causing the translation to be displayed as augmented reality content in a user interface of a view including the object. In one or more examples, in response to determining that the object is included in a group of one or more translated objects corresponding to a location, the translation is caused to be displayed as augmented reality content in the user interface.

[0106] Figure 9is a diagram including user interface 900 according to one or more examples, the user interface 900 including a view of a real-world scene and augmented reality content corresponding to a translation of an object included in the real-world scene. The user interface 900 may be displayed by user device 402. In one or more examples, the user interface 900 may be displayed in combination with a translated AR content item 408 executed within user application 406.

[0107] Figure 9 An illustrative example of indicates a first field of view 902 of a real-world scene. The first field of view 902 may include a first object 904. A first identifier 906 of the first object 904 in a first language may be displayed in the user interface 900 as augmented reality content such as a virtual object. Additionally, a first additional identifier 908 of the first object 904 in a second language may also be displayed in the user interface 900 as augmented reality content. Furthermore, the first identifier 906 and the first additional identifier 908 may be displayed near the first object 904. In response to determining that camera device data captured by camera device 410 corresponds to the first field of view 902, at least one of the first identifier 906 or the first additional identifier 908 may be displayed.

[0108] The field of view of the camera device 410 may be modified from the first field of view 902 to a second field of view 910. A second object 912 may be included in the second field of view 910. A second identifier 914 of the second object 912 in the first language may be displayed in the user interface 900 as augmented reality content. Additionally, a second additional identifier 916 of the second object 912 in the second language may also be displayed in the user interface 900 as augmented reality content. The second identifier 914 and the second additional identifier 916 may be displayed near the second object 912. A third object 918 may also be within the second field of view 910. A third identifier 920 of the third object 918 in the first language and a third additional identifier 922 of the third object 918 in the second language may be displayed in the user interface 900. The third identifier 920 and the third additional identifier 922 may be displayed near the third object 918.

[0109] In various examples, in response to determining that the camera device data captured by the camera device 410 corresponds to the second field of view 910, at least one of the second identifier 914, the second additional identifier 916, the third identifier 920, or the third additional identifier 922 may be displayed. In one or more examples, in response to determining that the field of view of the camera device 410 corresponds to the second field of view 910, each of the second identifier 914, the second additional identifier 916, the third identifier 920, and the third additional identifier 922 may be displayed in the user interface 900. In one or more additional examples, the gaze of the user 404 may be determined and estimated to be the center of the second field of view 910. In these scenarios, in response to determining that the second object 912 is at the center of the second field of view 910 of the camera device 410, the second identifier 914 and the second additional identifier 916 are displayed. Further, in response to determining that the third object 918 is at the center of the second field of view 910 of the camera device 410, the third identifier 920 and the third additional identifier 922 may be displayed.

[0110] In one or more examples, at least one of the identifiers 906, 908, 914, 916, 920, 922 may be displayed as at least one of text content, video content, image content, or animated content. Additionally, audio content corresponding to at least a portion of the pronunciation of the identifiers 906, 908, 914, 916, 920, 922 may be played. For example, in response to determining that the field of view of the camera device 410 corresponds to the first field of view 902, audio content corresponding to at least one of the first identifier 906 or the first additional identifier 908 may be displayed. Further, in response to determining that the field of view of the camera device 410 corresponds to the second field of view 910, audio content corresponding to at least one of the second identifier 914, the second additional identifier 916, the third identifier 920, or the third additional identifier 922 may be displayed.

[0111] Figure 10 is a block diagram of a networked system 1000 showing details of glasses 100 according to some examples. Figure 10 Shows a system 1000 that includes a head-wearable device 100 with a selector input device according to some examples. Figure 10 is a high-level functional block diagram of an example head-wearable device 100 that is communicatively coupled via various networks 1016 to a mobile device 1050 and various server systems 1004 (e.g., the Figure 11 described interaction server system 1110).

[0112] The head-wearable device 100 includes one or more camera devices, each of the one or more camera devices may be, for example, a visible light camera device 1006, an infrared emitter 1008, and an infrared camera device 1010.

[0113] The mobile device 1050 is connected to the head-wearable device 100 using both a low-power wireless connection 1012 and a high-speed wireless connection 1014. The mobile device 1050 is also connected to the server system 1004 and the network 1016.

[0114] The head-wearable device 100 also includes two image displays in an image display 1018 of an optical component. The two image displays 1018 of the optical component include one image display associated with the left lateral side of the head-wearable device 100 and one image display associated with the right lateral side of the head-wearable device 100. The head-wearable device 100 also includes an image display driver 1020, an image processor 1022, a low-power circuitry 1024, and a high-speed circuitry 1026. The image display 1018 of the optical component is used to present images and videos to a user of the head-wearable device 100, including images that may include a graphical user interface.

[0115] The image display driver 1020 commands and controls the image display 1018 of the optical component. The image display driver 1020 may deliver image data directly to the image display 1018 of the optical component for presentation or may convert the image data into a signal or data format suitable for delivery to an image display device. For example, the image data may be video data formatted according to a compression format such as H.264 (MPEG-4 Part 10), HEVC, Theora, Dirac, RealVideo RV40, VP8, VP9, etc., and the still image data may be formatted according to a compression format such as Portable Network Graphics (PNG), Joint Photographic Experts Group (JPEG), Tagged Image File Format (TIFF), or Exchangeable Image File Format (EXIF), etc.

[0116] The head-wearable device 100 includes a frame and a stem (or temple) extending from a lateral side of the frame. The head-wearable device 100 also includes a user input device 1028 (e.g., a touch sensor or a push button), including an input surface on the head-wearable device 100. The user input device 1028 (e.g., a touch sensor or a push button) is used to receive an input selection from a user to manipulate a graphical user interface of the presented image.

[0117] Figure 10The components for the head-wearable device 100 shown are located on one or more circuit boards (such as a PCB or a flexible PCB) in a frame or temple. Alternatively or additionally, the depicted components may be located in a block, frame, hinge, or nose bridge of the head-wearable device 100. The left visible light camera device and the right visible light camera device 1006 may include digital camera device elements, such as, for example, a complementary metal oxide semiconductor (CMOS) image sensor, a charge-coupled device, a camera device lens, or any other corresponding visible light capture element or light capture element that can be used to capture data, including an image of a scene with an unknown object).

[0118] The head-wearable device 100 includes a memory 1002 that stores instructions for performing a subset or all of the functions described herein. The memory 1002 may also include a storage device.

[0119] As Figure 10 shown, the high-speed circuitry 1026 includes a high-speed processor 1030, a memory 1002, and a high-speed wireless circuitry 1032. In some examples, an image display driver 1020 is coupled to the high-speed circuitry 1026 and is operated by the high-speed processor 1030 to drive the left image display and the right image display of the image display 1018 of the optical assembly. The high-speed processor 1030 may be any processor capable of managing the high-speed communication and operation of any general computing system required by the head-wearable device 100. The high-speed processor 1030 includes the processing resources required to manage high-speed data transmission over the high-speed wireless connection 1014 to a wireless local area network (WLAN) using the high-speed wireless circuitry 1032. In certain examples, the high-speed processor 1030 executes an operating system (such as, for example, the LINUX operating system) of the head-wearable device 100 or other such operating system, and the operating system is stored in the memory 1002 for execution. In addition to any other duties, the high-speed processor 1030 that executes the software architecture of the head-wearable device 100 manages data transmission with the high-speed wireless circuitry 1032. In certain examples, the high-speed wireless circuitry 1032 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standard, which is also referred to herein as WiFi. In some examples, other high-speed communication standards may be implemented by the high-speed wireless circuitry 1032.

[0120] The low-power wireless circuitry 1034 and the high-speed wireless circuitry 1032 of the head-wearable device 100 may include a short-range transceiver (Bluetooth TM) and a wireless wide area network transceiver, a wireless local area network transceiver, or a wide area network transceiver (e.g., cellular or WiFi). A mobile device 1050 including transceivers communicating via a low-power wireless connection 1012 and a high-speed wireless connection 1014 can implement using details of the architecture of the head-mounted device 100 (such as other elements that can be of network 1016).

[0121] The memory 1002 includes any storage device capable of storing various data and applications, including camera device data generated by the left visible light camera device and the right visible light camera device 1006, the infrared camera device 1010, and the image processor 1022, and images generated for display on the image display of the optical component by the image display driver 1020, etc. Although the memory 1002 is shown integrated with the high-speed circuitry 926, in some examples, the memory 1002 can be a separate stand-alone element of the head-mounted device 100. In certain such examples, electrical wiring can provide a connection from the image processor 1022 or the low-power processor 1036 to the memory 1002 through a chip including the high-speed processor 1030. In some examples, the high-speed processor 1030 can manage the addressing of the memory 1002 such that the low-power processor 1036 will initiate the high-speed processor 1030 whenever a read or write operation involving the memory 1002 is needed.

[0122] As Figure 10 shown, the low-power processor 1036 or the high-speed processor 1030 of the head-mounted device 100 can be coupled to a camera device (visible light camera device 1006, infrared emitter 1008, or infrared camera device 1010), an image display driver 1020, a user input device 1028 (e.g., a touch sensor or a push button), and the memory 1002.

[0123] The head-mounted device 100 is connected to a host computer. For example, the head-mounted device 100 is paired with the mobile device 1050 via the high-speed wireless connection 1014 or connected to the server system 1004 via the network 1016. The server system 1004 can be one or more computing devices that are part of a service or network computing system, e.g., including a processor, a memory, and a network communication interface to communicate with the mobile device 1050 and the head-mounted device 100 via the network 1016.

[0124] The mobile device 1050 includes a processor and a network communication interface coupled to the processor. The network communication interface allows communication via the network 1016, a low-power wireless connection 1012, or a high-speed wireless connection 1014. The mobile device 1050 may also store at least some portions of instructions for generating stereophonic audio content in the memory of the mobile device 1050 to implement the functions described herein.

[0125] The output components of the head-wearable device 100 include visual components such as a display such as a liquid crystal display (LCD), a plasma display panel (PDP), a light-emitting diode (LED) display, a projector, or a waveguide. The image display of the optical component is driven by an image display driver 1020. The output components of the head-wearable device 100 also include acoustic components (e.g., speakers), tactile components (e.g., vibration motors), other signal generators, and the like. The input components (e.g., the user input device 1028) of the head-wearable device 100, the mobile device 1050, and the server system 1004 may include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input components), point-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instruments), tactile input components (e.g., physical buttons, a touch screen that provides the position and force of a touch or touch gesture, or other tactile input components), audio input components (e.g., a microphone), and the like.

[0126] The head-wearable device 100 may also include additional peripheral device elements. Such peripheral device elements may include biometric sensors, additional sensors, or display elements integrated with the head-wearable device 100. For example, the peripheral device elements may include any I / O components, which include output components, motion components, positioning components, or any other such elements described herein.

[0127] For example, biometric components include components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biometric signals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying people (e.g., voice recognition, retina recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), and the like. Motion components include acceleration sensor components (e.g., accelerometers), gravity sensor components, rotational sensor components (e.g., gyroscopes), and the like. Positioning components include positioning sensor components for generating positioning coordinates (e.g., a global positioning system (GPS) receiver component), Wi-Fi or Bluetooth for generating positioning system coordinates TMTransceivers, altitude sensor components (e.g., altimeters or barometers that detect air pressure, from which altitude can be obtained), orientation sensor components (e.g., magnetometers), etc. Such positioning system coordinates can also be received from the mobile device 1050 via the low-power radio circuitry 1034 or the high-speed radio circuitry 1032 through the low-power wireless connection 1012 and the high-speed wireless connection 1014.

[0128] Figure 11 is a block diagram showing an example interaction system 1100 for facilitating interactions on a network (e.g., exchanging text messages, making text, audio, and video calls, or playing games). The interaction system 1100 includes a plurality of user systems 1102, each of the plurality of user systems 1102 hosting a plurality of applications including an interaction client 1104 and other applications 1106. Each interaction client 1104 is communicatively coupled via one or more communication networks including a network 1108 (e.g., the Internet) to other instances of interaction clients 1104 (e.g., hosted on corresponding other user systems 1102), an interaction server system 1110, and a third-party server 1112. The interaction client 1104 can also communicate with the locally hosted applications 1106 using an application programming interface (API).

[0129] Each user system 1102 can include a plurality of user devices, such as mobile devices 1114, head-wearable devices 1116, and computer client devices 1118, which are communicatively connected to exchange data and messages.

[0130] The interaction client 1104 interacts via the network 1108 with other interaction clients 1104 and with the interaction server system 1110. The data exchanged between the interaction clients 1104 (e.g., interaction 1120) and between the interaction client 1104 and the interaction server system 1110 includes functions (e.g., commands for activating functions) and payload data (e.g., text, audio, video, or other multimedia data).

[0131] The interaction server system 1110 provides server-side functions to the interaction clients 1104 via the network 1108. Although certain functions of the interaction system 1100 are described herein as being performed by the interaction client 1104 or by the interaction server system 1110, whether certain functions are located within the interaction client 1104 or within the interaction server system 1110 can be a design choice. For example, it may be technically preferred to initially deploy a particular technology and function within the interaction server system 1110, but then migrate the technology and function to the interaction client 1104 where the user system 1102 has sufficient processing power.

[0132] The interaction server system 1110 supports various services and operations provided to the interaction client 1104. Such operations include transmitting data to the interaction client 1104, receiving data from the interaction client 1104, and processing data generated by the interaction client 1104. The data can include message content, client device information, geolocation information, media enhancements and overlays, message content persistence conditions, social network information, and live event information. Data exchange within the interaction system 1100 is activated and controlled through functions available via the user interface (UI) of the interaction client 1104.

[0133] Now turning specifically to the interaction server system 1110, the application programming interface (API) server 1122 is coupled to the interaction server 1124 and provides a programming interface to the interaction server 1124, making the functions of the interaction server 1124 accessible to the interaction client 1104, other applications 1106, and third-party servers 1112. The interaction server 1124 is communicatively coupled to the database server 1126, thereby facilitating access to the database 1128, which stores data associated with interactions processed by the interaction server 1124. Similarly, the web server 1130 is coupled to the interaction server 1124 and provides a web-based interface to the interaction server 1124. To this end, the web server 1130 processes incoming network requests via the Hypertext Transfer Protocol (HTTP) and several other related protocols.

[0134] The application programming interface (API) server 1122 receives and transmits interaction data (e.g., commands and message payloads) between the interaction server 1124 and the client system 1102 (and, for example, the interaction client 1104 and other applications 1106) as well as third-party servers 1112. Specifically, the application programming interface (API) server 1122 provides a set of interfaces (e.g., routines and protocols) that can be invoked or queried by the interaction client 1104 and other applications 1106 to activate the functions of the interaction server 1124. The application programming interface (API) server 1122 exposes various functions supported by the interaction server 1124, including: account registration; login functionality; sending interaction data from a specific interaction client 1104 to another interaction client 1104 via the interaction server 1124; transmitting media files (e.g., images or videos) from the interaction client 1104 to the interaction server 1124; setting a collection of media data (e.g., a story); retrieving a friend list of a user of the user system 1102; retrieving messages and content; adding and deleting entities (e.g., friends) to and from an entity graph (e.g., a social graph); locating friends within the social graph; and opening application events (e.g., related to the interaction client 1104).

[0135] Figure 12FIG. 1200 is a block diagram showing a software architecture 1204 that can be installed on any one or more of the devices described herein. The software architecture 1204 is supported by hardware such as a machine 1202 that includes a processor 1220, a memory 1226, and I / O components 1238. In this example, the software architecture 1204 can be conceptualized as a stack of layers, where each layer provides a specific function. The software architecture 1204 includes layers such as an operating system 1212, libraries 1208, frameworks 1210, and applications 1206. In operation, an application 1206 makes API calls 1250 through the software stack and receives messages 1252 in response to the API calls 1250.

[0136] The operating system 1212 manages hardware resources and provides common services. The operating system 1212 includes, for example: a kernel 1214, services 1216, and drivers 1222. The kernel 1214 acts as an abstraction layer between the hardware and other software layers. For example, the kernel 1214 provides memory management, processor management (e.g., scheduling), component management, networking, and security settings, among other functions. The services 1216 can provide other common services to other software layers. The drivers 1222 are responsible for controlling or interfacing with the underlying hardware. For example, the drivers 1222 can include a display driver, a camera driver, or a low-power driver, a flash driver, a serial communication driver (e.g., a Universal Serial Bus (USB) driver), a driver, an audio driver, a power management driver, and so on.

[0137] Library 1208 provides low-level common infrastructure used by application 1206. Library 1208 may include system libraries 1218 (e.g., C standard library), and system libraries 1218 provide functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. Additionally, library 1208 may include API libraries 1224, such as media libraries (e.g., libraries for supporting the presentation and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), graphics libraries (e.g., OpenGL framework for presenting two-dimensional (2D) and three-dimensional (3D) image content on a display, GLMotif for implementing user interfaces), image feature extraction libraries (e.g., OpenIMAJ), database libraries (e.g., SQLite providing various relational database functions), web libraries (e.g., WebKit providing web browsing functions), etc. Library 1208 may also include various other libraries 1228 to provide many other APIs to application 1206.

[0138] Framework 1210 provides high-level common infrastructure used by application 1206. For example, framework 1210 provides various graphical user interface (GUI) functions, advanced resource management, and advanced location services. Framework 1210 may provide a wide range of other APIs that can be used by application 1206, and some of these APIs may be specific to a particular operating system or platform.

[0139] In an example, application 1206 may include a home application 1236, a contacts application 1230, a browser application 1232, a book reader application 1234, a location application 1242, a media application 1244, a messaging application 1246, a gaming application 1248, and various other applications such as third-party application 1240. Application 1206 is a program that executes functions defined in a program. One or more of applications 1206 can be created using various programming languages, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., C language or assembly language). In a specific example, third-party application 1240 (e.g., an application developed using an ANDROID TM or IOS TM software development kit (SDK)) can be an application on platforms such as IOS TM 、ANDROID TM 、 Mobile software running on the mobile operating system of a Phone or another mobile operating system. In this example, third-party application 1240 can activate API call 1250 provided by operating system 1212 to facilitate the functions described herein.

[0140] "Carrier signal" means any non-transitory medium that can store, encode, or carry instructions executable by a machine and includes digital or analog communication signals or other non-transitory media to facilitate the communication of such instructions. The instructions may be sent or received via a network interface device over a network using a transmission medium.

[0141] "Client device" means any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. A client device may be, but is not limited to, a mobile phone, desktop computer, laptop computer, portable digital assistant (PDA), smartphone, tablet computer, ultrabook, netbook, laptop, multiprocessor system, microprocessor-based or programmable consumer electronics, game console, set-top box, or any other communication device that a user may use to access a network.

[0142] "Communication network" means one or more portions of a network, which may be an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless WAN (WWAN), metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the public switched telephone network (PSTN), a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, a network, another type of network, or a combination of two or more such networks. For example, a network or a portion of a network may include a wireless network or a cellular network, and the coupling may be a code division multiple access (CDMA) connection, a global system for mobile communications (GSM) connection, or other type of cellular or wireless coupling. In this example, the coupling may implement any data transfer technology among various types of data transfer technologies, such as single carrier radio transmission technology (1xRTT), evolved data optimized (EVDO) technology, general packet radio service (GPRS) technology, GSM enhanced data rates for GSM evolution (EDGE) technology, the third generation partnership project (3GPP) including 3G, fourth generation wireless (4G) network, universal mobile telecommunications system (UMTS), high speed packet access (HSPA), worldwide interoperability for microwave access (WiMAX), long term evolution (LTE) standard, other data transfer technologies defined by various standards setting organizations, other long distance protocols, or other data transfer technologies.

[0143] "Component" refers to a device, physical entity, or logic that has boundaries defined by function or subroutine calls, branch points, APIs, or other technical definitions that provide for partitioning or modularizing a particular processing or control function. A component can be combined with other components via its interfaces to perform machine processing. A component can be an encapsulated functional hardware unit designed to be used with other components and a part of a program for a specific function that generally performs related functions. A component can be a software component (e.g., code implemented on a machine-readable medium) or a hardware component. A "hardware component" is a tangible unit capable of performing some operations and can be configured or arranged in a particular physical manner. In various examples, one or more computer systems (e.g., a stand-alone computer system, a client computer system, or a server computer system) or one or more hardware components of a computer system (e.g., a processor or a group of processors) can be configured by software (e.g., an application or a part of an application) to operate to perform some operations as described herein as a hardware component. A hardware component can also be implemented mechanically, electronically, or in any suitable combination thereof. For example, a hardware component can include dedicated circuitry or logic that is permanently configured to perform some operations. A hardware component can be a dedicated processor, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). A hardware component can also include programmable logic or circuitry that is temporarily configured by software to perform some operations. For example, a hardware component can include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a particular machine (or a particular component of a machine) customized to perform the configured function and is no longer a general-purpose processor. It will be understood that considerations of cost and time can drive the decision to implement a hardware component mechanically in dedicated and permanently configured circuitry or in circuitry that is temporarily configured (e.g., configured by software). Accordingly, the phrase "hardware component" (or "hardware-implemented component") should be understood to include a tangible entity, i.e., an entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a particular manner or to perform some operations described herein. Considering an example where a hardware component is temporarily configured (e.g., programmed), the hardware component may not be configured or instantiated at any given moment. For example, in the case where a hardware component includes a general-purpose processor that is configured by software to become a dedicated processor, the general-purpose processor can be configured at different times to be respective different dedicated processors (e.g., including different hardware components). The software accordingly configures one or more particular processors to, for example, constitute a particular hardware component at one moment and different hardware components at different moments. A hardware component can provide information to and receive information from other hardware components. Thus, the described hardware components can be considered communicatively coupled.In the presence of multiple hardware components, communication can be achieved through signal transmission between or among two or more hardware components (e.g., via appropriate circuits and buses). In examples where multiple hardware components are configured or instantiated at different times, communication between such hardware components can be achieved, for example, by storing information in a memory structure accessible to the multiple hardware components and retrieving the information from the memory structure. For example, one hardware component can perform an operation and store the output of the operation in a memory device communicatively coupled thereto. Then, another hardware component can access the memory device at a subsequent time to retrieve and process the stored output. Hardware components can also initiate communication with input or output devices and can operate on resources (e.g., collection of information). The various operations of the example methods described herein can be performed by one or more processors temporarily configured (e.g., via software) or permanently configured to perform the associated operations. Whether temporarily or permanently configured, such processors can constitute processor-implemented components that operate to perform one or more of the operations or functions described herein. As used herein, a "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the methods described herein can be partially implemented by a processor, where a particular one or more processors are examples of hardware. For example, some of the operations of the method can be performed by one or more processors or processor-implemented components. Additionally, one or more processors can also operate to support the execution of associated operations in a "cloud computing" environment or operate as "software as a service" (SaaS). For example, some of the operations can be performed by a group of computers (as an example of machines including processors), where the operations can be accessed via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs). The execution of some of the operations can be distributed among processors, reside within a single machine, and be deployed across multiple machines. In some examples, the processor or processor-implemented components can be located in a single geographical location (e.g., within a home environment, an office environment, or a server farm). In other examples, the processor or processor-implemented components can be distributed across multiple geographical locations.

[0144] "Computer-readable medium" refers to both machine storage media and transmission media. Thus, these terms include both storage devices / media and carrier / modulated data signals. The terms "machine-readable medium", "computer-readable medium", and "device-readable medium" mean the same thing and can be used interchangeably in this disclosure.

[0145] "Machine storage medium" means a single or multiple storage devices and / or media (e.g., centralized or distributed databases, and / or associated caches and servers) that store executable instructions, routines, and / or data. The term includes, but is not limited to, solid-state memory as well as optical and magnetic media, including memory internal or external to a processor. Specific examples of machine storage media, computer storage media, and / or device storage media include: non-volatile memory, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGA, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage medium", "device storage medium", and "computer storage medium" mean the same thing and may be used interchangeably in this disclosure. The terms "machine storage medium", "computer storage medium", and "device storage medium" expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are covered by the term "signal medium".

[0146] "Processor" means any circuit or virtual circuit (a physical circuit simulated by logic executed on an actual processor) that manipulates data values in accordance with control signals (e.g., "commands", "opcodes", "machine codes", etc.) and produces associated output signals that are applied to operate a machine. For example, a processor can be a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), or any combination thereof. A processor can also be a multi-core processor having two or more independent processors (sometimes called "cores") that can execute instructions simultaneously.

[0147] "Signal medium" means any intangible medium that is capable of storing, encoding, or carrying instructions executed by a machine, and "signal medium" includes digital or analog communication signals or other intangible media to facilitate the communication of software or data. The term "signal medium" can be considered to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal whose one or more characteristics are set or changed in such a manner as to encode information in the signal. The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure.

[0148] In view of the above implementations of the subject matter, the present application discloses the following list of examples, where an example that considers one feature of a single example or a combination of features, and optionally more than one feature of an example considered in combination with one or more features of one or more additional examples, is also an additional example that falls within the disclosure of the present application.

[0149] Example 1 is a computer-implemented method, including: obtaining camera device data by a computing system including one or more processors and a memory, the camera device data including at least one of image content or video content captured in a field of view of a camera device of a user device; analyzing the camera device data by the computing system to determine one or more objects indicated by the camera device data; identifying, by the computing system, a user input indicating a translation request to translate at least one of a word or a phrase corresponding to a first identifier of an object among the one or more objects, the at least one of the word or the phrase being in a first language; storing, by the computing system, the first identifier of the object in combination with an identifier of a user of the user device and in combination with a user's object translation list in a data store, the user's object translation list including at least one object having at least one of one or more words or one or more phrases that have been translated from the first language into a second language; determining, by the computing system, a translation of the first identifier of the object, the translation corresponding to a second identifier of the object in the second language; and causing, by the computing system, at least one of a word, a phrase, a character, or a symbol corresponding to the translation to be displayed as augmented reality content in a user interface that includes a view of a real-world scene including the object.

[0150] In Example 2, the subject matter of Example 1 includes: determining, by the computing system, a location of the user device; and storing, by the computing system, the identifier of the object in combination with the location in the object translation list.

[0151] In Example 3, the subject matter of Example 2 includes: determining the location of the object based on an arrangement of a plurality of objects in an environment including the object.

[0152] In Example 4, the subject matter of Example 2 includes: determining the location of the object based on Global Positioning System (GPS) data obtained from the user device.

[0153] In Example 5, the subject matter of any one of Examples 1 to 4 includes: generating, by the computing system, an audio file including an audible pronunciation of a second identifier of the object in the second language; and causing, by the computing system, the audio file to be sent to the user device for playback in combination with the augmented reality content including the translation being displayed.

[0154] In Example 6, the subject matter of any one of Examples 1 to 5 includes: in response to recognizing a user input indicating a translation request, generating, by a computing system, one or more application programming interface (API) calls to obtain a translation from a third-party translation service; sending, by the computing system, the one or more API calls to the third-party translation service; obtaining, by the computing system, the translation from the third-party translation service; and sending, by the computing system, the translation to a user device.

[0155] In Example 7, the subject matter of any one of Examples 1 to 6 includes: determining, by a computing system, that a field of view of a camera device has changed from a first field of view to a second field of view; obtaining, by the computing system, additional camera device data corresponding to the second field of view; analyzing, by the computing system, the additional camera device data to determine additional objects included in the second field of view; and determining, by the computing system, that the additional objects are included in a user's object translation list.

[0156] In Example 8, the subject matter of Example 7 includes: obtaining, by a computing system, an additional translation of at least one of one or more words, one or more phrases, one or more characters, or one or more symbols of an additional identifier of an additional object in a second language; causing, by the computing system, the additional translation to be displayed as additional augmented reality content in an additional user interface that includes an additional view of an additional real-world scene that includes the additional object; generating, by the computing system, an audio file of an audible pronunciation of the second language that includes the additional identifier of the additional object; and causing, by the computing system, the audio file to be sent to a user device for playback in combination with the additional augmented reality content that includes the additional translation being displayed.

[0157] In Example 9, the subject matter of any one of Examples 1 to 8 includes: a user input indicating a translation request corresponding to a touch input on a portion of a frame of a head-mounted device.

[0158] In Example 10, the subject matter of any one of Examples 1 to 9 includes: a user input indicating a translation request corresponding to at least one of an audio input captured by one or more microphones or one or more gestures captured by a camera device.

[0159] Example 11 is a computing device, comprising: one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform operations, the operations including: obtaining camera device data that includes at least one of image content or video content captured in a field of view of a camera device of a user device; analyzing the camera device data to determine one or more objects indicated by the camera device data; identifying user input indicating a translation request to translate at least one of a word or phrase corresponding to an object in the one or more objects, the at least one of the word or phrase being in a first language; storing the first identifier of the object in a data store in combination with an identifier of a user of the user device and in combination with a user's object translation list, the user's object translation list including at least one object having at least one of one or more words, one or more phrases, one or more characters, or one or more symbols that have been translated from the first language into a second language; determining a translation of the first identifier of the object that corresponds to a second identifier of the object in the second language; and causing at least one of the word, phrase, character, or symbol corresponding to the translation to be displayed as augmented reality content in a user interface that includes a view of a real-world scene including the object.

[0160] In example 12, the subject matter of example 11 includes: the object translation list includes a data structure that is stored in a data store in combination with an identifier of a user, an identifier of the user corresponding to a default language, one or more additional languages, and one or more locations, wherein each of the one or more locations corresponds to a group of one or more objects in which at least one of one or more words or one or more phrases of one or more identifiers of the one or more objects has been translated from the default language into at least one additional language of the one or more additional languages.

[0161] In example 13, the subject matter of example 11 or 12 includes: a memory storing additional instructions that, when executed by the one or more processors, cause the computing device to perform additional operations, the additional operations including: analyzing the camera device data to determine an arrangement of a plurality of objects in an environment including the object; determining a location of the object based on the arrangement of the plurality of objects; determining that the object is not present in the user's object translation list based on the location; and causing an option to be provided to the user to add the object to the user's object translation list.

[0162] Example 14 is a computer-implemented method, including: obtaining camera device data by a computing system including one or more processors and a memory, the camera device data including at least one of image content or video content captured in a field of view of a camera device of a user equipment; analyzing the camera device data by the computing system to determine one or more objects indicated by the camera device data; determining by the computing system that an object among the one or more objects is included in an object translation list of a user of the user equipment, the object translation list of the user including at least one object having an identifier: the identifier including at least one of one or more words or one or more phrases that have been translated from a first language into a second language; determining by the computing system a translation of a first identifier of the object, wherein the translation corresponds to a second identifier of the object in the second language, the second identifier including at least one of one or more words, one or more phrases, one or more characters or one or more symbols; and causing by the computing system the translation to be displayed as augmented reality content in a user interface, the user interface including a view of a real-world scene including the object.

[0163] In Example 15, the subject matter of Example 14 includes: one or more objects indicated by the camera device data include a plurality of objects, and includes: analyzing the camera device data by the computing system to determine an arrangement of the plurality of objects; determining by the computing system a position corresponding to the field of view based on the arrangement of the plurality of objects; and determining by the computing system that the object is included in a group of one or more translated objects corresponding to the position.

[0164] In Example 16, the subject matter of Example 15 includes: in response to determining that the object is included in a group of one or more translated objects corresponding to the position, causing the translation to be displayed as augmented reality content in the user interface.

[0165] In Example 17, the subject matter of any one of Examples 14 to 16 includes: determining by the computing system that a translation augmented reality (AR) content item is being executed within an instance of a user application being executed by the user equipment; wherein, in response to determining that the translation AR content item is being executed within an instance of a user application being executed by the user equipment, analyzing the camera device data to determine one or more objects indicated by the camera device data.

[0166] In Example 18, the subject matter of any one of Examples 14 to 17 includes: generating by the computing system an audio file including at least one of one or more words, one or more phrases, one or more characters or one or more symbols corresponding to the second identifier of the object in the second language; and causing by the computing system the audio file to be sent to the user equipment for playback in combination with the augmented reality content including the translation being displayed.

[0167] In Example 19, the subject matter of any one of Examples 14 to 18 includes: determining, by a computing system, that a field of view of a camera device changes from a first field of view to a second field of view; obtaining, by the computing system, additional camera device data corresponding to the second field of view; analyzing, by the computing system, the additional camera device data to identify additional objects included in the second field of view; determining, by the computing system, that the additional objects are not present in a user's object translation list; and causing, by the computing system, an option to be provided to the user to add the additional objects to the user's object translation list.

[0168] In Example 20, the subject matter of Example 19 includes: identifying, by a computing system, user input indicating a translation request to translate at least one of one or more first additional words, one or more first additional phrases, one or more first additional characters, or one or more first additional symbols of a first language of a first additional identifier of an additional object into a second language; causing, by the computing system, the first additional identifier of the additional object to be stored in a data store in combination with a user identifier in the user's object translation list; determining, by the computing system, an additional translation that includes at least one of one or more second additional words, one or more second additional phrases, one or more second additional characters, or one or more second additional symbols of a second language corresponding to a second additional identifier of the additional object; and causing, by the computing system, the second additional identifier to be displayed as augmented reality content in a user interface that includes a view of a real-world scene that includes the object.

[0169] Changes and modifications may be made to the disclosed examples without departing from the scope of the present disclosure. These and other changes or modifications are intended to be included within the scope of the present disclosure as expressed in the appended claims.

Claims

1. A computer-implemented method, comprising: obtaining, by a computing system including one or more processors and a memory, camera device data including at least one of image content or video content captured in a field of view of a camera device of a user device; analyzing, by the computing system, the camera device data to determine one or more objects indicated by the camera device data; identifying, by the computing system, user input indicating a translation request to translate at least one of a word or phrase corresponding to a first identifier of an object among the one or more objects, the at least one of the word or phrase being in a first language; storing, by the computing system, the first identifier of the object in a data store in combination with an identifier of a user of the user device and in combination with a user's object translation list of the user, the user's object translation list including at least one object having at least one of one or more words or one or more phrases that have been translated from the first language into a second language; determining, by the computing system, a translation of the first identifier of the object corresponding to a second identifier of the object in the second language; and causing, by the computing system, at least one of a word, phrase, character, or symbol corresponding to the translation to be displayed as augmented reality content in a user interface including a view of a real-world scene including the object.

2. The computer-implemented method according to claim 1, comprising: determining, by the computing system, a location of the user device; and storing, by the computing system, the identifier of the object in the object translation list in combination with the location.

3. The computer-implemented method according to claim 2, wherein, The location of the object is determined based on an arrangement of a plurality of objects in an environment including the object.

4. The computer-implemented method according to claim 2, wherein, The location of the object is determined based on global positioning system (GPS) data obtained from the user device.

5. The computer-implemented method according to claim 1, comprising: generating, by the computing system, an audio file including an audible pronunciation of the second identifier of the object in the second language; and causing, by the computing system, the audio file to be sent to the user device for playback in combination with the augmented reality content including the translation being displayed.

6. The computer-implemented method according to claim 1, comprising: in response to identifying the user input indicating the translation request, generating, by the computing system, one or more application programming interface (API) calls to obtain the translation from a third-party translation service; sending, by the computing system, the one or more API calls to the third-party translation service; obtaining, by the computing system, the translation from the third-party translation service; and sending, by the computing system, the translation to the user device.

7. The computer-implemented method according to claim 1, comprising: determining, by the computing system, that the field of view of the camera device changes from a first field of view to a second field of view; obtaining, by the computing system, additional camera device data corresponding to the second field of view; The additional camera device data is analyzed by the computing system to determine additional objects included in the second field of view; and the computing system determines that the additional objects are included in the user's object translation list.

8. The computer-implemented method according to claim 7, comprising: obtaining, by the computing system, an additional translation of at least one of one or more words, one or more phrases, one or more characters, or one or more symbols of an additional identifier of the additional object in the second language; causing, by the computing system, the additional translation to be displayed as additional augmented reality content in an additional user interface, the additional user interface including an additional view of an additional real-world scene including the additional object; generating, by the computing system, an audio file of an audible pronunciation of the second language including the additional identifier of the additional object; and causing, by the computing system, the audio file to be sent to the user device for playback in combination with the additional augmented reality content including the additional translation being displayed.

9. The computer-implemented method according to claim 1, wherein, The user input indicating the translation request corresponds to a touch input on a part of a frame of a head-mounted device.

10. The computer-implemented method according to claim 1, wherein, The user input indicating the translation request corresponds to at least one of an audio input captured by one or more microphones or one or more gestures captured by the camera device.

11. A computing device, comprising: one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the computing device to perform operations, the operations including: obtaining camera device data, the camera device data including at least one of image content or video content captured in a field of view of a camera device of a user device; analyzing the camera device data to determine one or more objects indicated by the camera device data; identifying a user input indicating a translation request to translate at least one of a word or phrase corresponding to a first identifier of an object among the one or more objects, the at least one of the word or phrase being in a first language; storing, in a data store, the first identifier of the object in combination with an identifier of a user of the user device and in combination with the user's object translation list, the user's object translation list including at least one object having at least one of one or more words, one or more phrases, one or more characters, or one or more symbols that have been translated from the first language into a second language; determining a translation of the first identifier of the object, the translation corresponding to a second identifier of the object in the second language; and causing at least one of a word, phrase, character, or symbol corresponding to the translation to be displayed as augmented reality content in a user interface, the user interface including a view of a real-world scene including the object.

12. The computing device according to claim 11, wherein, The object translation list includes a data structure that is stored in a data store in combination with an identifier of the user, an identifier of the user corresponding to a default language, one or more additional languages, and one or more locations, wherein each of the one or more locations corresponds to a group of one or more of the following objects: at least one of one or more words or one or more phrases of one or more identifiers of the one or more objects is translated from the default language into at least one additional language of the one or more additional languages.

13. The computing device according to claim 11, wherein the memory stores additional instructions that, when executed by the one or more processors, cause the computing device to perform additional operations, the additional operations including: Analyzing the camera device data to determine an arrangement of a plurality of objects in an environment including the object; Determining a location of the object based on the arrangement of the plurality of objects; Determining that the object does not exist in the user's object translation list based on the location; And Causing an option to be provided to the user to add the object to the user's object translation list.

14. A computer-implemented method, comprising: Obtaining, by a computing system including one or more processors and a memory, camera device data including at least one of image content or video content captured in a field of view of a camera device of a user device; Analyzing, by the computing system, the camera device data to determine one or more objects indicated by the camera device data; Determining, by the computing system, that an object among the one or more objects is included in a user's object translation list of the user device, the user's object translation list including at least one object having an identifier that includes at least one of one or more words or one or more phrases that have been translated from a first language into a second language; Determining, by the computing system, a translation of a first identifier of the object, wherein the translation corresponds to a second identifier of the object in the second language, the second identifier including at least one of one or more words, one or more phrases, one or more characters, or one or more symbols; And Causing, by the computing system, the translation to be displayed as augmented reality content in a user interface that includes a view of a real-world scene including the object.

15. The computer-implemented method according to claim 14, wherein, The one or more objects indicated by the camera device data include a plurality of objects, and the method includes: Analyzing, by the computing system, the camera device data to determine an arrangement of the plurality of objects; Determining, by the computing system, a location corresponding to the field of view based on the arrangement of the plurality of objects; and Determining, by the computing system, that the object is included in a group of one or more translated objects corresponding to the location.

16. The computer-implemented method according to claim 15, wherein, In response to determining that the object is included in a group of one or more translated objects corresponding to the location, causing the translation to be displayed as the augmented reality content in the user interface.

17. The computer-implemented method according to claim 14, comprising: determining, by the computing system, that a translated augmented reality (AR) content item is being executed within an instance of a user application being executed by the user device; wherein, in response to determining that the translated AR content item is being executed within the instance of the user application being executed by the user device, analyzing the camera device data to determine the one or more objects indicated by the camera device data.

18. The computer-implemented method according to claim 14, comprising: generating, by the computing system, an audio file including an audible pronunciation of at least one of the one or more words, one or more phrases, one or more characters, or one or more symbols including a second identifier corresponding to the object in the second language; and causing, by the computing system, the audio file to be sent to the user device for playback in combination with the augmented reality content including the translation being displayed.

19. The computer-implemented method according to claim 14, comprising: determining, by the computing system, that the field of view of the camera device changes from a first field of view to a second field of view; obtaining, by the computing system, additional camera device data corresponding to the second field of view; analyzing, by the computing system, the additional camera device data to determine additional objects included in the second field of view; determining, by the computing system, that the additional objects are not present in the user's object translation list; and causing, by the computing system, an option to be provided to the user to add the additional objects to the user's object translation list.

20. The computer-implemented method according to claim 19, comprising: identifying, by the computing system, a user input indicating a translation request to translate at least one of the one or more first additional words, one or more first additional phrases, one or more first additional characters, or one or more first additional symbols of a first language of a first additional identifier of the additional object into a second language; causing, by the computing system, the first additional identifier of the additional object to be stored in a data store in combination with the user's identifier in the user's object translation list; determining, by the computing system, an additional translation including at least one of the one or more second additional words, one or more second additional phrases, one or more second additional characters, or one or more second additional symbols of the second language corresponding to a second additional identifier of the additional object; and causing, by the computing system, the second additional identifier to be displayed in a user interface including a view of a real-world scene including the object.

Citation Information

Cited By

  • Generating augmented reality content including translations

    US12530544B2