Extended reality virtual assistant

CN122816449APending Publication Date: 2026-09-25QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610895247.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2017-07-20
Filing Date
2018-06-20
Publication Date
2026-09-25

Smart Images

  • Figure CN122816449A_ABST
    Figure CN122816449A_ABST
Patent Text Reader

Abstract

Methods, apparatuses, and devices are provided to facilitate positioning a virtual content item in an extended reality environment. For example, a first user can access the extended reality environment through a display of a mobile device, and in some instances, the method can determine locations and orientations of the first user and a second user in the extended reality environment. The method can also determine a placement location for the virtual content item in the extended reality environment based on the determined locations and orientations of the first user and the second user, and perform an operation to insert the virtual content item into the extended reality environment at the determined placement location.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to computer-implemented systems and processes for dynamically positioning virtual content in an extended reality environment. Background Technology

[0002] Mobile devices enable users to explore and immerse themselves in extended reality environments, such as augmented reality environments that provide a real-time view of a physically real-world environment combined with or enhanced by computer-generated graphical content. When immersed in an extended reality environment, many users experience a subtle connection to real-world information sources, which provides the ability to further enhance user interaction with and exploration of the extended reality environment. Summary of the Invention

[0003] The disclosed computer-implemented extended reality method may include determining the position and orientation of a first user in an extended reality environment by one or more processors. The method may further include determining the position and orientation of a second user in the extended reality environment by the one or more processors. The method may further include: determining, by the one or more processors, a placement position of a virtual content item in the extended reality environment based on the determined positions and orientations of the first user and the second user; and inserting the virtual content item into the extended reality environment at the determined placement position.

[0004] A disclosed device can be used in an extended reality environment. The device may include: a non-transitory machine-readable storage medium storing instructions; and at least one processor configured to execute the instructions to determine the position and orientation of a first user in the extended reality environment. The at least one processor is also configured to determine the position and orientation of a second user in the extended reality environment. The at least one processor is configured to execute instructions to determine the placement position of a virtual content item in the extended reality environment based on the determined positions and orientations of the first and second users, and to insert a virtual assistant into the extended reality environment at the determined placement position.

[0005] A disclosed device includes means for determining the position and orientation of a first user in an augmented reality environment. The device further includes means for determining the position and orientation of a second user in the augmented reality environment. The device includes means for determining a placement position of a virtual content item in the augmented reality environment, at least in part based on the determined positions and orientations of the first and second users. The device further includes means for inserting the virtual content item into the augmented reality environment at the determined placement position.

[0006] A disclosed non-transitory computer-readable storage medium is encoded with processor-executable program code, the processor-executable program code comprising: program code for determining the position and orientation of a first user in an extended reality environment; program code for determining the position and orientation of a second user in the extended reality environment; program code for determining a placement position of a virtual content item in the extended reality environment based at least in part on the determined position and orientation of the first user and the second user; and program code for inserting the virtual content item into the extended reality environment at the determined placement position. Attached Figure Description

[0007] Figure 1 is a block diagram of an exemplary network for an augmented reality environment based on some examples.

[0008] Figure 2 is a block diagram of an exemplary mobile device used in the augmented reality environment of Figure 1, based on some examples.

[0009] Figure 3 is a flowchart illustrating an exemplary process for dynamically locating a virtual assistant in an augmented reality environment using the mobile device shown in Figure 2, based on some examples.

[0010] Figures 4A and 4B are diagrams illustrating exemplary interactions between users using the mobile device in Figure 2 and the augmented reality environment, based on several examples.

[0011] Figure 5 shows example results of semantic scene analysis performed on the mobile device in Figure 2 based on some examples.

[0012] Figures 6A-6D, 7A, 7B, and 8 illustrate examples of virtual network computing using Figure 1. A graph illustrating various aspects of the scoring process for placing a hand at a candidate location in an augmented reality environment.

[0013] Figure 9 is a flowchart of an exemplary process for performing an operation in an augmented reality environment in response to detected gesture input, based on some examples.

[0014] Figures 10A and 10B are diagrams illustrating, based on some examples, how users interact with the augmented reality environment using the mobile device shown in Figure 2. Detailed Implementation

[0015] While the features, methods, apparatus, and systems described herein may be embodied in various forms, some exemplary and non-limiting embodiments are shown in the accompanying drawings and described below. Some components described in this disclosure are optional, and some embodiments may include additional, different, or fewer components compared to those expressly described in this disclosure.

[0016] Relative terms such as “lower,” “upper,” “horizontal,” “vertical,” “above,” “below,” “up,” “top,” “bottom,” and their derivatives (e.g., “horizontally,” “downward,” “upward,” etc.) refer to the orientation as described at the time or shown in the accompanying drawings discussed. Relative terms are provided for the reader's convenience. Relative terms do not limit the scope of the claims.

[0017] Many virtual assistant technologies, such as software pop-up assistants and voice call assistants, have been designed for use in telecommunications systems. This disclosure provides virtual assistants that utilize the potential of extended real-world environments, such as virtual reality environments, augmented reality environments, and augmented virtual environments.

[0018] The following describes examples of extended reality environments in the context of mobile computing. In some of the examples described below, extended reality generation and rendering tools are accessible via mobile devices and computing systems associated with extended reality platforms. These tools define extended reality environments based on elements of digital content such as captured digital video, digital images, digital audio content, or synthesized audiovisual content (e.g., computer-generated images and animations). The tools can deploy elements of the digital content on mobile devices to present them to users via displays incorporated into extended reality, virtual reality, or augmented reality headsets (such as head-mounted displays (HMDs)).

[0019] Mobile devices may include augmented reality (AR) eyewear (such as glasses, goggles, or any device covering a user's eyes) with one or more lenses or displays for displaying graphical elements of deployed digital content. For example, the eyewear may display the graphical elements as an augmented reality layer overlaid on real-world objects visible through the lenses. Additionally, portions of the digital content deployed for a mobile device—which establishes an augmented reality or other extended reality environment for the user of that mobile device—may also be deployed to other mobile devices. Users of other mobile devices can access these deployed portions of the digital content using their respective mobile devices to explore the augmented reality or other extended reality environments.

[0020] Users of mobile devices can also explore and interact with the augmented reality environment (e.g., through HMD) by using gestures or verbal input. For example, the mobile device can apply gesture recognition tools to gesture input to determine the context of the gesture input and perform additional operations corresponding to the determined context. Gesture input can be detected by a digital camera incorporated in or communicating with the mobile device, or by various sensors incorporated in or communicating with the mobile device. Examples of these sensors include, but are not limited to, inertial measurement units (IMUs) incorporated in the mobile device or IMUs incorporated in wearable devices (e.g., gloves) and communicating with the mobile device. In other instances, a microphone or other interface within the mobile device can capture spoken words. The mobile device can apply speech recognition tools or natural language processing algorithms to the captured words to determine the context of the spoken words and perform additional operations corresponding to the determined context. For example, additional operations can include presenting additional elements of digital content corresponding to the user's navigation through the augmented reality environment. Other instances include processes for obtaining information from one or more computing systems in response to user-spoken queries.

[0021] Gestures or verbal input can request virtual content items, such as virtual assistants, within the extended reality environment created by the mobile device (e.g., a virtual assistant can be "called"). For example, a virtual assistant can contain elements of animated digital content and synthetic audio content, presented in an appropriate and context-sensitive section of the extended reality environment. When rendered via an HMD or mobile device, the animated digital content and audio content elements facilitate enhanced interaction between the user and the extended reality environment and maintain the user's connection to the "real" world (outside the extended reality environment). In some cases, a virtual assistant may be associated with a library or collection of synthetic audio content responding to the user's spoken words. Such audio content can describe objects of interest arranged within the extended reality environment, utterances throughout the user's interaction with the augmented reality environment indicating actual real-world time or date, or indications of hazards or other dangerous situations present in the user's actual environment.

[0022] When rendered and presented within an augmented reality environment, the virtual assistant can also interact with the user and elicit further gestures or verbal queries from the user's mobile device. For example, the augmented reality environment could contain an augmented reality setting corresponding to a meeting attended by multiple geographically dispersed colleagues of the user, and the virtual assistant could prompt the user to make a verbal query to the mobile device, requesting the virtual assistant to record the meeting for later review. In other instances, the virtual assistant accesses elements of digital content (e.g., video content, images, etc.) and presents that digital content in a presentation area of ​​the augmented reality environment (e.g., an augmented reality whiteboard).

[0023] Mobile devices can capture spoken queries via microphones or other interfaces and can apply one or more speech recognition tools or natural language processing algorithms to the captured queries to determine the context of the spoken queries. Virtual assistants can perform additional operations corresponding to the determined context, either independently or by exchanging data with one or more other computing systems. By facilitating interaction between the user and both the augmented reality environment and the real world, virtual assistants create immersive augmented reality experiences for users. Virtual assistants can increase the use of augmented reality technologies and promote multi-user collaboration in augmented reality environments.

[0024] In response to captured gestures or verbal input—which invoke a virtual assistant within the extended reality environment—an extended reality computing system or mobile device can determine a portion (e.g., a “scene”) of the virtual environment currently visible to the user via an HMD or augmented reality glasses-wearing device. The extended reality computing system or mobile device can apply various image processing tools to generate data (e.g., a scene depth map) that establishes and assigns a value representing the depth of the pixel’s corresponding location within the extended reality environment to each pixel within the scene. In some instances, the extended reality computing system or mobile device also applies various semantic scene analysis processes to the visible portion of the extended reality environment and the scene depth map to identify and characterize objects arranged within the extended reality environment and to map the identified objects to their locations within the extended reality environment and the scene depth map.

[0025] Extended reality computing systems or mobile devices can also determine the position and orientation of the mobile device in an extended reality environment (e.g., the position and orientation of an HMD or augmented reality glasses wearer) and can further obtain data indicating the position and orientation of one or more other users in the extended reality environment. In one example, the user's position in the extended reality environment can be determined based on the latitude, longitude, or altitude values ​​of the mobile device operated or worn by the user. Similarly, extended reality computing systems or mobile devices can define the user's orientation in the extended reality environment based on the determined orientation of the mobile device operated or worn by the user. For example, the determined orientation of the mobile device, such as device posture, can be established based on one or more of the device's roll, pitch, and / or yaw values. Further, the position and orientation of the user and the mobile device can be based on one or more positioning signals and / or inertial sensor measurements obtained or received at the mobile device, as described in detail below.

[0026] In another instance, the user's orientation in the extended reality environment may correspond to the orientation of at least a portion of the user's body in the extended reality environment. For example, the user's orientation may be determined based on the orientation of a portion of the user's body relative to a portion of the mobile device the user is operating or wearing (e.g., the orientation of a portion of the user's head relative to the display surface of the mobile device, such as head posture). In other instances, the mobile device may include or be incorporated into a head-mounted display, and the extended reality computing system or mobile device may define the user's orientation based on a determined orientation of at least one of the user's eyes relative to a portion of the head-mounted display. For example, the head-mounted display may include augmented reality glasses, and the extended reality computing system or mobile device may define the user's orientation as the orientation of the user's left (or right) eye relative to the corresponding lens of the augmented reality glasses. Based on the generated scene depth map, the results of semantic processes, and data characterizing the position and orientation of the user or mobile device in the extended reality environment, the extended reality computing system or mobile device may establish multiple candidate locations for a virtual assistant in the extended reality environment and may calculate a placement score characterizing the feasibility of each candidate location for the virtual assistant.

[0027] Additionally, the extended reality computing system or mobile device can calculate a placement score for a specific candidate location that reflects the physical constraints imposed by the extended reality environment. For example, the presence of a hazard (e.g., a lake, cliff, etc.) at a specific candidate location may result in a low placement score. In another instance, the extended reality computing system or mobile device may calculate a low placement score for said candidate location if an object placed at that candidate location (e.g., a bookshelf or table placed at the candidate location) is unsuitable for supporting a virtual assistant, while a chair placed at the candidate location may result in a high placement score. In other instances, the extended reality computing system or mobile device calculates the placement score for a specific candidate location based at least in part on additional factors. These additional factors may include, but are not limited to, the viewpoint of the virtual assistant at the specific candidate location relative to other users positioned in the extended reality environment, the displacement between the specific candidate location and the positions of other users in the augmented reality environment, the determined visibility of each of the other users' faces to the virtual assistant positioned at the specific candidate location, and the interactions between all or some of the users positioned in the augmented reality environment.

[0028] An augmented reality computing system or mobile device can determine the minimum calculated placement score and identify candidate locations associated with the minimum calculated placement score. The augmented reality computing system or mobile device selects (establishes) the identified candidate locations as the positions of virtual content items, such as virtual assistants, in the augmented reality environment.

[0029] As described in detail below, an augmented reality computing system or mobile device can generate virtual content items (e.g., virtual assistants generated using animation and speech synthesis tools). The augmented reality computing system or mobile device can generate instructions that cause a display unit (e.g., an HMD) to present the virtual content item at a corresponding location in the augmented reality environment. The augmented reality computing system or mobile device can modify one or more visual characteristics of the virtual content item in response to additional gestures or verbal input (such as input representing user interaction with a virtual assistant).

[0030] Extended reality computing systems or mobile devices can also modify the position of virtual content items in the extended reality environment in response to changes in the state of the mobile device or display unit (e.g., changes in the position or orientation of the mobile device or display unit within the extended reality environment). Extended reality computing systems or mobile devices can also modify the position of virtual content items based on the detection of gesture input that guides the virtual content item to an alternative location in the extended reality environment (e.g., a location near an object of interest to the user in the augmented reality environment).

[0031] Figure 1 is a schematic block diagram of an exemplary network environment 100. Network environment 100 may include, for example, any number of mobile devices, such as mobile devices 102 and 104. Mobile devices 102 and 104 may establish and enable access to an extended reality environment by a corresponding user. As described herein, examples of extended reality environments include, but are not limited to, virtual reality environments, augmented reality environments, or augmented virtual environments. Mobile devices 102 and 104 may include any suitable mobile computing platform, such as, but not limited to, cellular phones, smartphones, personal digital assistants, low duty cycle communication devices, laptop computers, portable media player devices, personal navigation devices, and portable electronic devices including digital cameras.

[0032] Furthermore, in some instances, mobile devices 102 and 104 also include (or correspond to) wearable extended reality display units, such as HMDs that present stereoscopic graphic and audio content to create an extended reality environment for the corresponding user. In other instances, mobile devices 102 and 104 include augmented reality eyewear (e.g., glasses) containing one or more lenses for displaying graphic content (such as an augmented reality information layer) over real-world objects visible through such lenses to create an augmented reality environment. Mobile devices 102 and 104 can be operated by corresponding users, each of whom can access the extended reality environment using any of the procedures described below, and the mobile device can be positioned at a corresponding location within the accessed extended reality environment.

[0033] Network environment 100 may include extended reality (XR) computing system 130, positioning system 150, and one or more additional computing systems 160. Mobile devices 102 and 104 may wirelessly communicate with XR computing system 130, positioning system 150, and additional computing systems 160 across communication network 120. Communication network 120 may include one or more wide area networks (e.g., the Internet), local area networks (e.g., intranets), and / or personal area networks. For example, mobile devices 102 and 104 may wirelessly communicate with XR computing system 130 and additional computing systems 160 via any suitable communication protocol, including cellular communication protocols (such as Code Division Multiple Access (CDMA®), Global System for Mobile Communications (GSM®), or Wideband Code Division Multiple Access (WCDMA®)) and / or wireless LAN protocols (such as IEEE 802.11 (WiFi®) or Global Microwave Access Interoperability (WiMAX®)). Therefore, the communication network 120 may include one or more wireless transceivers. Mobile devices 102 and 104 may also use the wireless transceivers of the communication network 120 to obtain positioning information for estimating the location of the mobile devices.

[0034] Mobile devices 102 and 104 can estimate their corresponding geographic locations using trilateration-based methods. For example, techniques that mobile devices 102 and 104 can use include Advanced Forward Link Trilateration (AFLT) in CDMA®, Enhanced Observation Time Difference (EOTD) in GSM®, or Observed Time Difference of Arrival (OTDOA) in WCDMA®. OTDOA measures the relative time a radio signal arrives at the mobile device, where the radio signal is transmitted from each of several base stations equipped with transmitters. As another example, mobile devices 102 or 104 can estimate their own location by obtaining a Media Access Control (MAC) address or other suitable identifier associated with a radio transceiver and relating the MAC address or identifier to the known geographic location of the radio transceiver.

[0035] Mobile devices 102 or 104 can further obtain wireless positioning signals from positioning system 150 to estimate the corresponding mobile device location. For example, positioning system 150 may include a satellite positioning system (SPS) and / or a terrestrial positioning system. Satellite positioning systems may include, for example, Global Positioning System (GPS), Galileo, GLONASS, NAVSTAR, Global Navigation Satellite System (GNSS), systems using satellites from a combination of the aforementioned positioning systems, or any SPS developed in the future. As used herein, SPS may include pseudosatellite systems. The specific positioning techniques described herein are merely examples of positioning techniques and do not limit the claimed subject matter.

[0036] XR computing system 130 may include one or more servers and / or other suitable computing platforms. Therefore, XR computing system 130 may include non-transitory computer-readable storage media (“storage media”) 132 on which database 134 and instructions 136 are stored. XR computing system 130 may include one or more processors, such as processor 138 for executing instructions 136 or facilitating the storage and retrieval of data in database 134. XR computing system 130 may further include a communication interface 140 for facilitating communication with clients of communication network 120, said clients including mobile devices 102 and 104, positioning system 150, and additional computing system 160.

[0037] To facilitate understanding of the examples, instructions 136 are sometimes described in relation to one or more modules used to perform specific operations. As an example, instructions 136 may include a content management module 162 for managing elements of digital content, such as digital graphics and audio content, deployed to mobile devices 102 and 104. The graphics and audio content may include captured digital video, digital images, digital audio, or synthesized images or videos. Mobile devices 102 or 104 may present portions of the deployed graphics or audio content via corresponding display units (such as lenses of an HMD or augmented reality glasses) and may establish an extended reality environment at each mobile device in mobile devices 102 and 104. The established extended reality environment may contain graphics or audio content that allows users of mobile devices 102 or 104 to browse and explore various historical sites and locations within the extended reality environment or participate in virtual meetings attended by geographically dispersed participants.

[0038] Instruction 136 may also include an image processing module 164 for processing images representing portions of the extended reality environment visible to the user of mobile device 102 or mobile device 104 (e.g., via a corresponding HMD or via lenses of an augmented reality eye-wearing device). For example, image processing module 164 may, among other things, include a depth mapping module 166 for generating depth maps for each visible portion of the extended reality environment and a semantic analysis module 168 for identifying and characterizing objects arranged within the visible portions of the extended reality environment.

[0039] As an example, the depth mapping module 166 can receive images representing portions of the extended reality environment visible to a user of the mobile device 102 via a corresponding HMD. The depth mapping module 166 can generate a depth map that associates each pixel of one or more images with a corresponding location in the extended reality environment. The depth mapping module 166 can calculate a value characterizing the depth of each location in the corresponding location in the extended reality environment and associate the calculated depth with the corresponding location in the depth map. In some instances, the received images comprise stereo image pairs, each representing a visible portion of the extended reality environment from slightly different viewpoints (e.g., corresponding to the left and right lenses of the HMD or augmented reality glasses). Further, each location in the visible portion of the extended reality environment can be characterized by an offset (measured in pixels) between the two images. The offset is proportional to the distance between the position of the mobile device 102 in the extended reality environment and the user.

[0040] The depth mapping module 166 can further establish pixel offsets as depth values ​​representing positions within the generated depth map. For example, the depth mapping module 166 can establish values ​​proportional to the pixel offsets as depth values ​​representing positions within the depth map. The depth mapping module 166 can also establish a mapping function (e.g., a feature-depth mapping function) that correlates certain visual characteristics of the image of the visible portion of the extended reality environment with corresponding depth values ​​as illustrated in the depth map. The depth mapping module 166 can further process the data representing the image to identify the color value of each image pixel. The depth mapping module 166 can apply one or more suitable statistical techniques (e.g., regression) to the identified color values ​​and the depth values ​​illustrated in the depth map to generate the mapping function and correlate the pixel color values ​​with corresponding depths in the visible portion of the extended reality environment.

[0041] The subject matter of this invention is not limited to the examples of the depth mapping process described above, and the depth mapping module 166 may apply other or alternative image processing techniques to an image to generate a depth map representing the visible portion of an extended reality environment. For example, the depth mapping module 166 may process a portion of an image to determine its similarity to previous image data representing a previously visible portion of the extended reality environment (e.g., visible to a user of mobile device 102 or 104 via a corresponding HMD). In response to the determined similarity, the depth mapping module 166 may access database 134 and obtain data for a mapping function specifying the previously visible portion of the extended reality environment. The depth mapping module 166 may determine the color values ​​of the pixels representing said portion of the image, apply the mapping function to the determined color values, and directly generate a depth map of said portion of the image based on the output of the applied mapping function.

[0042] Referring back to Figure 1, the semantic analysis module 168 can process an image (e.g., the image representing a visible portion of an extended reality environment) and apply one or more semantic analysis techniques to identify and characterize objects arranged within the visible portion of the extended reality environment. For example, the semantic analysis module 168 can access image data associated with a corresponding image in the image and can apply one or more suitable computer vision or machine vision algorithms to the accessed image data. The computer vision or machine vision algorithm identifies the location of objects within the corresponding image in the image and the identified objects within the corresponding image in the image, thus identifying the location of the identified objects within the visible portion of the extended reality environment.

[0043] The applied computer vision or machine vision algorithms may rely on data stored locally by the XR computing system 130 (e.g., within database 134). The semantic analysis module 168 can obtain data supporting the application of computer vision or machine vision algorithms from one or more computing systems (such as another computing system 160) across the communication network 120. For example, the semantic analysis module 168 can perform operations via corresponding program interfaces to provide data to a corresponding computing system in another computing system 160 to facilitate image-based searches of various portions of the accessed image data. Examples are not limited to the semantic analysis techniques and image-based searches described above. The semantic analysis module 168 can, alone or in conjunction with another computing system 160, further apply additional or alternative algorithms or techniques to the obtained image data to identify and locate objects within the visible portion of the extended real-world environment.

[0044] Based on the results of applied computer vision or machine vision algorithms or image-based search results, the semantic analysis module 168 can generate data (e.g., metadata) for each identified object and its corresponding location within the visible portion of the augmented reality environment. The semantic analysis module 168 can also access a generated depth map of the visible portion of the augmented reality environment and correlate depth values ​​characterizing the visible portion with the identified objects and their corresponding locations.

[0045] Referring back to Figure 1, instruction 136 may further include a location determination module 170, a virtual content generation module 172, and a query processing module 174. The location determination module 170 can perform operations to establish multiple candidate locations for virtual content items, such as virtual assistants, within the extended reality environment established by mobile devices 102 and 104. The location determination module 170 provides means for determining the placement position of the virtual content item in the extended reality environment, at least in part, based on the user's determined location and orientation. The location determination module 170 can calculate a placement score characterizing the feasibility of each candidate location among candidate locations in the augmented or other reality environment. As described below, the placement score can be calculated based on generated depth map data specifying objects within the visible portion of the extended reality environment, data characterizing a portion and orientation of each user in the extended reality environment, or data characterizing the level of interaction between users in the extended reality environment.

[0046] The placement score calculated for a specific candidate location can reflect the physical constraints imposed by the extended reality environment. As mentioned above, the presence of hazards (e.g., lakes, cliffs, etc.) at a specific candidate location may result in a low placement score, and the extended reality computing system can calculate a low placement score for said candidate location if the object placed at the specific candidate location (e.g., a bookshelf or table placed at the candidate location of the virtual assistant) is not suitable for supporting the virtual content item. The placement score calculated for a specific candidate location can also reflect other factors (e.g., the virtual assistant's viewpoint relative to other users in the extended reality environment when the virtual assistant is at the specific candidate location, the displacement between the specific candidate location and the positions of other users in the extended reality environment, the determined visibility of each user's face to the virtual assistant placed at the specific candidate location, and / or the interaction between all or some users in the extended reality environment, or other augmented or other extended reality interactions). The location determination module 170 can determine the minimum calculated placement score, identify candidate locations associated with the minimum calculated placement score, and select and establish the identified candidate locations as the position of the virtual content item in the extended reality environment.

[0047] The virtual content generation module 172 can perform operations to generate and instantiate virtual content items, such as a virtual assistant, at an established location in the extended reality environment (e.g., visible to the user of mobile device 102 or mobile device 104). For example, the virtual content generation module 172 may include a graphics module 176 that generates an animated representation of the virtual assistant based on locally stored data (e.g., within database 134) specifying the visual characteristics of the virtual assistant, such as the visual characteristics of the avatar selected by the user of mobile device 102 or mobile device 104. The virtual content generation module 172 may also include a speech synthesis module 178 that generates audio content representing portions of an interactive dialogue spoken by the virtual assistant in the extended reality environment. In some instances, the speech synthesis module 178 generates the audio content and portions of the interactive spoken dialogue based on speech parameters locally stored by the XR computing system 130. For example, the speech parameters may specify a regional dialect or the language spoken by the user of mobile device 102 or 104.

[0048] The query processing module 174 can perform operations such as receiving query data from mobile device 102 or mobile device 104 specifying one or more queries. The query processing module 174 can obtain data in response to queries (e.g., data locally stored in storage medium 132 or obtained from another computing system 160 across communication network 120). The query processing module 174 can provide the obtained data to mobile device 102 or mobile device 104 in response to the received query data. For example, the extended reality environment established by mobile device 102 can include a virtual tour of the pyramid complex at Giza, Egypt, and in response to the synthesized voice of a virtual assistant, a user of mobile device 102 can issue a query requesting additional information about the construction practices used during the construction of the Great Pyramid. As described below, the voice recognition module of mobile device 102 can process the issued query—for example, using any suitable voice recognition algorithm or natural language processing algorithm—and generate text query data; the query module in mobile device 102 can package the text query data and transmit the text query data to XR computing system 130.

[0049] As described above, query processing module 174 can receive query data, generate or obtain data reflecting and responding to the received query data (e.g., based on data locally stored in storage medium 132 or obtained from another computing system 160). For example, query processing module 174 can execute a request and obtain information characterizing the construction techniques used during the construction of the Great Pyramid from one or more computing systems in the other computing system 160 (e.g., via a suitable program interface). Query processing module 174 can then transmit the obtained information to mobile device 102 as a response to the query data.

[0050] Before transmitting a response to mobile device 102, speech synthesis module 178 can access and process the acquired information to generate audio content that the virtual assistant can present to the user of mobile device 102 in an extended reality environment (e.g., the virtual assistant “speaks” in response to the user’s query).

[0051] The query processing module 174 can also transmit the obtained information to the mobile device 102 without further processing or speech synthesis. A local speech synthesis module maintained by the mobile device 102 can process the obtained information and generate synthesized speech that the virtual assistant can present using any of the processes described herein.

[0052] Database 134 may contain various types of data, such as media content data 180—e.g., captured digital video, digital images, digital audio, or synthetic images or videos suitable for deployment to mobile device 102 or mobile device 104—to establish a corresponding scenario for the extended reality environment (e.g., based on operations performed by content management module 162). Database 134 may also contain depth map data 182, which includes a depth map and data specifying a mapping function for the corresponding visible portion of the extended reality environment instantiated by mobile device 102 or 104. Database 134 may also contain object data 184, which includes metadata identifying objects and their positions within the corresponding visible portions (and further includes data relating the object's position to the corresponding portion of the depth map data 182).

[0053] Database 134 may also contain position and orientation data 186 identifying the position and orientation of users of mobile devices 102 and 104 within corresponding portions of the extended reality environment, and interaction data 188 characterizing the level or range of interaction between users in the extended reality environment. For example, the position of a mobile device may be represented as one or more latitude, longitude, or altitude values ​​measured relative to a reference reference, and the position of the mobile device may represent the position of the user of the mobile device. Further, and as described herein, the orientation of the mobile device may be represented by one or more of roll, pitch, and / or yaw values ​​measured relative to another or alternative reference reference, and the user's orientation may be represented as the orientation of the mobile device. In other instances, the user's orientation may be determined based on the orientation of at least a portion of the user's body relative to the extended reality environment or relative to a portion of the mobile device (such as the display surface of the mobile device). Further, in some instances, interaction data 188 may characterize the amount of audio communication between users in the augmented or other extended reality environment, as well as the amount of audio communication between the source user and the target user, as monitored and captured by the XR computing system 130.

[0054] Database 134 may further include graphic data 190 and voice data 192. Graphic data 190 may include data that facilitates and supports the generation of virtual content items, such as virtual assistants, by virtual content generation module 172. For example, graphic data 190 may include, but is not limited to, data specifying certain visual characteristics of the virtual assistant (such as the visual characteristics of an avatar selected by the user of mobile device 102 or mobile device 104). Further, for example, voice data 192 may include data that facilitates and supports the synthesis of speech suitable for presentation by the virtual assistant in an extended reality environment (such as, for example, a regional dialect or the language spoken by the user of mobile device 102 or 104).

[0055] Figure 2 is a schematic block diagram of an exemplary mobile device 200. The mobile device 200 is for at least some examples. Figure 1Non-limiting examples of mobile devices 102 and 104. Thus, mobile device 200 may include, for example, a communication interface 202 for facilitating communication with other computing platforms such as XR computing system 130, other mobile devices (e.g., mobile devices 102 and 104), positioning system 150, and / or additional computing system 160. Therefore, communication interface 202 can enable wireless communication with communication networks such as communication network 120. Mobile device 200 may also include a receiver 204 (e.g., a GPS receiver or SPS receiver) for receiving positioning signals from a positioning system such as positioning system 150 as shown in FIG. 1. Receiver 204 provides means for determining the location of mobile device 200—and therefore the user wearing or carrying the mobile device—in the extended reality environment.

[0056] Mobile device 200 may include one or more input units, such as input unit 206, for receiving input from a corresponding user. Examples of input unit 206 include, but are not limited to, one or more physical buttons, keyboards, controllers, microphones, pointing devices, and / or touch-sensitive surfaces.

[0057] Mobile device 200 may take the form of a wearable extended reality display unit, such as a head-mounted display (HMD) that presents stereoscopic graphic and audio content that establishes an extended reality environment. Mobile device 200 may also take the form of augmented reality glasses or eyeglasses, which include one or more lenses for displaying graphic content (such as an augmented reality information layer over real-world objects visible through the lenses and establishing the augmented reality environment). Mobile device 200 may include a display unit 208, such as a stereoscopic display that displays graphic content to a corresponding user, such that the graphic content establishes an augmented reality environment or other extended reality environment at mobile device 200. Display unit 208 may be incorporated into the augmented reality glasses or eyeglasses and may be further configured to display an augmented reality information layer overlaid on real-world objects visible through a single lens or alternatively through two lenses. Mobile device 200 may also include one or more output devices (not shown), such as audio speakers or headphone jacks for presenting audio content as part of the extended reality environment.

[0058] Mobile device 200 may include one or more inertial sensors 210 that collect inertial sensor measurements characterizing mobile device 200. The inertial sensors 210 provide means for determining the orientation of mobile device 200—and therefore, a user wearing or carrying mobile device 200—in an extended reality environment. Examples of suitable inertial sensors 210 include, but are not limited to, accelerometers, gyroscopes, or other suitable means for measuring the inertial state of mobile device 200. The inertial state of mobile device 200 may be measured by inertial sensors 210 along multiple axes in Cartesian and / or polar coordinate systems to provide an indication for determining the position or orientation of mobile device 200. Mobile device 200 may also process (e.g., integrate over time) data indicating the inertial sensor measurements obtained from inertial sensors 210 to generate an estimate of the position or orientation of the mobile device. As discussed above, the position of the mobile device 200 can be specified using latitude, longitude, or altitude values, and the orientation of the mobile device 200 can be specified using roll, pitch, or yaw values ​​measured relative to reference values.

[0059] Mobile device 200 may include a digital camera 212 configured to capture digital image data identifying one or more gestures or movements of a user (such as a predetermined gesture formed by the user's hand or fingers, or a pointing movement performed by the user's hand and arm). Digital camera 212 may include a digital camera having multiple optical elements (not shown). The optical elements may include one or more lenses for focusing light and / or one or more light-sensing elements for converting light into digital signals representing image and / or video data. As a non-limiting example, the light-sensing elements may include an optical pickup, a charge-coupled device, and / or an optoelectronic device for converting light into digital signals. As described below, mobile device 200 may be configured to detect one of a movement or gesture based on the digital image data captured by digital camera 212, identify an operation associated with the corresponding movement or gesture, and initiate the execution of said operation in response to the detected movement or gesture. For example, such an operation may invoke, undone, or reposition virtual content items such as virtual assistants in an extended reality environment.

[0060] Mobile device 200 may further include a non-transitory computer-readable storage medium (“storage medium”) 211 on which database 214 and instructions 216 are stored. Mobile device 200 may include one or more processors, such as processor 218, for executing instructions 216 and / or facilitating the storage and retrieval of data at database 214 to perform a computer-implemented extended reality method. Database 214 may contain various data, including some or all of the data elements described above with reference to database 134 of FIG1. ​​When displayed to a user via display unit 208, database 214 may also maintain media content data 220, which includes elements of captured digital video, digital images, digital audio, or synthesized images or videos used to establish an extended reality environment at mobile device 200. Mobile device 200 may further receive portions of media content data 220 from XR computing system 130 at regular predetermined intervals (e.g., a “push” operation) or in response to a request transmitted from mobile device 200 to XR computing system 130 (e.g., a “pull” operation) (e.g., via communication interface 202 using any suitable communication protocol).

[0061] Database 214 may contain depth map data 222, which includes a depth map and data specifying a mapping function for the corresponding visible portion of the extended reality environment instantiated by mobile device 200. Database 214 may also contain object data 224, which includes metadata identifying objects and their positions within the corresponding visible portions (and data relating the object's position to the corresponding portion of depth map data 222). Mobile device 200 may receive portions of depth map data 222 or object data 224 from XR computing system 130 (e.g., generated by depth mapping module 166 or semantic analysis module 168, respectively). In other cases, processor 218, described in more detail below, may execute portions of instructions 216 to generate local portions of depth map data 222 or object data 224.

[0062] Furthermore, similar to the portions of database 134 described above, database 214 may maintain a local copy of location and orientation data 226, which identifies the location and orientation of a user of mobile device 200 (and other mobile devices within network environment 100, such as mobile devices 102 and 104) in the corresponding portion of the extended reality environment. Database 214 may also maintain a local copy of interaction data 228 characterizing the level or range of interaction between users in the extended reality environment. Database 214 may further include local copies of graphics data 230 and voice data 232. In some instances, the local copy of graphics data 230 contains data that facilitates and supports a virtual assistant generated by mobile device 200 (e.g., by executing portions of instruction 216). The local copy of graphics data 230 may include, but is not limited to, data specifying the visual characteristics of the virtual assistant (such as the visual characteristics of an avatar selected by the user of mobile device 200). Furthermore, as described above, the local copy of the voice data 232 may contain data that facilitates and supports the synthesis of voice (such as a regional dialect or a language spoken by the user of the mobile device 200) suitable for presentation by the virtual assistant once instantiated in an extended reality environment.

[0063] Additionally, database 214 may also include a gesture library 234 and a verbal input library 236. Gesture library 234 may contain data identifying one or more candidate gesture inputs (e.g., gestures, pointing movements, facial expressions, etc.). The gesture input data associates the candidate gesture inputs with operations such as invoking a virtual assistant (or other virtual content item) in an extended reality environment, resuming the virtual assistant or other virtual content item, or repositioning a virtual assistant or virtual content item in an extended reality environment. Verbal input library 236 may further include text data representing one or more candidate verbal inputs and additional data associating the candidate verbal inputs with certain operations such as invoking a virtual assistant or virtual content item, revoking a virtual assistant or virtual content item, repositioning a virtual assistant or virtual content item, or requesting schedule data. The subject matter of this invention is not limited to the examples of the related operations described above, and in other cases, the gesture library 234 and the verbal input library 236 may contain data relating any additional or alternative gestures or verbal inputs detectable by the mobile device 200 to any additional or alternative operations performed by the mobile device 200 or the XR computing system 130.

[0064] Instruction 216 may include one or more of the modules and / or tools of instruction 136 described above with reference to FIG. 1. For the sake of brevity, the description of the common modules and / or tools included in both FIG. 1 and 2 will not be repeated. For example, instruction 216 may include an image processing module 164, which may further include a depth mapping module 166 and a semantic analysis module 168. Further, instruction 216 may also include a location determination module 170, a virtual content generation module 172, a graphics module 176, and a speech synthesis module 178, as described above with reference to the modules in instruction 136 of FIG. 1.

[0065] In another instance, instruction 216 also includes an extended reality building module, such as XR building module 238, for accessing portions of storage media 211 (e.g., media content data 220) and extracting elements of captured digital video, digital images, digital audio, or composite images. When executed by processor 218, XR building module 238 enables mobile device 200 to render and present portions of captured digital video, digital images, digital audio, or composite images to the user via display unit 208, thus establishing an extended reality environment for the user at mobile device 200.

[0066] Instruction 216 may also include a gesture detection module 240, a speech recognition module 242, and an operation module 244. In some instances, the gesture detection module 240 accesses digital image data captured by the digital camera 212 and may apply one or more image processing techniques, computer vision algorithms, or machine vision algorithms to detect human gestures arranged within a single frame of the digital image data. The digital image data may contain gestures established by the user's fingers. Alternatively or additionally, the digital image data may contain human movements occurring across multiple frames of the digital image data (e.g., pointing movements of the user's arm, hand, or fingers).

[0067] Operation module 244 can acquire data indicating detected gestures or movements and associate operations with the detected gestures or movements based on identifiers in gesture library 234. In response to an identifier, operation module 244 can initiate an associated operation performed by mobile device 200. As described above, the detected gestures or movements may correspond to a request from the user to reposition invoked digital content items, such as virtual assistants, to another part of the extended reality environment. That is, gesture detection module 240 can detect gestures that provide additional information about objects placed in the extended reality environment. Operation module 244 can cause mobile device 200 to perform operations that modify the placement position of the virtual assistant in the extended reality environment based on the detected gesture input.

[0068] In other cases, the speech recognition module 242 can access audio data containing utterances spoken by the user and captured by a microphone incorporated into the mobile device 200. The speech recognition module 242 can apply one or more speech recognition algorithms or natural language processing algorithms to parse the audio data and generate text data corresponding to the spoken utterances. The operation module 244 can identify operations associated with all or part of the generated text data based on various identifiers in the speech input library 236 and initiate operations performed by the mobile device 200. For example, the generated text data could correspond to the user's utterance "Open the virtual assistant." The operation module 244 can associate the utterance with a user's request to invoke the virtual assistant and place it at a location that conforms to constraints imposed by the extended reality environment and is contextually relevant to and enhances the user's exploration of the augmented reality environment or other extended reality environments and interaction with other users.

[0069] Figure 3 is a flowchart of an instance process 300 for dynamically locating a virtual assistant in an extended reality environment according to one embodiment. Process 300 can be performed by a mobile device (e.g., Figure 2 The mobile device 200 executes the instructions locally on one or more processors. Certain blocks of process 300 can be executed remotely by one or more processors of a server system or other suitable computing platform (such as the XR computing system 130 in Figure 1). Therefore, the various operations of process 300 can be represented by executable instructions stored in storage media of one or more computing platforms (such as storage media 132 of the XR computing system 130 and / or storage media 211 of the mobile device 200).

[0070] Referring to Figure 3, mobile device 200 can establish an extended reality environment, such as an augmented reality environment (e.g., in box 302). XR computing system 130 can access stored media content data 220, which includes elements of captured digital video, digital images, digital audio, or synthesized images or videos. When executed by processor 138, content management module 162 can package portions of captured digital video, digital images, digital audio, or synthesized images or videos into corresponding data packages, which XR computing system 130 can transmit to mobile device 200 across communication network 120. Mobile device 200 can receive the transmitted data packets and store portions of captured digital video, digital images, digital audio, or synthesized images or videos in storage media 211, such as media content data 220.

[0071] When executed by processor 218, XR creation module 238 enables mobile device 200 to render and present portions of captured digital video, digital images, digital audio, or synthesized images to the user via display unit 208, thus creating an augmented reality environment for the user at mobile device 200. As described above, the user can access the created augmented reality environment via head-mounted display (HMD), which can present a stereoscopic view of the created augmented reality environment to the user. For example, as shown in FIG4A, stereoscopic view 402 may include a pair of stereoscopic images, such as image 404 and image 406, which create the visible portion of the augmented reality environment when displayed to the user through corresponding lenses in the left and right lenses of the HMD, enabling the user to perceive depth in the augmented reality environment.

[0072] Referring back to Figure 3, mobile device 200 can receive user input that invokes and presents virtual content items, such as a virtual assistant, at a corresponding placement location within the visible portion of the augmented reality environment (e.g., in box 304). The received user input may include spoken words (e.g., “Open the virtual assistant”), captured via a microphone incorporated into mobile device 200. As described above, speech recognition module 242 can access the captured audio data and apply one or more speech recognition or natural language processing algorithms to parse the audio data and generate text data corresponding to the spoken words. Further, operation module 244 can access information associating the text data with corresponding operations (e.g., within the verbal input library 236 of Figure 2) and can confirm that the spoken words represent a request to invoke the virtual assistant. In response to the established request, mobile device 200 can execute any of the exemplary processes described herein to generate the virtual assistant, determine a placement location for the virtual assistant that conforms to any constraints that may be imposed by the augmented reality environment, and insert the virtual assistant into the augmented reality environment at the placement location.

[0073] In other cases, mobile device 102 may transmit data specifying a request to XR computing system 130 across network 120. XR computing system 130 may execute any of the exemplary procedures described herein to generate a virtual assistant, determine a placement location for the virtual assistant that conforms to any constraints that may be imposed by the augmented reality environment (or any other extended reality environment), and transmit the data specifying the generated virtual assistant and the determined placement location to mobile device 102 across network 120. Mobile device 102 may then display the virtual assistant as a virtual content item at the placement location in the augmented reality environment.

[0074] The subject matter of this invention is not limited to verbal input that invokes the placement of virtual content items in an augmented reality environment when detected by mobile device 200. In other instances, the digital camera 212 of mobile device 102 can capture digital image data containing gesture input, such as hand gestures or pointing movements, provided by the user to mobile device 200. Mobile device 200 can perform any of the processes described herein to identify the gesture input, calculate the correlation between the identified gesture input and the invocation of the virtual assistant (e.g., based on portions of a stored gesture library 234), and recognize based on the correlation that the gesture input represents a request to invoke the virtual assistant. Further, in other instances, the user can provide additional or alternative input to the input unit 206 of mobile device 200 (e.g., the one or more physical buttons, keyboard, controller, microphone, pointing device, and / or touch-sensitive surface) requesting the invocation of the virtual assistant.

[0075] In response to a detected request, mobile device 200 or XR computing system 130 can determine the location and orientation (e.g., in box 306) of one or more users accessing a visible portion of the augmented reality environment. These one or more users may include a first user operating mobile device 200 (e.g., a user wearing an HMD or augmented reality glasses) and one or more second users operating mobile device 102, mobile device 104, or other mobile devices in network environment 100. As described above, mobile device 200 can determine its own location and thus the location of its users based on positioning signals received from positioning system 150 (e.g., via receiver 204).

[0076] Furthermore, the mobile device 200 can determine its orientation (i.e., the orientation of the mobile device itself). In some instances (such as glasses), there is a fixed relationship between the orientation of the mobile device 200 and the orientation of the user. Therefore, the mobile device 200 can determine the user's orientation based on inertial sensor measurements obtained from one or more inertial sensors in the inertial sensor 210. The position of the mobile device 200 can be represented as one or more latitude, longitude, or altitude values. As shown in Figure 4B, the orientation of the mobile device 200 can be represented by a roll value (e.g., an angle about the longitudinal axis 414), a pitch value (e.g., an angle about the transverse axis 416), and a yaw value (e.g., an angle about the vertical axis 412). The orientation of the mobile device 200 establishes a vector 418 specifying the direction the user of the mobile device 200 faces when accessing the augmented reality environment. In other cases (not shown in Figure 4B), the mobile device 200 or the XR computing system 130 may determine the user's orientation based on the orientation of at least a part of the user's body (e.g., the user's head or eyes) relative to the augmented reality environment or a part of the mobile device 200 (e.g., the surface of the display unit 208).

[0077] Mobile device 200 can also receive data representing corresponding locations (e.g., latitude, longitude, or altitude) and corresponding orientations (e.g., roll, pitch, and yaw values) from mobile device 102, mobile device 104, or other mobile devices (not shown) via communication interface 202 across communication network 120. Mobile device 102 can further receive all or part of the data representing the location and orientation of mobile device 102, mobile device 104, or other mobile devices from XR computing system 130 (e.g., maintained in location and orientation data 186). Mobile device 200 can store the data representing the location and orientation of mobile device 200, mobile device 102 or 104, and other mobile devices, as well as additional data uniquely identifying the corresponding mobile device among mobile device 200, mobile device 102 or 104, and other mobile devices (e.g., MAC address, Internet Protocol (IP) address, etc.) in a corresponding portion of database 214 (e.g., in location and orientation data 226). In other cases, the XR computing system 130 can receive data characterizing the position or orientation of mobile device 200, mobile device 102 or 104 and other mobile devices and can store the received data in a corresponding part of the database 134 (e.g., in position and orientation data 186).

[0078] Mobile device 200 or XR computing system 130 may also identify user-visible portions of the augmented reality environment to mobile device 200 (e.g., in box 308) and obtain depth map data characterizing the visible portions of the augmented reality environment (e.g., in box 310). When executed by processor 138 or 218, depth mapping module 166 may receive images representing said portions of the augmented reality environment visible to the user of mobile device 200, such as stereoscopic images 404 and 406 of FIG4A, and may perform any of the example processes described herein to generate depth maps of the visible portions of the augmented reality environment. Mobile device 200 may further provide image data characterizing the visible portions of the augmented reality environment (e.g., stereoscopic images 404 and 406) to XR computing system 130. XR computing system 130 may then execute depth mapping module 166 to generate depth maps of the visible portions of the augmented reality environment and may perform operations to store the generated depth map data in a portion of database 134 (e.g., depth map data 182). In another scenario, the XR computing system 130 can transmit the generated depth map to the mobile device 102 across the communication network 120. The mobile device 200 can perform operations to store the generated depth map data, or alternatively, the received depth map data, in a portion of the database 214 (such as depth map data 222).

[0079] As described above, the generated or received depth map can associate calculated depth values ​​with corresponding locations in the visible portion of the augmented reality environment. This technique can locate common points in each of the stereo images 404 and 406, determine the offset (in pixels) between corresponding locations of points in each of the corresponding images 404 and 406, and set the depth value of the point in the depth map to be equal to the determined difference. Alternatively, the depth mapping module 166 can establish a value proportional to the pixel offset as a depth value characterizing a specific location within the depth map.

[0080] Referring back to Figure 3, the mobile device 200 or XR computing system 130 can identify one or more objects arranged within the visible portion of the augmented reality environment (e.g., in box 312). When executed by processor 138 or 218, the semantic analysis module 168 can access images representing the visible portion of the augmented reality environment (e.g., stereoscopic images 404 and 406). The semantic analysis module 168 can also apply one or more of the aforementioned semantic analysis techniques to the accessed images to identify physical objects within the accessed images, identify the location, type, and size of the identified objects within the accessed images, and thus identify the location and size of the identified physical objects within the visible portion of the augmented reality environment.

[0081] The semantic analysis module 168 can apply one or more of the semantic analysis techniques described above to stereoscopic images 404 and 406, which correspond to the portion of the augmented reality environment visible to the user of the mobile device 200 via the display unit 208 (e.g., HMD). Figure 5 shows a plan view of the visible portion of the augmented reality environment (e.g., derived from stereoscopic images 404 and 406 and generated depth map data). As shown in Figure 5, the applied semantic analysis techniques can identify multiple pieces of furniture, such as sofa 512, coffee tables 514A and 514B, chair 516, and table 518, within the visible portion of the augmented reality environment, as well as another user 520 (e.g., a user operating mobile device 102 or mobile device 104) positioned between sofa 512 and chair 516. Mobile device 200 or XR computing system 130 can perform operations that store semantic data representing identified physical objects (e.g., object types, such as furniture) and the position and size of the identified objects within the visible portion of the augmented reality environment in the corresponding portions of databases 214 or 134 (e.g., within the corresponding object data in object data 224 or 184). Figure 5 also shows (in dashed lines) multiple candidate locations 522, 524, and 526 for the virtual assistant described below.

[0082] In some instances, the XR computing system 130 can generate all or part of the data characterizing physical objects identified within the visible portion of the augmented reality environment, as well as the location and size of the identified objects. The mobile device 200 can provide the XR computing system 130 with image data characterizing the visible portion of the augmented reality environment (e.g., stereoscopic images 404 and 406), which can access stored depth map data 182 and perform semantic analysis module 168 to identify physical objects within the visible portion of the augmented reality environment, as well as their location and size. The XR computing system 130 can store the data characterizing the identified physical objects and their corresponding location and size in object data 184. Additionally, the XR computing system 130 can also transmit portions of the stored object data across communication network 120 to the mobile device 200, for example, as described above, for storage in object data 224.

[0083] When executed by processor 138 or processor 218, the location determination module 170 can establish multiple candidate locations (e.g., candidate locations 522, 524, and 526 for the virtual assistant in Figure 5) within the visible portion of the augmented reality environment based on the generated depth map data 222 and the identified physical objects (as well as their positions and dimensions) (e.g., in box 314). The location determination module 170 can also calculate a placement score characterizing the feasibility of each of the candidate locations 522, 524, and 526 for the virtual assistant within the augmented reality environment (e.g., in box 316).

[0084] As schematically shown in Figure 6A, the visible portion 600 of the augmented reality environment may include user 601 corresponding to the user of mobile device 200 and user 602 corresponding to the user of mobile device 102 or 104. Users 601 and 602 may be arranged within the visible portion 600 at corresponding positions and may be oriented in the direction specified by the corresponding vectors in vectors 601A and 602A. Additionally, an object 603 (such as sofa 512 or chair 516 as shown in Figure 5) may be arranged between users 601 and 602 within the augmented reality environment. Furthermore, the position determination module 170 may establish candidate positions for virtual assistants within the visible portion 600, as represented by candidate virtual assistants 604A, 604B, and 604C in Figure 6A. Although candidate virtual assistants 604A, 604B, and 604C are depicted as rectangular parallelepipeds in Figure 6A, the graphics module 176 can generate virtual assistants of any shape or use any avatar to depict the virtual assistants.

[0085] The location determination module 170 can calculate each candidate location of the virtual assistant within the visible portion of the augmented reality environment according to the following formula, for example, p. asst Placement score, for example, cost(p) asst ):

[0086]

[0087] For example, and for the corresponding candidate position among the candidate positions, c cons The value of reflects the physical constraints imposed by the augmented reality environment (such as hazards like lakes or cliffs) at the corresponding candidate location. Similarly, c supp The value represents the ability of the physical object placed at the corresponding candidate location to support the virtual assistant. If the corresponding candidate location is located in a lake or other natural hazard, the location determination module 170 can assign an infinite value to c in the absence of hazard. cons (i.e., c) cons = ∞) or any value much larger than any expected score, to indicate that the virtual assistant cannot be placed at the corresponding candidate position.

[0088] Similarly, if the corresponding candidate position coincides with a table or other physical object that cannot support the virtual assistant (e.g., candidate position 526 on coffee table 514A in Figure 5, or position on coffee table 514B or chair 516), then the position determination module 170 can assign an infinite value to c. supp (i.e., c) supp = ∞). Alternatively, if the corresponding candidate position matches a physical object capable of supporting the virtual assistant (e.g., candidate position 524 on sofa 512 or position on chair 516 in Figure 5), the position determination module 170 can assign a value of zero to c. sup (i.e., c) supp = 0), which indicates that the virtual assistant can be positioned at the corresponding candidate location and supported by a consistent object. The c assigned to each candidate location by the location determination module 170 cons value and c supp The value ensures that the final position of the virtual assistant, which has a visible portion of the augmented reality environment, corresponds to the real-world constraints imposed on each user in the augmented reality environment.

[0089] Referring to Figures 6B-6D, c angle c dis and c visThis represents the user-specific values ​​for each user 601 and the candidate virtual assistant 604A at each corresponding candidate location. Figure 6B shows the spatial relationship between user 601 and candidate virtual assistant 604A. The distance d between user 601 and candidate virtual assistant 604A is... au In Figure 6C, c angle V represents the viewpoint of the virtual assistant at the candidate location relative to the user at the given location. a With V u A function of the angle between them. Figure 6D In the middle, c dis The displacement between the candidate position of the candidate virtual assistant 604A and the determined position of the user 601 is represented. Based on the relationship shown in Figures 6B-6D, the position determination module 170 determines c. vis The value indicating the visibility of user 601's face to candidate virtual assistant 604A positioned at the candidate location.

[0090] As shown in Figure 6B, for each candidate position, the position determination module 170 can establish a corresponding orientation for the virtual assistant (e.g., corresponding values ​​based on roll, pitch, and yaw). Based on the established orientation for the virtual assistant, the position determination module 170 can calculate a vector indicating the direction the candidate virtual assistant 604A faces when the virtual assistant is located at the candidate position, for example, v. a Similarly, for each user 601 positioned within the visible portion of the augmented reality environment, the position determination module 170 can access the user 601's orientation data (e.g., data from position and orientation data 226 specifying roll, pitch, and yaw values) and calculate a vector indicating the direction the user 601 is facing at the corresponding position, such as v. u .

[0091] For each pair of candidate locations and user locations, the location determination module 170 can calculate v a With v u The value of the angle between them, such as measured in degrees (e.g., angle of view). For example, and referring to Figure 6B, user 601 can face the direction shown by direction vector 606, such as direction vector v. u The specified, and candidate virtual assistant 604A can face the direction indicated by direction vector 608, such as vector v. aThe specified location determination module 170 can calculate the difference between direction vectors 608 and 606 in degrees and can determine the c-position of user 601 and candidate positions associated with candidate virtual assistant 604A based on the predetermined empirical relationship 622 shown in Figure 6C. angle The corresponding value.

[0092] As shown in Figure 6C, c angle The value of v a With v u The difference is minimized when it is in the range of approximately 150° to approximately 210° (e.g., when user 601 and candidate virtual assistant 604A are roughly facing each other). Additionally, c angle The value of v a With v u The difference is greatest when it is close to 0° or 360° (e.g., when user 601 and candidate virtual assistant 604A are facing the same direction or back to back). For each combination of user and established candidate positions within the visible portion of the augmented reality environment, the c-value of user 601 and candidate virtual assistant 604A (e.g., if positioned at the corresponding candidate position) can be repeatedly calculated. angle The exemplary process.

[0093] Furthermore, for each pair of established candidate locations and user locations within the visible portion of the augmented reality environment, the location determination module 170 can calculate the displacement d between corresponding locations in the established candidate locations and user locations. au As shown in Figure 6B, the position determination module 170 can calculate the displacement 614 between the user 601 and the candidate virtual assistant 604A in the augmented reality environment. The position determination module 170 can also determine the c-axis of the candidate positions associated with the user 601 and the candidate virtual assistant 604A based on the predetermined empirical relationship 624 shown in Figure 6D. dis The corresponding value. As shown in Figure 6D, c dis The value is minimum when the displacement is about five to six feet and maximum when the displacement is close to zero or close to and exceeds fifteen feet.

[0094] In other cases, the location determination module 170 can adjust (or establish) the position of user 601 and candidate virtual assistant 604A. disThe value reflects the displacement between the candidate virtual assistant 604A (e.g., positioned at a candidate location) and one or more objects (e.g., object 603 in FIG. 6A) positioned in the augmented reality environment. For example, user 601 can interact with object 603 within the augmented reality environment, and the position determination module 170 can access the identifier of object 603 in the augmented reality environment. Location data within the territory (e.g., object data 224 in database 214). Location determination module 170 can calculate the displacement between the candidate virtual assistant 604A and object 603, and can adjust (or establish) the position of user 601 and candidate virtual assistant 604A. dis The value reflects the interaction between user 601 and the one or more objects. For each combination of user and candidate location within the visible portion of the augmented reality environment, the value of c for user 601 and candidate virtual assistant 604A (e.g., positioned at the corresponding candidate location) can be repeatedly calculated. dis The instance process.

[0095] Furthermore, for each pair of established candidate positions and user positions within the visible portion of the augmented reality environment, the position determination module 170 can calculate the user's field of view based on the user's orientation (e.g., roll, pitch, and yaw values). The position determination module 170 can also establish c vis The value corresponds to determining whether the virtual assistant will be visible or invisible to the user when placed at the corresponding established candidate location. As shown in Figure 7A, the location determination module 170 can calculate the user's field of view 702 and can determine that the candidate virtual assistant 604A does not exist within the field of view 702 and is therefore invisible to the user when placed at the corresponding established candidate location. Given the lack of visibility, the location determination module 170 can assign a value to the user 601 and the candidate location associated with the candidate virtual assistant 604A. vis The value assigned is 100.

[0096] As shown in Figure 7B, the location determination module 170 can calculate the user's additional field of view 704 based on the user's current orientation. Further, the location determination module 170 can determine that the candidate virtual assistant 604A exists within the additional field of view 704 and is visible to the user. Given this determined visibility, the location determination module 170 can determine the c-position of the user 601 and the candidate positions associated with the candidate virtual assistant 604A. visThe value assigned is zero. Additionally, as described above, for each combination of user and candidate location within the visible portion of the augmented reality environment, the c-value of user 601 and candidate virtual assistant 604A (e.g., positioned at the corresponding candidate location) can be repeatedly calculated. vis The exemplary process.

[0097] Furthermore, the exemplary process is not limited to c cons c supp or c vis Any specific value or specified c angle and c dis Any specific empirical relation to the value of c. The exemplary process described herein may assign any additional or alternative value to c. cons c supp or c vis Alternatively, c can be established based on any other or alternative empirical relationship. angle and c dis The value of the value will be suitable for the augmented virtual environment. Furthermore, although described as conical in Figures 7A and 7B, the user's field of view can be characterized by any suitable shape of the augmented reality environment and the display unit 208.

[0098] In the above example, in box 316, the mobile device 200 or XR computing system 130 executes the location determination module 170 to calculate the placement score p for each candidate location of the virtual assistant. ast The calculated placement score can reflect the structural and physical constraints imposed on candidate locations by the augmented reality environment, and further, it can reflect the visibility and proximity of the virtual assistant to each user in the augmented reality environment when the virtual assistant is placed at each candidate location. In other cases described below, the calculated placement score can also reflect the level of interaction between each user in the augmented reality environment.

[0099] For example, as shown in Figure 8, the visible portion 800 of the augmented reality environment may contain users 802, 806, 808, and 810, each of whom may be positioned at a corresponding location within the visible portion 800. A group of users (such as users 802, 804, and 806) may be positioned together within the visible portion 800 and may interact and converse closely while accessing the augmented reality environment. Further, the current and historical orientation of each user 802, 804, and 806 (e.g., maintained in location and orientation data 226) may indicate that users 802, 804, and 806 are positioned facing and toward each other, or alternatively, users 804 and 806 are positioned facing and toward user 802 (e.g., the "leader" of the group).

[0100] Other users (such as users 808 and 810) can be positioned away from users 802, 804, and 806. While other users 808 and 810 can monitor conversations between users 802, 804, and 806, they can choose not to participate in the monitored conversations and can experience limited interactions with each other and with users 802, 804, and 806. Furthermore, the current and historical orientation of each user 808 and 810 can indicate that users 808 and 810 are facing away from each other and from users 802, 804, and 806.

[0101] In box 316, the location determination module 170 can assign weights to each candidate location, such as p, for a virtual assistant or other virtual content item based on user-specific variations in the level of interaction within the augmented reality environment. asst Placement score, for example, cost(p) asst The location determination module 170 can calculate a "modified" placement score for each candidate location based on the following formula: (User-specific contribution).

[0102]

[0103] To account for variations in the level of interaction between each user, the location determination module 170 can use an interaction weighting factor s. inter Applied to C angle c dis and c vis Each user-specific combination. The location determination module 170 can determine the visible portion of the augmented reality environment for each user based on the analysis of stored audio data (e.g., stored in interaction data 228) of the captured interactions between users and based on each user's current and historical orientation.inter The value. For example, the position determination module 170 can determine a larger value s. inter Values ​​are assigned to users who interact closely within the visible portion of the augmented reality environment and smaller values ​​are assigned to those users. inter Values ​​are assigned to users who have experienced limited interaction. Additionally, based on analysis of stored audio and orientation data, the location determination module 170 can identify group leaders among interacting users and assign values ​​to... inter The value assigned to the group leader exceeded the value assigned to other interacting users. inter value.

[0104] By weighting the user-specific contribution of placement scores to candidate locations based on user-specific interaction levels, the disclosed implementation can bias the placement of virtual assistants in augmented reality environments to locations associated with significant levels of user interaction. For example, by using a larger s inter Values ​​are assigned to interactive users 802, 804, and 806 in Figure 8. The location determination module 170 can increase the likelihood that the virtual assistant's established location will be close to users 802, 804, and 806 within the visible portion of the augmented reality environment. Placing the virtual assistant close to highly interactive users can enhance the ability of highly interactive users to interact with the virtual assistant without modifying their corresponding posture or orientation. Exemplary arrangement and s inter The corresponding assignment value can bias any large changes in posture or orientation toward users with limited interaction experience, which can encourage users with lower interaction levels to further interact with the augmented reality environment.

[0105] Referring back to Figure 3, when performed by the mobile device 200 of the XR computing system 130, the location determination module 170 can establish one of the candidate locations as the placement position of the virtual content item within the visible portion of the augmented reality environment (e.g., in box 318). The location determination module 170 can identify the minimum calculated placement score (e.g., cost(p)) of the candidate locations. asst It also identifies the corresponding candidate positions associated with the minimum calculated placement score. In some instances, the position determination module 170 can establish the corresponding candidate positions as the placement positions for virtual assistants or other virtual content items. ∗ , As described below:

[0106]

[0107] For example, in box 318, the location determination module 170 can determine the minimum value of the placement cost representation placement score calculated for a candidate location associated with the candidate virtual assistant 604A (e.g., as shown in Figure 6A) and can establish the candidate location as the placement location for the virtual assistant or other virtual content item.

[0108] In other cases, the location determination module 170 may fail to identify the placement of a virtual assistant that aligns with the physical constraints imposed by the augmented reality environment and the user's position, orientation, and interaction level within the augmented reality environment. For example, a portion of the augmented reality environment visible to the user 601 (e.g., the user's gestures or verbal input invoking the virtual assistant) may contain ocean surrounded by cliffs. The location determination module 170 may perform any of the operations described herein to determine that each candidate location of the virtual assistant represents a hazard (e.g., relative to c). cons (associated with an infinite value), and therefore, the location determination module 170 is restricted to prevent the virtual assistant from being kept within the visible portion of the augmented reality environment.

[0109] Based on this determination, the location determination module 170 can access stored data identifying previously visible portions of the augmented reality environment to the user 601 (e.g., maintained by storage media 132 of the XR computing system 130 or storage media 211 of the mobile device 200). In other cases, the location determination module 170 can also determine portions of the augmented reality environment that are adjacent to the user 601 in the augmented reality environment (e.g., to the left, right, or behind the user 601). The location determination module 170 can then perform any of the operations described herein to determine a suitable placement position for the virtual assistant within the previously visible portion of the augmented reality environment or, alternatively, within an adjacent portion of the augmented reality environment.

[0110] Furthermore, in certain circumstances, the location determination module 170 may determine that a previously visible or adjacent arrangement of the augmented reality environment is unsuitable for placing the virtual assistant. In response to this determination, the location determination module 170 may generate an error signal that causes the mobile device 200 or the XR computing system 130 to maintain the virtual assistant's invisibility within the augmented reality environment, for example, for a specified future time period or until further gestures or verbal input from the user 601 is detected.

[0111] Referring back to Figure 3, mobile device 200 or XR computing system 130 can perform the operation of inserting a virtual content item (e.g., a virtual assistant) at a defined placement location within the visible portion of the augmented reality environment (e.g., in box 320). In some instances, when executed by processor 138, virtual content generation module 172 can execute any of the procedures described herein to generate a digital content item, such as an animated representation of a virtual assistant, and can transmit the generated virtual content item and the defined placement location across network 120 to mobile device 200, which can then perform the operation of displaying the animated representation and corresponding audio content at the defined placement location within the augmented reality environment, for example, via display unit 208. In other instances, when executed by processor 218, virtual content generation module 172 can execute any of the procedures described herein to locally generate a digital content item, such as an animated representation of a virtual assistant, and display the animated representation and corresponding audio content at the defined placement location within the augmented reality environment, for example, via display unit 208.

[0112] The virtual content generation module 172 provides means for inserting a virtual assistant into an augmented reality environment at a defined placement location. For example, the graphics module 176 of the virtual content generation module 172 can generate an animated representation based on stored data (e.g., graphics data 190 or local graphics data 230) specifying certain visual characteristics of the virtual assistant, such as the visual characteristics of the user-selected avatar. Further, the speech synthesis module 178 of the virtual content generation module 172 can generate audio content representing portions of an interactive dialogue spoken by the generated virtual assistant based on stored data (e.g., speech data 192 or local speech data 232). The simultaneous presentation of the animated representation and audio content establishes the virtual assistant within a visible portion of the augmented reality environment and facilitates interaction between the user of the mobile device 200 and the virtual assistant.

[0113] Mobile device 200 or XR computing system 130 can perform operations to determine whether a change in device state (such as a change in the position or orientation of mobile device 200) triggers a repositioning of the virtual assistant within the visible portion of the augmented reality environment (e.g., in box 322). A change in device state can trigger the virtual assistant's repositioning or modify objects or users within the visible portion of the augmented reality environment (e.g., to include new objects). For example, a change in device state can trigger the virtual assistant's repositioning when the magnitude of the change exceeds a predetermined threshold.

[0114] If mobile device 200 or XR computing system 130 determines that a change in device state triggers the repositioning of the virtual assistant (e.g., in box 322; yes), then exemplary process 300 may branch back to box 308. Mobile device 200 or XR computing system 130 may then identify the user-visible portion of the augmented reality environment for mobile device 200 based on the newly changed device state and execute any of the procedures described herein to move the virtual assistant to the location of the user-visible portion of the augmented reality environment. Alternatively, if mobile device 200 or XR computing system 130 determines that a change in device state does not trigger the repositioning of the virtual assistant (e.g., in box 324; no), then exemplary process 300 may branch back to box 324, and process 300 completes.

[0115] As described above, virtual content items may include animated representations of virtual assistants. In response to establishing a virtual assistant within a visible portion of the augmented reality environment, a user of mobile device 200 can interact with the virtual assistant and provide verbal or gestural input to mobile device 102, specifying a command, request, or query. In response to verbal input, mobile device 102 can execute any of the above-described processes to parse the verbal query and obtain text data corresponding to the command, request, or query associated with the verbal input, and perform an operation corresponding to the text data. In one instance, the verbal input corresponds to a verbal command to revoke the virtual assistant (e.g., "remove virtual assistant"), and operation module 244 can execute any of the above-described processes to remove the virtual assistant from the augmented reality environment.

[0116] Verbal input can further correspond to requests for additional information regarding objects positioned within the augmented reality environment or audio or graphical content provided by a virtual assistant. Operation module 244 can acquire text data corresponding to the issued query and package portions of the text data into query data, which mobile device 200 can transmit across communication network 120 to XR computing system 130. In one instance, XR computing system 130 can receive query data and execute any of the processes described above to obtain information in response to the query data (e.g., based on locally stored data or information received from another computing system 160). XR computing system 130 can also verify that the obtained information complies with optionally imposed security or privacy restrictions and transmit the obtained information back to mobile device 200 as a response to the query data. Mobile device 200 can also execute any of the processes described above to generate graphical or audio content (including synthesized speech) representing the obtained information and present the generated graphical or audio content through interaction with the user via a virtual assistant.

[0117] The subject matter of this invention is not limited to verbal input that facilitates interaction between a user and a virtual assistant within an augmented reality environment when detected by mobile device 200. The digital camera 212 of mobile device 200 can capture digital image data containing gestural input, such as hand gestures or pointing movements, provided by the user to mobile device 200. For example, an augmented reality environment could allow a user to explore the train sheds of a historic city train station, such as Pennsylvania Station in New York City, which is now demolished. The virtual assistant could be positioned on a platform adjacent to one or more carriages of a train and could provide synthesized audio content outlining the history of the station and the Pennsylvania Railroad. However, a user of mobile device 200 could be interested in steam locomotives and could point to them on the platform. Digital camera 212 could capture pointing movements in the corresponding image data. Mobile device 200 could identify gestural input within the image data or based on data received from various sensor units (such as IMUs or other sensors) incorporated into or communicating with mobile device 200 (e.g., incorporated into a glove worn by the user). As described below with reference to Figure 9, the mobile device 200 can associate the identified gesture input with the corresponding operation and perform the corresponding operation in response to the gesture input.

[0118] Figure 9 is a flowchart of an exemplary process 900 for performing operations in an extended reality environment in response to detected gesture input, according to some embodiments. Process 900 may be executed by one or more processors that execute instructions locally at a mobile device (e.g., mobile device 200 of Figure 2). Some blocks of process 900 may be executed remotely by one or more processors of a server system or other suitable computing platform (e.g., XR computing system 130 of Figure 1). Thus, the various operations of process 900 may be implemented by executable instructions held in storage media coupled to one or more computing platforms (e.g., storage media 132 of XR computing system 130 and / or storage media 211 of mobile device 200).

[0119] Mobile device 200 or XR computing system 130 can perform operations that detect user-provided gesture input (e.g., in box 902). As shown in FIG10A, a user of the mobile device (e.g., user 601) can access an extended reality environment via the corresponding mobile device 200 in the form of a head-mounted display (HMD) unit. User 601 can perform pointing motion 1002 in the real-world environment in response to content accessed in the extended reality environment. As described above, the extended reality environment may include an augmented reality environment that allows user 601 to explore the train sheds of Pennsylvania Station, and user 601 can perform pointing motion 1002 to attempt to obtain additional information about the steam locomotives on the platform from virtual assistant 1028.

[0120] As described above, the digital camera 212 of the mobile device 200 can capture and record digital image data of pointing motion 1002. When executed by the processor 218 of the mobile device 200, the gesture detection module 240 can access the digital image data and apply one or more of the aforementioned image processing techniques, computer vision algorithms, or machine vision algorithms to detect pointing motion 1002 within one or more frames of the digital image data. Further, when executed by the processor 218, the operation module 244 can obtain data indicating the detected pointing motion 1002. Based on the portions of the gesture library 234, the operation module 244 can determine that pointing motion 1002 corresponds to the repositioning virtual assistant 1028 and request additional information characterizing the object (e.g., a steam locomotive) associated with pointing motion 1002.

[0121] Additionally, the mobile device 200 can perform operations to obtain depth map data 222 and object data 224 representing the portion of the augmented reality environment visible to the user via the display unit 208 (e.g., in box 904). For example, the mobile device 200 can obtain depth map data of a depth map of a specified visible portion of the augmented reality environment from a corresponding portion of the database 214 (e.g., from the depth map data 222). Further, the mobile device 200 can obtain data identifying objects arranged within the visible portion of the augmented reality environment and data identifying the position or orientation of the identified objects from the object data 224 of the database 214.

[0122] In other cases, mobile device 200 may transmit data representing a request to reposition virtual assistant 1028 across network 120 to XR computing system 130, such as data indicating detected pointing motion 1002 and objects associated with pointing motion 1002. XR computing system 130 may obtain depth map data of a depth map specifying a visible portion of the augmented reality environment from a corresponding portion of database 134 (e.g., from depth map data 182). XR computing system 130 may obtain data identifying objects arranged within the visible portion of the augmented reality environment and data identifying the position or orientation of the identified objects from object data 184 in database 134.

[0123] The mobile device 200 or XR computing system 130 can also perform operations such as generating a gesture vector 1022 representing the detected direction of the motion 1002 and projecting the gesture vector 1022 onto a depth map established by depth map data 222 or depth map data 182 of the visible portion of the augmented reality environment (e.g., in box 906). For example, when executed by processor 218 or 138, the position determination module 170 can establish an origin for the motion 1002, such as the user's shoulder or elbow joint, and can generate the gesture vector 1022 based on a determined alignment of the user's extended arm, hand, and / or fingers. Further, the position determination module 170 can extend the gesture vector 1022 from the origin and project the gesture vector into a three-dimensional space established by the depth map of the visible portion of the augmented reality environment.

[0124] Mobile device 200 or XR computing system 130 can identify objects (e.g., in frame 908) within the visible portion of the augmented reality that correspond to the projection of a gesture vector. Based on object data 224, when executed by processor 218 of mobile device 200, position determination module 170 can determine the position or size of one or more objects arranged within the visible portion of the augmented reality. In other cases, when executed by processor 138 of XR computing system 130, position determination module 170 can determine the position or size of said one or more objects arranged within the visible portion of the augmented reality environment based on portions of object data 184. Position determination module 170 can then determine that the projected gesture vector intersects with the position or size of one of the arranged objects and can confirm that the arranged object corresponds to the projected gesture vector.

[0125] As shown in Figure 10B, the position determination module 170 can generate a gesture vector 1022 corresponding to the detected pointing motion 1002 and project the gesture vector (typically shown as 1024) onto a three-dimensional space 1020 constructed from depth map data 222 of the visible portion of the augmented reality environment. Furthermore, the position determination module 170 can also determine that the projected gesture vector 1024 intersects with an object 1026 arranged within the three-dimensional space 1020 and can access portions of the object data 224 to obtain information identifying the object 1026, such as semantic data identifying the object type.

[0126] Referring back to Figure 9, mobile device 200 or XR computing system 130 can obtain information characterizing and describing objects identified within the visible portion of the augmented reality environment (e.g., in box 910). As described above, the identified objects may correspond to a specific steam locomotive, and mobile device 200 can execute any of the processes described above to generate query data requesting additional information about the specific steam locomotive. Mobile device 200 can transmit the query data to XR computing system 130 and receive a response to the query data describing the specific steam locomotive and complying with applicable security and privacy restrictions.

[0127] Furthermore, by using any of the procedures described herein, mobile device 200 can generate animated representations of virtual assistant 1028 as well as graphical or audio content representing information about the descriptive objects. In other cases, XR computing system 130 can use any of the procedures described herein to generate animated representations of virtual assistant 1028 as well as graphical or audio content representing information about the descriptive objects and can transmit the animated representations and graphical or audio content to mobile device 200 across network 120.

[0128] In other cases, mobile device 200 or XR computing system 130 may execute any of the procedures described herein to generate a new placement location for virtual assistant 1028 in the augmented reality environment, such as the location of an identified object (e.g., in box 912). Mobile device 200 may render an animated representation of virtual assistant 1028 at the new placement location in the augmented reality environment and present graphical or audio content to provide an immersive experience that facilitates interaction between the user and the virtual assistant (e.g., in box 914). Exemplary procedure 900 is then completed in box 916.

[0129] In some of the examples described above, the extended reality environment established by mobile device 200, mobile device 102, or 104, or other devices operating in network environment 100, may correspond to an augmented reality environment that includes virtual tours of various historical sites or landmarks (such as the Giza Pyramids or Pennsylvania Station). The subject matter of this invention is not limited to exemplary environments, and in other cases, mobile device 200 and / or XR computing system 130 may operate together to establish virtual assistants or other virtual content items within any number of additional or alternative augmented reality environments.

[0130] For example, an augmented reality environment can correspond to a virtual meeting location (e.g., a conference room) containing multiple geographically dispersed participants. By using any of the processes described herein, mobile device 200 and / or XR computing system 130 can operate together to deploy virtual assistants at locations within the virtual meeting location and facilitate interaction with participants to support the meeting.

[0131] In some cases, participants may include the meeting's speaker or coordinator. When calculating the modified placement score for each candidate location within the virtual meeting venue, the mobile device 200 or XR computing system 130 may add additional or alternative weighting factors (e.g., as described above, in addition to s). inter (In addition to or as an alternative) applied to c associated with the speaker or coordinator of the meeting.angle c dis and c vis The combination of factors allows the mobile device 200 or XR computing system 130 to position the virtual assistant within the virtual meeting space closer to the speaker or coordinator, thereby enhancing interaction between the virtual assistant and the speaker or coordinator. For example, the mobile device 200 or XR computing system 130 can present the virtual assistant as a "talking head" or similar miniature object positioned in the center of a "virtual" meeting table within the virtual meeting space.

[0132] Participants can also provide verbal or gestural input to request the virtual assistant to display certain graphic content within a corresponding presentation area of ​​the augmented reality environment (such as a virtual whiteboard in a virtual meeting location). The input can also request the virtual assistant to perform actions, such as starting or stopping meeting recording or displaying the meeting agenda within the augmented reality environment. The virtual assistant can receive requests to set reminders for scheduled meetings on one or more participants' schedules, schedule meetings, track times, solicit feedback, or disappear for certain periods. In some cases, the mobile device 200 and / or the XR computing system 130 can process verbal or gestural input to identify requested graphic content and / or requested actions and perform operations consistent with the verbal or gestural input.

[0133] Additionally, when deployed in an augmented reality environment, virtual assistants can also present audio or graphical content to ensure the safety and awareness of one or more users in the real-world environment. For example, a virtual assistant can provide warnings about real-world hazards (such as obstacles, blind spots, mismatch between real and virtual objects, hot or cold objects, electrical equipment, etc.) or alert users to exit the augmented reality environment in response to a real-world emergency.

[0134] As described above, when executed by a mobile device (e.g., mobile device 200) or a computing system (e.g., XR computing system 130), augmented reality generation and presentation tools define an augmented reality environment based on certain elements of digital content, such as captured digital video, digital images, digital audio content, or synthesized audiovisual content (e.g., computer-generated images and animation content). These tools can deploy elements of the digital content to be presented to the user via a display unit 208 incorporated in mobile device 200, such as an augmented reality glasses wearable (e.g., glasses or goggles) with one or more lenses or displays for presenting the deployed digital content's graphic elements. For example, an augmented reality glasses wearable can display graphic elements as an augmented reality layer overlaid on real-world objects visible through the lenses, thereby enhancing the user's ability to interact with and explore the augmented reality environment.

[0135] Mobile device 200 can also capture gestures or verbal input requesting the placement of virtual content items, such as virtual assistants, within the augmented reality environment. For example, and in response to captured gestures or verbal input, mobile device 200 or XR computing system 130 can perform any of the exemplary processes described herein to generate a virtual content item, identify a placement location for the virtual content item that conforms to any constraints that may be imposed by the augmented reality environment, and insert the generated virtual content item into the augmented reality environment at the placement location. As mentioned above, the digital content item may include a virtual assistant that, when rendered for presentation in the augmented reality environment, can interact with the user and elicit further gestures or verbal queries from the user to mobile device 200.

[0136] In other instances, when executed by mobile device 200 or XR computing system 130, one or more of these tools may define a virtual reality environment based on certain elements of synthesized audiovisual content (e.g., individually or in combination with certain elements of captured digital audiovisual content). The defined virtual reality environment represents an artificial, computer-generated simulation or reconstruction of a real-world environment or situation, and the tools may deploy elements of the synthesized or captured audiovisual content via a head-mounted display (HMD) of mobile device 200 (such as a user-wearable virtual reality (VR) headset).

[0137] In some cases, a virtual reality environment can correspond to a virtual sporting event (such as a tennis match or a golf tournament). When presented to a user through a wearable VR headset, this virtual sporting event provides visual and auditory stimulation that immerses the user and allows them to experience the virtual sporting event firsthand. In other cases, a virtual environment can correspond to a virtual tour. When presented to a user through a wearable VR headset, this virtual tour allows the user to explore now-vanished historical landmarks or observe historically significant events, such as entire ancient military campaigns.

[0138] Mobile device 200 can also capture gestures or verbal input requesting the placement of virtual content items, such as virtual assistants, within a virtual reality or augmented virtual environment. For example, in response to a captured gesture or verbal input requesting a virtual assistant within the virtual reality or augmented virtual environment, mobile device 200 or XR computing system 130 can determine the portion of the virtual reality or augmented virtual environment currently visible to the user via the VR headset. Based on the determined visible portion, mobile device 200 or XR computing system 130 can execute any of the exemplary processes described herein to generate a virtual assistant, identify a placement location for the virtual assistant that conforms to any restrictions that may be imposed by the virtual reality or augmented virtual environment, and insert the generated virtual assistant into the virtual reality or augmented virtual environment at the placement location.

[0139] As described above, when rendered for presentation in an augmented reality, virtual reality, or augmented virtual environment, the virtual assistant can interact with the user and elicit further gestures or verbal queries from the user regarding the mobile device 200. For example, the virtual assistant can act as a virtual coach in a virtual sporting event, providing input reflecting the user's observations or participation in the virtual sporting event (and further, potentially translating into participation in a real-world sporting event). In other instances, the virtual assistant can act as a virtual tour guide in a virtual tour, providing additional information about certain objects or individuals within the virtual reality environment in response to user queries.

[0140] In some instances, the executed augmented reality tools, augmented virtual reality tools, and virtual reality tools can also define other virtual environments and perform operations to generate and locate virtual assistants within these defined virtual environments. For example, when executed by mobile device 200 or XR computing system 130, these tools can define one or more extended reality environments that combine elements of virtual reality and augmented reality environments and facilitate varying degrees of human-computer interaction.

[0141] For example, mobile device 200 or XR computing system 130 can select elements of captured or synthesized audiovisual content to be presented within an augmented reality environment based on data received from sensors integrated into or communicating with mobile device 200. Examples of such sensors include, but are not limited to, temperature sensors capable of establishing ambient temperature or biometric sensors capable of establishing the user's biometric characteristics (such as pulse or body temperature). These tools can then present elements of the synthesized or captured audiovisual content as an augmented reality layer overlaid on real-world objects visible through lenses, thereby enhancing the user's ability to interact with and explore the augmented reality environment in a way that adaptively reflects the user's environment or physical condition.

[0142] The methods and systems described herein can be embodied, at least in part, in the form of computer-implemented processes and devices for practicing the disclosed processes. The disclosed methods can also be embodied, at least in part, in the form of tangible, non-transitory, machine-readable storage media encoded with computer program code. The media can include, for example, random access memory (RAM), read-only memory (ROM), optical disc (CD)-ROM, digital versatile optical disc (DVD)-ROM, and "blue-ray disc". TM (BD)-ROM, hard disk drive, flash memory, or any other non-transitory machine-readable storage medium. When computer program code is loaded into and executed by the computer, the computer becomes a device for practicing the method. The method can also be embodied, at least in part, in the form of a computer, with computer program code loaded into or executed by the computer, making the computer a dedicated computer for practicing the method. When implemented on a general-purpose processor, computer program code segments configure the processor to create specific logic circuitry. The method can be embodied, at least in part, in an application-specific integrated circuit (ASIC) for performing the method.

[0143] The subject matter of the invention has been described with reference to exemplary embodiments. Since the exemplary embodiments are merely examples, the claimed invention is not limited to these embodiments. Changes and modifications may be made without departing from the spirit of the claimed subject matter. The claims are intended to cover such changes and modifications.

Claims

1. An apparatus comprising: Computer-readable storage media; as well as A processor coupled to the computer-readable storage medium, the processor being configured to: First image data is obtained from a first mobile device, wherein the first image data is based on at least one image of a real-world scene; Second image data is obtained from the second mobile device; An augmented reality (AR) environment is created based on the first image data and the second image data. Identify the objects placed in the real-world scene; Receive user input; Determine the placement location of virtual content items in the AR environment, the placement location being based on the AR environment, the identified objects, and the received user input; as well as The determined placement location is sent to one or more of the first mobile device and the second mobile device.

2. The device according to claim 1, wherein: The device is a server that communicates with the first mobile device and the second mobile device.

3. The device according to claim 1, wherein: The device is the first mobile device, and the placement position is sent to the second mobile device.

4. The device according to claim 3, wherein: The placement location is sent to the second mobile device via a server.

5. The device according to claim 3, wherein: The processor is also configured to invoke the placement of the virtual content item based on received user input.

6. The device of claim 5, wherein the user input is a gesture.

7. The device of claim 6, wherein the gesture is received via a touch-sensitive surface.

8. The device of claim 7, wherein the user input is associated with a location in the augmented reality environment.

9. The device of claim 8, wherein the placement position is further based on the posture of the first mobile device.

10. The device of claim 8, wherein the virtual content is associated with a real-world object in the AR environment.