Virtual and real object recording in mixed reality devices
By using an object selection device and control subsystem to select and record objects of interest to the user in a VR/AR system, the problem of noise interference is solved, and high-quality audio and video data recording and storage are achieved.
Patent Information
- Application Number
- CN202410676435.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-02-28
- Filing Date
- 2018-02-27
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2038-02-27
AI Technical Summary
Existing VR/AR systems suffer from reduced experience quality due to noise and unwanted sound interference when recording user experiences, making it difficult to effectively record the sounds of virtual or real objects that users are interested in.
An object selection device is used to receive input from the end user, select and record the object of interest, and generate and store audio and video data in combination with the display subsystem and the control subsystem. The sound of the real object is processed first, and the audio and video data is stored and rendered synchronously in the memory.
It enables the recording of only the sounds of virtual or real objects that users are interested in within VR/AR systems, improving recording quality and user experience, and ensuring the synchronization and efficient storage of audio and video data.
Smart Images

Figure CN118873933B_ABST
Abstract
Description
[0001] This application is a divisional application of application number 201880014388.7, filed on February 27, 2018, having the title “Virtual and real object recording in mixed reality devices”. TECHNICAL FIELD
[0002] The present invention relates generally to virtual reality and augmented reality systems. BACKGROUND
[0003] Modern computing and display technologies have facilitated the development of mixed reality (MR) systems that can present digital content, such as virtual objects, to users in a seemingly real or realistic fashion, in real time. As used herein, the term “mixed reality” refers generally to any technology that facilitates the real-time presentation of digital content, such as virtual objects, in a seemingly real or realistic fashion, and can be used interchangeably with the terms “virtual reality” and “augmented reality.” A virtual reality (VR) scenario typically involves presentation of digital or virtual image information without transparency to the actual real-world environment. An augmented reality (AR) scenario typically involves presentation of digital or virtual image information as an augmentation to the real-world environment. A mixed reality (MR) scenario can involve presentation of digital or virtual image information that has the appearance of being a part of the real-world environment.
[0004] For example, referring to Figure 1 , an augmented reality scenario is depicted in which a user of an AR technology sees a real-world parklike setting 6 featuring people, trees, buildings in the background, and a concrete platform 8. In addition to these items, the end user of the AR technology also perceives that he “sees” a robot statue 10 standing upon the real-world platform 8 and a cartoon-like avatar character 12, which seems to be a personification of a bumble bee, hovering above and beside the platform 8. These elements 10, 12 are of a digital or virtual
[0005] VR and AR systems typically employ a head-mounted display (or head-mounted display or smart glasses) that is at least loosely coupled to the user's head and moves therewith as the end user's head moves. If the display system detects head motion of the end user, the data being displayed can be updated to account for the change in head pose (i.e., the orientation and / or position of the user's head). Head-mounted displays that enable AR (i.e., the simultaneous viewing of virtual and real objects) can have several different types of configurations. In one such configuration, commonly referred to as a "video see-through" display, a camera captures elements of the real scene, a computing system superimposes virtual elements onto the captured real scene, and a non-transparent display presents the composite image to the eye. Another configuration is commonly referred to as an "optical see-through" display, in which the end user can see through transparent (or semi-transparent) elements in the display system to directly view light from real objects in the environment. The transparent elements, often referred to as "combiners," superimpose light from the display onto the end user's view of the real world.
[0006] Typically, a user of a VR / AR system can want to share his or her experience with others (e.g., while playing a game, a teleconference, or watching a movie) by recording and saving the experience on the VR / AR system for subsequent online posting. However, there can typically be noise and other unwanted or unexpected sounds in the recording due to a noisy environment or there can be too many sound sources, which can be distracting to the experience. Such unwanted / unexpected sounds can come from real objects, e.g., from children playing near the VR / AR system, or from virtual objects, e.g., from a virtual television being played back in the VR / AR system environment.
[0007] Accordingly, there remains a need to provide a simple and effective apparatus for recording only sounds from virtual or real objects that are of interest to the user. SUMMARY
[0008] According to a first aspect of the present application, a virtual image generation system for use by an end user includes a memory, a display subsystem, and an object selection apparatus configured to receive input from the end user and continuously select at least one object (e.g., a real object and / or a virtual object) in response to the end user input. In one embodiment, the display subsystem has a field of view, and the object selection apparatus is configured to continuously select objects in the field of view. In this case, the object selection apparatus can be configured to move a three-dimensional cursor in the field of view of the display subsystem and select objects in response to receiving the end user input. In another embodiment, the end user input includes one or more voice commands, and wherein the object selection apparatus includes one or more microphones configured to sense the voice commands. In yet another embodiment, the end user input includes one or more hand gestures, in which case the object selection apparatus can include one or more cameras configured to sense the hand gestures.
[0009] In the case of selecting multiple objects, the object selection apparatus can be configured to select the objects individually and / or globally in response to the end user input. If globally, the object selection apparatus can be configured to globally select all objects within an angular range of the field of view (which can be less than the entire angular range of the field of view or can be the entire angular range of the field of view) in response to the end user input. In one embodiment, the object selection apparatus is further configured to receive another input from the end user and continuously deselect previously selected objects in response to the other end user input.
[0010] The virtual image generation system further includes a control subsystem configured to generate video data originating from the at least one selected object, render a plurality of image frames in a three-dimensional scene from the video data, and transmit the image frames to the display subsystem. In one embodiment, the display subsystem is configured to be positioned in front of the end user's eyes. In another embodiment, the display subsystem includes a projection subsystem and a partially transparent display surface. In this case, the projection subsystem can be configured to project the image frames onto the partially transparent display surface, and the partially transparent display surface can be configured to be positioned in the field of view between the end user's eyes and the surrounding environment. The virtual image generation system can further include a frame structure configured to be worn by the end user and carry at least a portion of the display subsystem.
[0011] The control subsystem is further configured to generate audio data originating from the selected object and store the audio data in the memory. The virtual image generation system can further include a plurality of speakers, in which case the control subsystem is further configured to transmit the generated audio data to the speakers. In an optional embodiment, the control subsystem is further configured to store the video data in the memory synchronously with the audio data. In yet another embodiment, the virtual image generation system further includes at least one sensor configured to track a position of the selected object relative to a field of view of the display subsystem. In this case, the control subsystem can be configured to stop storing the audio data in the memory when the tracked position of the selected object moves out of the field of view of the display subsystem, or, optionally, the control subsystem is configured to continue storing the audio data in the memory when the tracked position of the selected object moves out of the field of view of the display subsystem.
[0012] If the selected object includes a real object, the virtual image generation system can further include a microphone assembly configured to generate audio output, in which case the control subsystem can be further configured to modify the directional audio output to preferentially sense sounds originating from the selected real object. The audio data can be derived from the modified audio output. The virtual image generation system can further include one or more cameras configured to capture video data originating from the selected real object, in which case the control subsystem can be further configured to store the video data in the memory synchronously with the audio data. The control subsystem can be configured to transform the captured video data into virtual content data for the selected real object and store the virtual content in the memory.
[0013] If the selected object includes a virtual object, the virtual image generation system can further include a database configured to store content data corresponding to sounds for a plurality of virtual objects, in which case the control subsystem can be further configured to retrieve content data corresponding to the selected virtual object from the database, and the audio data stored in the memory includes the retrieved content data. The control subsystem can be further configured to generate metadata corresponding to the selected virtual object (e.g., position, orientation, and volume data for the selected virtual object), in which case the audio data stored in the memory can include the retrieved content data and the generated metadata. In one embodiment, the virtual image generation system further includes one or more sensors configured to track a head pose of the end user, in which case the database can be configured to store absolute metadata for the plurality of virtual objects, and the control subsystem can be further configured to generate the metadata by retrieving absolute metadata corresponding to the selected virtual object, and to localize the absolute metadata to the end user based on the tracked head pose of the end user.
[0014] The virtual image generation system can also include at least one speaker, in which case the control subsystem can also be configured to retrieve stored audio data from the memory, derive audio from the retrieved audio data, and transmit the audio to the speaker. The audio data stored in the memory can include content data and metadata, in which case the control subsystem can also be configured to retrieve stored content data and metadata from the memory, render spatialized audio based on the retrieved content data and metadata, and transmit the rendered spatialized audio to the speaker.
[0015] According to a second aspect of the present invention, there is provided a method of operating a virtual image generation system by an end user. The method includes continuously selecting at least one object (e.g., a real object and / or a virtual object). In one method, selecting an object includes moving a three-dimensional cursor in a field of view of the end user and selecting the object with the three-dimensional cursor. In another method, selecting an object includes issuing one or more voice commands. In yet another method, selecting at least one object includes making one or more hand gestures. If multiple objects are selected, selecting the multiple objects can include selecting the objects individually and / or selecting the objects globally. If selected globally, the objects can be selected by defining an angular range of the field of view of the end user (which can be less than the entire angular range of the field of view or can be the entire angular range of the field of view) and selecting all of the objects within the defined angular range of the field of view of the end user. The method can also optionally include continuously deselecting previously selected objects.
[0016] The method also includes generating video data originating from the selected object, rendering a plurality of image frames in a three-dimensional scene from the generated video data, and displaying the image frames to the end user, generating audio data originating from the at least one selected object, and storing the audio data originating from the at least one selected object in a memory. One method can also include transforming the audio data originating from the selected object into sound perceived by the end user. The method can optionally include storing the video data and the audio data in the memory in synchronization. Yet another method can also include tracking a location of the selected object relative to a field of view of the end user. In this case, the method can also include stopping storing the audio data in the memory when the tracked location of the selected object moves out of the field of view of the end user, or alternatively, continuing to store the audio data in the memory when the tracked location of the selected object moves out of the field of view of the end user.
[0017] If the selected object includes a real object, the method can further include preferentially sensing sound originating from the selected real object relative to sound originating from other real objects, in which case the audio data can be derived from the preferentially sensed sound. The method can further include capturing video data originating from the selected real object and storing the video data in the memory in synchronization with the audio data. The captured video data can be transformed into virtual content data for storage in the memory.
[0018] If the selected object includes a virtual object, the method can further include storing content data corresponding to sound for a plurality of virtual objects and retrieving content data corresponding to the selected virtual object, in which case the audio data stored in the memory can include the retrieved content data. The method can further include generating metadata corresponding to the selected virtual object (e.g., position, orientation, and volume data for the selected virtual object), in which case the audio data stored in the memory can include the retrieved content data and the generated metadata. The method can further include tracking a head pose of the end user and storing absolute metadata for the plurality of virtual objects. In this case, generating the metadata can include retrieving absolute metadata corresponding to the selected virtual object and localizing the absolute metadata to the end user based on the tracked head pose of the end user.
[0019] The method can further include retrieving the stored audio data, deriving audio from the retrieved audio data, and transforming the audio into sound perceived by the end user. The stored audio data can include content data and metadata, in which case the method can further include retrieving the stored content data and metadata from the memory, rendering spatialized audio based on the retrieved content data and metadata, and transforming the spatialized audio into sound perceived by the end user.
[0020] According to a third aspect of the present invention, there is provided a virtual image generation system for use by a replay user. The virtual image generation system includes a memory configured to store audio content data and video content data originating from at least one object (e.g., a real object and / or a virtual object) in an original spatial environment, a plurality of loudspeakers, and a display subsystem. In one embodiment, the display subsystem is configured to be positioned in front of the eyes of an end user. In another embodiment, the display subsystem includes a projection subsystem and a partially transparent display surface. In this case, the projection subsystem can be configured to project image frames onto the partially transparent display surface, and the partially transparent display surface can be configured to be positioned in a field of view between the eyes of the end user and the surrounding environment. The virtual image generation system can further include a frame structure configured to be worn by the end user and to carry at least a portion of the display subsystem.
[0021] The virtual image generation system further includes a control subsystem configured to retrieve the audio content data and the video content data from the memory, render audio and video from the retrieved audio content data and the video content data, respectively, in a new spatial environment different from the original spatial environment, and synchronously transmit the rendered audio to the loudspeaker and the generated video data to the display subsystem.
[0022] In one embodiment, the control subsystem is configured to store the audio content data and the video content data in the memory. The virtual image generation system can further include an object selection device configured to receive input from an end user and continuously select an object in the original spatial environment in response to the end user input prior to storing the audio content data and the video content data in the memory.
[0023] If the object includes a real object, the virtual image generation system can further include a microphone assembly configured to capture the audio content data from the real object in the original spatial environment. The microphone assembly can be configured to generate an audio output, in which case the control subsystem can be further configured to modify the directional audio output to preferentially sense sounds originating from the selected real object. The audio content data can be derived from the modified audio output. The virtual image generation system can further include one or more cameras configured to capture video data from the selected real object in the original spatial environment. In an optional embodiment, the control subsystem can be configured to transform the captured video data into virtual content data for the selected real object, and store the virtual content data as the video content data in the memory.
[0024] If the object includes a virtual object, the virtual image generation system can further include a database configured to store content data corresponding to sounds for a plurality of virtual objects, in which case the control subsystem can be further configured to retrieve the content data corresponding to the virtual object from the database, and the audio data stored in the memory can include the retrieved content data.
[0025] In one embodiment, the control subsystem is configured to obtain absolute metadata corresponding to at least one object in the new spatial environment, and render audio in the new spatial environment in accordance with the retrieved audio content data and the absolute metadata. Obtaining absolute metadata corresponding to an object in the new spatial environment can include positioning the object in the new spatial environment. In this case, the virtual image generation system can further include a user input device configured to receive input from a replay user, in which case the control subsystem can be configured to position the object in the new spatial environment in response to the input from the replay user. The virtual image generation system can further include one or more sensors configured to track a head pose of the replay user, in which case the control subsystem can be further configured to localize the absolute metadata to the replay user based on the tracked head pose of the replay user such that the rendered audio is spatialized.
[0026] According to a fourth aspect of the application, there is provided a method of operating a virtual image generation system by a replay user to replay audio and video of at least one object (e.g., a real object and / or a virtual object) in an original spatial environment previously recorded as audio content data and video content data. The method includes retrieving the audio content data and the video content data from a memory. A method can further include storing the audio content data and the video content data in the memory. In this case, the method can further include continuously selecting the object in the original spatial environment prior to storing the audio content data and the video content data in the memory.
[0027] If the object includes a real object, the method can further include capturing the audio content data from the real object. In this case, the method can further include preferentially sensing sound originating from the selected real object relative to sound originating from other real objects. The audio content data is derived from the preferentially sensed sound. The method can further include capturing video data from the selected real object, and transforming the captured video data into virtual content data. If the object includes a virtual object, the method can further include storing content data corresponding to sound for a plurality of virtual objects, and retrieving the content data corresponding to the virtual object from the database. The audio content data stored in the memory can include the retrieved content data.
[0028] The method also includes rendering audio and video from the retrieved audio content data and video content data in a new spatial environment different from the original spatial environment, transforming the audio and video into sound and image frames, respectively, and synchronously delivering the sound and image frames to the replay user. A method also includes obtaining absolute metadata corresponding to objects in the new spatial environment, in which case the audio is rendered in the new spatial environment from the retrieved audio content data and the absolute metadata. The method can also include tracking a head pose of the replay user, and localizing the absolute metadata to the replay user based on the tracked head pose of the replay user, in which case the audio can be rendered in the new spatial environment from the retrieved audio content data and the localized metadata such that the rendered audio is spatialized. Obtaining absolute metadata corresponding to objects in the new spatial environment can include, for example, positioning the objects in the new spatial environment in response to input from the replay user.
[0029] Additional and other objects, features and advantages of the present application will be described in the detailed description which follows, and will be apparent to those of ordinary skill in the art. It is to be understood that both the foregoing information and the following detailed description are exemplary and intended to provide a detailed description of various embodiments of the application and are not intended to limit the scope of the application. BRIEF DESCRIPTION OF DRAWINGS
[0030] The accompanying drawings illustrate the design and utility of preferred embodiments of the present application, in which like reference numerals denote similar elements throughout the various figures included. For better understanding of how the above-mentioned and other advantages and objects of the application can be obtained, a more particular description of the application briefly described above will be rendered by reference to specific embodiments thereof, which are illustrated in the accompanying drawings. It is to be noted, that the detailed description filed herein is intended to be illustrative only and is presented by way of example to describe the application. Various other embodiments can be described and understood in view of the drawings, in which:
[0031] Figure 1 is a picture of a three-dimensional augmented reality scene that can be displayed to an end user by an augmented reality generation device of the prior art;
[0032] Figure 2 is a perspective view of an augmented reality system constructed in accordance with one embodiment of the present application;
[0033] Figure 3 is a block diagram of the augmented reality system of Figure 2
[0034] Figure 4 is a plan view of one embodiment of a spatialized speaker system used in the augmented reality system of Figure 2
[0035] Figure 5 is a plan view showing one technique used by the augmented reality system of Figure 2
[0036] Figure 6 is a plan view showing another technique used by the augmented reality system of Figure 2
[0037] Figure 7 is a plan view showing yet another technique used by the augmented reality system of Figure 2
[0038] Figure 8 is a plan view showing one technique used by the augmented reality system of Figure 2
[0039] Figure 9 is a plan view showing another technique used by the augmented reality system of Figure 2
[0040] Figure 10a is a plan view of one technique usable by the augmented reality system of Figure 2
[0041] Figure 10b is a plan view of another technique usable by the augmented reality system of Figure 2
[0042] Figure 10c is a plan view of yet another technique usable by the augmented reality system of Figure 2
[0043] Figure 10d is a plan view of still another technique usable by the augmented reality system of Figure 2
[0044] Figure 11 is a block diagram showing the augmented reality system of Figure 2
[0045] Figure 12 is a block diagram showing one embodiment of an audio processor used in the augmented reality system of Figure 2
[0046] Figure 13 is a diagram of a memory recording content data and metadata corresponding to virtual and real objects selected by the augmented reality system of Figure 2
[0047] Figure 14 is a schematic diagram of a microphone assembly and corresponding audio processing module used in the augmented reality system of Figure 2
[0048] Figure 15a is a plan view of a directional pattern generated by an audio processor of an augmented reality system of Figure 2 to preferentially receive sound from two objects having a first orientation relative to an end user;
[0049] Figure 15b is a plan view of a directional pattern generated by an audio processor of an augmented reality system of Figure 2 to preferentially receive sound from two objects having a second orientation relative to an end user;
[0050] Figure 16a is a block diagram of objects distributed in an original spatial environment relative to an end user;
[0051] Figure 16b is a block diagram of objects distributed in a new spatial environment relative to an end user of Figure 16a ; and
[0052] Figure 17 is a flowchart illustrating one method of operating an augmented reality system of Figure 2 to select and record audio and video of virtual and real objects; and
[0053] Figure 18 is a flowchart illustrating one method of operating an augmented reality system of Figure 2 to replay audio and video recorded in Figure 17 in a new spatial environment. DETAILED DESCRIPTION
[0054] The following description relates to display systems and methods to be used in an augmented reality system. It should be understood, however, that while the present application is well suited to application in an augmented reality system, the present application can not be limited to such use in its broadest aspects. For example, the present application can be applied in a virtual reality system. Thus, although often described herein in terms of an augmented reality system, the teachings should not be so limited. An augmented reality system can operate in the context of, for example, a video game, a teleconference with a combination of virtual and real people, or a movie viewing.
[0055] The augmented reality system described herein allows an end user to record audio data originating from at least one object (virtual or real) continuously selected by the end user. This recorded audio data can then be replayed by the same or a different end user. The sound originating from the recorded audio data can be replayed to the same or a different end user in the real environment in which the audio data was originally recorded. In addition to the content of the recorded audio data, metadata can be recorded in association with such audio data characterizing the environment and the head pose of the end user at which the audio content was originally recorded, so that during replay the audio can be re-rendered and transformed into spatialized sound that is aurally experienced in the same way as the end user aurally experienced the spatialized sound during the original recording. Optionally, the audio can be re-rendered and transformed into spatialized sound for perception by the same or a different end user in a new virtual or real environment, so that the same or a different end user can have an aural experience that is appropriate for the new environment. The audio data can be recorded in synchronization with video data originating from virtual objects and real objects in the surrounding environment.
[0056] The augmented reality system described herein can be operated to provide images of virtual objects mixed with real (or physical) objects in the field of view of the end user, as well as to provide virtual sound originating from virtual sources (inside or outside the field of view) mixed with real sound originating from real (or physical) sources (inside or outside the field of view). To this end, reference will now be made to Figure 2 and 3 An embodiment of an augmented reality system 100 constructed in accordance with the present application is described. The augmented reality system 100 includes a display subsystem 102 that includes a display screen 104 and a projection subsystem (not shown) that projects images onto the display screen 104.
[0057] In the illustrated embodiment, the display screen 104 is a partially transparent display screen through which the end user 50 can see real objects in the surrounding environment and on which images of virtual objects can be displayed. The augmented reality system 100 further includes a frame structure 106 worn by the end user 50 that carries the partially transparent display screen 104 so that the display screen 104 is positioned in front of the eyes 52 of the end user 50, in particular in the field of view of the end user 50 between the eyes 52 of the end user 50 and the surrounding environment.
[0058] The display subsystem 102 is designed to present to the eyes 52 of the end user 50 photo-based radiance patterns that can be comfortably perceived as augmentations to the physical reality with a high level of image quality and three-dimensional perception as well as to be able to present two-dimensional content. The display subsystem 102 presents a sequence of frames with a high frequency of perception that provides a single coherent scene.
[0059] In alternative embodiments, the augmented reality system 100 can use one or more imagers (e.g., cameras) to capture images of the surrounding environment and transform them into video data, which can then be mixed with video data representing virtual objects, in which case the augmented reality system 100 can display images representative of the mixed video data to the end user 50 on an opaque display surface.
[0060] Further details describing the display subsystem are described in U.S. Provisional Patent Application Serial No. 14 / 212,961 entitled "Display Subsystem and Method" and U.S. Provisional Patent Application Serial No. 14 / 331,216 entitled "Planar Waveguide Apparatus With Diffraction Element(s) and Subsystem Employing Same," which are expressly incorporated herein by reference.
[0061] The augmented reality system 100 also includes one or more speakers 108 for presenting only sounds from virtual objects to the end user 50 while allowing the end user 50 to hear directly sounds from real objects. In alternative embodiments, the augmented reality system 100 can include one or more microphones (not shown) to capture real sounds originating from the surrounding environment and transform them into audio data, which can be mixed with audio data from virtual sounds, in which case the speakers 108 can transmit sounds representative of the mixed audio data to the end user 50.
[0062] In any case, the speakers 108 are carried by the frame structure 106 so that the speakers 108 are positioned near (in or around) the ear canal of the end user 50, e.g., earbuds or headphones. The speakers 108 can provide stereophonic / shapeable sound control. While the speakers 108 are described as being positioned near the ear canal, other types of speakers that are not located near the ear canal can also be used to transmit sounds to the end user 50. For example, the speakers can be placed at a distance from the ear canal, e.g., using bone conduction technology. In Figure 4In the illustrated optional embodiment, the plurality of spatialized speakers 108 can be positioned around the head 54 of the end user 50 (e.g., four speakers 108-1, 108-2, 108-3, and 108-4) configured to receive sound from the left, right, front, and back of the head 54 and directed to the left and right ears 56 of the end user 50. Further details of spatialized speakers that can be used in an augmented reality system are described in U.S. Provisional Patent Application Serial No. 62 / 369,561, entitled "Mixed Reality System with Spatialized Audio," which is expressly incorporated herein by reference.
[0063] Importantly, the augmented reality system 100 is configured to allow the end user 50 to select one, several, or all of the objects (virtual or real) to record sound only from those selected objects. To this end, the augmented reality system 100 further includes an object selection device 110 configured to select one or more real objects (i.e., real objects from which real sound originates) and virtual objects (i.e., virtual objects from which virtual sound originates) to record sound therefrom in response to input from the end user 50. The object selection device 110 can be designed to select real objects or virtual objects in the field of view of the end user 50 individually and / or to select a subset or all of the real objects or virtual objects in the field of view of the end user 50 globally. The object selection device 110 can also be configured to deselect one or more previously selected real objects or virtual objects in response to additional input from the end user 50. In this case, the object selection device 110 can be designed to deselect real objects or virtual objects in the same manner as they were previously selected. In any case, the continuous selection of a particular object means that the particular object remains in the selected state until intentionally deselected.
[0064] In one embodiment, the display subsystem 102 can display a three-dimensional cursor in the field of view of the end user 50 that, in response to input into the object selection device 110, can be shifted in the field of view of the end user 50 for selecting a particular real object or virtual object in the augmented reality scene.
[0065] For example, as Figure 5As shown, four virtual objects (V1-V4) and two real objects (R1-R2) are located within the field of view 60 of the display 104. The display subsystem 102 can display a 3D cursor 62 in the field of view 60, which is shown in the figure in the form of a circle. In response to input by the end user 50 into the object selection device 110, the 3D cursor 62 can be moved over an object, and in this case, over the virtual object V3, thereby associating the 3D cursor 62 with the object. Then, in response to additional input by the end user 50 into the object selection device 110, the associated object can be selected. To provide visual feedback that a particular object, in this case, the virtual object V3, is associated with the 3D cursor 62 and ready for selection, the associated object or even the 3D cursor 62 itself can be highlighted (e.g., a change in color or shading). After selection, the object can remain highlighted until deselected. Of course, instead of or in addition to the virtual object V3, other objects in the augmented reality scene 4, including real objects, can be selected by placing the 3D cursor 62 over any of the other objects in the augmented reality scene 4 and selecting the object within the 3D cursor 62. It should also be understood that although the 3D cursor 62 is shown in the figure in the form of a circle, the 3D cursor 62 can be any shape, including an arrow, that the end user 50 can use to point to a particular object. Any of the previously selected objects in the field of view 60 can be deselected by moving the 3D cursor 62 over the previously selected object and deselecting the object. Figure 5
[0066] The object selection device 110 can take the form of any device that allows the end user 50 to move the 3D cursor 62 over a particular object and subsequently select the particular object. In one embodiment, the object selection device 110 takes the form of a conventional physical controller, such as a mouse, touchpad, joystick, directional button, etc., that can be physically manipulated to move the 3D cursor 62 over a particular object and "clicked" to select the particular object.
[0067] In another embodiment, the object selection device 110 can include a microphone and a corresponding speech interpretation module that, in response to speech commands, can move the 3D cursor 62 over a particular object and then select the particular object. For example, the end user 50 can speak directional commands, such as move left or move right, to continually move the 3D cursor 62 over a particular object and then speak a command such as "select" to select the particular object.
[0068] In yet another embodiment, object selection apparatus 110 can include one or more cameras (e.g., forward-facing camera 112) mounted to frame structure 106 and a corresponding processor (not shown) capable of tracking a physical gesture (e.g., finger movement) of end user 50 that correspondingly moves 3D cursor 62 over a particular object for selection of the particular object. For example, end user 50 can use a finger to "drag" 3D cursor 62 over the particular object within field of view 60 and then "tap" 3D cursor 62 to select the particular object. Or, for example, based at least in part on an orientation of head 54 of end user 50, forward-facing camera 112 can be used, for example, to detect or infer a center of attention of end user 50 that correspondingly moves 3D cursor 62 over the particular object for selection of the particular object. For example, end user 50 can move his or her head 50 to "drag" 3D cursor 62 over the particular object within field of view 60 and then snap his or her head 50 to select the particular object.
[0069] In yet another embodiment, object selection apparatus 110 can include one or more cameras (e.g., rear-facing camera 114 Figure 2corresponding processor that tracks the eyes 52 of the end user 50, particularly the direction and / or distance at which the end user 50 is focusing, which correspondingly moves the 3D cursor 62 over a particular object for selection of the particular object. The rear-facing camera 114 can track the angular position (direction in which the eyes are pointing), the blinking, and the depth of focus (by detecting eye convergence) of the eyes 52 of the end user 50. For example, the end user 50 can move his or her eyes 54 within the field of view to "drag" the 3D cursor over a particular object and then blink to select the particular object. Such eye tracking information can be discerned, for example, by projecting light at the end user's eyes and detecting the return or reflection of at least some of the projected light. Further details of eye tracking devices are discussed in U.S. Provisional Patent Application Serial No. 14 / 212,961 entitled "Display Subsystem and Method," U.S. Patent Application Serial No. 14 / 726,429 entitled "Methods and Subsystem for Creating Focal Planes in Virtual and Augmented Reality," and U.S. Patent Application Serial No. 14 / 205,126 entitled "Subsystem and Method for Augmented and Virtual Reality," which are expressly incorporated herein by reference.
[0070] In alternative embodiments, the object selection device 110 can combine a conventional physical controller, a microphone / voice interpretation module, and / or a camera to move and use the 3D cursor 62 to select objects. For example, a physical controller, a finger gesture, or eye movement can be used to move the 3D cursor 62 over a particular object, and a voice command can be used to select the particular object.
[0071] Rather than using the 3D cursor 62 to select an object in the field of view of the end user 50, a particular object can be selected by semantically recognizing the particular object or by selecting the object from a menu displayed to the end user 50, in which case the object need not be located in the field of view of the end user 50. In this case, if the particular object is semantically recognized, the object selection means 110 takes the form of a microphone and speech interpretation module that translates verbal commands provided by the end user 50. For example, if the virtual object V3 corresponds to a drum, the end user 50 can say "select drum," in response to which the drum V3 will be selected. To facilitate selection of an object corresponding to a verbal command, semantic information identifying all relevant objects in the field of view is preferably stored in a database, so that a verbal expression of an object by the end user 50 can be matched to a description of an object stored in the database. Metadata including semantic information can be pre-associated with virtual objects in the database, while real objects in the field of view can be pre-mapped and associated with semantic information in the manner described in U.S. Patent Application Serial No. 14 / 704,800, entitled "Method and System for Inserting Recognized Object Data into a Virtual World," which is expressly incorporated herein by reference.
[0072] Alternatively, a particular object can be selected without using the 3D cursor 62 simply by pointing or "clicking" on the particular object using a finger gesture. In this case, the object selection means 110 can include one or more cameras (e.g., the forward-facing camera 114) and a corresponding processor that tracks the finger gesture to select the particular object. For example, the end user 50 can simply select a particular object (in this case, the virtual object V3) by pointing at the particular object, as shown in Figure 6 In another embodiment, a particular object can be selected without using the 3D cursor 62 by forming a circle or partial circle using at least two fingers (e.g., the index finger and the thumb), as shown in Figure 7
[0073] Although the 3D cursor 62 has been described as being used to select only one object at a time, in an alternative or optional embodiment, the 3D cursor 62 can be used to select multiple objects at a time. For example, as shown in Figure 8 a line 64 can be drawn around a group of objects (e.g., around the real object Rl and the virtual objects V3 and V4) using the 3D cursor 62, thereby selecting the group of objects. The 3D cursor 62 can be controlled using, for example, the same means described above for selecting objects individually. Alternatively, a line can be drawn around a group of objects without using the 3D cursor 62, for example, by using a finger gesture.
[0074] In an optional embodiment, a set of objects within a predefined angular range of the field of view of the end user 50 can be selected, in which case the object selection means 110 can take the form of a single physical or virtual selection button that can be actuated by the end user 50 to select these objects, for example. The angular range of the field of view can be predefined by the end user 50 or can be preprogrammed into the augmented reality system 100. For example, as shown in Figure 9 the angular range 66 of 60 degrees (± 30 degrees from the center of the field of view) is shown in the context of a field of view 60 of 120 degrees. All objects within the angular range 64 of the field of view 60 (in this case, virtual objects VI, V2, and V3) can be selected globally upon actuation of the selection button, while all objects outside the angular range 64 of the field of view 60 (in this case, real objects Rl and R2 and virtual object V4) will not be selected upon actuation of the selection button. In one embodiment, the end user 50 can modify the angular range, for example, by dragging one or both edges of the defined angular range toward or away from the centerline of the field of view 60 (shown by the arrows). The end user 50 can adjust the angular range, for example, from a minimum of 0 degrees to the entire field of view (e.g., 120 degrees). Alternatively, the angular range 64 of the field of view 60 can be preprogrammed without the need for the end user 50 to be able to adjust it. For example, all objects in the entire field of view 60 can be selected in response to actuation of the selection button.
[0075] The augmented reality system 100 further includes one or more microphones configured to convert sound from real objects in the surrounding environment into audio signals. In particular, the augmented reality system 100 includes a microphone assembly 116 configured to preferentially receive sound at a particular direction and / or a particular distance corresponding to the direction and distance of one or more real objects selected by the end user 50 via the object selection means 110. The microphone assembly 116 includes an array of microphone elements 118 (e.g., four microphones) mounted to the frame structure 106, as shown in Figure 2 The details of the microphone assembly 116 will be described in further detail below. The augmented reality system 100 further includes a dedicated microphone 122 configured to convert speech of the end user 50 into audio signals, for example, for receiving commands or dictations from the end user 50.
[0076] The augmented reality system 100 tracks the position and orientation of selected real objects within a known coordinate system such that, with respect to unselected real objects, sounds originating from these real objects can be preferentially and continuously sensed by the microphone assembly 116 even if the position or orientation of the selected real objects relative to the augmented reality system changes. The positioning and location of all virtual objects in the known coordinate system are generally "known" (i.e., recorded in) the augmented reality system 100 and thus generally do not need to be actively tracked.
[0077] In the illustrated embodiment, the augmented reality system 100 employs a spatialized audio system that renders and presents spatialized audio corresponding to virtual objects having known virtual positions and orientations in a real and physical three-dimensional (3D) space such that, to the end user 50, the sounds appear to originate from the virtual positions of the real objects so as to affect the clarity or realism of the sounds. The augmented reality system 100 tracks the position of the end user 50 to more accurately render the spatialized audio such that the audio associated with various virtual objects appears to originate from their virtual positions. In addition, the augmented reality system 100 tracks the head pose of the end user 50 to more accurately render the spatialized audio such that directional audio associated with various virtual objects appears to propagate in virtual directions appropriate to the individual virtual objects (e.g., outside the mouth of a virtual character, rather than the back of the virtual character's head). In addition, the augmented reality system 100 takes into account other real physical and virtual objects when rendering the spatialized audio such that the audio associated with various virtual objects appears to be appropriately reflected or occluded or blocked by the real physical and virtual objects.
[0078] To this end, the augmented reality system 100 also includes a head / object tracking subsystem 120 for tracking the position and orientation of the head 54 of the end user 50 relative to the virtual three-dimensional scene and tracking the position and orientation of real objects relative to the head 54 of the end user 50. For example, the head / object tracking subsystem 120 can include one or more sensors configured to collect head pose data (position and orientation) of the end user 50 and a processor (not shown) configured to determine the head pose of the end user 50 in the known coordinate system based on the head pose data collected by the sensors 120. The sensors can include one or more of image capture devices (e.g., visible and infrared light cameras), inertial measurement units (including accelerometers and gyroscopes), compasses, microphones, GPS units, or radio devices. In the illustrated embodiment, the sensors include a forward-facing camera 112 (as shown in FIG. 1) and a microphone 114. The head / object tracking subsystem 120 can also include a GPS unit 122 and a radio device 124. The head / object tracking subsystem 120 can also include a processor (not shown) configured to determine the head pose of the end user 50 in the known coordinate system based on the head pose data collected by the sensors 120. Figure 2The forward-facing camera 120 is particularly suitable for capturing information indicative of the distance and angular position (i.e. the head pointing direction) of the head 54 of the end user 50 relative to the environment in which the end user 50 is located, when worn on the head in this way. The head direction can be detected in any direction (e.g. up / down, left, right relative to a reference frame of the end user 50). As will be described in further detail below, the forward-facing camera 114 is also configured to acquire video data of real objects in the surrounding environment, in order to facilitate the video recording function of the augmented reality system 100. Cameras can also be provided for tracking real objects in the surrounding environment. The frame structure 106 can be designed such that cameras can be mounted on the front and back of the frame structure 106. In this way, the array of cameras can surround the head 54 of the end user 50 to cover relevant objects in all directions.
[0079] The augmented reality system 100 further comprises a three-dimensional database 124 configured to store virtual three-dimensional scenes, including virtual objects (content data of the virtual objects and absolute metadata associated with these virtual objects, e.g. absolute position and orientation of these virtual objects in the 3D scene) and virtual objects (content data of the virtual objects, absolute metadata associated with these virtual objects (e.g. volume and absolute position and orientation of these virtual objects in the 3D scene), and spatial acoustics around each virtual object, including any virtual or real objects in the vicinity of the virtual source, room dimensions, wall / floor materials, etc.).
[0080] The augmented reality system 100 further comprises a control subsystem which, in addition to recording video data originating from virtual objects and real objects appearing in the field of view, also records audio data originating from only those virtual objects and real objects which have been selected by the end user 50 via the object selection means 110. The augmented reality system 100 can also record metadata associated with the video data and audio data, such that the synchronized video and audio can be accurately re-rendered during playback.
[0081] To this end, the control subsystem includes a video processor 126 configured to obtain video content and absolute metadata associated with the virtual objects from the three-dimensional database 124, head pose data of the end user 50 from the head / object tracking subsystem 120 (to be used to localize the absolute metadata for the video to the head 54 of the end user 50, as described in further detail below), and render video from the video content, and then pass the video to the display subsystem 102 for transformation into images mixed with images originating from real objects in the surrounding environment in the field of view of the end user 50. The video processor 126 is also configured to obtain video data originating from real objects in the surrounding environment from the forward-facing camera 112, which is then recorded with the video data originating from the virtual objects, as will be described further below.
[0082] Similarly, an audio processor 128 is configured to obtain audio content and metadata associated with the virtual objects from the three-dimensional database 124, head pose data of the end user 50 from the head / object tracking subsystem 120 (to be used to localize the absolute metadata for the audio to the head 54 of the end user 50, as described in further detail below), and render spatialized audio from the audio content, and then pass the spatialized audio to the speakers 108 for transformation into spatialized sound mixed with sound originating from real objects in the surrounding environment.
[0083] The audio processor 128 is also configured to obtain audio data from only selected real objects in the surrounding environment from the microphone assembly 116, which is then recorded with the spatialized audio data from the selected virtual objects, any resulting metadata for the head 54 of the end user 50 to which each virtual object is localized (e.g., position, orientation, and volume data), and global metadata (e.g., volume data set globally by the augmented reality system 100 or the end user 50), as will be described further below.
[0084] The augmented reality system 100 also includes a memory 130, a recorder 132 configured to store video and audio in the memory 130, and a player 134 configured to retrieve video and audio from the memory 130 for subsequent playback to the end user 50 or other end users. The recorder 132 obtains spatialized audio data (audio content audio data and metadata) corresponding to the selected virtual and real objects from the audio processor 128 and stores the spatialized audio data in the memory 130, and further obtains video data (video content data and metadata) corresponding to the selected virtual and real objects in accordance with the virtual and real objects that the selected virtual and real objects are consistent with. Although the player 134 is shown as being located in the same AR system 100 as the recorder 132 and the memory 130, it should be understood that the player can be located in a third-party AR system or even on a smartphone or computer that is replaying video and audio previously recorded by the AR system 100.
[0085] The control subsystems that perform the functions of the video processor 126, the audio processor 128, the recorder 132, and the player 134 can take any of a variety of forms and can include multiple controllers, such as one or more microcontrollers, microprocessors, or central processing units (CPUs), digital signal processors, graphics processing units (GPUs), other integrated circuit controllers, such as application specific integrated circuits (ASICs), programmable gate arrays (PGAs) (e.g., field PGAs (FPGAs)), and / or programmable logic controllers (PLUs).
[0086] The functions of the video processor 126, the audio processor 128, the recorder 132, and / or the player 134 can be performed by separate integrated devices, at least some of the functions of the video processor 126, the audio processor 128, the recorder 132, and / or the player 134 can be combined into a single integrated device, or the functions of each of the video processor 126, the audio processor 128, the recorder 132, or the player 134 can be distributed among several devices. For example, the video processor 126 can include a graphics processing unit (GPU) that obtains video data for virtual objects from the three-dimensional database 124 and renders composite video frames from the video data, and a central processing unit (CPU) that obtains video frames for real objects from the forward-facing camera 112. Similarly, the audio processor 128 can include a digital signal processor (DSP) that processes audio data obtained from the microphone assembly 116 and the user microphone 122, and a CPU that processes audio data obtained from the three-dimensional database 124. The recording functions of the recorder 132 and the playback functions of the player 134 can be performed by CPUs.
[0087] Further, various processing components of augmented reality system 100 can be physically contained in a distributed subsystem. For example, as shown in Figures 10a-10d augmented reality system 100 includes a local processing and data module 150 operatively coupled (e.g., by wired leads or wireless connectivity 152) to components mounted to the head 54 of the end user 50 (e.g., a projection subsystem of the display subsystem 102, a microphone assembly 116, speakers 104, and cameras 114, 118). The local processing and data module 150 can be mounted in various configurations, such as fixedly attached to the frame structure 106 Figure 10a ), fixedly attached to a helmet or hat 106a Figure 10b ), embedded in headphones, removably attached to the torso 58 of the end user 50 Figure 10c ), or removably attached to the hip 59 of the end user 50 in a belt-coupling configuration Figure 10d ). Augmented reality system 100 also includes a remote processing module 154 and a remote data repository 156 operatively coupled (e.g., by wired leads or wireless connectivity 158, 160) to the local processing and data module 150, so that these remote modules 154, 156 are operatively coupled to each other and available as sources of information to the local processing and data module 150.
[0088] The local processing and data module 150 can include a power-efficient processor or controller, as well as digital memory, such as flash memory, both of which can be used to help manage data
[0089] The couplings 152, 158, 160 between the various components described above can include one or more wired or wireless interfaces or ports to provide for optical or electrical communication, such as by RF, microwave, or IR. In some implementations, all communication can be wired, while in other implementations all communication can be wireless, except for the optical fiber used in the display subsystem 102. In still further implementations, a choice of wired and wireless communication can be made in dependence on the particular implementation and as desired. Figures 10a-10dThe differences are shown. Therefore, the specific choice between wired or wireless communication should not be regarded as a limitation.
[0090] In the illustrated embodiment, the light source and driving electronics (not shown) of the display subsystem 102, the processing components of the head / object tracking subsystem 120 and the object selection device 110, and the DSP of the audio processor 128 may be included in the local processing and data module 150. The GPU of the video processor 126 and the CPUs of the video processor 126 and the audio processor 128 may be included in the remote processing module 154, although in alternative embodiments, these components or portions thereof may be included in the local processing and data module 150. The 3D database 124 and the memory 130 may be associated with the remote data storage 156.
[0091] The processing and recording of audio data from virtual and real objects selected by end user 50 will be described in more detail. Figure 3 The audio processor 128 is shown. Figure 11 In the exemplary scenario shown, end user 50 (e.g., a parent) wants to record sounds from a four-piece band (including a virtual drummer V2 object, a real singer R2 (e.g., a child), a virtual guitarist V3, and a virtual bass guitarist V4), wants to monitor news or sports on a virtual TV V1 without recording sounds from the virtual TV, and also does not want to record sounds from a real kitchen R1 (e.g., someone cooking).
[0092] exist Figure 12 In the illustrated embodiment, the audio processor 128's functionality is distributed between a CPU 180 and a DSP 182. The CPU 180 processes audio originating from virtual objects, while the DSP 182 processes audio originating from real objects. The CPU 180 includes one or more effects modules 184 (in this case, effects modules 1-n) configured to generate spatialized audio data EFX-V1 to EFX-Vn corresponding to each virtual object V1-Vn. To this end, the effects modules 184 retrieve audio content data AUD-V1 to AUD-Vn and absolute metadata MD corresponding to the virtual objects V1-Vn from a 3D database 124. a -V1 to MD a -Vn and acquire head pose data from the head / object tracking subsystem 120, and convert absolute metadata MD based on the head pose data. a -V1 to MD a -Vn is localized to the head 54 of the end user 50, and localized metadata (e.g., location, orientation, and volume data) is applied to the audio content data to generate spatialized audio data for virtual objects V1-Vn.
[0093] The CPU 180 also comprises a mixer 186 configured to mix the spatialized audio data EFX-V1 to EFX-Vn received from the various effect modules 184 to obtain mixed audio data EFX, and a global effects module 188 configured to apply global metadata MD-OUT (e.g. global volume) to the mixed spatialized audio data to obtain final spatialized audio AUD-OUT EFX output through the plurality of channels to the loudspeakers 108.
[0094] Importantly, the effect modules 184 are configured to send audio content data originating from the virtual objects that have been selected by the end user 50 via the object selection means 110 and the metadata (localized and / or absolute) corresponding to these selected virtual objects to the logger 132 for storage in the memory 130 (as Figure 2 illustrated), and the global effects module 188 is configured to send the global metadata MD-OUT to the logger 132 for storage in the memory 130. In the exemplary embodiment, the audio content data AUD-V2 (i.e. virtual drummer), AUD-V3 (i.e. virtual guitarist), AUD-V4 (i.e. virtual bassist) are selected for recording, whereas the audio content data AUD-V1 (i.e. virtual TV) is not selected for recording. Accordingly, the audio content data AUD-V2, AUD-V3 and AUD-V4 and the corresponding localized metadata MD-V2, MD-V3 and MD-V4 are stored in the memory 130, as Figure 13 illustrated.
[0095] In an alternative embodiment, instead of or in addition to storing the audio content data from the selected virtual objects and the corresponding localized / absolute metadata and global metadata separately within the memory 130, the CPU 180 outputs spatialized audio generated by additionally mixing the spatialized audio data EFX-V2, EFX-V3, EFX-V4 corresponding only to the selected virtual objects AUD-V2, AUD-V3 and AUD-V4 and applying the global metadata MD-OUT to this mixed spatialized audio data to obtain spatialized audio comprising only audio from the selected virtual objects AUD-V2, AUD-V3 and AUD-V4. However, in this case, additional audio mixing functionality needs to be incorporated into the CPU 180.
[0096] The DSP 182 is configured to process audio signals acquired from the microphone assembly 116 and output audio signals that preferentially represent sound received by the microphone assembly 116 from a particular direction, and in this case, from the direction of each real object selected by the end user 50 via the object selection device 110. As the position and / or orientation of the real objects can move relative to the head 54 of the end user 50, real object tracking data can be received from the head / object tracking subsystem 120 so that any changes in the position and / or orientation of the real objects relative to the head 54 of the end user 50 can be taken into account so that the DSP 182 can dynamically modify the audio output to preferentially represent sound received by the microphone assembly 116 from the direction of the relatively moving real objects. For example, if the end user 50 moves his or her head 54 90 degrees counterclockwise relative to the orientation of the head 54 while a real object is selected, the preferential direction of the audio output from the DSP 182 can be dynamically shifted 90 degrees clockwise.
[0097] Reference is made to Figure 14 The microphone elements 118 of the microphone assembly 116 take the form of a phased array of microphone elements (in this case, microphone elements M1-Mn), each of which is configured to detect an ambient sound signal and convert the ambient sound signal into an audio signal. In the illustrated embodiment, the microphone elements 118 are digital in nature, thus converting the ambient sound signal into a digital audio signal, and in this case, a pulse density modulated (PDM) signal. Preferably, the microphone elements 118 are spaced apart from one another to maximize the directionality of the audio output. For example, as shown, two of the microphone elements 118 can be mounted to each arm of the frame structure 106, although more than two microphone elements (e.g., four microphone elements 118) can be mounted to each arm of the frame structure 106. Alternatively, the frame structure 106 can be designed so that the microphone elements 118 can be mounted to the front and rear of the frame structure 106. In this manner, the array of microphone elements 118 can surround the head 54 of the end user 50 to cover potential sound sources in all directions. Figure 2
[0098] The microphone assembly 116 also includes a plurality of digital microphone interfaces (DMICs) 190 (in this case, DMIC1 through DMICn, one for each microphone element M) that are configured to receive respective digital audio signals from the corresponding microphone elements 118 and perform a digital filter operation known as "decimation" to transform the digital audio signals from PDM format to pulse code modulation (PCM), which is easier to manipulate. Each of the DMICs 190 also performs fixed gain control on the digital audio signals.
[0099] DSP 182 includes multiple audio processing modules 200, each configured to process digital audio signals output from microphone assembly 116 and output directional audio signals AUD-R (one of directional audio signals AUD-R1 to AUD-Rm), which preferably represent sound received by microphone assembly 116 in the direction of a selected real object (one of R1 to Rm). The directional audio signals AUD-R1 to AUD-Rm output by each audio processing module 200 are combined into a directional audio output AUD-OUT MIC, which preferably represents sound originating from all selected real objects. In the illustrated embodiment, DSP 182 creates an instance of audio processing module 200 for each real object selected by end user 50 via object selection device 110.
[0100] To this end, each audio processing module 200 includes processing parameters in the following form: a plurality of delay elements 194 (in this case, delay elements D1-Dn, one for each microphone element M), a plurality of gain elements 196 (in this case, gain elements G1-Gn, one for each microphone element M), and an adder 198. The delay elements 194 apply delay factors to the amplified digital signals received from the corresponding gain amplifier 192 of the microphone assembly, and the gain elements 196 apply gain factors to the delayed digital signals. The adder 198(S) adds the gain-adjusted and delayed signals to generate corresponding directional audio signals AUD-R.
[0101] Microphone elements 118 are spatially arranged, and delay elements 194 and gain elements 196 of each audio processing module 200 are applied to the digital audio signals received from microphone elements 116 in a manner that generates received ambient sound according to a directional polarity pattern (i.e., sound arriving from a specific directional direction or angle is emphasized more than sound arriving from other directional directions). DSP 182 is configured to modify the directivity of the directional audio signals AUD-R1 to AUD-Rm by changing the delay factor of delay element 194 and the gain factor of gain element 196, thereby modifying the combined directional audio output AUD-OUTMIC.
[0102] Therefore, it is understandable that the directionality of the audio output AUD-OUT MIC can be modified based on the selected real object; for example, the direction of preferred sound reception can be set along the direction of the selected real object or source.
[0103] For example, refer to Figure 15a If we choose to follow two specific directions D respectively a and D b Two real objects R a and Rb then the DSP 182 will generate two instances of the audio processing module 200, and within each of these audio processing modules 200, select respective delay factors and gain factors for all of the delay elements 194 and gain elements 196 in each audio processing module 200 such that a receive gain pattern is generated having two lobes aligned with the directions D a and D b of the real objects R a and R b If the directions of the real objects R a and R b change relative to the head 54 of the end user 50, then the particular directions of the real objects R a and R b may change, in which case the DSP 182 can select different delay factors and gain factors for all of the delay elements 194 and gain elements 196 in each audio processing module 200 such that the receive gain pattern has two lobes aligned with the directions D c and D d as shown in Figure 15b .
[0104] To facilitate this dynamic modification of the directivity of the audio output AUD-OUT MIC, different sets of delay / gain values and corresponding preferred directions can be stored in the memory 130 for access by the DSP 182. That is, the DSP 182 matches the direction of each selected real object R to the closest direction value stored in the memory 130 and selects a corresponding set of delay / gain factors for that selected direction.
[0105] It should be noted that although the microphone elements 118 are described as digital, the microphone elements 118 are optionally analog. Further, although the delay elements 194, gain elements 196, and summer 198 are disclosed and shown as software components that exist inside the DSP 182, any one or more of the delay elements 194, gain elements 196, and summer 198 can comprise analog hardware components that exist outside of the DSP 182 but under the control of the DSP 182. However, the use of software-based audio processing modules 200 allows for sounds from several different real objects to be simultaneously preferentially received and processed.
[0106] Referring back to Figure 12DSP 182 also receives voice data from user microphone 122 and combines it with directional audio output AUD-OUT MIC. In an optional embodiment, DSP 182 is configured to perform acoustic echo cancellation (AEC) and noise suppression (NS) functions with respect to the sound from speaker 108 originating from the virtual object. That is, even though the direction in which the sound is preferentially received can not coincide with speaker 108, microphone assembly 116 can sense the sound emitted by speaker 108. To this end, the spatialized audio data output by global effects module 188 into speaker 108 is also input to DSP 182, which uses the spatialized audio data to suppress the resulting sound (treated as noise) output by speaker 108 into microphone assembly 116 and to cancel any echoes resulting from feedback from speaker 108 to microphone assembly 116.
[0107] Importantly, DSP 182 is also configured to send directional audio output AUD-OUT MIC and localization metadata (e.g., the position and orientation of the real object from which directional audio output AUD-OUT MIC originates) to recorder 132 for storage in memory 130 as audio content data (as shown in Figure 2 In the exemplary embodiment shown in Figure 11 Importantly, DSP 182 is also configured to send directional audio output AUD-OUT MIC and localization metadata (e.g., the position and orientation of the real object from which directional audio output AUD-OUT MIC originates) to recorder 132 for storage in memory 130 as audio content data (as shown in Figure 13
[0108] In an optional embodiment, directional audio output AUD-OUT MIC (which can be spatialized) can be input to speaker 108 or other speakers for playback to end user 50. Directional audio output AUD-OUT MIC can be spatialized in the same manner as the spatialized audio data originating from the virtual source so that end user 50 hears the sound as originating from the position of the real object, thereby affecting the clarity or realism of the sound. That is, localization metadata (e.g., the position and orientation of the real object from which directional audio output AUD-OUT MIC preferentially originates) can be applied to directional audio output AUD-OUT MIC to obtain spatialized audio data.
[0109] In another alternative embodiment, the sound originating from real objects or even virtual objects selected by the end user 50 can be profiled. Specifically, the DSP 182 can analyze the characteristics of the sound from the selected objects and compare them to the characteristics of the sound originating from other real objects in order to determine the type of the target sound. The DSP 182 can then include all the audio data originating from these real objects in the directional audio output AUD-OUT MIC if needed, to be recorded by the logger 132 into the memory 130 (as shown, for example). For example, if the end user 50 selects any of the music objects (AUD-V2, AUD-V3, AUD-V4, AUD-R2), the DSP 182 can control the microphone assembly 116 to preferentially sense all the music real objects. Figure 2
[0110] In the illustrated embodiment, the DSP 182 continues to output the directional audio output AUD-OUT MIC to the logger 130 for recording in the memory 130 even if the real object 198 selected by the end user 50 moves out of the field of view of the display subsystem 102 (as indicated by the real object tracking data received from the head / object tracking subsystem 120). In an alternative embodiment, the DSP 182 stops outputting the directional audio output AUD-OUT MIC to the logger 130 for recording in the memory 130 once the real object 198 selected by the end user 50 moves out of the field of view of the display subsystem 102, and resumes outputting the directional audio output AUD-OUT MIC to the logger 130 for recording in the memory 130 once the real object 198 selected by the end user 50 moves back into the field of view of the display subsystem 102.
[0111] In a similar manner to the audio processor 128 (CPU 180 and DSP 182 in the illustrated embodiment) sending the audio content data (in the example case, audio content data AUD-V2, AUD-V3 and AUD-V4 and AUD-MIC) and the localization metadata (in the example case, MD-V2, MD-V3, MD-V4 and MD-R2) originating from the selected virtual and real objects and the global metadata (MD-OUT) to the logger 132 for storage in the memory 130, the video processor 126 can send the video content data (in the example case, video content data VID-V2, VID-V3, VID-V4 and VID-R2) originating from the virtual and real objects to the logger 132 for storage in the memory 130. Figure 13 In the case of virtual objects, the video processor 126 simply retrieves the virtual objects from the 3D database 124 without further processing and sends these virtual objects to the recorder 132 for storage in the memory 130. In the case of real objects, the video processor 126 can extract or "cut out" any of the selected real objects from the video acquired from the camera 112 and store these real objects as virtual objects in the memory 130. In Figure 11 In the exemplary case shown, a video of the real singer R2 can be recorded as a virtual object VID-R2. In alternative embodiments, the video processor 126 sends the entire video acquired from the camera 112 (including video corresponding to unselected virtual and real objects) to the recorder 132 for storage in the memory 130.
[0112] The player 134 is configured to replay the video and / or audio recorded within the memory 130 to a replay user 50' (as shown in FIG. 1), which can be the original end user 50 who recorded the video / audio or a third party user. In response to commands given by the replay user 50', e.g., voice commands via the user microphone 122, the player 134 can selectively replay the audio / video. For example, the replay user 50' can use a "virtual audio on / off" command to turn on or off virtual audio replay, or a "display on / off" command to turn on or off virtual video replay, or a "real audio on / off" command to turn on or off real audio replay. Figure 16a
[0113] In the embodiment shown, the audio processor 128 retrieves the audio content data and metadata (corresponding to the selected virtual and real objects) from the memory 130, renders spatialized audio from the audio content data and metadata, and communicates the spatialized audio to the player 134 for replay to the replay user 50' through the speakers 108. In alternative embodiments where mixed spatialized audio data (rather than content and metadata) is stored, the player 134 can simply retrieve the audio data from the memory 130 for replay to the replay user 50' without re-rendering the audio data or otherwise further processing the audio data.
[0114] Further, in the illustrated embodiment, the video processor 126 retrieves the video content data and metadata (corresponding to the selected virtual and real objects), renders a video from the video content data and metadata, and passes the video to the player 134 for playback to the playback user 50' via the display subsystem 102 in synchronization with the audio being played back via the speakers 108. Alternatively, where all of the video data captured by the camera 112 is stored, the player 134 can simply retrieve the video data from the memory 130 to play back to the playback user 50' without rendering or otherwise further processing the video data. The augmented reality system 10 can provide the playback user 50' with the option to play back only the video corresponding to the selected virtual objects and real objects or to play back the full video captured by the camera 112.
[0115] In one embodiment, the current head pose of the playback user 50' is not taken into account during playback of the video / audio. However, the video / audio is played back to the playback user 50' using the head pose originally detected during the recording of the video / audio data, which will be reflected in the localization metadata stored with the audio / video content data in the memory 130 or, if the mixed spatialized audio is recorded without metadata, the head pose will be reflected within the mixed spatialized audio stored in the memory 130. In this case, the playback user 50' will experience the video / audio in the same way that the original end user 50 experienced the video / audio, except that only the audio and optionally only the video from the sources of the virtual and real objects selected by the original end user 50 will be played back. In this case, the playback user 50' can not be immersed in the augmented reality since the head pose of the playback user 50' will be taken into account. However, the playback user 50' can experience the audio playback using headphones (so the audio will not be affected by the environment) or the playback user 50' can experience the audio playback in a quiet room.
[0116] In an alternative embodiment, the current head pose of the replay user 50' can be taken into account during the replay of the video / audio. In this case, the head pose of the replay user 50' during the recording of the video / audio does not need to be incorporated into the metadata stored in the memory 130 along with the video / audio content data, since the current head pose of the replay user 50' detected during the replay will be used to re-render the video / audio data. However, the absolute metadata stored in the memory 130 (e.g., the volume and absolute position and orientation of these virtual objects in the 3D scene, and the spatial acoustics around each virtual object including any virtual or real objects in the vicinity of the virtual sources, room dimensions, wall / floor materials, etc.) will be localized to the head pose of the replay user 50' using the current head pose of the replay user 50', and then this absolute metadata is used to render the audio / video. Thus, during the replay of the video / audio, the replay user 50' will be immersed in an augmented reality.
[0117] The replay user 50' can experience the augmented reality in the original spatial environment (e.g., "same physical space") in which the video / audio was recorded or can experience the augmented reality in a new physical or virtual spatial environment (e.g., "different physical or virtual room").
[0118] If the replay user 50' experiences the augmented reality in the original spatial environment in which the video / audio was recorded, then there is no need to modify the absolute metadata associated with the selected objects in order to accurately replay the spatialized audio. Conversely, if the replay user 50' experiences the augmented reality in a new spatial environment, then it can be necessary to modify the absolute metadata associated with the objects in order to accurately render the audio / video in the new spatial environment.
[0119] For example, in an exemplary embodiment, the audio / video content from the virtual objects AUD-V2 (i.e., virtual drummer), AUD-V3 (i.e., virtual guitarist), AUD-V4 (i.e., virtual bass guitarist), and the real object (i.e., real singer) can be recorded in a small room 250, as shown in Figure 16a The previously recorded audio from the virtual objects AUD-V2 (i.e., virtual drummer), AUD-V3 (i.e., virtual guitarist), AUD-V4 (i.e., virtual bass guitarist), and the real object (i.e., real singer) can be replayed in a concert hall 252, as shown in Figure 16bAs shown. Augmented reality system 10 can reposition objects to any location in concert hall 252 and can generate or otherwise acquire absolute metadata including the new location of each object in concert hall 252 and the spatial acoustics around each object in concert hall 252. This absolute metadata can then be localized using the current head pose of replay user 50', and then used to render audio and video in concert hall 252 for replay to replay user 50'.
[0120] The layout and functions of the augmented reality system 100 have already been described; now, referring to... Figure 17 A method 300 is described that uses an augmented reality system 100 to select at least one object and record audio and video from the selected object.
[0121] First, the end user 50 continuously selects at least one object (e.g., real and / or virtual) in the spatial environment via the object selection device 110 (step 302). This can be achieved, for example, by moving a 3D cursor 62 within the end user 50's field of view 60 and selecting objects using the 3D cursor 62 (e.g., ...). Figure 5 (As shown), while selecting objects in the field of view 60 of the end user 50. Alternatively, hand gestures (such as...) can be used. Figure 6 (As shown in Figure 7) or by using voice commands to select objects. Multiple objects can be selected individually, or multiple objects can be selected globally, for example, by drawing lines around the objects (as shown in Figure 64). Figure 8 (as shown), or by limiting the angular range 66 of the field of view 60 of the terminal user 50 (which may be smaller than the entire angular range of the field of view 60 of the terminal user 50) and selecting all objects (such as...) within the limited angular range 66 of the field of view 60 of the terminal user 50. Figure 9 (As shown).
[0122] Next, the audio and video content of all virtual objects within the spatial environment, as well as the absolute metadata associated with the virtual objects, are acquired (step 304). Next, the current head pose of the end user 50 is tracked (step 306), and the absolute metadata is localized to the head 54 of the end user 50 using the current head pose data (step 308). This absolute metadata is then applied to the audio and video content of the virtual objects to obtain video data and spatialized audio data for all virtual objects within the corresponding virtual objects (step 310). The spatialized audio data of all virtual objects within the corresponding virtual objects in the 3D scene is blended (step 312), and global metadata is applied to the blended spatialized audio data to obtain the final spatialized audio for all virtual objects in the 3D scene (step 314). This final spatialized audio is then transformed into sound perceived by the end user 50 (step 316). Next, the video data obtained in step 310 is transformed into image frames perceived by the end user 50 (step 318). Next, record the audio / video content and all associated metadata (absolute and localized metadata) of all virtual objects selected by the end user 50 in step 302 (step 320).
[0123] In parallel with steps 304-320, the position and / or orientation of the selected real object relative to the head 54 of the end user 50 is tracked (step 322), and sound from the selected real object is sensed preferentially based on the tracked position and orientation of the real object (step 324). Next, an image of the selected real object is captured (step 326), and optionally transformed into virtual video content. Next, audio content associated with the preferentially sensed sound from the selected real object and video content associated with the captured image of the selected real object, along with all associated metadata (position and orientation of the real object), are recorded for each of the selected real objects (step 328).
[0124] Now refer to Figure 18 A method 400 is described that uses an augmented reality system 100 to replay audio and video of at least one previously recorded object to a replay user 50'. Such audio and video may have been recorded as described above. Figure 17 The method described in method 300 is recorded as audio content data and video content data. The objects can be real and / or virtual and can be continuously selected by the end user 50. In exemplary method 400, the audio and video have previously been recorded in an original spatial environment (e.g., small room 250) and are replayed in a new spatial environment different from the original spatial environment (e.g., concert hall 252), as per [the relevant information]. Figure 16a and 16b As stated above.
[0125] First, previously recorded audio content data and video content data are retrieved (step 402). If the new spatial environment is at least partially virtual, additional virtual content (audio or video) associated with the new spatial environment can also be retrieved. Then, objects can be repositioned in the new spatial environment in response to input from the replay user 50' (step 404). Then, absolute metadata corresponding to the objects positioned in the new spatial environment is retrieved (step 406), the head pose of the replay user 50' is tracked in the new spatial environment (step 408), and the absolute metadata is localized to the replay user 50' based on the tracked head pose of the replay user 50' (step 410). Next, audio and video from the retrieved audio content data and video content data are rendered based on the localized metadata in the new spatial environment (step 412). Then, the rendered audio and video are transformed into sound and image frames, respectively, that are synchronously perceived by the replay user 50' (step 414).
[0126] In the foregoing specification, the application has been described with reference to specific embodiments thereof. It is evident, however, that various modifications and changes can be made thereto without departing from the broader spirit and scope of the application. For example, the order of process actions can be changed, and other processes can be added or factors altered. As an example, process flows have been described with reference to the certain ordering of process actions. However, many of the process actions described can take place at the same time or in other sequences that are not described in this specification. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.
Claims
1. A method of operating a virtual image generation system by an end user, comprising: continually selecting at least one object including a virtual object; generating video data originating from the at least one selected object; rendering a plurality of image frames in a three-dimensional scene from the generated video data; displaying the image frames to the end user; generating audio data originating from the at least one selected object; and storing the audio data originating from the at least one selected object in a memory, generating metadata corresponding to the selected virtual object, wherein generating the metadata includes retrieving absolute metadata corresponding to the selected virtual object and localizing the absolute metadata to the end user based on a tracked head pose of the end user.
2. The method of claim 1, further comprising storing the video data in the memory in synchronization with the audio data.
3. The method of claim 1, further comprising transforming the audio data originating from the at least one selected object into sound perceived by the end user. selecting the at least one object in a field of view of the end user.
4. The method of claim 1, wherein, selecting the at least one object includes moving a three-dimensional cursor in the field of view of the end user and selecting the at least one object with the three-dimensional cursor.
5. The method of claim 4, wherein, selecting the at least one object includes issuing one or more voice commands.
6. The method of claim 1, wherein, 7. The method of claim 1, selecting the at least one object includes making one or more hand gestures. the at least one object includes a plurality of objects, and selecting the plurality of objects includes selecting the objects individually.
8. The method of claim 1, wherein, the at least one object includes a plurality of objects, and selecting the plurality of objects includes selecting the objects globally.
9. The method of claim 1, wherein, globally selecting the objects includes defining an angular range of a field of view of the end user and selecting all of the objects within the defined angular range of the field of view of the end user.
10. The method of claim 9, wherein, the defined angular range is less than an entire angular range of the field of view of the end user.
11. The method of claim 10, wherein, the defined angular range is an entire angular range of the field of view of the end user.
12. The method of claim 10, wherein, 13. The method of claim 1, further comprising continually deselecting the at least one previously selected object.
14. The method of claim 1, further comprising tracking a location of the at least one selected object relative to a field of view of the end user. stopping storing the audio data originating from the at least one selected object in the memory when the tracked location of the at least one selected object moves out of the field of view of the end user.
15. The method of claim 14, further comprising: continuing storing the audio data originating from the at least one selected object in the memory when the tracked location of the at least one selected object moves out of the field of view of the end user.
16. The method of claim 14, further comprising: the at least one selected object further includes a real object.
17. The method of claim 1, wherein, preferentially sensing sound originating from the selected real object relative to sound originating from other real objects, wherein the audio data is derived from the preferentially sensed sound.
18. The method of claim 17, further comprising: 19. The method of claim 17, further comprising: capturing video data originating from the selected real object; and storing the video data in the memory in synchronization with the audio data.
20. The method of claim 19, further comprising: transforming the captured video data into virtual content data, and storing the virtual content data in the memory.
21. The method of claim 1, further comprising: storing content data corresponding to sound for a plurality of virtual objects; and retrieving the content data corresponding to the selected virtual object, wherein the audio data stored in the memory includes the retrieved content data.
22. The method of claim 21, wherein, the audio data stored in the memory includes the retrieved content data and the generated metadata.
23. The method of claim 22, wherein, the metadata includes position, orientation, and volume data for the selected virtual object.
24. The method of claim 22, further comprising: tracking the head pose of the end user; and storing the absolute metadata for the plurality of virtual objects.
25. The method of claim 1, further comprising: retrieving the stored audio data, deriving audio from the retrieved audio data, and transforming the audio into sound perceived by the end user.
26. The method of claim 1, wherein, the stored audio data includes content data and metadata, the method further comprising: retrieving the stored content data and metadata from the memory; rendering spatialized audio based on the retrieved content data and metadata; and transforming the spatialized audio into sound perceived by the end user.
Citation Information
Patent Citations
Display system and method
US20140267420A1
Firearm trigger pull training system and methods
US9728095B1
Methods and system for creating focal planes in virtual and augmented reality
US9857591B2
Method for displaying image combined with playing audio in an electronic device
CN104065869A
System and method for augmented and virtual reality
CN105188516A