Method and apparatus for presenting audio and synthetic reality experiences

By correlating the timeline of audio files with user environment analysis, synchronous synthetic reality content is acquired and presented in real time, solving the problem of non-immersive audiovisual experience in existing technologies, realizing an immersive audio/synthetic reality experience, and enhancing the user's interaction and perception with virtual objects.

CN112189183BActive Publication Date: 2026-04-17APPLE INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
APPLE INC
Filing Date
2019-05-29
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing audiovisual experiences, such as music videos and algorithmic audio visualizations, are often not truly immersive and are not tailored to the user's environment.

Method used

By correlating the timeline of audio files with the analysis of the user's environment, synchronized synthetic reality content is acquired and presented in real time. The display of virtual objects is adjusted in real time using head-mounted devices and controllers to match the user's physical environment.

Benefits of technology

It achieves an immersive audio/synthetic reality experience, enhances user interaction and perception with virtual objects, and improves environmental adaptability and interactivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112189183B_ABST
    Figure CN112189183B_ABST
Patent Text Reader

Abstract

In various implementations, methods for presenting an audio / SR experience are disclosed. In one embodiment, when an audio file is played in an environment, the SR content event is displayed in association with the environment in response to determining a corresponding time criterion and a corresponding environmental criterion that satisfy the SR content event. In one embodiment, SR content is acquired based on a 3D point cloud of the audio file and the environment and displayed in association with the environment. In one embodiment, SR content is acquired based on spoken words from real sounds in the environment and displayed in association with the environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates in general to audio and synthetic reality experiences, and more particularly to systems, methods, and apparatus for presenting synthetic reality experiences accompanied by audio. Background Technology

[0002] A physical setting refers to a world that an individual can perceive and / or interact with without the aid of electronic systems. A physical setting (e.g., a physical forest) includes physical elements (e.g., physical trees, physical structures, and physical animals). Individuals can directly interact with and / or perceive a physical setting, such as through touch, sight, smell, hearing, and taste.

[0003] In contrast, synthetic reality (SR) scenes refer to computer-created scenes, wholly or partially, that can be perceived and / or interacted with by an individual via an electronic system. In SR, a subset of an individual's movement is monitored, and in response, one or more properties of one or more virtual objects within the SR scene are changed in a manner consistent with one or more physical laws. For example, an SR system might detect that an individual has taken a few steps forward and, in response, adjust the graphics and audio presented to the individual in a manner similar to how such scenes and sounds would change in a physical environment. Modifications to one or more properties of one or more virtual objects within the SR scene may also be made in response to representations of movement (e.g., audio commands).

[0004] An individual can interact with and / or perceive SR objects using any of their senses, including touch, smell, vision, taste, and sound. For example, an individual can interact with and / or perceive auditory objects that create multidimensional (e.g., three-dimensional) or spatial auditory settings and / or achieve auditory transparency. Multidimensional or spatial auditory settings provide an individual with the perception of discrete auditory sources in a multidimensional space. Auditory transparency selectively combines sound from the physical setting, with or without computer-generated audio. In some SR settings, an individual can interact with and / or perceive only the auditory object.

[0005] An example of SR is Virtual Reality (VR). A VR scene is a simulated scene designed to include only computer-created sensory input for at least one sense. A VR scene includes multiple virtual objects that an individual can interact with and / or perceive. An individual can interact with and / or perceive the virtual objects in a VR scene by simulating a subset of their own actions within the computer-created scene and / or by simulating the individual or their presence within the computer-created scene.

[0006] Another example of SR is Mixed Reality (MR). An MR set refers to a simulated set designed to integrate computer-generated sensory input (e.g., virtual objects) with sensory input from a physical set or its representation. In the reality spectrum, mixed reality sets lie between VR sets at one end and fully physical sets at the other, and do not include either of these sets.

[0007] In some MR scenes, computer-generated sensory input can adapt to changes in sensory input from the physical scene. Additionally, some electronic systems used to present MR scenes can monitor orientation and / or position relative to the physical scene, enabling virtual objects to interact with real objects (i.e., physical elements from the physical scene or their representations). For example, the system can monitor motion so that virtual plants appear stationary relative to physical buildings.

[0008] An example of mixed reality is augmented reality (AR). An AR scene refers to a simulated scene in which at least one virtual object is superimposed on a physical scene or its representation. For example, an electronic system may have an opaque display and at least one imaging sensor for capturing images or videos of the physical scene, which are representations of the physical scene. The system combines the images or videos with virtual objects and displays the combination on the opaque display. An individual uses the system to indirectly view the physical scene via images or videos of the physical scene and observes the virtual objects superimposed on the physical scene. When the system uses one or more image sensors to capture images of the physical scene and uses those images to present the AR scene on an opaque display, the displayed images are referred to as video pass-through. Alternatively, the electronic system for displaying the AR scene may have a transparent or semi-transparent display through which an individual can directly view the physical scene. The system may display virtual objects on the transparent or semi-transparent display, allowing an individual to observe the virtual objects superimposed on the physical scene using the system. As another example, the system may include a projection system that projects virtual objects onto the physical scene. Virtual objects can be projected, for example, onto a physical surface or as holograms, allowing individuals to use the system to observe virtual objects superimposed on a physical setting.

[0009] Augmented reality scenes can also refer to simulated scenes in which the representation of a physical scene is altered by sensory information created by a computer. For example, a portion of the representation of a physical scene can be graphically altered (e.g., magnified) such that the altered portion still represents one or more initially captured images, but is not a faithful reproduction. As another example, when providing video pass-through, the system can alter at least one of the sensor images to impose a specific viewpoint different from the viewpoint captured by one or more image sensors. Furthermore, the representation of a physical scene can be altered by graphically blurring or removing portions of it.

[0010] Another example of mixed reality is augmented virtual (AV). An AV scene refers to a computer-created or virtual scene incorporating at least one sensory input from a physical scene. The one or more sensory inputs from the physical scene can be a representation of at least one feature of the physical scene. For example, virtual objects can exhibit colors of physical elements captured by one or more imaging sensors. Similarly, virtual objects can exhibit features consistent with actual weather conditions in the physical scene, such as those identified via weather-related imaging sensors and / or online weather data. In another example, an augmented reality forest can have virtual trees and structures, but the animals can have features accurately reproduced from images taken of physical animals.

[0011] Many electronic systems enable individuals to interact with and / or perceive various SR (Real-Time) scenes. One example includes a head-mounted system. The head-mounted system may have an opaque display and one or more speakers. Alternatively, the head-mounted system may be designed to receive external displays (e.g., smartphones). The head-mounted system may have one or more imaging sensors and / or microphones for capturing images / videos of the physical scene and / or capturing audio of the physical scene. The head-mounted system may also have a transparent or semi-transparent display. The transparent or semi-transparent display may be combined with a substrate through which light representing the image is directed to the individual's eyes. The display may incorporate LEDs, OLEDs, digital light projectors, laser scanning light sources, liquid crystal on silicon, or any combination of these technologies. The substrate transmitting light may be an optical waveguide, an optical combiner, a light reflector, a holographic substrate, or any combination of these substrates. In one embodiment, the transparent or semi-transparent display may selectively switch between an opaque state and a transparent or semi-transparent state. As another example, the electronic system may be a projection-based system. A projection-based system may use retinal projection to project images onto the individual's retina. Alternatively, projection systems can project virtual objects onto a physical set (e.g., onto a physical surface or as a hologram). Other examples of SR systems include head-up displays, car windshields capable of displaying graphics, windows capable of displaying graphics, lenses capable of displaying graphics, headphones or earbuds, speaker arrangements, input mechanisms (e.g., controllers with or without haptic feedback), tablets, smartphones, and desktop or laptop computers. While music is typically an audio experience, lyrics, sound dynamics, or other features are suited to complement visual experiences. Previously available audiovisual experiences, such as music videos and / or algorithmic audio visualizations, are not truly immersive and / or not tailored to the user's environment. Attached Figure Description

[0012] Therefore, this disclosure will be understood by those skilled in the art, and a more detailed description can be made with reference to aspects of some exemplary embodiments, some of which are shown in the accompanying drawings.

[0013] Figure 1A It is a block diagram of an exemplary operating architecture according to some implementation methods.

[0014] Figure 1B It is a block diagram of an exemplary operating architecture according to some implementation methods.

[0015] Figure 2 This is a block diagram of an exemplary controller according to some implementation methods.

[0016] Figure 3 This is a block diagram of an exemplary head-mounted device (HMD) according to some implementation methods.

[0017] Figure 4A-4G The SR volume environment during playback of a first audio file is shown according to some embodiments.

[0018] Figure 5A-5G Another SR volume environment during playback of a first audio file is shown according to some implementations.

[0019] Figure 6 An audio / SR experience data object is shown according to some implementations.

[0020] Figure 7 This is a flowchart representation of a first method for presenting an audio / SR experience according to some implementation methods.

[0021] Figures 8A-8B The following is illustrated during playback of a second audio file according to some embodiments. Figure 4A SR volume environment.

[0022] Figures 9A-9C The following describes the playback of a third audio file according to some embodiments. Figure 4A SR volume environment.

[0023] Figures 10A-10E The following is illustrated according to some embodiments during the playback of the fourth audio file. Figure 4A SR volume environment.

[0024] Figure 11 This is a flowchart representation of a second method for presenting an audio / SR experience according to some implementation methods.

[0025] Figures 12A-12E This illustrates, according to some embodiments, the process during a story told by a storyteller. Figure 4A SR volume environment.

[0026] Figure 13 This is a flowchart representation of a third method for presenting an audio / SR experience according to some implementation methods.

[0027] As is customary, the various features shown in the accompanying drawings may not be drawn to scale. Therefore, for clarity, the dimensions of various features may be arbitrarily expanded or reduced. Additionally, some drawings may not depict all components of a given system, method, or apparatus. Finally, similar reference numerals may be used throughout the specification and drawings to denote similar features. Summary of the Invention

[0028] The various embodiments disclosed herein include devices, systems, and methods for presenting an audio / SR experience. In various embodiments, a first method is performed by a device including one or more processors, non-transitory memory, a speaker, and a display. The method includes: storing an audio file having an associated timeline in the non-transitory memory. The method includes: storing a plurality of SR content events in the non-transitory memory in association with the audio file, wherein each of the plurality of SR content events is associated with a corresponding time standard and a corresponding environmental standard. When playing the audio file via the speaker, the method includes: using the processor to determine, based on the current position on the timeline of the audio file, the corresponding time standard that satisfies a particular SR content event among the plurality of SR content events. When playing the audio file via the speaker, the method includes: using the processor to determine, based on environmental data of the environment, the corresponding environmental standard that satisfies the particular SR event among the plurality of SR events. When playing the audio file via the speaker, the method includes: in response to determining the corresponding time standard and the corresponding environmental standard that satisfy the particular SR content event among the plurality of SR content events, displaying, in association with the environment, the particular SR content event among the plurality of SR content events on the display.

[0029] In various implementations, the second method is performed by a device including one or more processors, non-transitory memory, a speaker, and a display. The method includes: acquiring a three-dimensional (3D) point cloud of an environment. The method includes: acquiring SR content based on an audio file and the 3D point cloud of the environment. The method includes: simultaneously playing the audio file via the speaker and displaying the SR content on the display in association with the environment.

[0030] In various implementations, the third method is performed by a device including one or more processors, non-transitory memory, a speaker, and a display. The method includes: recording real sound generated in the environment via the microphone. The method includes: detecting one or more spoken words in the real sound using the one or more processors. The method includes: acquiring SR content based on the one or more spoken words. The method includes: displaying the SR content on the display in association with the environment.

[0031] According to some embodiments, an apparatus includes one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors. The one or more programs include instructions for performing or causing to perform any of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors of the apparatus, cause the apparatus to perform or cause to perform any of the methods described herein. According to some embodiments, an apparatus includes one or more processors, non-transitory memory, and means for performing or causing to perform any of the methods described herein. Detailed Implementation

[0032] Numerous details have been described to provide a thorough understanding of the exemplary embodiments illustrated in the accompanying drawings. However, the drawings illustrate only some exemplary aspects of this disclosure and should not be considered limiting. Those skilled in the art will understand that other effective aspects and / or variations do not include all the specific details set forth herein. Furthermore, well-known systems, methods, components, devices, and circuits have not been described exhaustively so as not to obscure further relevant aspects of the exemplary embodiments described herein.

[0033] As mentioned above, previously available audiovisual experiences, such as music videos and / or algorithmic audio visualizations, are not truly immersive and / or tailored to the user's environment. Therefore, in the various embodiments described herein, an audio / SR experience is presented. In the various embodiments described herein, the timeline of the audio file is associated with curated SR content events displayed based on analysis of the user's environment. In the various embodiments, SR content is acquired and presented in the user's environment in real-time based on analysis of the audio file (e.g., audio or metadata [lyrics, title, artist, etc.]) during playback. In the various embodiments, SR content is presented based on spoken words detected in the audio heard by the user.

[0034] Figure 1A This is a block diagram of an exemplary operating architecture 100A according to some embodiments. Although relevant features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure further relevant aspects of the exemplary embodiments disclosed herein. Therefore, as a non-limiting example, operating architecture 100A includes electronic device 120A.

[0035] In some embodiments, electronic device 120A is configured to present CGR content to a user. In some embodiments, electronic device 120A includes a suitable combination of software, firmware, and / or hardware. According to some embodiments, electronic device 120A presents SR content to a user via display 122 when the user is physically present within physical environment 103, the physical environment including table 107 within the field of view 111 of electronic device 120A. In some embodiments, the user holds electronic device 120A in one or both of his / her hands. In some embodiments, while providing augmented reality (AR) content, electronic device 120A is configured to display AR objects (e.g., AR cube 109) and implement video pass-through of physical environment 103 (e.g., representation 117 including table 107) on display 122.

[0036] Figure 1B This is a block diagram of an exemplary operating architecture 100B according to some embodiments. Although relevant features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and to avoid obscuring further relevant aspects of the exemplary embodiments disclosed herein. Therefore, as a non-limiting example, the operating environment 100B includes a controller 110 and a head-mounted device (HMD) 120B.

[0037] In some implementations, controller 110 is configured to manage and coordinate the presentation of SR content to the user. In some implementations, controller 110 includes a suitable combination of software, firmware, and / or hardware. Reference is made below. Figure 2 The controller 110 is described in more detail. In some embodiments, the controller 110 is a computing device located locally or remotely relative to scene 105. For example, the controller 110 is a local server located within scene 105. Alternatively, the controller 110 is a remote server (e.g., a cloud server, central server, etc.) located outside scene 105. In some embodiments, the controller 110 is communicatively coupled to the HMD 120B via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). Alternatively, the controller 110 is included within the housing of the HMD 120B.

[0038] In some implementations, the HMD 120B is configured to present SR content to a user. In some implementations, the HMD 120B includes a suitable combination of software, firmware, and / or hardware. References are provided below. Figure 3 HMD 120B is described in more detail. In some implementations, the functionality of controller 110 is provided by and / or combined with HMD 120B.

[0039] According to some implementations, HMD 120B provides SR content to the user when the user is virtually and / or physically present in scene 105.

[0040] In some implementations, the user wears the HMD 120B on his / her head. Therefore, the HMD 120B includes one or more SR displays provided for displaying SR content. For example, in various implementations, the HMD 120B surrounds the user's field of view. In some implementations, such as Figure 1A As shown, a handheld device (such as a smartphone or tablet) configured to display SR content is used instead of the HMD 120B, and the user no longer wears the HMD 120B but holds the device, with the display facing the user's field of view and the camera facing scene 105. In some embodiments, the handheld device may be placed inside a housing that can be worn on the user's head. In some embodiments, an SR pod, housing, or chamber configured to display SR content is used instead of the HMD 120B, wherein the user no longer wears or holds the HMD 120B.

[0041] Figure 2 This is a block diagram of an example controller 110 according to some embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and in order not to obscure further relevant aspects of the embodiments disclosed herein. Therefore, as a non-limiting example, in some implementations, controller 110 includes one or more processing units 202 (e.g., microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZigBee, and / or similar type interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these components and various other components.

[0042] In some embodiments, the one or more communication buses 204 include circuitry for communication between interconnecting system components and control system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0043] Memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220 or the non-transitory computer-readable storage medium of memory 220 stores programs, modules, and data structures or subsets thereof, including optional operating system 230 and SR experience module 240.

[0044] Operating system 230 includes processes for handling various basic system services and for performing hardware-related tasks. In some embodiments, SR experience module 240 is configured to manage and coordinate single or multiple SR experiences for one or more users (e.g., single SR experiences for one or more users, or multiple SR experiences for corresponding groups of one or more users). To this end, in various embodiments, SR experience module 240 includes a data acquisition unit 242, a tracking unit 244, a coordination unit 246, and a data transmission unit 248.

[0045] In some implementations, the data acquisition unit 242 is configured to acquire at least data from the HMD 120B (e.g., presentation data, interaction data, sensor data, location data, etc.). To this end, in various implementations, the data acquisition unit 242 includes instructions and / or logic components for instructions, as well as heuristics and metadata for heuristics.

[0046] In some embodiments, the tracking unit 244 is configured to map scene 105 and at least track the position / location of HMD 120B relative to scene 105. To this end, in various embodiments, the tracking unit 244 includes instructions and / or logic components for instructions, as well as heuristics and metadata for the heuristics.

[0047] In some implementations, the coordination unit 246 is configured to manage and coordinate the SR experience presented to the user by the HMD 120B. To this end, in various implementations, the coordination unit 246 includes instructions and / or logic components for the instructions, as well as heuristics and metadata for the heuristics.

[0048] In some embodiments, the data transmission unit 248 is configured to transmit at least data (e.g., presentation data, location data, etc.) to the HMD 120B. To this end, in various embodiments, the data transmission unit 248 includes instructions and / or logic components for instructions, as well as heuristics and metadata for heuristics.

[0049] Although the data acquisition unit 242, tracking unit 244, coordination unit 246 and data transmission unit 248 are shown residing on a single device (e.g., controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 242, tracking unit 244, coordination unit 246 and data transmission unit 248 may reside in a separate computing device.

[0050] also, Figure 2 This is used more as a functional description of various features that may exist in a particular embodiment, and differs from the structural schematic diagrams of the embodiments described herein. As those skilled in the art will recognize, individually shown items may be combined, and some items may be separate. For example, Figure 2 Some functional modules shown individually can be implemented in a single module, and the various functions of a single functional block can be implemented through one or more functional blocks in various implementations. The actual number of modules and the division of specific functions, as well as how features are allocated therein, will vary depending on the implementation and, in some implementations, depend in part on the specific combination of hardware, software, and / or firmware selected for a particular embodiment.

[0051] Figure 3This is a block diagram of an example of an HMD 120B according to some embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and in order not to obscure further relevant aspects of the embodiments disclosed herein. Therefore, as a non-limiting example, in some implementations, the HMD 120B includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BlueTooth, ZigBee, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more SR displays 312, one or more internal and / or external image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these components and various other components.

[0052] In some embodiments, one or more communication buses 304 include circuitry for interconnecting and communicating between system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones 307A, one or more speakers 307B (e.g., headphones or loudspeakers), a haptic engine, and / or one or more depth sensors (e.g., structured light, time-of-flight, etc.).

[0053] In some embodiments, one or more SR displays 312 are configured to provide an SR experience to a user. In some embodiments, the one or more SR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, the one or more SR displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. For example, the HMD 120B includes a single SR display. Alternatively, the HMD 120B may include SR displays for each of the user's eyes. In some embodiments, the one or more SR displays 312 are capable of displaying AR and VR content.

[0054] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of a user's face (including the user's eyes) (and thus may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward in order to acquire image data corresponding to the scene that the user would see when the HMD 120B is not present (and thus may be referred to as a scene camera). The one or more image sensors 314 may include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.

[0055] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320 or the non-transitory computer-readable storage medium of memory 320 stores programs, modules, and data structures, or subsets thereof, including optional operating system 330 and SR rendering module 340.

[0056] Operating system 330 includes processes for handling various basic system services and for performing hardware-related tasks. In some embodiments, SR presentation module 340 is configured to present SR content to a user via one or more SR displays 312. For this purpose, in various embodiments, SR presentation module 340 includes a data acquisition unit 342, an audio / SR presentation unit 344, and a data transmission unit 348.

[0057] In some implementations, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from one or more of the controller 110 (e.g., via communication interface 308), I / O devices, and sensors 306 or one or more image sensors 314. To this end, in various implementations, the data acquisition unit 342 includes instructions and / or logic components for instructions, as well as heuristics and metadata for heuristics.

[0058] In some embodiments, the audio / SR presentation unit 344 is configured to present an audio / SR experience via one or more SR displays 312 (and, in various embodiments, speakers 307B and / or microphones 307A). To this end, in various embodiments, the SR presentation unit 344 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.

[0059] In some embodiments, the data transmission unit 346 is configured to transmit at least data (e.g., presentation data, location data, etc.) to the controller 110. To this end, in various embodiments, the data transmission unit 346 includes instructions and / or logic components for instructions, as well as heuristics and metadata for heuristics.

[0060] Although the data acquisition unit 342, the audio / SR presentation unit 344, and the data transmission unit 346 are shown residing on a single device (e.g., HMD 120B), it should be understood that in other embodiments, any combination of the data acquisition unit 342, the audio / SR presentation unit 344, and the data transmission unit 346 may reside in a separate computing device.

[0061] also, Figure 3 This is more of a functional description of various features that may exist in a particular embodiment, and differs from the structural schematic diagrams of the embodiments described herein. As those skilled in the art will recognize, individually shown items may be combined, and some items may be separate. For example, Figure 3Some functional modules shown individually can be implemented in a single module, and the various functions of a single functional block can be implemented through one or more functional blocks in various implementations. The actual number of modules and the division of specific functions, as well as how features are allocated therein, will vary depending on the implementation and, in some implementations, depend in part on the specific combination of hardware, software, and / or firmware selected for a particular embodiment.

[0062] Figure 4A An SR volumetric environment 400 based on a real-world environment surveyed by a scene camera of a device is illustrated. In various embodiments, the scene camera is part of a device worn by a user and including a display showing the SR volumetric environment 400. Thus, in various embodiments, the user is physically present in the environment. In various embodiments, the scene camera is part of a remote device (such as a drone or robot avatar) that transmits images from the scene camera to a local device worn by the user and including a display showing the SR volumetric environment 400.

[0063] Figure 4A The SR volume environment 400 is shown at the first moment during the playback of the first audio file (e.g., a song named "SongName1" by an artist named "ArtistName1").

[0064] The SR volume environment 400 includes multiple objects, including one or more real objects (e.g., a photograph 411, a table 412, a television 413, a lamp 414, and a window 415) and one or more virtual objects (e.g., an audio playback indicator 420). In various embodiments, each object is displayed at a location within the first SR volume environment 400 (e.g., at a location defined by three coordinates in a three-dimensional (3D) SR coordinate system). Thus, when a user moves within the SR volume environment 400 (e.g., changes position and / or orientation), the object moves on the HMD's display but maintains its position within the SR volume environment 400. In various embodiments, certain virtual objects (such as the audio playback indicator 420) are displayed at locations on the display such that when the user moves within the SR volume environment 400, the object remains stationary on the HMD's display.

[0065] The audio playback indicator 420 includes information about the playback of an audio file. In various embodiments, the audio file is associated with a timeline, such that various portions of the audio file are played at various times. In various embodiments, the audio playback indicator 420 includes text, such as the artist associated with the audio file and / or the title associated with the audio file. In various embodiments, the audio playback indicator 420 includes an audio progress bar that indicates the current position on the timeline of the audio file being played. In various embodiments, the audio playback indicator 420 includes event markers that indicate the time standard of SR content events. Although in Figure 4A The audio playback indicator 420 is shown in the figure, but in various embodiments, the audio playback indicator 420 is not shown, even when an audio file is being played.

[0066] Figure 4B The second time during the playback of the first audio file is shown. Figure 4A The SR volume environment 400 includes the first SR content event in response to determining that a first time criterion and a first environment criterion for the first SR content event are satisfied. The first time criterion for the first SR content event is satisfied when the current position on the timeline of the first audio file matches the trigger time (e.g., indicated by a first event marker of the audio playback indicator 420). The first environment criterion for the first SR content event is satisfied when the SR volume environment 400 includes a square object with a specific reflectivity. Figure 4B In this embodiment, the first environmental criterion is met because the SR volume environment 400 includes a photograph 411. In other embodiments, the first environmental criterion is met because the SR volume environment includes a digital photo frame or a framed graduation certificate. Displaying the first SR content event includes displaying a virtual object (e.g., broken glass 421) on top of a square object with a specific reflectivity.

[0067] Figure 4C The third time during the playback of the first audio file is shown. Figure 4A The SR volume environment 400 includes a second SR content event in response to determining a second time criterion and a second environment criterion for satisfying the second SR content event. The second time criterion for the second SR content event is satisfied when the current position on the timeline of the first audio file matches a trigger time (e.g., indicated by a second event marker of the audio playback indicator 420). The second environment criterion for the second SR content event is satisfied when the SR volume environment 400 includes an object with a specific shape having a long, slender portion with a larger portion at the top. Figure 4CIn this embodiment, the second environment criterion is met because the SR volume environment 400 includes a light 414. In other embodiments, the second environment criterion is met because the SR volume environment includes a tree, a sculpture, or a person wearing a large hat. Displaying the second SR content event includes displaying a virtual object (e.g., an alien 422) that exits from or behind a larger portion of an object of a particular shape and retracts and hides within or behind the larger portion.

[0068] Figure 4D The fourth time during the playback of the first audio file is shown. Figure 4A The SR volume environment 400 includes a third SR content event in response to determining a third time criterion and a third environment criterion for satisfying the third SR content event. The third time criterion for the third SR content event is satisfied when the current position on the timeline of the first audio file is within a trigger window (e.g., indicated by a third event marker of the audio playback indicator 420). The third environment criterion for the third SR content event is satisfied when the SR volume environment 400 includes a dynamic square object with a specific reflectivity. Figure 4D In this embodiment, the SR volume environment 400 satisfies the third environment standard because it includes a television 413. In other embodiments, the third environment standard is satisfied because the SR volume environment includes a digital picture frame or a computer monitor. Displaying third SR content events includes displaying virtual objects (e.g., a video clip 423 of ArtistName1 playing a part of a song) on ​​top of a dynamic square object with a specific reflectivity.

[0069] Figure 4E The fifth time point during the playback of the first audio file is shown. Figure 4A The SR volume environment 400 includes a fourth SR content event in response to determining that a fourth time criterion and a fourth environment criterion for the fourth SR content event are satisfied. The fourth time criterion for the fourth SR content event is satisfied when the current position on the timeline of the first audio file matches a trigger time (e.g., indicated by a fourth event marker on the audio playback indicator 420). The fourth environment criterion for the fourth SR content event is satisfied when the SR volume environment 400 includes a table. Figure 4E In this embodiment, the fourth environment criterion is satisfied because the SR volume environment 400 includes table 412. In other embodiments, the fourth environment criterion is satisfied because the SR volume environment includes a different table or another object classified as a table. Events that display fourth SR content include virtual objects moving on the table (e.g., another alien 424).

[0070] Figure 4F The sixth time during the playback of the first audio file is shown. Figure 4A The SR volume environment 400. In response to determining a fifth time criterion and a fifth environment criterion for satisfying a fifth SR content event, and further in response to determining another playback criterion for satisfying the fifth SR content event, the SR volume environment 400 includes the fifth SR content event. The fifth time criterion for the fifth SR content event is satisfied when the current position on the timeline of the first audio file matches the trigger time (e.g., indicated by a fifth event marker on the audio playback indicator 420). The fifth environment criterion for the fifth SR content event is satisfied (like the first environment criterion for the first SR content event) when the SR volume environment 400 includes a square object with a specific reflectivity. Figure 4F In this embodiment, the fifth environment criterion is met because the SR volume environment 400 includes photograph 411. In other embodiments, the fifth environment criterion is met because the SR volume environment includes a digital photo frame or a framed graduation certificate. Another playback criterion for the fifth SR content event is met when the first SR content event is previously displayed. Therefore, in various embodiments, the fifth SR content event is not displayed even if the fifth time criterion and the fifth environment criterion are met because the first SR content event is not displayed (e.g., because photograph 411 is not in the field of view of the scene camera at the corresponding trigger time). Displaying the fifth SR content event includes displaying a virtual object (e.g., broken glass 425) falling from the position of a square object with a specific reflectivity.

[0071] Figure 4G The seventh time during the playback of the first audio file is shown. Figure 4A The SR volume environment 400. In response to determining that the sixth time criterion and the sixth environment criterion of the sixth SR content event are satisfied, the SR volume environment 400 includes the sixth SR content event. The sixth time criterion of the sixth SR content event is satisfied when the current position on the timeline of the first audio file matches the trigger time (e.g., indicated by the sixth event marker of the audio playback indicator 420). The sixth environment criterion of the sixth SR content event is satisfied when the SR volume environment 400 is classified as internal. Figure 4G In this context, because the SR volume environment 400 is within the room, it meets the sixth environment criterion. Displaying sixth SR content events includes displaying virtual objects that have intruded into the room (e.g., another alien 426). Therefore, in various implementations, SR environments that include windows (such as inside a car or outside a house) but are classified as external will not trigger the display of sixth SR content events.

[0072] Figure 5A It shows a real-world environment surveyed by the device's scene camera (different from the real-world environment surveyed by the device). Figure 4A-4GThe SR volume environment 500 is a real-world environment. In various embodiments, a scene camera is part of a device worn by a user and including a display showing the first SR environment 500. Thus, in various embodiments, the user is physically present in the environment. In various embodiments, the scene camera is part of a remote device (such as a drone or robot avatar) that transmits images from the scene camera to a local device worn by the user and including a display showing the SR environment 500.

[0073] Figure 5A Another SR volume environment 500 is shown during the first moment of playback of the first audio file (e.g., a song named "SongName1" by an artist named "ArtistName1").

[0074] The SR volume environment 500 includes multiple objects, including one or more real objects (e.g., sky 511, tree 512, table 513, and beach 514) and one or more virtual objects (e.g., audio playback indicator 420). In various embodiments, each object is displayed at a location within the first SR volume environment 400 (e.g., at a location defined by three coordinates in a three-dimensional (3D) SR coordinate system). Thus, as the user moves within the SR volume environment 500 (e.g., changes position and / or orientation), the object moves on the HMD's display but maintains its position within the SR volume environment 500. In various embodiments, certain virtual objects (such as the audio playback indicator 420) are displayed at locations on the display such that when the user moves within the SR volume environment 400, the object appears stationary on the HMD's display.

[0075] Figure 5B The second time during the playback of the first audio file is shown. Figure 5A The SR volume environment 500. Because the SR volume environment 500 does not include square objects with a specific reflectivity, the first SR content event is not displayed (in Figure 4B (As shown in the image). However, in response to determining that the seventh time criterion and the seventh environment criterion of the seventh SR content event are satisfied, the SR volume environment 500 includes the seventh SR content event. The seventh time criterion of the seventh SR content event is satisfied when the current position on the timeline of the first audio file matches the trigger time (e.g., indicated by the first event marker of the audio playback indicator 420, e.g., the same trigger time of the first SR content event). The seventh environment criterion of the seventh SR content event is satisfied when the SR volume environment 500 includes the sky. Figure 5BIn this context, because the SR volume environment 500 includes the sky 511, it meets the fifth environment standard. Displaying the seventh SR content event includes displaying virtual objects moving in the sky (e.g., the spaceship 521).

[0076] Figure 5C The third time during the playback of the first audio file is shown. Figure 5A The SR volume environment is 500. Similar to... Figure 4C In response to determining that a second time criterion and a second environment criterion are satisfied for a second SR content event, the SR volume environment 500 includes the second SR content event. As described above, the second time criterion for the second SR content event is satisfied when the current position on the timeline of the first audio file matches the trigger time (e.g., indicated by a second event marker of the audio playback indicator 420), and the second environment criterion for the second SR content event is satisfied when the SR volume environment 500 includes an object with a specific shape having a long, slender portion with a larger portion at the top. Figure 5C In this context, because the SR volume environment 500 includes tree 512, it meets the second environment criterion. As described above, the event of displaying the second SR content includes displaying a virtual object (e.g., alien 422) that exits from or after the larger portion and retracts and hides within or after the larger portion.

[0077] Figure 5D The fourth time during the playback of the first audio file is shown. Figure 5A The SR volume environment 500. Because the SR volume environment 500 does not include dynamic square objects with a specific reflectivity, the third SR content event is not displayed (in Figure 4D (as shown in the image). Figure 5D In this implementation, the SR content event is not displayed in the fourth time period.

[0078] Figure 5E The fifth time point during the playback of the first audio file is shown. Figure 5A The SR volume environment is 500. Similar to... Figure 4E In response to determining that a fourth time criterion and a fourth environment criterion for a fourth SR content event are satisfied, the SR volume environment 500 includes the fourth SR content event. As described above, the fourth time criterion for the fourth SR content event is satisfied when the current position on the timeline of the first audio file matches the trigger time (e.g., indicated by the fourth event marker of the audio playback indicator 420), and the fourth environment criterion for the fourth SR content event is satisfied when the SR volume environment 400 includes a table. Figure 5EIn this context, because the SR volume environment 500 includes table 513, it meets the fourth environment standard. As mentioned above, events that display the fourth SR content include virtual objects moving on the table (e.g., another alien 424).

[0079] Figure 5F The sixth time during the playback of the first audio file is shown. Figure 5A The SR volume environment 500. In response to determining an eighth time criterion and an eighth environment criterion for satisfying the eighth SR content event, and further in response to determining another playback criterion for satisfying the eighth SR content event, the SR volume environment 500 includes the eighth SR content event. The eighth time criterion for the eighth SR content event is satisfied when the current position on the timeline of the first audio file matches the trigger time (e.g., indicated by the fifth event marker of the audio playback indicator 420). The eighth environment criterion for the eighth SR content event is satisfied when the SR volume environment 500 includes a sky (like the seventh environment criterion for the seventh SR content event). Figure 5F In this context, because the SR volume environment 500 includes the sky 511, the eighth environment criterion is met. Another playback criterion for the eighth SR content event is met when the seventh SR content event was previously displayed. Therefore, in various implementations, the eighth SR content event is not displayed even if the eighth time criterion and the eighth environment criterion are met, because the seventh SR content event is not displayed (e.g., because the sky 511 is not in the scene camera's field of view at the corresponding trigger time (e.g., when the user was inside but has moved outside)). Displaying the eighth SR content event includes displaying a virtual object (e.g., the spaceship 521 from the seventh SR content event) moving in the sky along with another virtual object (e.g., a fighter jet 525).

[0080] Figure 5G The seventh time during the playback of the first audio file is shown. Figure 5A The SR volume environment 500. Because SR volume environment 500 is not classified as internal, the sixth SR content event is not displayed (in Figure 4G (as shown in the image). However, in response to determining the ninth time criterion and the ninth environment criterion for satisfying the ninth SR content event, the SR volume environment 500 includes the ninth SR content event. The ninth time criterion for the ninth SR content event is satisfied when the current position on the timeline of the first audio file matches the trigger time (e.g., indicated by the sixth event marker of the audio playback indicator 420). The ninth environment criterion for the fifth SR content event is satisfied when the SR volume environment 500 is classified as external. Figure 5GIn this context, because the SR volume environment 500 is located outdoors, on a beach, it meets the Ninth Environment Standard. Displaying Ninth SR content events includes displaying multiple virtual objects excavated from the ground (e.g., multiple aliens 526A-526B).

[0081] Figure 6 An audio / SR experience data object 600 according to some embodiments is shown. The audio / SR experience data object 600 includes audio / SR experience metadata 610. In various embodiments, the audio / SR experience metadata 610 includes a header for the audio / SR experience data object 600. In various embodiments, the audio / SR experience metadata 610 includes data indicating the creator or provider of the audio / SR experience. In various embodiments, the audio / SR experience metadata 610 includes data indicating the creation time of the audio / SR experience (e.g., date or year).

[0082] Audio / SR Experience Data Object 600 includes an audio file 620. In various embodiments, the audio file is, for example, an MP3 file, an AAC file, or a WAV file. Audio file 620 includes audio file metadata 622. In various embodiments, audio file metadata 622 includes data indicating title, artist, album, release year, lyrics, etc. Audio file 620 includes audio data 624. In various embodiments, audio data 624 includes data indicating audio that can be played by a speaker. In various embodiments, audio data 624 includes data indicating a song or spoken words.

[0083] The audio / SR experience data object 600 includes an SR content package 630. In various embodiments, the SR content package 630 includes SR content package metadata 632 and multiple SR content events 634. The multiple SR content events 634 include a first SR content event 640A, a second SR content event 640B, and a third SR content event 640C. In various embodiments, the SR content package 630 may include any number of SR content events 640A-640C.

[0084] The first SR content event 640A includes metadata (such as an event identifier or the name of the SR content event). The first SR content event 640A includes data indicating a time standard. In various embodiments, the time standard is satisfied when the current position on the timeline of the audio file matches the trigger time of the SR content event. In various embodiments, the time standard is satisfied when the current position on the timeline of the audio file is within the trigger time range of the SR content event.

[0085] The first SR content event 640A includes data indicating environmental criteria. In various embodiments, environmental criteria are met when the environment of the scene camera surveying the environment is a specific environment category. For example, in various embodiments, the environment is categorized as large room, small room, interior, exterior, bright, or dark. Therefore, in various embodiments, environmental criteria include data indicating the environment category. In various embodiments, environmental criteria are met when the environment of the scene camera surveying the environment includes objects of a specific shape. For example, in various embodiments, a specific shape is a generally circular object within a certain size range or an object above a first threshold but narrower than a second threshold. Therefore, in various embodiments, environmental criteria include a definition of shape or shape category, such as shape matching criteria. In various embodiments, environmental criteria are met when the environment of the scene camera surveying the environment includes objects of a specific type. For example, in various embodiments, a specific type could be a window, table, mirror, display screen, etc. Therefore, in various embodiments, environmental criteria include data indicating the object type.

[0086] The first SR content event 640A includes data indicating other criteria. In various implementations, these other criteria are related to one or more of the following: user settings, other previously played SR content events, user triggers / actions, time of day, user biometrics, ambient sounds, random numbers, etc.

[0087] The first SR content event 640A includes SR content displayed when time and environmental criteria (and any other criteria) are met. In various embodiments, the SR content is displayed on top of a representation of a real object in the environment. In various embodiments, the SR content includes a supplementary audio file played simultaneously with the audio file.

[0088] The second SR content event 640B and the third SR content event 640C include fields substantially similar to those of the first SR content event 640A. In various embodiments, the plurality of SR content events 634 include two events having the same time standard but different environmental standards. In various embodiments, the plurality of SR content events 634 include two events having the same environmental standard but different time standards.

[0089] In various implementations, the audio / SR experience data object 600 is stored as a file. In various implementations, the audio / SR experience data object 600 is created by a human designer using, for example, a programming interface of a computing device, such as design, engineering, and / or programming.

[0090] In various implementations, multiple audio / SR experience data objects are stored in non-transitory memory. In various implementations, multiple audio / SR experience data objects are stored by a content provider. Furthermore, the content provider may provide a virtual storefront for selling audio / SR experience data objects to consumers. Therefore, in various implementations, one or more of the multiple audio / SR experience data objects are being transmitted and / or downloaded (e.g., streamed by a consumer or stored on different non-transitory memories).

[0091] In various implementations, the SR content package is stored, sold, transmitted, downloaded, and / or streamed separately from the audio files. Therefore, in various implementations, the SR content package metadata includes an indication of the audio files associated with it.

[0092] Figure 7 This is a flowchart illustrating a first method 700 for presenting an audio / SR experience according to some embodiments. In various embodiments, method 700 is performed by a device having one or more processors, non-transitory memory, speakers, and a display (e.g., Figure 3 The method 700 is executed by the HMD 120B. In some embodiments, method 700 is executed by processing logic components (including hardware, firmware, software, or a combination thereof). In some embodiments, method 700 is executed by a processor that executes instructions (e.g., code) stored in a non-transitory computer-readable medium (e.g., memory).

[0093] Method 700 begins at block 710, wherein the device stores an audio file with an associated timeline. In various embodiments, the audio file is an MP3 file, an AAC file, a WAV file, etc. In various embodiments, the audio file includes audio data representing music and / or spoken words (e.g., audiobooks).

[0094] Method 700 continues in block 720, wherein the device stores multiple SR content events in association with an audio file. Each of the multiple SR content events is associated with a corresponding time standard and a corresponding ambient standard. In various embodiments, the multiple SR content events are stored as a single data object associated with the audio file, such as... Figure 6 Audio / SR experience data object 600. In various implementations, multiple SR content events are stored separately from the audio files in association with them, but together with metadata indicating the audio files associated with the multiple SR content events.

[0095] Method 700 continues in block 730, wherein, while playing an audio file, the device determines a corresponding time criterion that satisfies a specific SR content event among multiple SR content events based on the current position on the audio file's timeline. In various implementations, the corresponding time criterion is satisfied when the current position on the audio file's timeline matches the trigger time of a specific SR content event among multiple SR content events. For example, in Figure 4B In this context, when the current position on the timeline of the first audio file matches a trigger time (e.g., indicated by a first event marker of the audio playback indicator 420), a first time criterion for the first SR content event is satisfied. In various embodiments, when the current position on the timeline of the audio file falls within the trigger time range of a specific SR content event among multiple SR content events, the corresponding time criterion for that specific SR content event among the multiple SR content events is satisfied. For example, in... Figure 4D In this context, when the current position on the timeline of the first audio file is within the trigger window (e.g., indicated by the third event marker of the audio playback indicator 420), the third time criterion of the third SR content event is satisfied. In various embodiments, the SR content events have an SR content timeline that coexists with the time window, and when the corresponding environmental criterion is satisfied, the corresponding portion of a particular SR content event among a plurality of SR content events is displayed (as described below with respect to box 750). For example, in various embodiments, a particular SR content event among a plurality of SR content events includes a video segment, and whenever the field of view of the scene camera includes a television during the time window, the corresponding portion of the video segment is displayed on the television.

[0096] Method 700 continues at block 740, wherein, while playing an audio file, the device determines, based on environmental data (e.g., the environment surveyed by the device's scene camera), a corresponding environmental criterion satisfying a specific SR content event among multiple SR content events. In various embodiments, the scene camera is part of a device worn by the user and includes a display showing the specific SR content event among the multiple SR content events (as described below with respect to block 750). Thus, in various embodiments, the user is physically present in the environment. In various embodiments, the scene camera is part of a remote device (such as a drone or robot avatar) that transmits images from the scene camera to a local device worn by the user and including a display showing the specific SR content event among the multiple SR content events.

[0097] In various implementations, when the environment is a specific environment category, the corresponding environment criteria for a specific SR content event among multiple SR content events are met. In various implementations, the device classifies the environment as internal, external, large room, small room, bright, dark, etc. For example, in... Figure 4GIn this context, because the SR volume environment 400 is located within a room and is therefore classified as an "internal" environment, it meets the sixth environment criterion for the sixth SR content event. For example, in... Figure 5G In the case of SR volume environment 500, since it is not classified as an "internal" environment, it does not meet the sixth environmental criterion. However, since SR volume environment 500 is external, on the beach, and is therefore classified as an "external" environment, it meets the ninth environmental criterion for the ninth SR content event.

[0098] In various implementations, the device determines the appropriate environmental criteria for a specific SR content event among multiple SR content events by performing image analysis on images of the environment (e.g., the environment captured by the device's scene camera). Therefore, in various implementations, environmental data of the environment includes images of the environment. In various implementations, performing image analysis on images of the environment includes performing object detection and / or classification.

[0099] In various implementations, when the image of the environment includes objects of a specific shape, the corresponding environment criteria for a specific SR content event among multiple SR content events are satisfied. For example, in Figure 4C In this context, because the SR volume environment 400 includes objects of a specific shape (e.g., lamp 414) with a long, slender portion where the top is the larger part, it satisfies the second environment criterion for the second SR content event. For example, in... Figure 5C In this context, because the SR volume environment 500 also includes objects of a specific shape with a long, slender portion at the top that is the larger part (e.g., a tree 512), it also meets the second environment standard.

[0100] In various embodiments, the specific shape is a generally circular object within a certain size range. Therefore, in various embodiments, displaying a specific SR content event among multiple SR content events (as described below with respect to box 750) includes displaying a disco ball on top of the generally circular object. In various embodiments, the specific shape is a flat surface of at least a threshold size. Therefore, in various embodiments, displaying a specific SR content event among multiple SR content events (as described below with respect to box 750) includes displaying video content on the flat surface.

[0101] In various implementations, when the image of the environment includes objects of a specific type, the corresponding environment criteria for a specific SR content event among multiple SR content events are satisfied. For example, in Figure 4E In this context, because the SR volume environment 400 includes objects categorized as "tables" (specifically, table 412), it satisfies the fourth environment criterion for the fourth SR content event. For example, in... Figure 5EIn this context, because the SR volume environment 500 also includes objects classified as "tables" (specifically, table 513), it also meets the fourth environment standard.

[0102] In various implementations, the specific type is a "mirror". Therefore, in various implementations, displaying a specific SR content event among multiple SR content events (as described below with respect to box 750) includes displaying SR content on top of a mirror, such as virtual objects that exist only in the "mirror world". In various implementations, the specific type is a "window". Therefore, in various implementations, displaying a specific SR content event among multiple SR content events (as described below with respect to box 750) includes displaying SR content on top of a window, such as virtual objects that appear to be outside the window. In various implementations, the specific type is a "display screen". Therefore, in various implementations, displaying a specific SR content event among the multiple SR content events (as described below with respect to box 750) includes displaying SR content on a display screen, such as video clips of an artist on a television. In various implementations, the specific type is a "photograph". Therefore, in various implementations, displaying a specific SR content event among multiple SR content events (as described below with respect to box 750) includes displaying SR content on top of a photograph, such as replacing a family photo with a photograph of an artist or replacing a poster of a cat with a poster promoting a concert starring an artist.

[0103] As described above, in various embodiments, environmental data includes images of the environment. In various embodiments, environmental data includes the device's GPS location. Therefore, in various embodiments, the corresponding environmental criteria are met when the GPS location indicates that the device is located in a specific location. In various embodiments, environmental data includes network connectivity information. Therefore, in various embodiments, the corresponding environmental criteria are met when the device is connected to a user's home Wi-Fi. In various embodiments, the corresponding environmental criteria are met when the device is connected to a public Wi-Fi network.

[0104] In various implementations, environmental data includes ambient sound recordings. Therefore, in various implementations, a corresponding environmental standard is met when a specific sound is detected or when the environment is quiet.

[0105] Method 700 continues in block 750, wherein while playing an audio file and in response to determining that corresponding time criteria and corresponding environmental criteria are met, the device displays a specific SR content event among a plurality of SR content events in association with the environment. In various embodiments, displaying a specific SR content event among a plurality of SR content events includes displaying content over an object detected in the environment. In various embodiments, displaying a specific SR content event among a plurality of SR content events includes displaying content over a detected object. Therefore, in various embodiments, displaying a specific SR content event among a plurality of SR content events includes replacing a real object in the SR volume environment with a virtual object in the SR volume environment. In various embodiments, displaying a specific SR content event among a plurality of SR content events includes displaying content adjacent to or attached to a detected object. Therefore, in various embodiments, displaying a specific SR content event among a plurality of SR content events includes displaying a virtual object attached to a real object in the SR volume environment.

[0106] Figure 8A This shows the first moment during playback of the second audio file (e.g., a song titled "SongName2" by an artist named "ArtistName2"). Figure 4A The SR volume environment is 400.

[0107] Figure 8B The second time during the playback of the second audio file is shown. Figure 4A The SR volume environment is 400. In Figure 8B In the SR content (based on audio files and the real environment), the content is displayed in the SR volume environment 400.

[0108] exist Figure 8B In the second audio file, photo 411 is replaced with album cover 811 associated with the album of the second audio file. Television 413 is changed from displaying a news program to displaying a concert segment 813 of the artist performing the song in the second audio file. Window 415 is changed to make rain 815 appear on the outside (e.g., based on the title of the second audio file which includes the word "rain"). Lamp 414 is replaced with Greek column 814 (e.g., based on the genre of the second audio file being "Greek"). Fireplace 816 is displayed on the wall of SR volume environment 400 (e.g., based on the lyrics of the second audio file which includes the phrase "sitting by the fire"). Candle 812 is displayed on table 412 (e.g., in response to the lyrics / music analysis of the second audio file indicating that it is classified as a love song).

[0109] Figure 9AThis shows the first moment during playback of a third audio file (e.g., a song titled "SongName3" by an artist named "ArtistName2"). Figure 4A The SR volume environment is 400.

[0110] Figure 9B The second time during the playback of the third audio file is shown. Figure 4A The SR volume environment is 400. In Figure 9B In the SR content (based on audio files and the real environment), the content is displayed in the SR volume environment 400.

[0111] exist Figure 9B In this configuration, the first set of audio / SR lines 911A is displayed on the room boundaries (e.g., ceiling, floor, and walls) of the SR volume environment 400. The audio / SR lines 911A are based on audio data from a third audio file. For example, in various embodiments, the audio / SR lines 911A are based on the volume and / or frequency of the audio data from the third audio file at a second time. Therefore, the audio / SR lines 911A are generated using an audio visualization algorithm. The audio / SR lines 911A are based on the real environment because they are only displayed on the room boundaries of the SR volume environment 400. Therefore, the audio / SR lines 911A are obscured by the television 413. Similarly, the audio / SR lines 911A are distorted due to the location of the room boundaries, for example, bending at the corners of the room in the SR volume environment 400.

[0112] Figure 9C The third time was shown during the playback of the third audio file. Figure 4A The SR volume environment is 400. In Figure 9C In the SR content (based on audio files and the real environment), the content is displayed in the SR volume environment 400.

[0113] exist Figure 9C In this configuration, the second set of audio / SR lines 911B is displayed on the room boundaries (e.g., ceiling, floor, and walls) of the SR volume environment. The audio / SR lines 911B are based on audio data from a third audio file at a third time. For example, in various embodiments, the audio / SR lines 911B are based on the volume and / or frequency of the audio data from the third audio file at a third time. Therefore, the audio / SR lines 911B are generated using an audio visualization algorithm. The audio / SR lines 911B are based on the real environment because they are only displayed on the room boundaries of the SR volume environment 400. Therefore, the audio / SR lines 911B are obscured by the television 413, table 412, and lamp 414. Similarly, the audio / SR lines 911B are distorted due to the location of the room boundaries, for example, bending at the corners of the room in the SR volume environment 400.

[0114] Figure 10A This shows the first moment during playback of the fourth audio file (e.g., a song titled "SongName4" by an artist named "ArtistName2"). Figure 4A The SR volume environment is 400.

[0115] Figure 10B The second time during the playback of the fourth audio file is shown. Figure 4A The SR volume environment is 400. In Figure 10B In the SR content (based on audio files and the real environment), the content is displayed in the SR volume environment 400.

[0116] exist Figure 10B In the image, the first set of audio / SR magic balls 1011A is displayed in the SR volume environment 400. The audio / SR magic balls 1011A are based on audio data from a fourth audio file. For example, in various embodiments, the audio / SR magic balls 1011A are based on the volume and / or frequency of the audio data from the fourth audio file at a second time. For example, in various embodiments, the size of the audio / SR magic balls 1011A at various locations is based on the volume at various frequencies of the audio data from the fourth audio file at a second time. Therefore, in various embodiments, the audio / SR magic balls 1011A are generated using an audio visualization algorithm.

[0117] The audio / SR Magic Balls 1011A are based on the real environment because they are virtual objects located at various locations within the SR volume environment 400 that interact with the SR virtual environment 400 (as described below) and are affected by the SR virtual environment 400 (e.g., illuminated by light emitted from lamp 414 or through window 415).

[0118] Figure 10C The third time during the playback of the fourth audio file is shown. Figure 4A The SR volume environment is 400. In Figure 10C In the middle, the first set of audio / SR magic ball 1011A has moved (e.g., fallen) in the SR volume environment 400.

[0119] Figure 10D The fourth time point during the playback of the fourth audio file is shown. Figure 4A The SR volume environment is 400. In Figure 10D In the middle, the first set of audio / SR magic balls 1011A has already moved within the SR volume environment 400 (e.g., further falling and interacting with table 412). As Figure 10DAs shown, several of the audio / SR magic balls 1011A have been affected by the SR volume environment 400. Specifically, the paths of three audio / SR magic balls in the audio / SR magic ball 1011A have been altered due to the presence of the table 412 in the SR volume environment 400.

[0120] exist Figure 10D In the image, the second set of audio / SR magic balls 1011B is displayed in the SR volume environment 400. The audio / SR magic balls 1011B are based on audio data of a fourth audio file at a fourth time. For example, in various embodiments, the audio / SR magic balls 1011B are based on the volume and / or frequency of the audio data of the fourth audio file at a fourth time. For example, in various embodiments, the size of the audio / SR magic balls 1011B at various locations is based on the volume at various frequencies of the audio data of the fourth audio file at a fourth time. Therefore, in various embodiments, the audio / SR magic balls 1011A are generated using an audio visualization algorithm.

[0121] Figure 10E The fifth time point during the playback of the fourth audio file is shown. Figure 4A The SR volume environment is 400. In Figure 10E In the middle, the first set of audio / SR magic balls 1011A has already moved within the SR volume environment 400 (e.g., further falling and interacting with the table 412 and the floor). As Figure 10E As shown, several of the audio / SR magic balls 1011A have been affected by the SR volume environment 400. Specifically, the paths of many of the audio / SR magic balls in the audio / SR magic balls 1011A have been altered by the base plate of the SR volume environment 400. Furthermore, two of the audio / SR magic balls 1011A are blocked by the lamp 414. Similarly, a second set of audio / SR magic balls 1011B has moved (e.g., fallen) within the SR volume environment 400.

[0122] Figure 11 This is a flowchart illustrating a second method 1100 for presenting an audio / SR experience according to some embodiments. In various embodiments, method 1100 is performed by a device having one or more processors, non-transitory memory, speakers, and a display (e.g., Figure 3 The method is executed by the HMD 120B. In some embodiments, method 1100 is executed by a processing logic unit (including hardware, firmware, software, or a combination thereof). In some embodiments, method 1100 is executed by a processor that executes instructions (e.g., code) stored in a non-transitory computer-readable medium (e.g., memory).

[0123] Method 1100 begins at box 1110, wherein the device acquires a three-dimensional (3D) point cloud of the environment. In various embodiments, the point cloud is based on images of the environment acquired by a scene camera and / or other hardware. In various embodiments, the device includes a scene camera and / or other hardware. In various embodiments, the scene camera and / or other hardware is part of a device worn by a user and includes a display for displaying SR content (as described below with respect to box 1130). Thus, in various embodiments, the user is physically present in the environment. In various embodiments, the scene camera is part of a remote device (such as a drone or robot avatar) that transmits images from the scene camera to a local device worn by the user and including a display for displaying SR content.

[0124] In various implementations, the point cloud comprises a plurality of 3D points within a 3D SR coordinate system. In various implementations, the 3DSR coordinate system is gravity-aligned such that one of the coordinates (e.g., the z-coordinate) extends in the opposite direction to the gravity vector. The gravity vector can be obtained from the device's accelerometer. Each point in the point cloud represents a point on the surface of the environment, e.g., relative to... Figure 4A The SR volumetric environment 400 comprises points on walls (or photographs 411, televisions 413, or windows 415), floors, lamps 414, or tables 412. In various embodiments, the point cloud is acquired using VIO (Visual Inertial Odometry) and / or depth sensors. In various embodiments, the point cloud is based on images of the environment and previous images of the environment taken at different angles to provide stereoscopic imaging. In various embodiments, the points in the point cloud are associated with metadata, which may be confidence levels regarding the color, texture, reflectivity, or transmissivity of points on surfaces in the environment, or the location of points on surfaces in the environment.

[0125] Therefore, in various implementations, a 3D point cloud of the environment is obtained from transparent image data (e.g., images captured by a scene camera) characterized by multiple poses in the environment (e.g., at various orientations and / or locations), wherein each of the multiple poses in the environment is associated with a corresponding field of view of an image sensor (e.g., a scene camera).

[0126] In various implementations, the device generates representation vectors for points in a 3D point cloud, where each representation vector includes one or more labels. One or more labels of a specific representation vector for a particular point in the 3D point cloud are associated with the type and / or characteristics of a physical object. For example, a representation vector may include labels indicating that a particular point in the 3D point cloud corresponds to a surface of a room boundary (such as a ceiling, floor, or wall), a table, or a lamp in a real-world environment. In various implementations, a representation vector includes multiple labels associated with macro and micro information. For example, a representation vector may include a first label indicating that a particular point in the 3D point cloud corresponds to a surface of a room boundary and a second label indicating that a particular point in the 3D point cloud corresponds to a surface of a wall. As another example, a representation vector may include a first label indicating that a particular point in the 3D point cloud corresponds to a surface of a table and a second label indicating that a particular point in the 3D point cloud corresponds to a surface of a table leg.

[0127] In various implementations, the representation vectors are generated by a machine learning process (e.g., by a neural network). In some implementations, generating the representation vectors involves eliminating point clusters from a 3D point cloud, where the representation vectors of the point clusters satisfy an object confidence criterion. In various implementations, a full-error threshold is used if the machine learning-assigned labels included in the representation vectors of the corresponding points are sufficiently similar to each other. In various implementations, multiple point clusters of multiple candidate objects are identified.

[0128] In various implementations, generating the representation vector includes determining a volume region of a group of points, where the volume region corresponds to a 3D representation of an object in space. In various implementations, an object confidence criterion is met when the 3D point cloud comprises a sufficient number of points relative to a specific candidate object to satisfy a threshold confidence level regarding object identification and / or a threshold confidence level regarding the accuracy of the calculated volume. In various implementations, an object confidence criterion is met when the points of the 3D point cloud are sufficiently close to each other relative to a specific candidate object to satisfy a threshold confidence level regarding object identification and / or a threshold confidence level regarding the accuracy of the calculated volume. Therefore, in various implementations, the device detects objects of a specific object type in the environment based on the representation vectors of the points in the 3D point cloud.

[0129] In various implementations, the device detects one or more surfaces in the environment. In various implementations, the surfaces are labeled (e.g., based on representation vectors). In various implementations, the surfaces are not labeled and are defined only by their location and boundaries. The device can employ various methods to detect surfaces (e.g., planar surfaces) from a point cloud. For example, in various implementations, the RANSAC (Random Sample Consensus) method is used to detect surfaces based on the point cloud. In a RANSAC method for detecting planar surfaces, iterations include selecting three random points in the point cloud, determining the plane defined by these three random points, and determining the number of points in the point cloud within a preset distance (e.g., 1 cm) of the plane. The number of points forms a score (or confidence level) for the plane, and after several iterations, the plane with the highest score is selected as the detected planar surface. If points on the detected plane are removed from the point cloud, the method can be repeated to detect another planar surface.

[0130] Method 1100 continues in box 1120, where the device acquires SR content based on the audio file and the 3D point cloud of the environment.

[0131] In various implementations, the device acquires SR content based on the audio data of the audio file. For example, in Figures 9B-9C In this context, the device acquires the SR content, including audio / SR lines 911A-911B, based on audio data (e.g., volume and / or frequency) from a third audio file. For example, in... Figure 10B-10E In this context, the device acquires SR content, including audio / SR magic balls 1011A-1011B, based on audio data (e.g., volume and / or frequency) of a fourth audio file. Therefore, in various implementations, the SR content includes an abstract visualization based on volume dynamics and / or frequency dynamics. For example, in... Figure 8B In this process, the device acquires the SR content, including the candle 812, based at least in part on the analysis of the audio data of the second file (e.g., rhythm or mood [along with optional lyrics] indicating that the song corresponding to the second file is a love song).

[0132] In various implementations, the device retrieves SR content based on metadata of the audio file, such as title, artist, album, genre, lyrics, etc. For example, in Figure 8B In this context, the device retrieves SR content, including the album art 811, based on the metadata of the second audio file, including an album field with a value indicating a specific album. For example, in... Figure 8B In this context, the device uses the metadata of the second audio file, including an artist field with an artist field value indicating the artist, to retrieve the SR content of a concert segment 813 featuring the artist performing the song in the second audio file. For example, in... Figure 8BIn this context, the device retrieves the SR content containing "rain 815" based on the metadata of the second audio file, including a title field with a value containing the word "rain". For example, in... Figure 8B In this process, the device retrieves SR content including Greek column 814 based on the metadata of the second audio file, including the genre field with the value of "Greek".

[0133] In various implementations, the device obtains SR content based on the lyrics of an audio file, through audio analysis that determines the lyrics, or based on the lyrics field value of the lyrics field in the metadata of the audio file. For example, in Figure 8B In this context, the device retrieves the SR content, including fireplace 816, based on the metadata of the second audio file, including a lyrics field with a value containing the phrase "sitting by the fire." Similarly, in... Figure 8B In this example, the device acquires SR content including candle 812 based at least in part on the lyrics of the second file indicating that the song in the second audio file is a love song. Alternatively, the device acquires SR content that shifts the SR volumetric environment into space based on lyrics that evoke a spatial theme (e.g., “The stars look very different today…far above the moon / the planet Earth is blue”).

[0134] In various implementations, the device acquires SR content based on a set of surfaces detected in a 3D point cloud of the environment satisfying rendering criteria. For example, in various implementations, the device acquires SR content based on detecting a flat, vertical surface that is sufficiently large for rendering the SR content. Figure 8B In this embodiment, the device acquires SR content including fireplace 816 based on detecting a flat vertical surface (e.g., a wall) at least large enough to display a threshold size for the fireplace. Alternatively, in various embodiments, the device acquires SR content based on detecting multiple surfaces defining an object of a specific shape. For example, in... Figure 8B In this example, the device acquires SR content including Greek column 814 based on the detection of a tall, slender object (e.g., lamp 414). Similarly, the device acquires SR content including a disco ball based on the detection of a circular object with a diameter of approximately one foot.

[0135] In various implementations, the device acquires SR content based on representation vectors of a 3D point cloud of the environment. For example, in various implementations, the device acquires SR content based on detecting objects of a specific object type in the environment. Figure 8B In this context, the device acquires SR content, including the album cover 811, based on the detected photo 411. For example, in... Figure 8B In this context, the device acquires SR content, including a concert segment 813 of an artist performing a song based on the detection of television 413. For example, in... Figure 8BIn this context, the device acquires SR content, including raindrop 815, based on the detection of window 415. For example, in... Figure 8B In the process, the device acquires SR content, including candle 812, based on the detection of table 412.

[0136] In various implementations, the device retrieves SR content by selecting SR content from a library of tagged SR content elements based on one or more characteristics of the audio file and the environment. In various implementations, the library is located remotely from device storage, for example, it may be accessible via the Internet.

[0137] Method 1100 continues at block 1130, wherein the device simultaneously plays an audio file via a speaker and displays SR content on a display in relation to the environment. In various embodiments, displaying SR content includes displaying the SR content over an object detected in the environment (in a particular embodiment, the object on which the SR content is based). Therefore, in various embodiments, displaying SR content includes replacing a real object in the SR volume environment with a virtual object in the SR volume environment. In various embodiments, displaying SR content includes displaying content that is adjacent to or attached to the detected object (specifically, the object on which the SR content is based). Therefore, in various embodiments, displaying SR content includes displaying a virtual object attached to a real object in the SR volume environment.

[0138] In various implementations, displaying SR content in relation to the environment includes simultaneously playing a supplementary audio file associated with the SR content via a speaker. For example, in Figure 8B In the concert segment 813, audio of the audience cheering may be included; audio of rain 815 may be included of rain falling; or audio of a fireplace 816 may be included of a flame burning.

[0139] Figure 12A This shows the first moment during the story told by storyteller 490. Figure 4A The SR volume environment 400. In various embodiments, a storyteller 490 exists within the SR volume environment. In various embodiments, the storyteller 490 is a real object, such as a person or an audio generating device. In various embodiments, the storyteller 490 is a virtual object displayed by a device.

[0140] Figure 12B Showing the second time period during the story Figure 4A The SR volume environment 400. In various embodiments, the storyteller 490 generates realistic sounds in the environment. In various embodiments, the realistic sounds include one or more spoken words. Figure 12B In the SR volume environment 400, SR content based on spoken words and optionally real-world context is displayed.

[0141] exist Figure 12B In response to the spoken words in the first part of the story 491, which includes the word "dog", the SR volume environment 400 includes other virtual objects, such as the dog 441.

[0142] Figure 12C It shows the third time period during the story. Figure 4A The SR volume environment is 400. In Figure 12C In response to the spoken words in the second part 492 of the story, including the phrase "poor eyesight", the SR volume environment 400 is displayed through virtual objects such as optical filters 442.

[0143] Figure 12D It shows the fourth time period during the story. Figure 4A The SR volume environment is 400. In Figure 12D In response to determining that the spoken words in the third part 493 of the story include the words “food” and the modifier phrase “on the table”, and further in response to detecting the table 412 in the SR volume environment 400, which includes another virtual object such as food 443.

[0144] Figure 12E It shows the fifth time period during the story. Figure 4A The SR volume environment is 400. In Figure 12E In response to determining that the spoken words of the fourth part 494 of the story include the word "rain" and the modifier phrase "external", and further in response to detecting window 415 in the SR volume environment 400, which includes another virtual object, such as rain 444.

[0145] Figure 13 This is a flowchart illustrating a third method 1300 for presenting an audio / SR experience according to some embodiments. In various embodiments, method 1300 is performed by a device having one or more processors, non-transitory memory, and one or more SR displays (e.g., Figure 3 The method is executed by the HMD 120B. In some embodiments, method 1300 is executed by processing logic components (including hardware, firmware, software, or a combination thereof). In some embodiments, method 1300 is executed by a processor that executes instructions (e.g., code) stored in a non-transitory computer-readable medium (e.g., memory).

[0146] Method 1300 begins at block 1310, wherein the device records real-world sounds generated in the environment via a microphone. In various embodiments, the device includes a microphone. In various embodiments, the microphone is part of a device worn by a user and includes a display for showing SR content (as described below with respect to block 1340). Thus, in various embodiments, the user is physically present in the environment. In various embodiments, a scene camera is part of a remote device (such as a drone or robot avatar) that transmits images from the scene camera to a local device worn by the user and including a display for showing SR content.

[0147] Method 1300 continues in block 1320, wherein the device detects one or more spoken words in the real sound. In various embodiments, the device employs one or more speech recognition algorithms to detect spoken words in the real sound.

[0148] Method 1300 continues in block 1330, wherein the device acquires SR content based on one or more spoken words. In various embodiments, the device detects a trigger word among one or more spoken words and acquires SR content based on the trigger word. For example, in Figure 12B In response to detecting the trigger word "dog," the device acquires SR content including "dog" 441. In various implementations, the device further detects modifier words associated with the trigger word and acquires SR content based on these modifier words. For example, in... Figure 12D In response to the detection of the trigger word "rain" and the modifier word "external", the device acquires SR content including rain 444.

[0149] In various implementations, the acquisition of SR content is further based on one or more spatial characteristics of the environment (e.g., characteristics other than the presence of sound in the environment).

[0150] For example, in various implementations, SR content is obtained based on whether the environment is a specific environment category. In various implementations, SR content is obtained based on whether the environment includes objects of a specific shape. For example, in Figure 12E In response to detecting an "outside" phase and further in response to detecting window 415, the SR content includes rain 444 outside window 415. In various implementations, the SR content is acquired based on the environment including objects of a specific type. For example, in... Figure 12D In response to the detection of the phrase "on the table" and further in response to the detection of table 412 in the SR volume environment 400, the SR content includes food 443 on table 412.

[0151] In various implementations, acquiring SR content involves selecting SR content from a library of tagged SR content elements based on at least one of one or more spoken words. In various implementations, the library is stored remotely to the device, for example, via the Internet.

[0152] Method 1300 continues at block 1340, wherein the device displays SR content on a display in relation to the environment. In various embodiments, displaying SR content includes displaying the SR content over an object detected in the environment (in a particular embodiment, the object on which the SR content is based). Therefore, in various embodiments, displaying SR content includes replacing a real object in the SR volume environment with a virtual object in the SR volume environment. In various embodiments, displaying SR content includes displaying content adjacent to or attached to the detected object (specifically, the object on which the SR content is based). Therefore, in various embodiments, displaying SR content includes displaying a virtual object attached to a real object in the SR volume environment.

[0153] In various implementations, displaying SR content in context includes playing supplementary audio files associated with the SR content via speakers. For example, in Figure 12E In the context, dog 441 may include audio of a dog barking, or rain 444 may include audio of rain.

[0154] While various aspects of embodiments within the scope of the appended claims have been described above, it should be apparent that the various features of the above embodiments can be embodied in a wide variety of forms, and any particular structure and / or function described above are merely illustrative. Based on this disclosure, those skilled in the art will understand that the aspects described herein can be implemented independently of any other aspects, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement an apparatus and / or to practice a method. Furthermore, other structures and / or functions, other than or different from those set forth herein, can be used to implement such an apparatus and / or to practice such a method.

[0155] It will also be understood that while terms such as "first," "second," etc., may be used in this document to describe various elements, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, a first node can be called a second node, and similarly, a second node can be called a first node, changing the meaning of the description, provided that all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. First nodes and second nodes are both nodes, but they are not the same node.

[0156] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of these embodiments and the appended claims, the singular forms “a” and “the” are intended to also cover the plural forms unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the term “comprising” as used in this specification specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0157] As used herein, the term "if" can be interpreted as meaning "when the prerequisite is true" or "when the prerequisite is true" or "in response to determination" or "according to determination" or "in response to detection" that the prerequisite is true, depending on the context. Similarly, the phrases "if it is determined [the prerequisite is true]" or "if [the prerequisite is true]" or "when [the prerequisite is true]" are interpreted as meaning "when it is determined that the prerequisite is true" or "in response to determination" or "according to determination" that the prerequisite is true or "when the prerequisite is detected" or "in response to detection" that the prerequisite is true, depending on the context.

Claims

1. A method for synthesizing reality, comprising: In devices that include processors, non-transitory memory, image sensors, speakers, and displays: Audio files with associated timelines are stored in the non-transitory memory; Multiple synthetic reality content events are stored in the non-transitory memory in association with the audio file, wherein each of the multiple synthetic reality content events is associated with a corresponding time standard and a corresponding environmental standard; as well as When the audio file is played via the speaker: The processor determines, based on the current position on the timeline of the audio file, the corresponding time standard for a specific synthetic reality content event among the plurality of synthetic reality content events is satisfied; The image sensor captures images of the physical environment of the device; Based on image analysis of the image, and based on the image indicating the physical environment of the device including physical objects of a specific shape or type, it is determined that the corresponding environmental criteria for the specific synthetic reality content event among the plurality of synthetic reality content events are met; and In response to the satisfaction of both the corresponding time criterion and the corresponding environmental criterion for determining a specific synthetic reality content event among the plurality of synthetic reality content events, the specific synthetic reality content event among the plurality of synthetic reality content events is displayed on the display in association with the physical environment, wherein displaying the specific synthetic reality content event among the plurality of synthetic reality content events includes (i) initiating the display of a virtual object adjacent to the physical object or initiating the display of a virtual object on top of the physical object and (ii) initiating the playback of a supplementary audio file, wherein the supplementary audio file is played simultaneously with the audio file via the speaker, and the supplementary audio file is associated with the specific synthetic reality content event among the plurality of synthetic reality content events. If the corresponding environmental standard or the corresponding time standard is not met, the specific synthetic reality content event among the plurality of synthetic reality content events will not be displayed. The failure to display the specific synthetic reality content event among the plurality of synthetic reality content events includes not displaying the virtual object and not playing the supplementary audio file.

2. The method of claim 1, wherein determining the corresponding time standard for the specific synthetic reality content event among the plurality of synthetic reality content events comprises: Determine the current position on the timeline of the audio file to match the trigger time of the specific synthetic reality content event among the plurality of synthetic reality content events.

3. The method of claim 1, wherein determining the corresponding time standard for the specific synthetic reality content event among the plurality of synthetic reality content events comprises: The current position of the audio file on the timeline is determined to be within the triggering time range of the specific synthetic reality content event among the plurality of synthetic reality content events.

4. The method according to any one of claims 1 to 3, wherein determining the corresponding environmental criteria for a specific synthetic reality content event among the plurality of synthetic reality content events comprises: The environment is determined to be a specific environment category.

5. The method according to any one of claims 1 to 3, wherein displaying a particular synthetic reality content event among the plurality of synthetic reality content events on the display in association with the physical environment is further performed in response to determining that one or more additional criteria are met.

6. The method according to any one of claims 1 to 3, wherein: The first synthetic reality content event among the plurality of synthetic reality content events is associated with a first time standard and a first environment standard; The second synthetic reality content event among the plurality of synthetic reality content events is associated with a second time standard and a second environmental standard; The first time standard is the same as the second time standard; The first environmental standard differs from the second environmental standard; as well as The first synthetic reality content event is different from the second synthetic reality content event; The method includes displaying the first synthetic reality content event on the display in association with the physical environment, without displaying the second synthetic reality content event.

7. The method according to any one of claims 1 to 3, wherein: The first synthetic reality content event among the plurality of synthetic reality content events is associated with a first time standard and a first environment standard; The second synthetic reality content event among the plurality of synthetic reality content events is associated with a second time standard and a second environmental standard; The first time standard is different from the second time standard; The first environmental standard is the same as the second environmental standard; The first synthetic reality content event is different from the second synthetic reality content event; The method includes: displaying the first synthetic reality content event on the display according to the first time standard and displaying the second synthetic reality content event according to the second time standard.

8. The method according to any one of claims 1 to 3, comprising: Multiple audio files are stored in the non-transitory memory, each audio file having an associated timeline; as well as Multiple synthetic reality content packages are encapsulated and stored in the non-transitory memory in association with corresponding audio files among the multiple audio files, each synthetic reality content package including multiple synthetic reality content events associated with a corresponding time standard and a corresponding environmental standard.

9. The method of claim 1, wherein determining that a particular synthetic reality content event among a plurality of synthetic reality content events is satisfied includes determining that the physical object is a square object with a specific reflectivity.

10. The method of claim 1, wherein displaying a specific synthetic reality content event among a plurality of synthetic reality content events includes displaying a virtual object on a display to provide the appearance of the virtual object being adjacent to the physical object.

11. An electronic device, comprising: speaker; monitor; Non-transitory memory; as well as One or more processors, said one or more processors being used for: Audio files with associated timelines are stored in the non-transitory memory; Multiple synthetic reality content events are stored in the non-transitory memory in association with the audio file, wherein each of the multiple synthetic reality content events is associated with a corresponding time standard and a corresponding environmental standard; as well as When the audio file is played via the speaker: Based on the current position on the timeline of the audio file, the corresponding time standard for determining a specific synthetic reality content event among the plurality of synthetic reality content events is satisfied; The image sensor captures images of the physical environment of the device; The image based on the physical environment indicates that the physical environment of the device includes physical objects with a specific shape or type, and determines that the corresponding environmental criteria for the specific synthetic reality content event among the plurality of synthetic reality content events are met; and In response to the satisfaction of both the corresponding time criterion and the corresponding environmental criterion for determining a specific synthetic reality content event among the plurality of synthetic reality content events, the specific synthetic reality content event among the plurality of synthetic reality content events is displayed on the display in association with the physical environment, wherein displaying the specific synthetic reality content event among the plurality of synthetic reality content events includes (i) initiating the display of a virtual object adjacent to the physical object or initiating the display of a virtual object on top of the physical object and (ii) initiating the playback of a supplementary audio file, wherein the supplementary audio file is played simultaneously with the audio file via the speaker, and the supplementary audio file is associated with the specific synthetic reality content event among the plurality of synthetic reality content events. If the corresponding environmental standard or the corresponding time standard is not met, the specific synthetic reality content event among the plurality of synthetic reality content events will not be displayed. The failure to display the specific synthetic reality content event among the plurality of synthetic reality content events includes not displaying the virtual object and not playing the supplementary audio file.

12. The electronic device according to claim 11, wherein, The one or more processors are further configured to determine, by determining that the physical environment belongs to a specific environment category, that the corresponding environment criterion for the specific synthetic reality content event among the plurality of synthetic reality content events is met.

13. The electronic device according to claim 11, wherein: The first content event in the plurality of synthetic reality content events is associated with a first time standard and a first environment standard; The second content event in the plurality of synthetic reality content events is associated with a second time standard and a second environmental standard; The first time standard is the same as the second time standard; The first environmental standard differs from the second environmental standard; The first content event is different from the second content event; as well as The one or more processors are configured to display the first content event on the display in association with the physical environment, but not the second content event.

14. The electronic device according to claim 11, wherein: The first content event in the plurality of synthetic reality content events is associated with a first time standard and a first environment standard; The second content event in the plurality of synthetic reality content events is associated with a second time standard and a second environmental standard; The first time standard is different from the second time standard; The first environmental standard is the same as the second environmental standard; The first content event is different from the second content event; as well as The one or more processors are configured to display the first content event on the display according to the first time standard, and to display the second content event according to the second time standard.

15. The electronic device according to claim 11, wherein, The specific synthetic reality content event among the plurality of synthetic reality content events includes: in response to determining that the physical object has a specific shape with a long, slender portion at the top being a larger portion, displaying a virtual object on the display to provide the appearance of the virtual object exiting from or behind the larger portion of the physical object and retracting and hiding within or behind the larger portion of the physical object.

16. The electronic device according to claim 11, wherein, Determining that the corresponding time criterion of a specific synthetic reality content event among the plurality of synthetic reality content events is met includes: determining that the current position on the timeline of the audio file matches the trigger time of the specific synthetic reality content event among the plurality of synthetic reality content events.

17. A non-transitory computer-readable medium having instructions encoded thereon, the instructions, when executed by one or more processors of a device including a speaker, an image sensor, and a display, causing the device to: Store audio files with associated timelines; Multiple synthetic reality content events are stored in association with the audio file, wherein each of the multiple synthetic reality content events is associated with a corresponding time standard and a corresponding environmental standard; as well as When the audio file is played via the speaker: Based on the current position on the timeline of the audio file, it is determined that the corresponding time standard of a specific synthetic reality content event among the plurality of synthetic reality content events is satisfied; The image sensor captures images of the physical environment of the device; The image based on the physical environment indicates that the physical environment of the device includes physical objects with a specific shape or type, and determines that the corresponding environmental criteria for the specific synthetic reality content event among the plurality of synthetic reality content events are met; and In response to the satisfaction of both the corresponding time criterion and the corresponding environmental criterion for determining a specific synthetic reality content event among the plurality of synthetic reality content events, the specific synthetic reality content event is displayed on the display in association with the physical environment, wherein displaying the specific synthetic reality content event among the plurality of synthetic reality content events includes (i) initiating the display of a virtual object adjacent to the physical object or initiating the display of a virtual object on top of the physical object and (ii) initiating the playback of a supplementary audio file, wherein the supplementary audio file is played simultaneously with the audio file via the speaker, and the supplementary audio file is associated with the specific synthetic reality content event among the plurality of synthetic reality content events. If the corresponding environmental standard or the corresponding time standard is not met, the specific synthetic reality content event among the plurality of synthetic reality content events will not be displayed. The failure to display the specific synthetic reality content event among the plurality of synthetic reality content events includes not displaying the virtual object and not playing the supplementary audio file.

18. The non-transitory computer-readable medium according to claim 17, wherein, Determining that the corresponding environmental criteria for a specific synthetic reality content event among the plurality of synthetic reality content events are met includes: determining that the physical object is a square object with a specific reflectivity.

19. The non-transitory computer-readable medium according to claim 17, wherein, The specific synthetic reality content event among the plurality of synthetic reality content events includes: in response to determining that the physical object has a particular shape with a long, slender portion at the top being a larger portion, displaying a virtual object on the display to provide the appearance of the virtual object exiting from or behind the larger portion of the physical object and shrinking back into or hiding behind the larger portion of the physical object.

20. The non-transitory computer-readable medium according to claim 17, wherein, Determining that the corresponding time criterion of a specific synthetic reality content event among the plurality of synthetic reality content events is satisfied includes: determining that the current position on the timeline of the audio file is within the trigger time range of the specific synthetic reality content event among the plurality of synthetic reality content events.

21. The non-transitory computer-readable medium according to claim 17, wherein, Determining that the corresponding environmental criteria for a specific synthetic reality content event among the plurality of synthetic reality content events are met includes: determining that the physical environment belongs to a specific environment category based on the image.

Citation Information

Patent Citations

  • Audio making and decoding method and device

    CN106448687A

  • Handmade product and virtual reality experience system and method based on mobile terminal

    CN108062796A

  • Head pose mixing of audio files

    US20170078825A1

  • Categorized and tagged video annotation

    US8984405B1