Adaptive playback of extended reality media content based on real-world conditions
The media playback system dynamically adapts virtual scenes to match real-world acoustics, addressing XR audio challenges by enhancing immersion and reducing resource usage.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SONOS INC
- Filing Date
- 2026-01-26
- Publication Date
- 2026-07-30
AI Technical Summary
Existing extended reality (XR) audio systems face challenges in accurately simulating sound propagation in virtual environments, requiring high computational resources and room correction, while maintaining user awareness of real-world sounds and social interactions.
A media playback system that analyzes the user's real-world environment to adapt virtual scenes to match real-world acoustics, adjusting virtual objects and sound reflections for a personalized and immersive experience.
Enhances immersion and realism by tailoring XR experiences to individual environments, improving accuracy and reducing computational demands.
Smart Images

Figure 00000096_0000 
Figure 00000097_0000 
Figure 00000098_0000
Abstract
Description
PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO ADAPTIVE PLAYBACK OF EXTENDED REALITY MEDIA CONTENT BASED ON REAL-WORLD CONDITIONSCROSS-REFERENCE TO RELATED APPLICATION(S)
[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 750,012, filed January 27, 2025, which is hereby incorporated by reference in its entirety.FIELD OF THE DISCLOSURE
[0002] The present disclosure is related to consumer goods and, more particularly, to methods, systems, products, features, services, and other elements directed to media playback or some aspect thereof.BACKGROUND
[0003] Options for accessing and listening to digital audio in an out-loud setting were limited until in 2002. when SONOS, Inc. began development of anew type of playback system. Sonos then filed one of its first patent applications in 2003, entitled “Method for Synchronizing Audio Playback between Multiple Networked Devices,” and began offering its first media playback systems for sale in 2005. The Sonos Wireless Home Sound System enables people to experience music from many sources via one or more networked playback devices. Through a software control application installed on a controller (e g., smartphone, tablet, computer, voice input device), one can play what she wants in any room having a networked playback device. Media content (e.g., songs, podcasts, video sound) can be streamed to playback devices such that each room with a playback device can play back corresponding different media content. In addition, rooms can be grouped together for synchronous playback of the same media content, and / or the same media content can be heard in all rooms synchronously.BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Features, aspects, and advantages of the presently disclosed technology may be better understood with regard to the following description, appended claims, and accompanying drawings, as listed below. A person skilled in the relevant art will understand that the features shown in the drawings are for purposes of illustrations, and variations, including different and / or additional features and arrangements thereof, are possible.
[0005] Figure 1A is a partial cutaway view of an environment having a media playback system configured in accordance with aspects of the disclosed technology.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0006] Figure IB is a schematic diagram of the media playback system of Figure 1 A and one or more networks.
[0007] Figure 1C is a block diagram of a playback device.
[0008] Figure ID is a block diagram of a playback device.
[0009] Figure IE is a block diagram of a bonded playback device.
[0010] Figure IF is a block diagram of a network microphone device.
[0011] Figure 1G is a block diagram of a playback device.
[0012] Figure 1H is a partially schematic diagram of a control device.
[0013] Figures II through IL show schematic diagrams of corresponding media playback system zones.
[0014] Figure IM shows a schematic diagram of media playback system areas.
[0015] Figure 2 is a schematic diagram of an extended reality (XR) system including in accordance with aspects of the disclosed technology7.
[0016] Figure 3 is a schematic diagram illustrating a blockchain-capable playback device with generative media components and generative context and control components in accordance with aspects of the disclosed technology.
[0017] Figure 4 illustrates an example modification of XR media content based on real-world conditions in accordance with aspects of the disclosed technology.
[0018] Figure 5 illustrates an example modification of XR media content based on two different real-world environments in accordance with aspects of the disclosed technology.
[0019] Figure 6 is a diagram of an example environment showing a user wearing a w earable playback device in proximity to an out-loud playback device, in accordance with aspects of the disclosed technology.
[0020] Figure 7 is a swim lane diagram showing the sequence of operations performed by the wearable XR playback device and the out-loud playback device to achieve synchronization through data stream analysis of embedded reference signatures, in accordance with aspects of the disclosed technology.
[0021] Figure 8 is a swim lane diagram showing the interactions between audio transducers, internal microphones, and external microphones within the wearable playback device for dualmicrophone detection synchronization, in accordance with aspects of the disclosed technology.
[0022] Figures 9-11 are flow diagrams illustrating example methods in accordance with aspects of the disclosed technology.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0023] The drawings are for the purpose of illustrating examples of the present technology', but those of ordinary skill in the art will understand that the technology disclosed herein is not limited to the arrangements and / or instrumentality shown in the drawings.DETAILED DESCRIPTIONI. Overview
[0024] The present technology relates to media content, including spatialized media content and / or extended reality' (XR) media content. As used herein, ‘‘extended reality” encompasses virtual reality (VR), mixed reality (MR) and augmented reality (AR) experiences, and / or other perceptually immersive technologies. “VR” may encompass experiences that fully visually immerse users in computer-generated environments, while AR may encompass experiences that overlay digital content onto views of the real world. XR experiences may encompass experiences that combine visual content, presented through wearable devices such as headmounted displays (HMDs) or smart glasses, with audio content for a multi-sensory' experience. In some examples, XR experiences may comprise holographic projection devices presented without any wearable devices. While many traditional XR systems utilize headphones or earbuds for personal audio playback, implementations described herein may additionally or alternatively provide audio content through out-loud playback devices such as speakers, soundbars, or other audio transducers positioned in the real-world environment. This approach enables users to experience spatialized audio content while maintaining awareness of real-world sounds and facilitating social interactions with others in the physical space.
[0025] The current state of the art in extended reality (XR) audio involves techniques such as acoustic simulation software to model sound propagation in virtual environments and generate impulse responses that capture the environment's unique acoustic signature. This data is then translated to the real world through techniques such as convolution, binaural rendering, and multi-channel audio, often with dynamic adaptation based on head tracking and changes in the virtual environment. However, this approach presents several challenges, including the accuracy of the simulation, computational resource demands, speaker placement limitations, and the need for room correction to compensate for real-world listening environments, among other challenges.
[0026] The techniques described herein address these and other challenges by enabling a dynamic interplay betw een the virtual and real w orlds. By analyzing the acoustic properties of the user's real-world environment (such as the user’s real-world room), for instance, the system can adapt a virtual scene to match the real-world acoustics of the user’s real-worldPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO environment. This creates a more immersive and realistic experience by adjusting the size, shape, and materials of virtual objects and / or spaces to create realistic sound reflections and reverberations, as well as personalizing the experience by adjusting the virtual scene to match the user's real-world and / or preferred acoustic environments. This approach allows for a more personalized and realistic XR experience that is tailored to the user's specific environment and preferences.
[0027] While some examples described herein may refer to functions performed by given actors such as ‘‘users,” “listeners,” and / or other entities, it should be understood that this is for purposes of explanation only. The claims should not be interpreted to require action by any such example actor unless explicitly required by the language of the claims themselves.
[0028] In the Figures, identical reference numbers identify generally similar, and / or identical, elements. To facilitate the discussion of any particular element, the most significant digit or digits of a reference number refers to the Figure in which that element is first introduced. For example, element 110a is first introduced and discussed with reference to Figure 1 A. Many of the details, dimensions, angles and other features shown in the Figures are merely illustrative of particular examples of the disclosed technology. Accordingly, other examples can have other details, dimensions, angles and features without departing from the spirit or scope of the disclosure. In addition, those of ordinary skill in the art will appreciate that further examples of the various disclosed technologies can be practiced without several of the details described below.II. Suitable Operating Environment
[0029] Figure 1A is a partial cutaway view of a media playback system 100 distributed in an environment 101 (e.g., a house). The media playback system 100 comprises one or more playback devices 110 (identified individually as playback devices HOa-n), one or more network microphone devices (“NMDs”) 120 (identified individually as NMDs 120a-c), and one or more control devices 130 (identified individually as control devices 130a and 130b).
[0030] As used herein the term “playback device” can generally refer to a network device configured to receive, process, and / or output data of a media playback system. For example, a playback device can be a network device that receives and processes audio content. In some examples, a playback device includes one or more transducers or speakers powered by one or more amplifiers. In other examples, however, a playback device includes one of (or neither of) the speaker and the amplifier. For instance, a playback device can comprise one or morePATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO amplifiers configured to drive one or more speakers external to the playback device via a corresponding wire or cable.
[0031] Moreover, as used herein the term NMD (z. e. , a “network microphone device’’) can generally refer to a network device that is configured for audio detection. In some examples, an NMD is a stand-alone device configured primarily for audio detection. In other examples, an NMD is incorporated into a playback device (or vice versa).
[0032] The term “control device” can generally refer to a network device configured to perform functions relevant to facilitating user access, control, and / or configuration of the media playback system 100.
[0033] Each of the playback devices 110 is configured to receive audio signals or data from one or more media sources (e.g., one or more remote servers or one or more local devices) and play back the received audio signals or data as sound. The one or more NMDs 120 are configured to receive spoken word commands, and the one or more control devices 130 are configured to receive user input. In response to the received spoken word commands and / or user input, the media playback system 100 can play back audio via one or more of the playback devices 110. In certain examples, the playback devices 110 are configured to commence playback of media content in response to a trigger. For instance, one or more of the playback devices 110 can be configured to play back a morning playlist upon detection of an associated trigger condition (e.g., presence of a user in a kitchen, detection of a coffee machine operation). In some examples, for instance, the media playback system 100 is configured to play back audio from a first playback device (e.g., the playback device 110a) in synchrony with a second playback device (e.g., the playback device 110b). Interactions between the playback devices 110, NMDs 120, and / or control devices 130 of the media playback system 100 configured in accordance with the various examples of the disclosure are described in greater detail below with respect to Figures 1B-1H.
[0034] In the illustrated example of Figure 1 A, the environment 101 comprises a household having several rooms, spaces, and / or playback zones, including (clockwise from upper left) a master bathroom 101a. a master bedroom 101b, a second bedroom 101c, a family room or den lOld, an office lOle, a living room 10 If, a dining room 101g, a kitchen lOlh, and an outdoor patio lOli. While certain examples and examples are described below in the context of a home environment, the technologies described herein may be implemented in other types of environments. In some examples, for instance, the media playback system 100 can be implemented in one or more commercial settings (e.g., a restaurant, mall, airport, hotel, a retailPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO or other store), one or more vehicles (e g., a utility vehicle, bus, car, a ship, a boat, an airplane), multiple environments (e.g., a combination of home and vehicle environments), and / or another suitable environment where multi-zone audio may be desirable.
[0035] The media playback system 100 can comprise one or more playback zones, some of which may correspond to the rooms in the environment 101. The media playback system 100 can be established with one or more playback zones, after which additional zones may be added, or removed to form, for example, the configuration shown in Figure 1 A. Each zone may be given a name according to a different room or space such as the office 101 e, master bathroom 101a, master bedroom 101b, the second bedroom 101c, kitchen lOlh, dining room 101g, living room lOlf, and / or the outdoor patio lOli. In some aspects, a single playback zone may include multiple rooms or spaces. In certain aspects, a single room or space may include multiple playback zones.
[0036] In the illustrated example of Figure 1A, the master bathroom 101a, the second bedroom 101c, the office lOle, the living room lOlf, the dining room 101g, the kitchen lOlh, and the outdoor patio lOli each include one playback device 110, and the master bedroom 101b and the den 101 d include a plurality of playback devices 110. In the master bedroom 101b, the playback devices 1101 and 110m may be configured, for example, to play back audio content in synchrony as individual ones of playback devices 110, as a bonded playback zone, as a consolidated playback device, and / or any combination thereof. Similarly, in the den lOld, the playback devices HOh-j can be configured, for instance, to play back audio content in synchrony as individual ones of playback devices 110, as one or more bonded playback devices, and / or as one or more consolidated playback devices. Additional details regarding bonded and consolidated playback devices are described below with respect to Figures IB and IE.
[0037] In some aspects, one or more of the playback zones in the environment 101 may each be playing different audio content. For instance, a user may be grilling on the patio lOli and listening to hip hop music being played by the playback device 110c while another user is preparing food in the kitchen lOlh and listening to classical music played by the playback device 110b. In another example, a playback zone may play the same audio content in synchrony with another playback zone. For instance, the user may be in the office lOle listening to the playback device 1 lOf playing back the same hip hop music being played back by playback device 110c on the patio lOli. In some aspects, the playback devices 110c and 11 Of play back the hip hop music in synchrony such that the user perceives that the audioPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO content is being played seamlessly (or at least substantially seamlessly) while moving between different playback zones. Additional details regarding audio playback synchronization among playback devices and / or zones can be found, for example, in U.S. PatentNo. 8,234,395 entitled, “System and method for synchronizing operations among a plurality of independently clocked digital data processing devices,’' which is incorporated herein by reference in its entirety. a. Suitable Media Playback System
[0038] Figure IB is a schematic diagram of the media playback system 100 and a cloud network 102. For ease of illustration, certain devices of the media playback system 100 and the cloud network 102 are omitted from Figure IB. One or more communication links 103 (referred to hereinafter as “the links 103”) communicatively couple the media playback system 100 and the cloud network 102.
[0039] The links 103 can comprise, for example, one or more wired networks, one or more wireless networks, one or more wide area networks (WAN), one or more local area networks (LAN), one or more personal area networks (PAN), one or more telecommunication networks (e.g., one or more Global System for Mobiles (GSM) networks, Code Division Multiple Access (CDMA) networks, Long-Term Evolution (LTE) networks, 5G communication network networks, 6G communication networks and / or other suitable data transmission protocol networks), etc. The cloud network 102 is configured to deliver media content (e.g., audio content, video content, photographs, social media content) to the media playback system 100 in response to a request transmitted from the media playback system 100 via the links 103. In some examples, the cloud network 102 is further configured to receive data (e.g. voice input data) from the media playback system 100 and correspondingly transmit commands and / or media content to the media playback system 100.
[0040] The cloud network 102 comprises computing devices 106 (identified separately as a first computing device 106a, a second computing device 106b, and a third computing device 106c). The computing devices 106 can comprise individual computers or servers, such as, for example, a media streaming service server storing audio and / or other media content, a voice service server, a social media server, a media playback system control server, etc. In some examples, one or more of the computing devices 106 comprise modules of a single computer or server. In certain examples, one or more of the computing devices 106 comprise one or more modules, computers, and / or servers. Moreover, while the cloud network 102 is described above in the context of a single cloud network, in some examples the cloud network 102 comprises a plurality of cloud networks comprising communicatively coupled computing devices.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO Furthermore, while the cloud network 102 is shown in Figure IB as having three of the computing devices 106, in some examples, the cloud network 102 comprises fewer (or more than) three computing devices 106.
[0041] The media playback system 100 is configured to receive media content from the networks 102 via the links 103. The received media content can comprise, for example, a Uniform Resource Identifier (URI) and / or a Uniform Resource Locator (URL). For instance, in some examples, the media playback system 100 can stream, download, or otherwise obtain data from a URI or a URL corresponding to the received media content. A network 104 communicatively couples the links 103 and at least a portion of the devices (e.g., one or more of the playback devices 110, NMDs 120, and / or control devices 130) of the media playback system 100. The network 104 can include, for example, a wireless network (e.g., a WiFi network, a Bluetooth, a Z-Wave network, a ZigBee, and / or other suitable wireless communication protocol network) and / or a wired network (e.g., a network comprising Ethernet, Universal Serial Bus (USB), and / or another suitable wired communication). As those of ordinary skill in the art will appreciate, as used herein, “WiFi” can refer to several different communication protocols including, for example. Institute of Electrical and Electronics Modules (IEEE) 802.11 a, 802.1 lb, 802.11 g, 802.1 In, 802.11 ac, 802.11 ac, 802.11 ad, 802.11 af, 802.11ah, 802.11ai, 802.11aj, 802.11aq, 802.11ax, 802.11ay, 802.15, etc. transmitted at 2.4 Gigahertz (GHz), 5 GHz, and / or another suitable frequency.
[0042] In some examples, the network 104 comprises a dedicated communication network that the media playback system 100 uses to transmit messages between individual devices and / or to transmit media content to and from media content sources (e.g., one or more of the computing devices 106). In certain examples, the network 104 is configured to be accessible only to devices in the media playback system 100, thereby reducing interference and competition with other household devices. In other examples, however, the network 104 comprises an existing household communication network (e.g., a household WiFi network). In some examples, the links 103 and the network 104 comprise one or more of the same networks. In some aspects, for example, the links 103 and the network 104 comprise a telecommunication network (e.g., an LTE network, a 5G network). Moreover, in some examples, the media playback system 100 is implemented without the network 104, and devices comprising the media playback system 100 can communicate with each other, for example, via one or more direct connections, PANs, telecommunication networks, and / or other suitable communication links.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0043] In some examples, audio content sources may be regularly added or removed from the media playback system 100. In some examples, for instance, the media playback system 100 performs an indexing of media items when one or more media content sources are updated, added to, and / or removed from the media playback system 100. The media playback system 100 can scan identifiable media items in some or all folders and / or directories accessible to the playback devices 110, and generate or update a media content database comprising metadata (e.g., title, artist, album, track length) and other associated information (e.g., URIs, URLs) for each identifiable media item found. In some examples, for instance, the media content database is stored on one or more of the playback devices 110, NMDs 120, and / or control devices 130.
[0044] In the illustrated example of Figure IB, the playback devices 1101 and 110m comprise a group 107a. The playback devices 1101 and 110m can be positioned in different rooms in a household and be grouped together in the group 107a on a temporary or permanent basis based on user input received at the control device 130a and / or another control device 130 in the media playback system 100. When arranged in the group 107a, the playback devices 1101 and 110m can be configured to play back the same or similar audio content in synchrony from one or more audio content sources. In certain examples, for instance, the group 107a comprises a bonded zone in which the playback devices 1101 and 110m comprise left audio and right audio channels, respectively, of multi-channel audio content, thereby producing or enhancing a stereo effect of the audio content. In some examples, the group 107a includes additional playback devices 110. In other examples, however, the media playback system 100 omits the group 1 7a and / or other grouped arrangements of the playback devices 110.
[0045] The media playback system 100 includes the NMDs 120a and 120d, each comprising one or more microphones configured to receive voice utterances from a user. In the illustrated example of Figure IB. the NMD 120a is a standalone device and the NMD 120d is integrated into the playback device HOn. The NMD 120a, for example, is configured to receive voice input 121 from a user 123. In some examples, the NMD 120a transmits data associated with the received voice input 121 to a voice assistant service (VAS) configured to (i) process the received voice input data and (ii) transmit a corresponding command to the media playback system 100. In some aspects, for example, the computing device 106c comprises one or more modules and / or servers of a VAS (e.g., a VAS operated by one or more of SONOS®, AMAZON®, GOOGLE® APPLE®, MICROSOFT®). The computing device 106c can receive the voice input data from the NMD 120a via the network 104 and the links 103. In response to receiving the voice input data, the computing device 106c processes the voice inputPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO data (z.e., “Play Hey Jude by The Beatles”), and determines that the processed voice input includes a command to play a song (e.g.. “Hey Jude”). The computing device 106c accordingly transmits commands to the media playback system 100 to play back “Hey Jude” by the Beatles from a suitable media service (e.g., via one or more of the computing devices 106) on one or more of the playback devices 110.b. Suitable Playback Devices
[0046] Figure 1C is a block diagram of the playback device 110a comprising an input / output 111. The input / output 111 can include an analog I / O Illa (e.g., one or more wires, cables, and / or other suitable communication links configured to cany7analog signals) and / or a digital I / O 11 lb (e.g., one or more wires, cables, or other suitable communication links configured to carry digital signals). In some examples, the analog I / O Illa is an audio line-in input connection comprising, for example, an auto-detecting 3.5mm audio line-in connection. In some examples, the digital I / O 111b comprises a Sony / Philips Digital Interface Format (S / PDIF) communication interface and / or cable and / or a Toshiba Link (TOSLINK) cable. In some examples, the digital I / O 111b comprises a High-Definition Multimedia Interface (HDMI) interface and / or cable. In some examples, the digital I / O 111b includes one or more wireless communication links comprising, for example, a radio frequency (RF), infrared, WiFi, Bluetooth, or another suitable communication protocol. In certain examples, the analog I / O Illa and the digital 111b comprise interfaces (e.g.. ports, plugs, jacks) configured to receive connectors of cables transmitting analog and digital signals, respectively, without necessarily including cables.
[0047] The playback device 110a, for example, can receive media content (e.g., audio content comprising music and / or other sounds) from a local audio source 105 via the input / output 111 (e.g.. a cable, a wire, a PAN. a Bluetooth connection, an ad hoc wired or wireless communication network, and / or another suitable communication link). The local audio source 105 can comprise, for example, a mobile device (e.g., a smartphone, a tablet, a laptop computer) or another suitable audio component (e.g., a television, a desktop computer, an amplifier, a phonograph, a Blu-ray player, a memory storing digital media files). In some aspects, the local audio source 105 includes local music libraries on a smartphone, a computer, a networked-attached storage (NAS), and / or another suitable device configured to store media files. In certain examples, one or more of the playback devices 110, NMDs 120, and / or control devices 130 comprise the local audio source 105. In other examples, however, the media playback system omits the local audio source 105 altogether. In some examples, the playbackPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO device 110a does not include an input / output 111 and receives all audio content via the network 104.
[0048] The playback device 110a further comprises electronics 112, a user interface 113 (e.g., one or more buttons, knobs, dials, touch-sensitive surfaces, displays, touchscreens), and one or more transducers 114 (referred to hereinafter as “the transducers 114”). The electronics 112 is configured to receive audio from an audio source (e.g., the local audio source 105) via the input / output 111, one or more of the computing devices 106a-c via the network 104 (Figure IB)), amplify the received audio, and output the amplified audio for playback via one or more of the transducers 114. In some examples, the playback device 110a optionally includes one or more microphones 115 (e.g., a single microphone, a plurality of microphones, a microphone array) (hereinafter referred to as “the microphones 115”). In certain examples, for instance, the playback device 110a having one or more of the optional microphones 115 can operate as an NMD configured to receive voice input from a user and correspondingly perform one or more operations based on the received voice input.
[0049] In the illustrated example of Figure 1C, the electronics 112 comprise one or more processors 112a (referred to hereinafter as “the processors 112a”), memory 112b, software components 112c, a network interface 112d, one or more audio processing components 112g (referred to hereinafter as “the audio components 112g”), one or more audio amplifiers 112h (referred to hereinafter as “the amplifiers 112h”), and power 112i (e.g., one or more power supplies, power cables, power receptacles, batteries, induction coils, Power-over Ethernet (POE) interfaces, and / or other suitable sources of electric power). In some examples, the electronics 112 optionally include one or more other components 112j (e.g., one or more sensors, video displays, touchscreens, battery charging bases).
[0050] The processors 112a can comprise clock-driven computing component(s) configured to process data, and the memory 112b can comprise a computer-readable medium (e.g., a tangible, non-transitory computer-readable medium, data storage loaded with one or more of the software components 112c) configured to store instructions for performing various operations and / or functions. The processors 112a are configured to execute the instructions stored on the memory 112b to perform one or more of the operations. The operations can include, for example, causing the playback device 110a to retrieve audio data from an audio source (e.g., one or more of the computing devices 106a-c (Figure IB)), and / or another one of the playback devices 110. In some examples, the operations further include causing the playback device 110a to send audio data to another one of the playback devices 110a and / orPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO another device (e.g., one of the NMDs 120). Certain examples include operations causing the playback device 110a to pair with another of the one or more playback devices 110 to enable a multi-channel audio environment (e.g., a stereo pair, a bonded zone).
[0051] The processors 112a can be further configured to perform operations causing the playback device 110a to synchronize playback of audio content with another of the one or more playback devices 110. As those of ordinary skill in the art will appreciate, during synchronous playback of audio content on a plurality of playback devices, a listener will preferably be unable to perceive time-delay differences between playback of the audio content by the playback device 110a and the other one or more other playback devices 110. Additional details regarding audio playback synchronization among playback devices can be found, for example, in U.S. Patent No. 8,234,395, which was incorporated by reference above.
[0052] In some examples, the memory 112b is further configured to store data associated with the playback device 110a, such as one or more zones and / or zone groups of which the playback device 110a is a member, audio sources accessible to the playback device 110a, and / or a playback queue that the playback device 110a (and / or another of the one or more playback devices) can be associated with. The stored data can comprise one or more state variables that are periodically updated and used to describe a state of the playback device 110a. The memory 112b can also include data associated with a state of one or more of the other devices (e.g., the playback devices 110, NMDs 120, control devices 130) of the media playback system 100. In some aspects, for example, the state data is shared during predetermined intervals of time (e.g., every 5 seconds, every 10 seconds, every 60 seconds) among at least a portion of the devices of the media playback system 100, so that one or more of the devices have the most recent data associated with the media playback system 100.
[0053] The network interface 112d is configured to facilitate a transmission of data between the playback device 110a and one or more other devices on a data netw ork such as, for example, the links 103 and / or the netw ork 104 (Figure IB). The netw ork interface 112d is configured to transmit and receive data corresponding to media content (e.g., audio content, video content, text, photographs) and other signals (e.g., non-transitory signals) comprising digital packet data including an Internet Protocol (IP)-based source address and / or an IP-based destination address. The network interface 112d can parse the digital packet data such that the electronics 112 properly receives and processes the data destined for the playback device 110a.
[0054] In the illustrated example of Figure 1C, the netw ork interface 112d comprises one or more wireless interfaces 112e (referred to hereinafter as "The wireless interface 112e?’). ThePATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO wireless interface 112e (e.g., a suitable interface comprising one or more antennae) can be configured to wirelessly communicate with one or more other devices (e.g.. one or more of the other playback devices 110, NMDs 120, and / or control devices 130) that are communicatively coupled to the network 104 (Figure IB) in accordance with a suitable wireless communication protocol (e.g., WiFi, Bluetooth, LTE). In some examples, the network interface 112d optionally includes a wired interface 112f (e.g.. an interface or receptacle configured to receive a network cable such as an Ethernet, a USB-A, USB-C, and / or Thunderbolt cable) configured to communicate over a wired connection with other devices in accordance with a suitable wired communication protocol. In certain examples, the network interface 112d includes the wired interface 112f and excludes the wireless interface 112e. In some examples, the electronics 112 excludes the network interface 112d altogether and transmits and receives media content and / or other data via another communication path (e.g., the input / output 111).
[0055] The audio components 112g are configured to process and / or filter data comprising media content received by the electronics 112 (e g., via the input / output 111 and / or the network interface 112d) to produce output audio signals. In some examples, the audio processing components 112g comprise, for example, one or more digital-to-analog converters (DAC), audio preprocessing components, audio enhancement components, a digital signal processors (DSPs), and / or other suitable audio processing components, modules, circuits, etc. In certain examples, one or more of the audio processing components 112g can comprise one or more subcomponents of the processors 112a. In some examples, the electronics 112 omits the audio processing components 112g. In some aspects, for example, the processors 112a execute instructions stored on the memory 112b to perform audio processing operations to produce the output audio signals.
[0056] The amplifiers 112h are configured to receive and amplify the audio output signals produced by the audio processing components 112g and / or the processors 112a. The amplifiers 112h can comprise electronic devices and / or components configured to amplify audio signals to levels sufficient for driving one or more of the transducers 114. In some examples, for instance, the amplifiers 112h include one or more switching or class-D power amplifiers. In other examples, however, the amplifiers include one or more other types of power amplifiers (e.g., linear gain power amplifiers, class-A amplifiers, class-B amplifiers, class-AB amplifiers, class-C amplifiers, class-D amplifiers, class-E amplifiers, class-F amplifiers, class-G and / or class H amplifiers, and / or another suitable type of power amplifier). In certain examples, the amplifiers 112h comprise a suitable combination of two or more of the foregoing types ofPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO power amplifiers. Moreover, in some examples, individual ones of the amplifiers 112h correspond to individual ones of the transducers 114. In other examples, however, the electronics 112 includes a single one of the amplifiers 112h configured to output amplified audio signals to a plurality of the transducers 114. In some other examples, the electronics 112 omits the amplifiers 112h.
[0057] The transducers 114 (e.g., one or more speakers and / or speaker drivers) receive the amplified audio signals from the amplifier I I2h and render or output the amplified audio signals as sound (e.g., audible sound waves having a frequency between about 20 Hertz (Hz) and 20 kilohertz (kHz)). In some examples, the transducers 114 can comprise a single transducer. In other examples, however, the transducers 114 comprise a plurality of audio transducers. In some examples, the transducers 114 comprise more than one type of transducer. For example, the transducers 114 can include one or more low frequency transducers (e.g., subwoofers, woofers), mid-range frequency transducers (e.g., mid-range transducers, midwoofers), and one or more high frequency transducers (e.g., one or more tweeters). As used herein, “low frequency"’ can generally refer to audible frequencies below about 500 Hz, “midrange frequency” can generally refer to audible frequencies between about 500 Hz and about 2 kHz, and “high frequency” can generally refer to audible frequencies above 2 kHz. In certain examples, however, one or more of the transducers 114 comprise transducers that do not adhere to the foregoing frequency ranges. For example, one of the transducers 114 may comprise a mid-woofer transducer configured to output sound at frequencies between about 200 Hz and about 5 kHz.
[0058] By way of illustration, SONOS, Inc. presently offers (or has offered) for sale certain playback devices including, for example, a “SONOS ONE,” “PLAY:1,” “PLAY:3,” “PLAY:5,” “PLAYBAR,” “PLAYBASE.” “CONNECT: AMP,” “CONNECT,” and “SUB.” Other suitable playback devices may additionally or alternatively be used to implement the playback devices of example examples disclosed herein. Additionally, one of ordinary skilled in the art will appreciate that a playback device is not limited to the examples described herein or to SONOS product offerings. In some examples, for instance, one or more playback devices 110 comprises wired or wireless headphones (e.g., over-the-ear headphones, on-ear headphones, in-ear earphones). In other examples, one or more of the playback devices 110 comprise a docking station and / or an interface configured to interact with a docking station for personal mobile media playback devices. In certain examples, a playback device may be integral to another device or component such as a television, a lighting fixture, or some otherPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO device for indoor or outdoor use. In some examples, a playback device omits a user interface and / or one or more transducers. For example, Figure ID is a block diagram of a playback device IlOp comprising the input / output 111 and electronics 112 without the user interface 113 or transducers 114.
[0059] Figure IE is a block diagram of a bonded playback device HOq comprising the playback device 110a (Figure 1C) sonically bonded with the playback device HOi (e.g., a subwoofer) (Figure 1A). In the illustrated example, the playback devices 110a and HOi are separate ones of the playback devices 110 housed in separate enclosures. In some examples, however, the bonded playback device HOq comprises a single enclosure housing both the playback devices 110a and HOi. The bonded playback device HOq can be configured to process and reproduce sound differently than an unbonded playback device (e.g., the playback device 110a of Figure 1C) and / or paired or bonded playback devices (e.g., the playback devices 1101 and 110m of Figure IB). In some examples, for instance, the playback device 110a is fullrange playback device configured to render low frequency, mid-range frequency, and high frequency audio content, and the playback device 1 lOi is a subwoofer configured to render low frequency audio content. In some aspects, the playback device 110a, when bonded with the first playback device, is configured to render only the mid-range and high frequency components of a particular audio content, while the playback device HOi renders the low frequency component of the particular audio content. In some examples, the bonded playback device 1 lOq includes additional playback devices and / or another bonded playback device. c. Suitable Network Microphone Devices (NMDs)
[0060] Figure IF is a block diagram of the NMD 120a (Figures 1A and IB). The NMD 120a includes one or more voice processing components 124 (hereinafter “the voice components 124”) and several components described with respect to the playback device 110a (Figure 1C) including the processors 112a, the memory 112b, and the microphones 115. The NMD 120a optionally comprises other components also included in the playback device 110a (Figure 1C), such as the user interface 113 and / or the transducers 114. In some examples, the NMD 120a is configured as a media playback device (e.g., one or more of the playback devices 110), and further includes, for example, one or more of the audio components 112g (Figure 1C), the amplifiers, and / or other playback device components. In certain examples, the NMD 120a comprises an Internet of Things (loT) device such as, for example, a thermostat, alarm panel, fire and / or smoke detector, etc. In some examples, the NMD 120a comprises the microphones 115, the voice processing 124, and only a portion of the components of the electronics 112PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO described above with respect to Figure IB. In some aspects, for example, the NMD 120a includes the processor 112a and the memory 112b (Figure IB), while omitting one or more other components of the electronics 112. In some examples, the NMD 120a includes additional components (e.g., one or more sensors, cameras, thermometers, barometers, hygrometers).
[0061] In some examples, an NMD can be integrated into a playback device. Figure 1G is a block diagram of a playback device 1 lOr comprising an NMD 120d. The playback device 1 lOr can comprise many or all of the components of the playback device 110a and further include the microphones 115 and voice processing 124 (Figure IF). The playback device HOr optionally includes an integrated control device 130c. The control device 130c can comprise, for example, a user interface (e.g., the user interface 113 of Figure IB) configured to receive user input (e.g., touch input, voice input) without a separate control device. In other examples, however, the playback device 1 lOr receives commands from another control device (e.g., the control device 130a of Figure IB).
[0062] Referring again to Figure IF, the microphones 115 are configured to acquire, capture, and / or receive sound from an environment (e.g., the environment 101 of Figure 1A) and / or a room in which the NMD 120a is positioned. The received sound can include, for example, vocal utterances, audio played back by the NMD 120a and / or another playback device, background voices, ambient sounds, etc. The microphones 115 convert the received sound into electrical signals to produce microphone data. The voice processing 124 receives and analyzes the microphone data to determine whether a voice input is present in the microphone data. The voice input can comprise, for example, an activation word followed by an utterance including a user request. As those of ordinary' skill in the art will appreciate, an activation word is a word or other audio cue that signifying a user voice input. For instance, in query ing the AMAZON® VAS, a user might speak the activation word “Alexa / ’ Other examples include “Ok. Google” for invoking the GOOGLE® VAS and “Hey, Siri” for invoking the APPLE® VAS.
[0063] After detecting the activation word, voice processing 124 monitors the microphone data for an accompanying user request in the voice input. The user request may include, for example, a command to control a third-party device, such as a thermostat (e.g.. NEST® thermostat), an illumination device (e.g., a PHILIPS HUE ® lighting device), or a media playback device (e.g., a Sonos® playback device). For example, a user might speak the activation word “Alexa” followed by the utterance “set the thermostat to 68 degrees” to set a temperature in ahome (e g., the environment 101 of Figure 1 A). The user might speak the same activation word followed by the utterance “turn on the living room” to turn on illuminationPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO devices in a living room area of the home. The user may similarly speak an activation word followed by a request to play a particular song, an album, or a playlist of music on a playback device in the home.d. Suitable Control Devices
[0064] Figure 1H is a partially schematic diagram of the control device 130a (Figures 1A and IB). As used herein, the term ‘“control device” can be used interchangeably with ■‘controller” or '‘control system.” Among other features, the control device 130a is configured to receive user input related to the media playback system 100 and, in response, cause one or more devices in the media playback system 100 to perform an action(s) or operation(s) corresponding to the user input. In the illustrated example, the control device 130a comprises a smartphone (e.g., an iPhone™, an Android phone) on which media playback system controller application software is installed. In some examples, the control device 130a comprises, for example, a tablet (e.g., an iPad™), a computer (e.g., a laptop computer, a desktop computer), and / or another suitable device (e.g., a television, an automobile audio head unit, an loT device). In certain examples, the control device 130a comprises a dedicated controller for the media playback system 100. In other examples, as described above with respect to Figure 1G, the control device 130a is integrated into another device in the media playback system 100 (e.g., one more of the playback devices 110, NMDs 120, and / or other suitable devices configured to communicate over a network).
[0065] The control device 130a includes electronics 132, a user interface 133, one or more speakers 134, and one or more microphones 135. The electronics 132 comprise one or more processors 132a (referred to hereinafter as “the processors 132a”), a memory 132b, software components 132c, and a network interface 132d. The processor 132a can be configured to perform functions relevant to facilitating user access, control, and configuration of the media playback system 100. The memory 132b can comprise data storage that can be loaded with one or more of the software components executable by the processor 112a to perform those functions. The software components 132c can comprise applications and / or other executable software configured to facilitate control of the media playback system 100. The memory 112b can be configured to store, for example, the software components 132c, media playback system controller application software, and / or other data associated with the media playback system 100 and the user.
[0066] The network interface 132d is configured to facilitate network communications between the control device 130a and one or more other devices in the media playback systemPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO 100, and / or one or more remote devices. In some examples, the network interface 132d is configured to operate according to one or more suitable communication industry standards (e.g., infrared, radio, wired standards including IEEE 802.3, wireless standards including IEEE 802.11a, 802.11b, 802.11g, 802.1 In, 802.1 lac, 802.15, 4G, LTE). The network interface 132d can be configured, for example, to transmit data to and / or receive data from the playback devices 110, the NMDs 120, other ones of the control devices 130, one of the computing devices 106 of Figure IB, devices comprising one or more other media playback systems, etc. The transmitted and / or received data can include, for example, playback device control commands, state variables, playback zone and / or zone group configurations. For instance, based on user input received at the user interface 133, the network interface 132d can transmit a playback device control command (e.g.. volume control, audio playback control, audio content selection) from the control device 130 to one or more of the playback devices 110. The network interface 132d can also transmit and / or receive configuration changes such as, for example, adding / removing one or more playback devices 110 to / from a zone, adding / removing one or more zones to / from a zone group, forming a bonded or consolidated player, separating one or more playback devices from a bonded or consolidated player, among others. Additional description of zones and groups can be found below with respect to Figures II through IM.
[0067] The user interface 133 is configured to receive user input and can facilitate 'control of the media playback system 100. The user interface 133 includes media content art 133a (e.g., album art, lyrics, videos), a playback status indicator 133b (e.g., an elapsed and / or remaining time indicator), media content information region 133c, a playback control region 133d, and a zone indicator 133e. The media content information region 133c can include a display of relevant information (e.g., title, artist, album, genre, release year) about media content currently playing and / or media content in a queue or playlist. The playback control region 133d can include selectable (e g., via touch input and / or via a cursor or another suitable selector) icons to cause one or more playback devices in a selected playback zone or zone group to perform playback actions such as, for example, play or pause, fast forward, rewind, skip to next, skip to previous, enter / exit shuffle mode, enter / exit repeat mode, enter / exit cross fade mode, etc. The playback control region 133d may also include selectable icons to modify equalization settings, playback volume, and / or other suitable playback actions. In the illustrated example, the user interface 133 comprises a display presented on a touch screen interface of a smartphone (e.g., an iPhone™ an Android phone). In some examples, however, user interfaces of varyingPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO formats, sty les, and interactive sequences may alternatively be implemented on one or more network devices to provide comparable control access to a media playback system.
[0068] The one or more speakers 134 (e g., one or more transducers) can be configured to output sound to the user of the control device 130a. In some examples, the one or more speakers comprise individual transducers configured to correspondingly output low frequencies, midrange frequencies, and / or high frequencies. In some aspects, for example, the control device 130a is configured as a playback device (e.g., one of the playback devices 110). Similarly, in some examples the control device 130a is configured as an NMD (e.g., one of the NMDs 120), receiving voice commands and other sounds via the one or more microphones 135.
[0069] The one or more microphones 135 can comprise, for example, one or more condenser microphones, electret condenser microphones, dynamic microphones, and / or other suitable types of microphones or transducers. In some examples, two or more of the microphones 135 are arranged to capture location information of an audio source (e.g., voice, audible sound) and / or configured to facilitate filtering of background noise. Moreover, in certain examples, the control device 130a is configured to operate as playback device and an NMD. In other examples, however, the control device 130a omits the one or more speakers 134 and / or the one or more microphones 135. For instance, the control device 130a may comprise a device (e.g., a thermostat, an loT device, a network device) comprising a portion of the electronics 132 and the user interface 133 (e.g., a touch screen) without any speakers or microphones.e. Suitable Playback Device Configurations
[0070] Figures II through IM show example configurations of playback devices in zones and zone groups. Referring first to Figure IM, in one example, a single playback device may belong to a zone. For example, the playback device 110g in the second bedroom 101c (Figure 1A) may belong to Zone C. In some implementations described below, multiple playback devices may be '‘bonded” to form a “bonded pair” which together form a single zone. For example, the playback device 1101 (e.g., a left playback device) can be bonded to the playback device 1101 (e.g., a left playback device) to form Zone A. Bonded playback devices may have different playback responsibilities (e.g.. channel responsibilities). In another implementation described below, multiple playback devices may be merged to form a single zone. For example, the playback device 1 lOh (e.g., a front playback device) may be merged with the playback device HOi (e.g., a subwoofer), and the playback devices HOj and 110k (e.g., left and right surround speakers, respectively) to form a single Zone D. In another example, the playback devices 110g and 11 Oh can be merged to form a merged group or a zone group 108b. ThePATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO merged playback devices 110g and 11 Oh may not be specifically assigned different playback responsibilities. That is, the merged playback devices 1 lOh and 1 lOi may, aside from playing audio content in synchrony, each play audio content as they would if they were not merged.
[0071] Each zone in the media playback system 100 may be provided for control as a single user interface (UI) entity. For example, Zone A may be provided as a single entity named Master Bathroom. Zone B may be provided as a single entity named Master Bedroom. Zone C may be provided as a single entity named Second Bedroom.
[0072] Playback devices that are bonded may have different playback responsibilities, such as responsibilities for certain audio channels. For example, as shown in Figure 1 -I, the playback devices 1101 and 110m may be bonded so as to produce or enhance a stereo effect of audio content. In this example, the playback device 1101 may be configured to play a left channel audio component, while the playback device 110k may be configured to play a right channel audio component. In some implementations, such stereo bonding may be referred to as “pairing.'’
[0073] Additionally, bonded playback devices may have additional and / or different respective speaker drivers. As shown in Figure 1J, the playback device 1 lOh named Front may be bonded with the playback device HOi named SUB. The Front device IlOh can be configured to render a range of mid to high frequencies and the SUB device 1 lOi can be configured render low frequencies. When unbonded, however, the Front device 1 lOh can be configured render a full range of frequencies. As another example. Figure IK shows the Front and SUB devices 11 Oh and HOi further bonded with Left and Right playback devices HOj and 110k, respectively. In some implementations, the Right and Left devices HOj and 102k can be configured to form surround or “satellite” channels of a home theater system. The bonded playback devices 1 lOh, 1 lOi, 1 lOj, and 110k may form a single Zone D (Figure IM).
[0074] Playback devices that are merged may not have assigned playback responsibilities, and may each render the full range of audio content the respective playback device is capable of. Nevertheless, merged devices may be represented as a single UI entity (i.e., a zone, as discussed above). For instance, the playback devices 110a and 1 lOn the master bathroom have the single UI entity of Zone A. In one example, the playback devices 110a and 1 lOn may each output the full range of audio content each respective playback devices 110a and 11 On are capable of, in synchrony.
[0075] In some examples, an NMD is bonded or merged with another device so as to form a zone. For example, the NMD 120b may be bonded with the playback device HOe, whichPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO together form Zone F, named Living Room. In other examples, a stand-alone network microphone device may be in a zone by itself. In other examples, however, a stand-alone network microphone device may not be associated with a zone. Additional details regarding associating network microphone devices and playback devices as designated or default devices may be found, for example, in previously referenced U.S. Patent Application No. 15 / 438,749.
[0076] Zones of individual, bonded, and / or merged devices may be grouped to form a zone group. For example, referring to Figure IM, Zone A may be grouped with Zone B to form a zone group 108a that includes the two zones. Similarly, Zone G may be grouped with Zone H to form the zone group 108b. As another example, Zone A may be grouped with one or more other Zones C-I. The Zones A-I may be grouped and ungrouped in numerous ways. For example, three, four, five, or more (e.g., all) of the Zones A-I may be grouped. When grouped, the zones of individual and / or bonded playback devices may play back audio in synchrony with one another, as described in previously referenced U.S. Patent No. 8,234,395. Playback devices may be dynamically grouped and ungrouped to form new or different groups that synchronously play back audio content.
[0077] In various implementations, the zones in an environment may be the default name of a zone within the group or a combination of the names of the zones within a zone group. For example, Zone Group 108b can have be assigned a name such as “Dining + Kitchen”, as shown in Figure IM. In some examples, a zone group may be given a unique name selected by a user.
[0078] Certain data may be stored in a memory of a playback device (e.g., the memory 112b of Figure 1C) as one or more state variables that are periodically updated and used to describe the state of a playback zone, the playback device(s), and / or a zone group associated therewith. The memory may also include the data associated with the state of the other devices of the media system, and shared from time to time among the devices so that one or more of the devices have the most recent data associated with the system.
[0079] In some examples, the memory may store instances of various variable types associated with the states. Variables instances may be stored with identifiers (e.g., tags) corresponding to type. For example, certain identifiers may be a first type “al” to identify playback device(s) of a zone, a second type “bl” to identify playback device(s) that may be bonded in the zone, and a third type “cl” to identify a zone group to which the zone may belong. As a related example, identifiers associated with the second bedroom 101c may indicate that the playback device is the only playback device of the Zone C and not in a zone group. Identifiers associated with the Den may indicate that the Den is not grouped with otherPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO zones but includes bonded playback devices 110h-l 10k. Identifiers associated with the Dining Room may indicate that the Dining Room is part of the Dining + Kitchen zone group 108b and that devices 110b and HOd are grouped (Figure IL). Identifiers associated with the Kitchen may indicate the same or similar information by virtue of the Kitchen being part of the Dining + Kitchen zone group 108b. Other example zone variables and identifiers are described below.
[0080] In yet another example, the media playback system 100 may variables or identifiers representing other associations of zones and zone groups, such as identifiers associated with Areas, as show n in Figure IM. An area may involve a cluster of zone groups and / or zones not within a zone group. For instance, Figure IM shows an Upper Area 109a including Zones A-D, and a Low er Area 109b including Zones E-I. In one aspect, an Area may be used to invoke a cluster of zone groups and / or zones that share one or more zones and / or zone groups of another cluster. In another aspect, this differs from a zone group, w hich does not share a zone with another zone group. Further examples of techniques for implementing Areas may be found, for example, in U.S. Application No. 15 / 682,506 filed August 21, 2017 and titled “Room Association Based on Name / ’ and U.S. Patent No. 8.483,853 filed September 11, 2007, and titled “Controlling and manipulating groupings in a multi-zone media system.” Each of these applications is incorporated herein by reference in its entirety. In some examples, the media playback system 100 may not implement Areas, in which case the system may not store variables associated with Areas.111. Example Systems for Modifying XR Media Content Based on Real-World Conditionsa. Overview'
[0081] Various technologies for adapting extended reality (XR) experiences to real-world acoustic environments are described. A media playback system and / or an XR system may analyze acoustic characteristics of a physical space where a user is located and modify virtual content accordingly. The modifications may account for acoustic capabilities and limitations of the physical space, including properties of available audio equipment. In some implementations, the system (or multiple systems) supports shared XR experiences between multiple users in different physical locations by generating customized audio content for each location while maintaining a synchronized virtual experience. In various examples, the system may store environmental data and user preferences using distributed ledger technologies (DLT), enabling secure and persistent access to information used in generating the experiences such as acoustic signatures of different physical spaces. A generative media module may createPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO customized virtual environments based on the stored environmental and preference data. These technologies enable more immersive and personalized XR experiences by dynamically adapting virtual content to match real-world acoustic conditions.
[0082] Figure 2 is a schematic diagram of an example cloud-local distributed system, local hub system, control system, media playback system, and / or XR system 200. As illustrated, the system 200 includes components that enable adaptive audio experiences based on real-world acoustic environments. The system includes an XR device 202 that may incorporate various components for environmental analysis and content modification. These components may include environmental analysis component(s) 216 that work in conjunction with microphone(s) 222, sensor(s) 218, and audio transducer(s) 226 to analyze acoustic properties of real-world environments. A content generation component 214 may adjust virtual content based on the analyzed environmental characteristics. In some examples, the system 200 includes one or more cloud servers (e.g., one or more of the computing devices 106 of Figure IB). In other examples, the system 200 is a local network that lacks any cloud or remote computing devices. In some examples, the system 200 operates as a local or edge network with a capability’ to connect to one or more additional local networks and / or cloud servers.
[0083] In some implementations, the system 200 may utilize various types of remote computing infrastructure to support different operational needs. A conventional cloud server infrastructure, such as those provided by cloud service providers, may handle traditional data storage and remote computation tasks. Additionally or alternatively, the system may interact with public blockchain networks for storing and accessing characteristic data, usage patterns, and other relevant information in a distributed manner. The system 200 may further incorporate specialized cloud-based generative inference resources, such as inference-as-a-service platforms, specifically for producing generative audio content and modifications. While these different infrastructure types may all be considered cloud-based resources, in various implementations the local systems may communicate with and utilize these distinct physical resources at different times and for different purposes. For example, the system 200 may use conventional cloud storage for maintaining user profiles, access blockchain resources for verifying acoustic characteristics, and leverage specialized inference services for generating dynamic audio content. The system 200 may additionally implement different communication protocols, security measures, and access patterns appropriate to each type of remote resource while maintaining coordinated operation across the various infrastructure elements.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0084] The system 200 additionally includes various other components, including generative media components 204 and XR media content source(s) 206. The system 200 optionally includes one or more XR display devices 208 (e.g., wearable display devices such as headmounted displays, smartglasses, etc., and / or non- wearable display devices (e.g., conventional long-throw, medium- throw, or ultra short-throw projectors, holographic projection devices / systems, 3D screens, etc.). In some implementations, multiple audio playback devices 210a and 210b as well as playback networks 210c may be supported, enabling shared XR experiences across different physical environments.
[0085] The system 200 may additionally incorporate and / or interact with distributed ledger(s) 212 that optionally store various types of data including environmental data 228, XR media content 230, and user preference data 232. In some implementations, the distributed ledger(s) 212 may implement smart contract(s) 234, token(s) 236, and DAO(s) 238 to manage access rights, content distribution, and collaborative features. This architecture enables the system to maintain persistent environmental and user data while supporting secure and personalized XR experiences that adapt to the acoustic properties of different physical spaces. In the illustrated example, while the distributed ledgers 212 are represented as an independent feature of the system 200, it should not be understood that distributed ledgers 212 are necessarily implemented independent from other devices within the system 200. In one example, distributed ledger 212 may be implemented separate from other devices within system 200, including remote from the playback network and / or other playback devices. However, as those of ordinary skill in the art will appreciate, in other examples the distributed ledgers 212 can alternatively or additionally be distributed among some or all of the devices comprising the system 200, including local to and / or on the playback network and / or other playback devices. Additional details regarding the use of blockchain technologies, in addition to other useful features and techniques, are described in International Patent Application No. PCT / US2024 / 039870, filed July 26, 2024, titled “Systems and Methods for Maintaining Distributed Media Content History and Preferences,” which is hereby incorporated by reference in its entirety for all purposes. In particular, Figure 5 of the ‘870 application describes a distributed ledger system that may be used with some implementations of the present technology.
[0086] Communication among these various entities and devices can be carried out via network(s) 102, which as noted above can include any suitable wired or wireless network connections or combinations thereof (e.g., WiFi network, a Bluetooth, a Z-Wave network, aPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO ZigBee, an Ethernet connection, a Universal Serial Bus (USB) connection, a cellular telecommunications network connection (e.g. via a connection associated with 4G, 5G, 6G. 7G (and so on) technologies), etc ).a. Example Content Generation components
[0087] In some implementations, content generation component(s) 214 may include various modules and subcomponents configured to modify media content, such as XR media content, based on real-world environmental data. The content generation component(s) 214 may receive environmental data from environmental analysis component(s) 216, which may include acoustic measurements, room dimensions, material properties, audio equipment specifications, environment type (e.g.. vehicle, residence, commercial space, etc.), and auxiliary zone availability data from the real-world environment where an XR device 202 is located. The environmental data may additionally include temporal variations showing how acoustic characteristics change over time or under different conditions.
[0088] In some implementations, the system may modify extended reality content based on non-physical characteristics of the real-world environment, including temporal, social, and contextual factors. The environmental analysis components may consider time of day when determining appropriate acoustic modifications. For instance, during late night hours between 11 PM and 6 AM, the system may automatically adjust audio levels downward, reduce low frequency content that could propagate through walls, and implement more subtle acoustic effects to avoid disturbing others. Conversely, during typical waking hours, the system may allow fuller dynamic range and more impactful acoustic elements. The system may additionally analyze the type and social context of physical spaces to inform content modification. For example, in private single-family residences, the system may permit higher volume levels and more expansive sound fields, while in multi-unit dwellings, educational environments, or public spaces, the system may implement more constrained acoustic presentations with reduced sound propagation characteristics. These non-physical factors may be combined with physical acoustic analysis to determine appropriate modifications. In some implementations, users may provide additional context about their environments via distributed ledgers, enabling the system to build sophisticated models of temporal and social patterns that inform automated content adaptation.
[0089] The content generation component(s) 214 may include an acoustic adaptation module configured to analyze the received environmental data and determine modifications needed for virtual content. In some implementations, the acoustic adaptation module may evaluatePATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO frequency response limitations, spatial audio reproduction capabilities, and reverberation characteristics of the real-world environment. The module may analyze relationships between different acoustic profiles, enabling pattern recognition and prediction of acoustic behavior in similar environments. Based on this evaluation, the module may generate a set of modification parameters for adjusting the virtual content. These parameters may be stored locally and / or remotely. In an example, these parameters may be stored as blockchain data, including as one or more tokens such non-fungible tokens (NFTs) with associated metadata describing their acoustic properties and intended use cases.
[0090] In some implementations, the content generation component(s) 214 may implement and / or modify visual characteristics of virtual objects to maintain consistency with acoustic limitations of the real-world environment. As used herein, content generation components 214 may perform initial generation of content, modify pre-existing content, mix or combine preexisting content, or any combination thereof to produced output content that differs from any pre-existing content (e.g.. XR media content from XR media content sources 206). For example, if the real-world environment has limited low-frequency reproduction capabilities, the content generation component(s) 214 may adjust the size or material properties of virtual objects that would typically generate low-frequency sounds. A large virtual concert hall might be modified to appear as a smaller chamber music venue, with corresponding adjustments to material textures and architectural features to match the achievable acoustic properties. The system 200 may store multiple versions or variations of assets optimized for different acoustic environments, allowing selection of the most appropriate version based on real-world conditions.
[0091] The content generation component(s) 214 may include a spatial audio processing module configured to adjust positions and orientations of virtual sound sources. This module may analyze the placement of audio playback devices 210 in the real-world environment and modify the virtual scene to optimize sound source positioning. In vehicle-based implementations, the module may account for vehicle speed, cabin acoustics, and operation mode when determining spatial audio modifications. In some implementations, the module may identify auxiliary zones in the real-world environment that provide additional acoustic opportunities and adjust the virtual environment to utilize these zones effectively. For autonomous vehicles, the system may reconfigure the cabin space when transitioning between manual and autonomous operation modes to optimize the acoustic environment. Additional details regarding determining placement of playback devices, users, and / or other componentsPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO within a space, as well as determining environmental characteristics, can be found in the following patents, each of which is hereby incorporated by reference in its entirety for all purposes: (1) U.S. Patent No. 9,084,058, issued July 14, 2015, titled ‘'Sound Field Calibration Using Listener Localization”; (2) U.S. Patent No. 9,949,054, issued April 17, 2018, titled “Spatial Mapping of Audio Playback Devices in a Listening Environment”; (3) International Patent Application No. PCT / US2024 / 037012, filed July 8, 2024, titled “Height Audio Adjustment Based on Listening Environment Characteristics”; and (4) U.S. Patent No.11,800,318, issued October 24, 2023, titled “Systems and Methods for Playback Device Management.”
[0092] A psychoacoustic compensation module may be included in the content generation component(s) 214. This module may implement various psychoacoustic techniques to create perceptual effects that match the intended experience despite hardware limitations. For example, in implementations where low-frequency reproduction is limited, the module may generate harmonics or employ other psychoacoustic principles to create the impression of lower frequencies than the audio playback devices 210 can physically reproduce. In certain examples, the psychoacoustic compensation module includes one or more local psychoacoustic models that incorporate and / or are trained with data associated with user biological characteristics (e.g., biometric data, physiological data, and / or anatomical data). Additional details regarding use and determination of biological characteristics can be found in U.S. Patent No. 10,318,233, issued June 11, 2019, titled “Multimedia Experience According to Biometrics,” which is hereby incorporated by reference in its entirety for all purposes.
[0093] The content generation component(s) 214 may incorporate a real-time adaptation engine that continuously monitors changes in the real-world environment and updates modifications accordingly. This engine may process data from environmental analysis component(s) 216 to detect changes in room acoustics, user position, or audio system status. The engine may additionally monitor vehicle motion parameters in vehicle-based implementations, adjusting the acoustic environment based on speed, road conditions, and cabin noise levels. In some implementations, the engine may implement smooth transitions between different acoustic states (e.g., the user(s) move(s) from one room to another room, or one area of a room to another area of the room) to maintain user immersion. In various implementations, the various acoustic profiles associated with different real-world environments may be stored locally, on the cloud, via a distributed ledger (e.g., blockchain),PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO etc. For multi-vehicle scenarios, the engine may coordinate acoustic environments across multiple vehicles while maintaining synchronized experiences.
[0094] In multi-user scenarios, the content generation component(s) 214 may include synchronization modules that coordinate modifications across different XR devices 202 (including, in some examples, geographically remote XR devices 202) or XR display devices 208 (including, in some examples, geographically remote XR display devices 208). These modules may ensure that while each user receives content optimized for their specific environment, temporal alignment and spatial relationships between virtual elements are maintained across all participants. The synchronization modules may account for network latency and processing delays to maintain coherent shared experiences. In implementations involving asymmetric capabilities, the modules may provide different audio feeds based on each environment's capabilities while preserving the overall shared experience. For example, users experiencing a virtual concert may receive different acoustic representations based on their available speaker configurations while maintaining synchronized timing and spatial relationships.
[0095] In some implementations, the system may employ different strategies for handling varying environmental constraints in multi-user scenarios. In some instances, the system 200 may generate shared experiences based on the most restrictive environmental limitations among participating users. For example, if one user's environment permits full-range audio reproduction while another user's environment requires quieter operation, the system may adapt the shared virtual environment to accommodate the more restrictive quiet requirement, effectively implementing a lowest-common-denominator approach to maintain consistent experiences across users. Additionally or alternatively, the system 200 may generate asymmetric presentations of the same underlying content tailored to each user's environmental capabilities while maintaining temporal and narrative synchronization. For instance, during a shared virtual concert experience, a user in an environment supporting high-volume playback might experience the event in a large virtual arena with full concert acoustics, while another user in a noise-sensitive environment might simultaneously experience an acoustic version of the same performance in a smaller virtual venue. Despite these different presentations, the system 200 can maintain core experiential elements such as song selection, timing, and basic spatial relationships to preserve meaningful shared aspects of the experience. Smart contracts may govern how content can be modified across these different presentation modes while maintaining artistic integrity and appropriate synchronization between participants.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0096] The content generation component(s) 214 may also include an emotional impact preservation module. This module may analyze the intended emotional characteristics of virtual content and enable modifications that maintain or enhance these characteristics within the constraints of the real-world environment. For example, if a virtual scene is intended to create a sense of awe through massive reverberant spaces, but the real-world environment has limited reverberation capabilities, the module may employ alternative acoustic and visual techniques to achieve similar emotional impact. The module may incorporate personalized audio training simulator capabilities, adjusting content based on user responses and learning objectives while maintaining emotional engagement.
[0097] In some implementations, the content generation component(s) 214 may interface with distributed ledger(s) 212 to access and store modification parameters and user preferences. This integration may enable the system to leam from previous modifications and build personalized adjustment profiles for different users and environments. Smart contracts may govern how modifications are applied and shared across different XR experiences. The system 200 may implement token economics to for instance, incentivize high-quality environmental data collection and content creation, while DAOs may manage certain aspects of the modification ecosystem through community governance.
[0098] In some implementations, the system 200 may store generated modification parameters along with detailed contextual metadata describing the conditions under which they were created. This metadata may include temporal information (such as time and date), user identifiers, device specifications, environmental characteristics, and other relevant contextual factors that influenced the parameter generation. The modification parameters and associated metadata may be recorded on distributed ledgers as tokens that maintain verifiable relationships between content, modifications, and usage context. When these stored parameters are subsequently accessed or reused, the system 200 may analyze the recorded contextual metadata to validate appropriateness for new implementations, verity- modification provenance, determine parameter relevance for similar scenarios, and enable proper attribution or compensation for parameter reuse. Smart contracts may govern how stored parameters can be accessed and applied while maintaining appropriate relationships between original generation context and new usage scenarios.
[0099] The content generation component(s) 214 may include biometric response modules that incorporate user physiological data in the modification process. These modules may analyze biometric data from sensor(s) 218 to assess user engagement and comfort, adjustingPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO modifications accordingly. For example, if biometric data indicates user discomfort with certain acoustic properties, the module may automatically adjust the modifications to improve the experience. In vehicle implementations, the module may additionally monitor driver alertness and adjust acoustic properties to maintain optimal awareness levels. Additional details regarding obtaining and / or using such biometric data for multimedia playback can be found in U.S. Patent No. 10,318,233. issued June 11, 2019, titled “Multimedia Experience According to Biometrics." which is hereby incorporated by reference in its entirety for all purposes.
[0100] A machine learning engine may be incorporated within, and / or may be utilized by, the content generation component(s) 214 to optimize modification strategies over time. This engine may analyze relationships between environmental characteristics, applied modifications, and user responses to develop increasingly sophisticated adaptation techniques. The engine may utilize transformer models, diffusion models, or other suitable architectures for generating and modifying audio content. In some implementations, the engine may generate entirely new modification approaches based on patterns identified in historical data. The engine may additionally learn optimal transition strategies for different types of acoustic environment changes.
[0101] The content generation component(s) 214 may include head-related transfer function (HRTF) adaptation modules that adjust spatial audio rendering based on individual user characteristics. These modules may process personal HRTF data to create customized spatial audio experiences that account for both the user's unique hearing characteristics and the limitations of the real-world environment. In character perspective implementations, the modules may generate and map character-specific HRTFs to user characteristics, enabling perception of spatial audio from different virtual entities' perspectives.
[0102] In some implementations, the content generation component(s) 214 may employ generative Al techniques to create novel audio content that fits within environmental constraints while maintaining the intended user experience. For example, if certain sound effects cannot be adequately reproduced, the system may generate alternative sound designs that achieve similar narrative or functional purposes while working within the available acoustic capabilities. The system may utilize various generative models including audiospecific architectures like MusicLM, AudioGen, or other suitable approaches for creating and adapting content. In some examples, the generative Al techniques may involve the use of general generative Al models, including one or more large language models (LLMs), multimodal models, transformer models, world models, and / or other suitable models.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO Additional details regarding generative media content and associated devices, systems, and methods can be found in the following patents and applications, each of which is hereby incorporated by reference in its entirety for all purposes: (1) U.S. PatentNo. 11,985,376, issued May 14, 2024, titled “Playback of Generative Media Content”; (2) U.S. PatentNo. 12,175,161, issued December 24, 2024, titled “Generative Audio Playback via Wearable Playback Devices”; (3) International Patent Application No. PCT / US2024 / 026459. filed April 26, 2024, titled “Providing Moodscapes and Other Media Experiences”; and (4) International Patent Application No. PCT / US2024 / 039870, titled “Systems and Methods for Maintaining Distributed Media Content History and Preferences.”
[0103] The content generation component(s) 214 may also include modules for handling dynamic events and interactive content. These modules may predict potential acoustic requirements based on possible user interactions and prepare modification strategies in advance. This predictive capability may enable smooth transitions and consistent audio quality even during rapid changes in virtual content or user behavior. In vehicle implementations, the modules may anticipate acoustic needs based on navigation data and prepare appropriate modifications for upcoming route segments or vehicle state changes.b. Example Environmental Analysis Components
[0104] In some implementations, environmental analysis component(s) 216 may include various modules and subcomponents configured to analyze acoustic and physical properties of real-world environments where XR experiences take place. These components may be integrated into XR device(s) 202, implemented in wearable XR display devices 208 such as head-mounted displays (HMDs) or smartglasses, or distributed across various devices in the environment including playback devices 210, mobile devices, or dedicated environmental sensors.
[0105] The environmental analysis component(s) 216 may include acoustic measurement modules that work in conjunction with microphone(s) 222, and / or other sensors, to capture and analyze sound characteristics of the physical space. These modules may perform various types of acoustic analysis by measuring room impulse responses (IR) to characterize how sound propagates through the space. The modules may determine reverberation times across different frequency bands and analyze early reflection patterns and timing. In some examples, for instance, the environmental analysis component(s) 216 may cause the output of an impulse into the room (e g., a loud short burst of sound or noise) via one or more transducers and correspondingly use data acquired the microphone(s) 222 to determine the RT60 reverberationPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO time or another suitable reverberation metric of the physical space. Additionally, the modules may evaluate frequency response characteristics of the room, measure background noise levels and spectral content, determine sound absorption coefficients of different surfaces, and analyze the spatial distribution of acoustic properties.
[0106] In some implementations, the environmental analysis component(s) 216 may incorporate spatial mapping capabilities that work alongside the XR device's tracking systems to create detailed three-dimensional models of the physical space. These capabilities may use LiDAR or depth sensor scanning to capture room dimensions and geometry. The system may employ computer vision analysis to identify surface materials and textures, while tracking movable objects that might affect room acoustics. The spatial mapping capabilities may also detect potential auxiliary acoustic zones and map speaker and listener positions within the space. Additional details regarding mapping and characterizing a space can be found in International Patent Application No. PCT / US2024 / 037012, filed July 8, 2024, titled “Height Audio Adjustment Based on Listening Environment Characteristics,’' which is hereby incorporated by reference in its entirety for all purposes.
[0107] The environmental analysis component(s) 216 may include audio equipment analysis modules configured to evaluate the capabilities and limitations of available playback devices. These modules may measure frequency response characteristics of speakers and headphones and determine maximum output levels and distortion thresholds. The modules may assess spatial audio reproduction capabilities and evaluate latency and synchronization characteristics. Furthermore, they may monitor real-time performance and status of audio devices and detect connection and configuration changes in the audio system.
[0108] In some implementations, the environmental analysis component(s) 216 may implement distributed sensing networks that combine data from multiple devices in the environment. These networks may aggregate acoustic measurements from different spatial positions and cross-reference data from various sensor types. The system may combine measurements from different time periods, synthesize data from fixed and mobile sensors, and coordinate analysis across multiple rooms or spaces. Additional details regarding distributed sensing networks and associated devices can be found in the following patents, each of which is hereby incorporated by reference in its entirety for all purposes: (1) U.S. Patent No.9,084,058, issued July 14, 2015, titled “Sound Field Calibration Using Listener Localization”; and (2) U.S. Patent No. 10,318,233, issued June 11, 2019. titled “Multimedia Experience According to Biometrics.”PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0109] The environmental analysis component(s) 216 may include real-time monitoring modules that continuously track changes in the acoustic environment. These modules may detect movement of people or objects that affect acoustics and monitor changes in background noise levels. The system may identify when doors and windows are opened or closed and detect activation of HVAC systems or other noise sources. Additionally, the modules may track variations in humidity or temperature that might affect sound propagation.
[0110] A machine learning analysis engine may be incorporated within the environmental analysis component(s) 216 to identify patterns in acoustic behavior over time and predict how environmental changes might affect acoustics. The engine may classify different acoustic scenarios or conditions and optimize sensor fusion and data processing. In some implementations, the engine may generate acoustic models from incomplete measurement data.[OHl] In some implementations, the environmental analysis component(s) 216 may include crowd-sourced data collection modules that aggregate environmental data from multiple users and sessions. These modules may build databases of acoustic signatures for different spaces and track changes in environmental characteristics over time. The system may identify common acoustic challenges in similar environments and share optimized analysis parameters across devices. Furthermore, the modules may generate statistical models of acoustic variability.
[0112] The environmental analysis component(s) 216 may incorporate user interaction analysis modules that correlate environmental characteristics with user behavior and preferences. These modules may track user movement patterns and preferred listening positions while monitoring user adjustments to audio settings. The system 200 may analyze user feedback and comfort levels and detect patterns in user engagement with different acoustic environments. Additionally, the modules may identify environmental factors that impact user experience.
[0113] In some implementations, the environmental analysis component(s) 216 may include calibration and verification modules that facilitate accurate and consistent measurements. These modules may perform regular self-calibration of sensors and compare measurements across different devices. The system may validate data against known acoustic standards and detect and compensate for sensor drift or degradation. Furthermore, the modules may maintain measurement accuracy over time.
[0114] The environmental analysis component(s) 216 optionally integrates with blockchain systems through data management modules that store environmental measurements securelyPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO on distributed ledgers 212. These modules may track the provenance and validity of acoustic data and implement smart contracts for data sharing and access. The system may enable tokenization of valuable acoustic signatures and maintain audit trails of environmental analysis. Additional details regarding integration with blockchain technologies can be found in U.S. Patent No. 12,167,062, issued December 10, 2024, titled “Blockchain Data Based on Synthetic Content Generation,” which is hereby incorporated by reference in its entirety for all purposes.
[0115] Network analysis modules may be included in the environmental analysis component(s) 216 to evaluate available bandwidth for real-time data sharing and network latency between devices. The modules may assess quality of service for audio streaming and reliability of device synchronization. Additionally, the system may manage resource allocation for distributed processing.
[0116] The environmental analysis component(s) 216 may include biometric sensing modules that correlate user physiological responses with environmental characteristics. These modules may analyze heart rate variability in different acoustic conditions and stress indicators related to sound exposure. The system may track movement patterns in response to acoustic stimuli and measure attention and focus metrics. Furthermore, the modules may monitor overall comfort and well-being indicators.
[0117] In some implementations, the environmental analysis component(s) 216 may- incorporate predictive modeling modules that anticipate potential acoustic changes based on environmental factors. These modules may generate probabilistic models of acoustic behavior and forecast maintenance needs for acoustic treatment. The system may predict user responses to environmental conditions and model long-term acoustic trends. In some implementations, the system may utilize distributed computing capabilities across networked playback devices to optimize processing tasks. A playback device that is not actively rendering content may perform computational operations on behalf of other devices in the network that are actively playing content. For example, while a first playback device renders current audio content, a second idle playback device may execute predictive calculations to prepare acoustic modifications for upcoming content segments or anticipated user movements. The system may dynamically recruit available processing resources from inactive devices in the playback network to assist with computationally intensive tasks such as real-time acoustic analysis, modification parameter generation, or environmental modeling. Smart contracts may govern how processing resources can be shared between devices while maintaining appropriate performance priorities and compensation mechanisms for resource utilization.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0118] The environmental analysis component(s) 216 may include error detection and compensation modules designed to identify and filter out measurement artifacts and detect sensor malfunctions or interference. The modules may compensate for missing or corrupted data and maintain analysis quality during device failures. Additionally, the system may provide graceful degradation of analysis capabilities when needed.
[0119] These various components and modules may work together to provide comprehensive analysis of real-world environments, enabling the system 200 to optimize content delivery and user experience based on actual physical conditions and constraints.c. Example Generative Media Components
[0120] According to some aspects, generative media components 204 can be used to output generative media for use via the system 200. Generative media content can include any media content (e.g., XR media content, VR media content, AR media content, MR media content, audio, video, image, audio-visual output, tactile output (e.g., haptic or vibratory), text, olfactory, or any other suitable media content) that is dynamically created, synthesized, and / or modified by a non-human. via one or more algorithms, models (e.g., neural networks such as, for instance, a transformer models or similar suitable models), and / or agents (e.g., an Al agent running on one or more of the devices). This creation or modification can occur for playback in real-time or near real-time. Additionally or alternatively, generative media content can be produced or modified asynchronously (e.g.. ahead of time before playback is requested), and the particular item of generative media content may then be selected for playback at a later time. As used herein, a “generative media module” includes any system, whether implemented in software, a physical model, or combination thereof, that can produce generative media content based on one or more inputs and / or perform actions, operations, etc. based on one or more inputs, commands, and / or instructions. In some examples, such generative media content includes novel synthetic media content that can be created as wholly new or can be created by mixing, combining, manipulating, or otherwise modifying one or more pre-existing pieces of media content. In some examples, the generative media module comprises one or more autonomous agents to cany’ out one or more tasks based on an input(s) or command(s) received from one or more users and / or one or more other generative media modules. In some examples, the autonomous agents perform tasks absent input from one or more users.
[0121] In some implementations, the generative media module(s) are configured to produce a generative output based on one or more inputs using a generative media content model (or multiple generative models). The inputs can include sensor data, user input (e.g., as receivedPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO from a control device or via direct user interaction with the device, and / or media other content source(s). For example, a generative media module can produce and continuously modify generative media by adjusting various characteristics of the generative audio based on one or more input parameters (e.g., sensor data relating to one or more users).
[0122] Any suitable algorithm or combination of algorithms can be used to produce generative media content. Examples of such algorithms include those using machine learning techniques (e.g., generative adversarial networks, neural networks, etc.), formal grammars, Markov models, finite-state automata, and / or any other suitable algorithms. Transformer models, originally developed for natural language processing tasks, have been successfully adapted for media generation. These models utilize self-attention mechanisms to capture long-range dependencies in sequential data, making them well-suited for modeling complex audio structures. In the context of audio generation, transformer-based architectures like OpenAI's Jukebox have shown the ability to generate multi-instrumental music with coherent long-term structure, even incorporating sty listic elements and rudimentary lyrics. These models can be conditioned on various inputs, including genre, artist style, and even textual descriptions, allowing for fine-grained control over the generated media content.
[0123] As another example, diffusion models can be used high-fidelity audio synthesis. In many examples, these models work by learning to reverse a gradual noising process, effectively reconstructing media content from pure noise. Notable examples in the audio domain include Google's AudioLM and Meta Al's AudioGen. Diffusion models can generate realistic environmental sounds, speech, and music, and capture fine-grained audio details and maintaining consistency over extended durations. Moreover, diffusion models can be utilized in tasks such as text-to-audio generation and audio inpainting, where missing segments of audio are convincingly reconstructed.
[0124] Generative models may also be categorized based on their data type and modeling approach. Continuous latent models like AudioLDM and RAVE can employ iterative refinement techniques, while discrete latent models such as MusicGen and LLMs may utilize next-token prediction or autoregressive methods. Models working directly with raw audio data, like MaskGIT and WaveGAN, use masked prediction and one-step generation approaches respectively. Each of these models offers unique capabilities and trade-offs in generating audio content. For instance, MusicLM and MusicGen, introduced in 2023, represent significant advancements in connecting text and music modalities. Other notable models in the evolution of this technology include SoundStream, Encodec, CLIP, CLAP, and AudioLM from earlierPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO years, as well as more recent developments like Stable Audio, Suno, and UDio in 2024. In addition to audio-specific models, more general purpose (and optionally multimodal) models may be utilized (e.g., ChatGPT from OpenAI, Claude.ai from Anthropic, Llama from Meta, Gemini from Google, Deepseek, or other suitable models). As those of ordinary skill in the art will appreciate, in some examples, the generative models may be stored and / or ran locally (e.g., on one or more devices of the system 200), in the cloud (e.g., on one or more of the computing devices 106 of Figure IB), and / or distributed among multiple devices. Additional details regarding the distribution of generative modules can be found in U.S. Patent No. 11,812,240, issued November 7, 2023, titled “Playback of Generative Media Content,” which is hereby incorporated by reference in its entirety for all purposes.
[0125] In various examples, the generative media module(s) can utilize adaptive or deterministic models. Each approach offers distinct advantages in the production of media content. Deterministic models follow a fixed set of rules or parameters to generate output, producing consistent results given the same inputs. These models are predictable and can be precisely controlled, making them suitable for applications where reproducibility is desired. In contrast, adaptive models dynamically adjust their behavior based on real-time feedback or changing contextual inputs. These models can leam and evolve their output strategies over time, potentially producing more varied and context-aware content. In the realm of generative audio, adaptive models might adjust musical elements like tempo, harmony, or instrumentation in response to user interactions, environmental factors, or biometric data. This adaptability allows for a more responsive and personalized audio experience, potentially better suited to the dynamic nature of blockchain-capable playback devices. However, the choice between adaptive and deterministic approaches often depends on the specific use case, with some applications benefiting from a hybrid approach that combines elements of both model types. In various examples, the generative media module(s) can utilize any suitable generative algorithms now existing or developed in the future.
[0126] In various implementations, producing the generative media content (e.g., XR media content) can involve changing various characteristics of the media content in real time and / or algorithmically generating novel media content in real-time or near real-time.d. Example Distributed Ledger(s)
[0127] As noted above, the system 200 can include and / or communicate with one or more distributed ledger(s) 212, which in turn can store, instantiate, or communicate with one or more smart contracts 234, tokens 236, and / or decentralized autonomous organizations (DAOs) 238.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO Additional details regarding the use of or integration with distributed ledgers can be found in the following patents and applications, each of which is hereby incorporated by reference in its entirety for all purposes: (1) International Patent Application No. PCT / US2024 / 039870, filed July 26, 2024, titled “Systems and Methods for Maintaining Distributed Media Content History and Preferences'’; and (2) U.S. Patent No. 12,167,062, issued December 10, 2024, titled “Blockchain Data Based on Synthetic Content Generation.”
[0128] In some implementations, distributed ledger(s) 212 may provide decentralized storage and management capabilities for various types of data used in the system 200. The distributed ledger(s) 212 may be implemented using blockchain technology7or other distributed consensus mechanisms that facilitate data integrity, transparency, data provenance, trust, and secure access control. In some examples, for instance, the distributed ledger(s) 212 comprise one or more directed acyclic graphs (DAGs), hashgraphs, and / or other suitable distributed ledgers.
[0129] Environmental data 228 stored via the distributed ledger(s) 212 may include detailed acoustic signatures of physical spaces where XR experiences occur. This data may comprise any of the information obtained via the environmental analysis component(s) 216, for instance room impulse responses, frequency response characteristics, reverberation times, and other acoustic measurements. The environmental data 228 may be organized into acoustic profiles that can be referenced and retrieved when users return to previously analyzed spaces. In some implementations, the environmental data 228 may include temporal variations, showing how acoustic characteristics change over time or under different conditions. The system may store relationships between different acoustic profiles, enabling pattern recognition and prediction of acoustic behavior in similar environments.
[0130] XR media content 230 stored via the distributed ledger(s) 212 may include virtual environment data, sound effects, music, spatial audio recordings, and other media assets used in XR experiences. In some implementations, these assets may be stored as non-fungible tokens (NFTs) with associated metadata describing their acoustic properties, intended use cases, and modification parameters. The XR media content 230 may include multiple versions or variations of assets optimized for different acoustic environments, allowing the system to select the most appropriate version based on real-world conditions.
[0131] User preference data 232 may store individual users' acoustic preferences, listening histories, and interaction patterns. This data may include preferred audio settings for different types of content, sensitivity thresholds for various frequencies, and spatial audio preferences.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO In some implementations, the user preference data 232 may incorporate biometric response data, correlating physiological measurements with acoustic experiences to build personalized comfort profiles. The system may store user preferences for specific environments or content types, enabling automatic optimization of XR experiences based on individual user characteristics.
[0132] Smart contract(s) 234 implemented in the distributed ledger(s) 212 may govern various aspects of the system's operation. These contracts may manage access rights to environmental data and media content, automatically executing licensing agreements when content is used. In some implementations, smart contracts may control the modification and adaptation of XR content based on environmental conditions, ensuring that adjustments comply with content creators' specifications and user preferences. The smart contracts may also facilitate automated payments or rewards for users w ho contribute valuable environmental data or content modifications to the system.
[0133] Token(s) 236 may represent various digital assets and rights within the system 200. These may include utility tokens that grant access to specific features or content, governance tokens that enable participation in system decision-making, and asset tokens that represent ownership of virtual acoustic spaces or sound effects. In some implementations, tokens may be used to create acoustic marketplaces where users can trade or lease acoustic profiles, custom modifications, or specialized audio content. The system may implement token economics to incentivize high-quality environmental data collection and content creation.
[0134] Decentralized autonomous organizations (DAOs) 238 may manage certain aspects of the system 200 through community governance. These organizations may establish standards for acoustic data quality, vote on system improvements, and manage shared resources. In some implementations, DAOs may curate libraries of acoustic profiles and content modifications, ensuring that contributed data meets quality standards and serves the community's needs. The DAOs may also coordinate collaborative development of new acoustic features or experience types.
[0135] The distributed ledger(s) 212 may implement sophisticated access control mechanisms that protect sensitive data while enabling appropriate sharing and collaboration. These mechanisms may use hierarchical permission structures, allowing different levels of access for various user types and use cases. In some implementations, the system may employ zero-knowledge proofs or other privacy-preserving techniques to enable acoustic optimization without exposing detailed environmental or user data.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0136] The system may include data validation and verification modules that ensure the integrity of stored information. These modules may implement consensus mechanisms to validate new environmental measurements, content modifications, or user preference updates before they are added to the distributed ledger. In some implementations, the system may use reputation systems to track the reliability of data sources and content contributors.
[0137] Synchronization mechanisms within the distributed ledger(s) 212 may ensure consistency across multiple XR devices and environments. These mechanisms may manage real-time updates to acoustic data and content modifications while maintaining temporal consistency in shared XR experiences. The system 200 may implement efficient data distribution protocols that minimize latency while ensuring all participants have access to necessary information.
[0138] The distributed ledger(s) 212 may include analytics modules that process stored data to identify patterns, trends, and opportunities for optimization. These modules may analyze relationships between environmental characteristics, content modifications, and user responses to develop improved adaptation strategies. In some implementations, the analytics may inform the development of new smart contracts or DAO governance proposals.
[0139] Version control and history tracking capabilities within the distributed ledger(s) 212 may maintain records of changes to environmental data, content modifications, and user preferences. This historical data may enable rollback of changes if needed and provide valuable information for system improvement. The system may implement efficient storage mechanisms that balance the need for historical data with practical storage limitations.
[0140] The distributed ledger(s) 212 may include integration interfaces that enable interaction with external systems and services. These interfaces may allow import of acoustic data from other measurement systems, integration with content creation tools, and export of anonymized data for research or analysis. In some implementations, the system may implement cross-chain communication to interact with other blockchain networks or decentralized sen-ices.
[0141] Recovery and backup mechanisms may facilitate the resilience of stored data and system functionality. These mechanisms may include distributed backup strategies, redundant storage of critical data, and procedures for recovering from network or node failures. The system may implement automatic backup scheduling and verification to maintain data availability and integrity.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO e. Example Blockchain-Capable Playback Devices with Generative Components
[0142] In some implementations, the system 200 can include a device with distributed ledger functionality such as a blockchain-capable playback device that includes generative media components and / or generative context and control components. In some implementations, such a device can be an XR device 202, an audio playback device 210, or any other suitable network device. Figure 3 illustrates a schematic diagram of such a blockchain-capable playback device 300, which can include three component groups: distributed ledger components 302, generative media components 304, and generative context and control components 310. These components can work in concert, for instance sharing information with one another and with one or more blockchains, to produce a personalized and responsive media playback ecosystem that leverages both on-chain data and real-time contextual information. Additional details regarding blockchain-capable devices with generative media components and / or generative context and control components can be found in International Patent Application No. PCT / US24 / 39870, filed July 26, 2024 and titled “Systems and Methods for Maintaining Distributed Media Content History and Preferences / ’ which is hereby incorporated by reference in its entirety for all purposes.
[0143] As illustrated in Figure 3, the distributed ledger components 302 facilitate interaction with one or more distributed ledger networks (e.g. blockchain networks) (as described elsewhere herein), allowing the device to access, store, and update data on distributed ledgers. This capability enables the playback device to maintain and / or utilize decentralized records of user preferences, listening history, and other relevant data. In various implementations, the distributed ledger components 302 can be used to interact with one or more generative components, such as an artificial intelligence inference model that is implemented by the playback device or by the media playback system including the playback device. Such an artificial intelligence inference model can obtain inputs from and / or write outputs to, a blockchain or other distributed ledger technology using the distributed ledger components 302. For instance, the inputs and / or outputs can be obtained from or written to distributed ledger(s) 212, which as illustrated can include a sequence of blocks 320a-320h. which may store records such as Content Experience Record Sets (CERS), Content Network Record Sets (CNRS), or other suitable data regarding system, device, or user history and / or preferences.
[0144] The generative media components 304 include a generative media module 306 capable of producing novel, synthetic generative media content 308. These components can leverage artificial intelligence algorithms and models to create or modify media content in real-PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO time, providing users with unique and dynamically generated listening experiences. Additional details regarding suitable generative media components can be found in International Patent Application No. PCT / US2021 / 072454, filed November 17, 2021, titled “Playback of Generative Media Content,” which is hereby incorporated by reference in its entirety for all purposes.
[0145] The generative context and control components 310 encompass various applications and services that inform and guide the behavior of the playback device or other components of a media playback system including the playback device. These include a positioning system application 312 with common positioning API 1014 and a localization application 316. The generative context and control components 310 can further include a personalization service 318. These components work together to analyze the device's environment and user behavior, enabling context-aware adjustments to both the generative media output and the overall playback experience. In various examples, the generative context and control components 310 can provide an output in the form of a system recommendation, for instance relating to media playback, configuration of a media playback system, or for other aspects of the user's environment, media playback system, or components thereof. Additional details regarding suitable generative context and control components can be found in International Patent Application No. PCT / US2024 / 039698, filed July 26, 2024, titled “Personalization Techniques for Media Playback Systems," which is hereby incorporated by reference in its entirety for all purposes.
[0146] The interaction between these component groups allows the blockchain-capable playback device 300 to deliver personalized and dynamic media playback. By combining secure, decentralized data management with advanced generative capabilities and contextual awareness, the device can create audio experiences that are tailored to each user's preferences, environment, and current situation. The blockchain-capable playback device 300, which may be equipped with generative components, can use blockchain data as both inputs and outputs for its generative models, creating a dynamic and interconnected ecosystem of personalized audio experiences. As inputs, the generative media components 304 can utilize blockchain-stored data such as user preferences, listening history, and collaborative filtering data from the content experience record sets (CERS) to inform the generation of novel audio content. For example, a generative model might create a unique musical composition based on the user's preferred genres and artists as recorded on the blockchain.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0147] Similarly, the generative context and control components 310 can use DLT data (e.g., blockchain) data about device configurations and environmental preferences to adapt playback settings in real-time. As for outputs, the device 300 can record generated content metadata, user interactions with generated content, and performance metrics of the generative models back to the blockchain. For instance, if the device 300 generates a personalized playlist, it could store the playlist structure and user engagement data on the blockchain for future reference or sharing. Additionally, the device could contribute improvements to the generative models themselves, such as updated parameters or new training data, by recording these advancements on the blockchain for other devices to access and incorporate. This bidirectional flow7of data between the blockchain and the generative components can create a continuously evolving and improving system of personalized media content experiences.
[0148] The integration of both DLT and / or blockchain technology with generative media techniques can provide highly personalized and dynamic media content experiences. For instance, by leveraging data stored on distributed ledgers as input parameters for Al models, such as generative Al and / or generative media models, blockchain-capable playback devices can create tailored content and draw from a wide variety of data sources that may be available via one or more blockchains. This approach can also allow7for a more individualized understanding of user preferences, listening history, and environmental factors, all of which can inform the generative process.
[0149] Blockchain data, such as content experience record sets (CERS) and content network record sets (CNRS) described above, can provide a rich source of information for generative media models. As noted previously, these decentralized records can include detailed histories of user interactions, preferred audio characteristics, and even collaborative filtering data from similar users. When fed into generative algorithms, this blockchain-sourced data enables the creation of media content that not only reflects individual tastes but can also incorporate broader trends and patterns observed across the network.
[0150] Additionally7or alternatively, the outputs of generative media models and the resulting generative media content can be recorded or otherwise stored, in whole or in part, on decentralized ledgers. This process can, in some cases, create a feedback loop, in which the generated content and associated playback data become part of the user's blockchain-recorded history. By writing this information to the distributed ledger, the system can facilitate a transparent and immutable record of the generative media experience, which can then be used to further refine future content generation. This cyclical flow of data between blockchain andPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO generative systems allows for continuous improvement and personalization of the media content experience.
[0151] According to some aspects, generative media content can include any media content (e.g., audio, video, image, lighting, audio-visual output, tactile output, text, olfactory, or any other suitable media content) that is dynamically created, synthesized, and / or modified by a non-human, via one or more algorithms or models (e.g.. neural networks such as. for instance, a transformer models, world models, or similar suitable models). This creation or modification can occur for playback in real-time or near real-time. Additionally or alternatively, generative media content can be produced or modified asynchronously (e.g., ahead of time before playback is requested), and the particular item of generative media content may then be selected for playback at a later time. As used herein, a ‘"generative media module” includes any system, whether implemented in software, a physical model, or combination thereof, that can produce generative media content based on one or more inputs. In some examples, such generative media content includes novel synthetic media content that can be created as wholly new or can be created by mixing, combining, manipulating, or otherwise modifying one or more preexisting pieces of media content. Although several examples throughout this discussion refer to audio content (e.g., music, spoken work, and / or other sound(s)), the techniques disclosed herein can be applied in some examples to other types of media content, e.g., video, audiovisual, tactile, text, or otherwise.
[0152] The generative context and control components 310, which include localization applications and positioning system applications, can both retrieve data from and contribute data to blockchain sources. This bidirectional flow of information allows for a dynamic interplay between historical, persistent data stored on-chain and real-time, dynamic data generated by the playback device and its environment.
[0153] Localization and positioning system application data may be sourced from the blockchain, forming its own distinct “record set” within the broader blockchain ecosystem. This data can be used in combination with information stored locally on the device, on other devices within the same household / network, devices associated with a community and / or a DAO, or in traditional cloud storage. For example, a user's playback and listening history or media environment configuration might be retrieved from the blockchain (e.g., by accessing a CERS or CNRS), providing a comprehensive and portable record of preferences and behaviors. Meanwhile, current device positioning information, which may be more sensitive and requirePATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO real-time accuracy, might be stored locally within the household network for improved privacy and reduced latency.
[0154] In some implementations, even sensitive or private information can be stored on-chain when appropriate privacy -preserving techniques are employed. For instance, zeroknowledge proofs can be utilized to verify certain attributes or conditions without revealing the underlying data, allowing for the benefits of blockchain-based data management while maintaining user privacy.
[0155] The interaction between blockchain data and generative context and control components enables a more personalized approach to media playback. In an example approach, historical data from the blockchain can inform long-term preferences and patterns, while realtime contextual data allows for immediate responsiveness to the user's current environment and situation. In some examples, contextual and / or input data can be collected and shared among devices within the MPS 100 to support generative context and control operations. Examples of suitable contextual data can be found in International Application No. PCT / US2022 / 077185, filed September 28, 2022, titled “Spatial Mapping of Media Playback System Components,” which is hereby incorporated by reference in its entirety for all purposes.
[0156] In some implementations, generative models are used for purposes of context and control. Additional details regarding suitable generative models can also be found in International Patent Application No. PCT / US2024 / 029969, filed May 17, 2024, titled “Learned Device Targeting,” which is hereby incorporated by reference in its entirety for all purposes.
[0157] Additional details regarding mapping contextual or environmental data with respect to subscriber devices can be found in International Patent Application No. PCT / US2024 / 026459, filed April 26, 2024, titled “Providing Moodscapes and Other Media Experiences.” which is hereby incorporated by reference in its entirety for all purposes. Among examples, such data regarding a map of provider contextual or environmental data with respect to subscriber devices can be stored via a distributed ledger (e.g., as part of a CERS or CNRS). The use of a blockchain allows for a verified experience such that individual subscribers can confirm and trust that the moodscape simulated at their location corresponds (optionally in real-time) to the provider moodscape. Additionally or alternatively, to facilitate production of a suitable moodscape as described in the above-referenced application, a personalization model may be used to tailor the moodscape for the listener’s context and environmental conditions.
[0158] According to certain aspects, a positioning system can be implemented to determine relative positioning of devices within the environment and optionally to control or modifyPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO behavior of one or more devices based on the relative positions. Positioning or localization information can be acquired through various techniques, optionally using sensors in some instances, examples of which are discussed below. In certain examples, one or more devices in the MPS 100, such as one or more blockchain-capable playback devices 300, playback devices 110, NMDs 120, or controller devices 130 may host a localization application that may implement features that process localization information to enhance user experiences with the MPS 100. Examples of such features include sophisticated acoustic manipulation (e.g., features directed to psychoacoustic effects during audio playback) and autonomous device configuration / reconfiguration (e.g., features directed to detection and configuration of new7devices or devices that have moved or otherwise been changed in some way), among others. The requirements that these features place on localization information vary, with some features requiring low latency, high precision localization information and other features being able to operate using high latency, low precision localization information.
[0159] According to certain examples, a positioning system can be implemented using a variety of different devices to generate the localization information utilized by certain application features. However, the number, arrangement, and configuration of these devices can vary between examples. Additionally, or alternatively, the communications technology and / or sensors employed by the devices can vary. Given the number of variables in play within any particular MPS and the concomitant inefficiencies that this variability7imposes on MPS application feature development and maintenance, some examples disclosed herein utilize one or more blockchain-capable playback devices 300, playback devices 110, NMDs 120, or controller devices 130 to implement a positioning system using a common positioning application programming interface (API) that decouples the positioning / localization information from specific devices or underlying enabling technologies.IV. Example Methods
[0160] As noted previously, the system 200 described above may modify virtual environments based on real-world conditions by analyzing characteristics of physical spaces and adjusting virtual content accordingly. The system may adapt virtual scenes for implementations ranging from residential rooms to vehicle cabins, modifying content based on factors such as room acoustics, available playback devices, user location, and environmental conditions. These modifications may account for vary ing hardware capabilities, from simple stereo setups to complex multi-channel arrays, as well as different acoustic conditions such as reverberation times, frequency response limitations, and background noise levels. The systemPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO may additionally consider user preferences, biometric responses, and shared experience requirements when determining appropriate modifications. Through the various example methods described below, the system 200 enables creation of optimized extended reality experiences across diverse physical environments while maintaining intended perceptual and emotional characteristics of the virtual content. The following scenarios illustrate examples of such implementations, though it should be understood that these are provided for illustration only and do not limit the scope of the technology.
[0161] Figure 4 illustrates operation of content generation component(s) 214 in an example extended reality (XR) system. In some implementations, XR media content source(s) 206 may provide XR media content to content generation component(s) 214. The XR media content may include virtual environment data such as three-dimensional models, textures, audio content, and associated metadata defining a virtual scene or experience.
[0162] Environmental analysis component(s) 216 may provide environmental parameter(s) to the content generation component(s) 214. The environmental parameter(s) may characterize acoustic properties of a real-world environment where an XR experience takes place. In some implementations, these parameters may include room impulse responses, reverberation times, frequency response characteristics, background noise levels, and other acoustic measurements obtained through analysis of the physical space.
[0163] In some implementations, the environmental analysis components 216 may collect and process a comprehensive set of system parameters including detailed information about all available playback devices in the network. These parameters may include specifications, capabilities, and precise positional data not only for speakers within the active playback zone where XR content is being experienced, but also for all playback devices throughout the broader environment that could potentially supplement audio reproduction. For example, the system 200 may maintain awareness of auxiliary subwoofers, satellite speakers, or full-range playback devices in adjacent rooms or other floors of a building that could be temporarily recruited to enhance audio presentation. The environmental analysis may include data about each device's frequency response range, power handling capabilities, current operational status, physical orientation, distance from the primary playback zone, and potential acoustic contributions through walls, floors, or other architectural elements. Smart contracts may record and manage this comprehensive device capability data to enable dynamic resource allocation and optimization across the extended playback network.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0164] In some implementations, distributed ledger(s) 212 may provide additional data to the content generation component(s) 214. This additional data may include user preferences, historical acoustic profiles, content modification parameters, or other information stored in the distributed ledger system. The content generation component(s) 214 may use this data to inform how the XR media content should be modified.
[0165] Based on the received inputs, the content generation component(s) 214 may generate two types of modified content. First, the content generation component(s) 214 may output modified XR visual content to an XR display device, such as a head-mounted display (HMD) worn by a user. The modified XR visual content may include adjustments to the virtual environment that maintain consistency with acoustic limitations or capabilities of the real-world space.
[0166] Additionally or alternatively, the content generation component(s) 214 may output modified XR audio content to one or more playback devices in the real-world environment. The modified XR audio content may be adjusted based on the environmental parameter(s) to optimize audio reproduction given the acoustic properties and limitations of the physical space. In some implementations, the modifications to both visual and audio content may be coordinated to provide a coherent and immersive XR experience that accounts for real-world environmental conditions.
[0167] Figure 5 illustrates operation of an extended reality (XR) system configured to support multiple users in different physical environments. The figure shows how XR media content may be modified differently for each environment while maintaining a coordinated experience between users. In the illustrated implementation, XR media content source(s) 206 provide XR media content to content generation component(s) 214. This content may include virtual environment data such as three-dimensional models, textures, audio content, and associated metadata defining a shared virtual experience.
[0168] A first environment, shown in the upper portion of Figure 5, includes environmental analysis component(s) 216a configured to analyze acoustic properties of the first physical space. These component(s) 216a may generate first environmental parameter(s) characterizing the acoustic capabilities and limitations of the first environment. The first environment may include an XR display device (e.g., head-mounted display) worn by a first user and an audio playback device for providing audio output.
[0169] A second environment, shown in the lower portion of Figure 5, includes environmental analysis component(s) 216b configured to analyze acoustic properties of thePATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO second physical space. These component(s) 216b may generate second environmental parameter(s) characterizing the acoustic capabilities and limitations of the second environment. The second environment may similarly include an XR display device worn by a second user and an audio playback device, which may have different capabilities than those in the first environment.
[0170] The content generation component(s) 214 may receive both the first and second environmental parameter(s) and generate customized modifications for each environment. In some implementations, the content generation component(s) 214 may output first modified XR content optimized for the acoustic properties of the first environment, and second modified XR content optimized for the acoustic properties of the second environment.
[0171] For example, if the first environment has limited low-frequency reproduction capabilities while the second environment has full-range audio capabilities, the content generation component(s) 214 may generate different audio modifications for each space. The first environment might receive modified content that employs psychoacoustic techniques to simulate low frequencies, while the second environment receives content that takes full advantage of its superior bass response.
[0172] In some implementations, the visual aspects of the virtual environment may also be modified differently for each user while maintaining spatial and temporal consistency between the environments. For instance, if the first environment has significant reverberation that limits directional audio cues, the visual representation might be adjusted to provide additional visual feedback for sound source locations. Meanwhile, the second environment might maintain standard visual representations if its acoustic properties allow for accurate spatial audio reproduction.
[0173] The content generation component(s) 214 may implement synchronization mechanisms that, despite different modifications, enable both users experience a coherent shared virtual space. This may involve maintaining consistent timing of events, preserving relative spatial relationships between virtual objects, and ensuring that interactive elements remain coordinated between environments. For example, in some implementations, the system may distinguish between core synchronization elements that must remain identical across users and adaptive elements that can vary while maintaining meaningful shared experiences. The system may implement a multi-tier synchronization framework where certain interactive elements are designated as invariant across all participants, while other aspects of the experience may be modified based on local environmental capabilities and constraints. ForPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO example, in a shared virtual concert experience, the system may maintain strict synchronization of fundamental elements such as song selection, timing, and basic performer positions, enabling users to experience substantially the same or similar musical progressions. However, the system may simultaneously allow other aspects of the experience to vary based on local playback capabilities and environmental constraints. A user with a basic stereo system might experience a simplified two-channel mix of the performance, while another user with a full home theater array might receive an expanded spatial audio presentation incorporating additional elements such as spatialized crowd reactions, detailed venue acoustics, and immersive environmental effects. Similarly, users in noise-sensitive environments might experience the same musical performance rendered as an intimate acoustic arrangement, while users in environments permitting higher volume levels might experience a full arena rock presentation, with both versions maintaining synchronized timing and core musical elements. The system may employ smart contracts to govern which aspects of experiences can be modified while ensuring minimum synchronization requirements are maintained. These contracts may define tiered experience levels based on available playback capabilities while preserving essential shared elements that enable meaningful social interaction and shared reference points between users experiencing different versions of the same content.
[0174] In some implementations, the content generation component(s) 214 may dynamically adjust modifications as users move or interact within their respective environments. The system may continuously monitor environmental parameters and update modifications accordingly, while maintaining the synchronized experience between users.
[0175] The system may also consider network conditions between environments, adjusting modifications to account for latency or bandwidth limitations. This may include predictive modifications that anticipate user actions or environmental changes to maintain smooth coordination between spaces.
[0176] When audio content includes virtual sound sources, the content generation component(s) 214 may position these sources differently in each environment based on the available playback devices and acoustic properties. However, the relative positioning and movement of sound sources may be maintained to preserve the spatial relationships of the shared experience.
[0177] In some implementations, the system may include mechanisms for users to perceive aspects of each other's environments while maintaining individual optimizations. For example,PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO visual or audio cues might indicate when one user's environment has different acoustic capabilities, helping users understand and adapt to their shared but distinct experiences.
[0178] In some implementations, the system may dynamically modify real-world speaker configurations in response to virtual content requirements and opportunities. The system may analyze available playback devices across different zones of a playback network to identify auxiliary audio resources that could enhance content presentation. For example, when a user experiences virtual content in a home theater zone, the system may detect an available subwoofer in a basement zone that could supplement low-frequency reproduction capabilities. Upon determining that upcoming virtual content would benefit from enhanced bass response, such as during a virtual explosion effect, the system may temporarily bond the auxiliary subwoofer with the home theater playback devices to create an expanded speaker configuration. The content modification components may then generate audio content that specifically utilizes this enhanced speaker array, taking advantage of the additional low-frequency capabilities while maintaining appropriate acoustic balance. After the content segment requiring enhanced capabilities concludes, the system may automatically dissolve the temporary bond and return devices to their original configurations. Smart contracts may govern how auxiliary playback devices can be dynamically recruited and bonded while maintaining appropriate access controls and usage tracking across different zones of the playback network.a. Methods for Synchronizing Out-Loud and Wearable Audio Playback
[0179] In some implementations, a user may experience XR media content using both a wearable audio playback device (e.g., headphones) and out-loud audio playback devices (e.g., a stationary soundbar or other playback device). In such scenarios, it can be useful to enable audio content from the various devices to be synchronized with respect to the user. Additional details regarding such synchronization techniques can be found in U.S. Provisional Patent Application No. 63 / 891,022, filed September 30, 2025 and titled “Synchronization of Wearable and Out-Loud Audio Playback,” which is hereby incorporated by reference in its entirety.
[0180] In some implementations, the system can achieve synchronization by embedding reference signature components into audio data streams and comparing known reference timing with detected ultrasonic signals to calculate timing offsets. Additionally or alternatively, the system can utilize internal and external microphones on the wearable device to detect ultrasonic synchronization signals from both the wearable device itself and the out-loud speakers, enabling precise measurement of signal arrival times at the listening position. The synchronization process operates efficiently on resource-constrained devices and can bePATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO triggered by various events including device activation, user movement, device changes (e.g., dropouts in wireless communication links) or environmental changes to maintain proper temporal alignment of audio content.
[0181] Figure 6 illustrates an example scenario with a user wearing a wearable playback device 602a in proximity to an out-loud playback device 602b. The wearable playback device 602a takes the form of a headphone device that is positioned over or around the user's ears, though in other implementations the wearable playback device 602a may comprise other forms of wearable audio devices such as earbuds, in-ear monitors, audio glasses, or open-ear hearable devices. The user may also wear a separate XR display device, or such a device may be integrated with the wearable playback device 602a. The out-loud playback device 602b is shown as a soundbar positioned nearby the user, though in various implementations the out-loud playback device 602b may comprise other forms of speakers such as loudspeakers, portable speakers, or combinations of multiple speakers arranged in the listening environment.
[0182] The wearable playback device 602a includes various components that enable both audio playback and acoustic monitoring functions. One or more audio transducers 718 are positioned within the wearable playback device 602a to provide audio output to the user. These audio transducers 718 may include drivers, speakers, or other acoustic output devices configured to generate both audible audio content and ultrasonic reference signals as described herein. The audio transducers 718 may be configured to reproduce a full range of audio frequencies, including frequencies extending into the near-ultrasonic range for synchronization purposes.
[0183] The wearable playback device 602a further includes a microphone system comprising both external microphones 722a and internal microphones 722b positioned to capture different acoustic signals. The one or more external microphones 722a are positioned on or within the wearable playback device 602a such that they are configured to detect audio signals from the out-loud playback device 602b and other external sound sources in the environment. In some implementations, the external microphones 722a may be positioned on an outer surface of the wearable playback device 602a or in locations that provide acoustic access to the ambient environment while the device is worn.
[0184] The one or more internal microphones 722b are positioned within the wearable playback device 602a such that they are configured to primarily detect audio signals generated by the audio transducers 718 of the wearable playback device 602a itself, particularly at the near-ultrasonic frequency range where acoustic isolation is stronger. In some implementations,PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO the internal microphones 722b may be positioned in or near the acoustic chamber of the wearable playback device 602a, in ear cups of headphone implementations, or in other locations where they can effectively monitor the audio output that reaches the user's ears. This positioning allows the internal microphones 722b to capture the self-generated audio signals, including ultrasonic reference signatures, at or near the listening position.
[0185] The out-loud playback device 602b is configured to generate audio content and ultrasonic reference signals that propagate through the acoustic environment to reach both the user's ears directly and the external microphones 722a of the wearable playback device 602a. In operation, the out-loud playback device 602b may include components similar to those described above with respect to the playback devices 102 of the MPS 100, including audio processing components, amplifiers, speakers, and network interfaces for communication with other devices in the system.
[0186] Both the wearable playback device 602a and the out-loud playback device 602b may include processing circuitry, memory, and communication interfaces to enable the synchronization techniques described herein. These devices may communicate with each other directly or through intermediate network devices to coordinate the generation of ultrasonic reference signals and the exchange of timing information necessary for synchronization adjustments. In some implementations, the wearable playback device 602a may perform the primary analysis of captured acoustic data and timing calculations, while in other implementations these functions may be distributed across multiple devices in the system. In various implementations, the devices may also communicate known delays due to signal processing and communication interfaces, which may provide a coarse estimate of overall offset which is then refined with acoustic analysis as described herein.
[0187] Figure 7 illustrates a swim diagram 700 showing the sequence of operations performed by the wearable playback device 602a and the out-loud playback device 602b to achieve synchronization through data stream analysis of embedded reference signatures. The diagram demonstrates how both devices coordinate to embed, detect, and analyze ultrasonic reference signatures within audio data streams to determine and correct timing offsets.
[0188] The process begins with the out-loud playback device 602b performing an embedding operation 702 to embed an ultrasonic signature into the audio data stream. This embedding operation 702 involves inserting reference signature components into the data stream that includes the audio content intended for playback. The reference signature component may bePATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO embedded as metadata, as a separate data channel, or as frequency components within the audio data itself. In some implementations, the embedding occurs at predetermined intervals during audio playback, while in other implementations the embedding may be triggered by specific events such as initiation of a synchronization procedure or detection of timing drift or detection of events that may cause a timing drift (e.g., packet loss on wireless communications interface, etc.). In various implementations, the ultrasonic signature can be embedded by another playback device, by an audio source directly, or by any other components of the media playback system, and need not be embedded directly by the out-loud playback device 602b itself.
[0189] Following the embedding operation 702, the audio data stream with the embedded ultrasonic signature 704 is transmitted from the out-loud playback device 602b to the wearable playback device 602a. This transmission may occur through various communication channels, including wireless networks, wired connections, or other data transmission methods available within the media playback system. The audio data stream includes both the intended audio content for playback and the embedded reference signature components that will enable the synchronization analysis. In some implementations, the wearable playback device 602a and the out-loud playback device 602b each separately receive audio data streams with embedded signatures (e g., from an audio source directly, from other devices within the media playback system, etc.), with neither one transmitting the stream to the other.
[0190] Both devices then proceed to playback operations, with the wearable playback device 602a performing a playback operation 706 to play back audio data with the embedded ultrasonic signature, and the out-loud playback device 602b performing a corresponding playback operation 708 to play back audio data with the ultrasonic signature. These playback operations can occur substantially simultaneously, with each device generating both the audible audio content and the ultrasonic reference signatures based on the same reference signature components embedded in the data stream. The ultrasonic signatures may comprise frequencies in a range between approximately 16 kHz and 24 kHz, to ensure they remain above the A pical range of human hearing while still being reproducible by standard audio equipment. Alternatively, the playback operations can occur at known temporal offsets from one another, which may facilitate disambiguation of the two ultrasonic reference signals.
[0191] The wearable playback device 602a then performs a detection operation 710 to detect the ultrasonic signature in the data stream for self-playback. This detection operation 710PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO involves analyzing the data stream in a buffer maintained by the wearable playback device 602a to identify the reference signature components and determine the reference timing (according to its local clock) of the ultrasonic reference signature that the device is generating. The analysis may involve parsing the embedded metadata, extracting frequency components, or otherwise identifying the timing information associated with the reference signature within the data stream. This step enables the wearable playback device 602a to establish a precise temporal reference for its own ultrasonic output (relative to its own local clock without factoring in any downstream delays to acoustic output on the wearable playback device 602a).
[0192] The data stream analysis for reference timing determination may utilize multiple sources of timing information depending on the system configuration and accuracy requirements. In some implementations, the wearable playback device 602a analyzes the data stream directly from audio buffers containing data scheduled for playback, providing timing information based on the device's internal processing timeline. Additionally or alternatively, the system may utilize acoustic feedback by capturing the device's own audio output through internal microphones after the signal has been rendered through the complete playback chain, including digital-to-analog conversion, amplification, and acoustic transduction. This microphone-based approach becomes particularly valuable when variable or unknown downstream delays exist in the playback path after the point where digital signal analysis can be performed. The system may employ a combination of both buffered data analysis and acoustic self-monitoring to account for the complete signal path from digital processing through acoustic output, thereby providing more accurate timing references that reflect the actual acoustic signal timing at the listening position.
[0193] Concurrently, the wearable playback device 602a performs another detection operation 712 to detect the ultrasonic signature in the out-loud playback. This detection operation 712 utilizes the microphones of the wearable playback device 602a, particularly the external microphones 622a described with respect to Figure 6, to capture acoustic data from the environment and identify the ultrasonic reference signature generated by the out-loud playback device 602b. The detection may involve signal processing techniques such as frequency analysis, pattern recognition, or watermark detection (similar to the techniques described above with respect to operation 710) to identify the specific ultrasonic signature pattern within the captured acoustic data and determine its time-of-arrival.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0194] The wearable playback device 602a then performs a comparison operation 814 to compare timing data for the out-loud signature and self-signature. This comparison operation 714 involves analyzing the reference timing determined from the data stream analysis in operation 710 and comparing it with the time-of-arrival of the ultrasonic signature detected from the out-loud playback device 602b in operation 712. The comparison takes into account the known timing relationship between the reference signatures generated by both devices, accounting for any intentional offsets or delays that may be part of the system design. The result of this comparison is a calculated time offset that represents the timing difference between the playback paths of the two devices.
[0195] Based on the results of the comparison operation 714, the wearable playback device 602a generates timing adjustment data 716 that is optionally transmitted to the out-loud playback device 602b. This timing adjustment data 716 includes information about the calculated time offset and may specify the adjustments needed to achieve synchronization between the devices. In some implementations, the timing adjustment data 716 may include specific delay values, timing correction factors, or other control parameters necessary for the out-loud playback device 602b to adjust its playback timing. Additionally or alternatively, the timing adjustment data may be calculated on another device (e.g., the out-loud playback device 602b or another network device) which is then transmitted to the wearable playback device 602a which can in turn calculate its own timing offset accordingly.
[0196] The timing adjustment calculation process can incorporate statistical filtering techniques to produce stable synchronization corrections. In some implementations, the system employs Kalman filtering algorithms that combine current timing offset measurements with historical data and system models to estimate optimal timing corrections while accounting for measurement uncertainty and system noise. Additionally or alternatively, the system may utilize other statistical filtering approaches such as exponential smoothing, moving averages, or adaptive filtering techniques that weight recent measurements according to their associated confidence levels.
[0197] The system may implement hysteresis mechanisms to prevent excessive or frequent timing adjustments that could degrade the listening experience. These mechanisms establish perceptually relevant thresholds that determine when timing corrections should be applied, ensuring that adjustments occur only when timing offsets exceed levels that would be noticeable to users. In some implementations, the hysteresis control may require that timingPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO offset measurements consistently exceed a threshold value across multiple measurement cycles before triggering an adjustment, thereby avoiding oscillatory behavior or unnecessary corrections due to temporary measurement variations. The perceptually relevant thresholds may be configured based on psychoacoustic research regarding human sensitivity to timing differences in audio playback, typically in the range of several milliseconds for most listening scenarios.
[0198] Finally, one or both devices can perform timing adjustment operations to synchronize their subsequent audio playback. The wearable playback device 602a performs an adjustment operation 718 to adjust playback timing for subsequent audio playback, while the out-loud playback device 602b performs a corresponding adjustment operation 720 to adjust playback timing for subsequent audio playback. These adjustments may involve modifying delay buffers in the signal processing chains of the respective devices, adjusting clock timing, or implementing other timing correction mechanisms to ensure that audio content from both devices reaches the user's ears with proper temporal alignment.
[0199] The process illustrated in Figure 7 may be performed at various intervals during audio playback to maintain synchronization. Orchestration of this process can be performed by any suitable device in the network, including a home theatre primary device, the out-loud playback device, the wearable playback device, or any other suitable device. In some implementations, the synchronization process occurs at predetermined intervals determined based on detected user movement, detected environmental changes, or a predetermined schedule configured to compensate for timing drift. Additionally or alternatively, the process may be triggered by specific events such as detection of the wearable playback device being donned by a user, detection of user movement using an inertial measurement unit, or detection of an additional wearable audio playback device joining an audio playback session. The process can be designed to operate efficiently on resource-constrained wearable devices while providing accurate synchronization between multiple playback devices in the system.
[0200] Figure 8 illustrates a swim-lane diagram of a synchronization process utilizing a dualmicrophone detection approach within the wearable playback device 602a. The diagram shows the specific interactions between the various hardware components of the wearable playback device 602a and their coordination with the out-loud playback device 602b to achieve synchronized audio playback through detection of self-emitted and externally-emitted ultrasonic synchronization signals.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0201] The wearable playback device 602a comprises three primary components involved in the synchronization process: audio transducers 618, internal microphones 622b. and external microphones 622a. These components work in coordination to generate ultrasonic synchronization signals and detect both self-generated and external ultrasonic signals for timing analysis.
[0202] The audio transducers 618 of the wearable playback device 602a perform a playback operation 802 to play back audio data with ultrasonic signature. During this operation 802, the audio transducers 618 generate both the intended audible audio content and a first ultrasonic synchronization signal based on reference signature components. The ultrasonic signature may comprise frequencies in a near-ultrasonic frequency range (e.g., between approximately 16 kHz and 24 kHz). The ultrasonic signature may take the form of a sequence of tones, a chirp pattern, or other signal patterns configured to remain identifiable in the presence of background noise.
[0203] Simultaneously, the out-loud playback device 602b performs a corresponding playback operation 804 to play back audio data with ultrasonic signature. This operation 804 generates a second ultrasonic synchronization signal that propagates through the acoustic environment toward the wearable playback device 602a. The ultrasonic signatures generated by both devices may be identical or may have known timing relationships that enable comparison for synchronization purposes.
[0204] The internal microphones 622b of the wearable playback device 602a are positioned to primarily capture audio signals generated by the audio transducers 618 of the same device. These internal microphones 622b perform a detection operation 806 to detect ultrasonic signature in captured audio data. During this detection operation 806, the internal microphones 622b capture first acoustic data that includes the first ultrasonic synchronization signal generated by the audio transducers 618. The positioning of the internal microphones 622b allows them to detect the self-generated ultrasonic signal at or near the listening position, providing a reference measurement for the timing of the wearable device's audio output at the user's ear.
[0205] Concurrently, the external microphones 622a of the wearable playback device 602a are positioned to primarily capture audio signals from external sources, including the out-loud playback device 602b. These external microphones 622a perform a detection operation 808 to detect ultrasonic signature in captured audio data. During this detection operation 808, the external microphones 622a capture second acoustic data that includes the second ultrasonicPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO synchronization signal generated by the out-loud playback device 602b. The external microphones 622a may be positioned on an outer surface of the wearable playback device 602a or in other locations that provide acoustic access to the ambient environment.
[0206] The wearable playback device 602a then analyzes the first acoustic data from the internal microphones 622b to determine a first time-of-arrival of the first ultrasonic synchronization signal at the listening position. Similarly, the device analyzes the second acoustic data from the external microphones 622a to determine a second time-of-arrival of the second ultrasonic synchronization signal at the listening position. The analysis may involve identifying watermark patterns within the respective ultrasonic synchronization signals, where the watermark patterns have a know n target time relationship.
[0207] Based on a comparison of the first time-of-arrival and the second time-of-arrival, the wearable playback device 602a determines a time offset between the wearable audio device and the out-loud playback device. This time offset calculation accounts for the different acoustic paths from the internal microphones 622b and external microphones 622a to the estimated listening position. In some implementations, the system applies acoustic path compensation to account for signal propagation delays from the microphones to an estimated ear drum position of the user.
[0208] The system may combine acoustic data from multiple microphones within the wearable playback device 602a to improve accuracy of the time-of-arrival determinations. The analysis is designed to operate with a constrained computational footprint to enable real-time operation on the wearable audio device. In some implementations, the analysis may be performed without using cross-correlation analysis to maintain computational efficiency.
[0209] Following the time offset calculation, the system generates timing adjustment data 810 that may optionally be transmitted from the wearable playback device 602a to the out-loud playback device 602b. This timing adjustment data 810 includes information about the calculated time offset and may specify adjustments needed to achieve synchronization between the devices. The transmission of timing adjustment data 810 enables coordinated timing corrections across both devices in the system.
[0210] Based on the determined time offset, the system performs timing adjustment operations to synchronize subsequent audio playback. The wearable playback device 602a performs an adjustment operation 812 to adjust playback timing for subsequent audio playback. This adjustment may involve modifying a delay buffer in the signal processing chain of the wearablePATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO audio device, adjusting clock timing, or implementing other timing correction mechanisms specific to the wearable device.
[0211] Additionally or alternatively, the out-loud playback device 602b may perform an adjustment operation 814 to adjust playback timing for subsequent audio playback. This adjustment operation 814 may be triggered by receipt of the timing adjustment data 810 from the wearable playback device 602a. The adjustment may similarly involve modifying delay buffers, timing parameters, or other playback characteristics of the out-loud playback device 602b to achieve temporal alignment with the wearable device output.
[0212] The process illustrated in Figure 8 may be performed periodically during audio playback to compensate for timing drift and physical changes in positioning of the wearable audio device. The synchronization process may be triggered by various events including detection of user movement, environmental changes, initiation of an audio playback session, or other conditions that may affect the timing relationship between the devices. The dualmicrophone approach enables precise measurement of signal arrival times at the actual listening position while maintaining efficient operation suitable for resource-constrained wearable devices.
[0213] In some implementations, the wearable playback device may achieve synchronization without generating its own ultrasonic reference signature. Instead of embedding ultrasonic components in the audio output, the system may utilize auxiliary timing data transmitted alongside the audio data stream to provide reference timing information for the wearable device. This auxiliary data may include timestamps, sequence identifiers, or other temporal markers that establish a timing reference relative to the wearable device's local clock. The wearable device determines its internal playback latency through calibration or real-time measurement and uses this information along with the auxiliary timing data to establish a precise temporal reference. The external microphones of the wearable device continue to detect ultrasonic synchronization signals from the out-loud playback device, and the system calculates timing offsets by comparing the detected arrival times with the timing reference derived from the auxiliary data and known internal delays. This approach eliminates the need for ultrasonic signal generation by the wearable device, potentially reducing power consumption and avoiding any ultrasonic content in the user's audio experience.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO b. Methods for Modifying XR Media Content
[0214] Figures 911 are flow diagrams of example methods 900, 1000, 1100 relating to modification of XR media content based on real-world conditions. The methods 900, 1000, 1100 can be implemented by any of the devices or systems described herein (e.g., system 200, XR device 202, device 300, wearable device 602a, etc.), or any other devices or systems now known or later developed. Various examples of the methods 900, 1000, 1100 include one or more operations, functions, or actions illustrated by blocks. Although the blocks are illustrated in sequential order, these blocks may also be performed in parallel, and / or in a different order than the order disclosed and described herein. Also, the various blocks may be combined into fewer blocks, divided into additional blocks, and / or removed based upon a desired implementation.
[0215] In addition, for the methods 800, 900, 1000 and for other processes and methods disclosed herein, the flowcharts show functionality and operation of possible implementations of some examples. In this regard, each block may represent a module, a segment, or a portion of program code, which includes one or more instructions executable by one or more processors for implementing specific logical functions or steps in the process. The program code may be stored on any type of computer readable medium, for example, such as a storage device including a disk or hard drive. The computer readable medium may include non-transitory computer readable media, for example, such as tangible, non-transitory computer-readable media that stores data for short periods of time like register memory, processor cache, and Random-Access Memory (RAM). The computer readable medium may also include non-transitory media, such as secondary' or persistent long-term storage, like read only memory' (ROM), optical or magnetic disks, compact disc read only memory (CD-ROM), for example. The computer readable media may also be any other volatile or non-volatile storage systems. The computer readable medium may be considered a computer readable storage medium, for example, or a tangible storage device. In addition, for the methods and for other processes and methods disclosed herein, each block in Figures 9-11 may represent circuitry that is w ired to perform the specific logical functions in the process. In some examples, the methods 900, 1000, 1100 are not encoded deterministically in software or code, but instead performed spontaneously or extemporaneously' by one or more generative media modules and / or one or more Al agents, such as those described above with respect to Figure 2.
[0216] Figure 9 illustrates a flowchart of an example method 900 for adapting virtual content based on real-world acoustic characteristics. At block 902, the method includes obtainingPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO acoustic characteristics of a real-world environment. This may involve analyzing sound reflections, reverberation time, and frequency response characteristics of the space. In some implementations, the analysis may utilize microphones and sensors to capture room impulse responses, measure background noise levels, and determine spatial distribution of acoustic properties. The analysis may also evaluate audio equipment specifications to determine frequency response limitations and spatial audio reproduction capabilities.
[0217] At block 904, the method includes receiving virtual environment data defining acoustic properties of a virtual environment. This may include three-dimensional models, textures, audio content, and associated metadata defining the acoustic behavior of virtual objects and spaces. In some implementations, the virtual environment data may include specifications for virtual sound sources, their positions, and intended acoustic interactions.
[0218] At block 906, the method includes modifying the virtual environment data based on the acoustic characteristics of the real-world environment. This may involve adjusting sizes, shapes, or materials of virtual objects to create sound reflections and reverberations that match the real-world capabilities. The modifications may include repositioning virtual sound sources based on playback device locations and implementing psychoacoustic techniques to compensate for hardware limitations. In some implementations, the method may incorporate biometric data from user sensors to further optimize the modifications. The system may also continuously monitor and dynamically update modifications based on real-time changes in the environment.
[0219] At block 908, the method includes providing the modified virtual environment to an extended reality (XR) display device for playback. This may involve coordinating visual and audio modifications to maintain consistency while accounting for environmental limitations. The method may utilize head-related transfer function (HRTF) data to personalize spatial audio rendering and may adjust emotional characteristics of the content while working within acoustic constraints.
[0220] Figure 10 illustrates a flowchart of an example method 1000 for providing shared XR experiences across different environments. At block 1002, the method includes obtaining first acoustic data from a first XR device in a first real-world environment and second acoustic data from a second XR device in a second real-world environment. This may involve collecting comprehensive acoustic measurements and equipment specifications from each location.
[0221] At block 1004, the method includes analyzing the first and second acoustic data to determine respective acoustic capabilities. This analysis may evaluate the specific limitationsPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO and opportunities presented by each environment's acoustic properties and available playback hardware. The analysis may consider factors such as room acoustics, equipment specifications, and potential interference sources.
[0222] At block 1006, the method includes generating first modified virtual environment data for the first XR device and second modified virtual environment data for the second XR device. These modifications may be customized for each environment's capabilities while maintaining synchronized timing and spatial relationships. The method may create different audio mixes, adjust virtual room properties, and implement environment-specific optimizations while preserving the coherence of the shared experience. In some implementations, the system may monitor network conditions and adjust modifications to account for latency between environments.
[0223] At block 1008, the method includes providing the first and second modified virtual environment data to the respective XR devices. This may involve coordinating delivery of the modified content while maintaining temporal and spatial synchronization between environments. The method may include mechanisms for users to understand aspects of each other's environmental constraints while enjoying their individually optimized experiences.
[0224] Figure 11 illustrates a flow chart of an example method 1100 for utilizing blockchain-stored environmental data in XR experiences. At block 1102, the method includes receiving a request to generate a virtual environment for an XR device in a physical space. This request may include information about the intended experience and any specific requirements or preferences.
[0225] At block 1104, the method includes retrieving, via a blockchain (or another suitable distributed ledger), stored real-world environment data characterizing acoustic properties of the physical space. This data may be stored as non-fungible tokens (NFTs) containing acoustic signatures, user preferences, and historical measurements. The blockchain may implement smart contracts governing access rights and usage terms for the stored data.
[0226] At block 1106, the method includes generating customized virtual environment data based on the retrieved real-world environment data. This may involve using machine learning to analyze historical interaction patterns and create personalized modifications. The method may incorporate biometric data and emotional characteristics stored on the blockchain to optimize the experience. In some implementations, the system may use collaborative audio experiences stored as NFTs to inform the customization process.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0227] At block 1108. the method includes providing the customized virtual environment data to the XR device for playback. The method may utilize smart contracts to manage content licensing and ensure appropriate use of stored acoustic profiles and modifications. The blockchain may maintain audit trails of how environmental data and modifications are used while protecting user privacy through appropriate access controls.c. Vehicle-Based XR Implementations
[0228] In some examples, the XR system can be integrated within vehicle environments. A vehicle, such as an autonomous car, may sen e as a dynamic extended reality environment where audio playback characteristics adapt based on vehicle state and motion parameters. The vehicle's cabin can function as an acoustically controlled space for experiencing XR content.
[0229] In some implementations, one or more sensors in the vehicle monitor various vehicle parameters including, but not limited to, vehicle speed, acceleration, road conditions, cabin noise levels, and vehicle operation mode. These parameters can be provided to the content generation components to dynamically adjust the virtual environment characteristics. For example, in some implementations, the system may adjust audio spatialization based on vehicle speed. At higher speeds, where cabin noise may increase, the system can modify frequency response curves and spatial audio rendering to maintain optimal audio perception. The content generation components may implement active noise cancellation in coordination with the virtual audio content to preserve the intended acoustic experience.
[0230] Additionally or alternatively, the system may modify the virtual environment scale or acoustic properties based on vehicle motion. For example, when the vehicle is stationary, the system may present a larger virtual concert hall environment. As vehicle speed increases, the system may gradually adjust the virtual environment to a more intimate acoustic space that better matches the vehicle's noise profile and acoustic limitations at speed.
[0231] In some implementations, the system can detect when the vehicle enters autonomous operation mode and automatically reconfigure the cabin for optimal extended reality experiences. This may include adjusting seat positions, deploying or retracting acoustic treatment elements, and modifying climate control settings to create ideal acoustic conditions.
[0232] The vehicle's audio system may include multiple playback devices positioned throughout the cabin. In some implementations, these playback devices can be dynamically reconfigured based on the virtual content being presented. For example, certain speakers may be reassigned from typical vehicle audio functions (like navigation prompts) to extended reality content playback when an XR experience is active.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO d. Digital Twinning
[0233] In some implementations, the system creates and maintains digital twins of physical spaces where extended reality experiences take place. A digital twin can comprise detailed virtual replica of a real-world environment, including precise acoustic characteristics, object positions, and material properties. The digital twin can serve as a reference model for adapting virtual content to match real-world constraints.
[0234] The digital twin creation process may include multiple stages of environmental analysis. Initially, the system may perform detailed acoustic measurements including, but not limited to, room impulse responses, frequency response analysis, and spatial acoustic mapping. Additionally or alternatively, the system may utilize computer vision and depth sensing to capture physical dimensions, object positions, and surface characteristics of the space.
[0235] In some implementations, the digital twin maintains a bi-directional relationship with the physical space through continuous monitoring and updates. Environmental analysis components may detect changes in the physical environment, such as furniture movement, door / window states, or occupancy changes, and automatically update the digital twin accordingly. This synchronization can occur in real-time or at predetermined intervals depending on system requirements.
[0236] The digital twin may include multiple lay ers of data representation. A geometric layer can capture physical dimensions and object positions. An acoustic layer may model sound propagation characteristics and frequency response properties. A material properties layer might define surface reflectivity, absorption coefficients, and other acoustic characteristics. Additional layers may represent dynamic elements like movable objects, variable acoustic treatments, or temporary modifications to the space.
[0237] In some implementations, the system stores digital twin data using distributed ledger technologies. Smart contracts may govern how digital twin data is updated and validated. For example, a smart contract might require consensus from multiple environmental sensors before accepting a significant change to the acoustic model. The blockchain can maintain a verifiable history of environmental changes, enabling analysis of acoustic variations over time.
[0238] The system may implement different synchronization strategies for various aspects of the digital twin. Time-critical parameters that directly affect user experience may be updated in real-time, while less critical data might be updated asynchronously to optimize system performance. In some implementations, the system may predict and pre-compute likely environmental changes to reduce latency in digital twin updates.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO e. Dynamic Playback Device Bonding
[0239] In some implementations, the system can dynamically modify audio playback device bonding configurations when transitioning between standard audio playback and extended reality experiences. Playback device bonding refers to the logical grouping of multiple playback devices to create specific audio effects or channel configurations. The system may break or reconfigure these bonds to optimize spatial audio reproduction for XR content. Additional details regarding dynamic bonding of playback devices can be found in U.S. Patent No. 9,864,571, issued January 9, 2018, titled “Dynamic Bonding of Playback Devices,” which is hereby incorporated by reference in its entirety for all purposes.
[0240] For example, two playback devices originally bonded as a stereo pair for music playback may be unbonded and individually controlled when presenting XR content. This allows for more precise positioning of virtual sound sources and better adaptation to the specific acoustic requirements of the virtual environment.
[0241] The system may implement a state management architecture to handle playback device bonding transitions. This architecture can track current bonding configurations, maintain speaker capabilities and characteristics, and manage the transition process between different bonding states. In some implementations, the state management system may utilize distributed ledger technologies to maintain consistent playback device configurations across multiple playback devices and control points.
[0242] When transitioning between bonded and unbonded states, the system may implement smooth crossfading or other audio blending techniques to prevent abrupt changes in the listening experience. In some implementations, the system may temporarily maintain both bonded and unbonded signal paths during transitions, gradually shifting the audio balance to the new configuration.
[0243] The system may store different bonding profiles for various types of content or usage scenarios. These profiles can define optimal playback device configurations for different virtual environments or acoustic requirements. Smart contracts may govern how bonding profiles are selected and applied based on content type, user preferences, and environmental conditions.
[0244] In some implementations, the system dynamically optimizes playback device bonding configurations based on real-time analysis of the virtual content and physical environment. For example, if the virtual environment includes multiple distinct sound sources, the system may unbond playback devices to provide more granular spatial control. Conversely,PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO if the virtual content requires specific channel relationships, the system may create temporary bonds to maintain these relationships.f. Asymmetric Experience Management
[0245] In some implementations, the system supports asymmetric shared experiences where users in different physical environments receive modified versions of the same virtual content optimized for their specific acoustic capabilities while maintaining a coherent shared experience. This asymmetric adaptation goes beyond basic audio quality adjustments to potentially present significantly different but conceptually aligned experiences.
[0246] For example, in a shared virtual concert experience, a user in a space with full-range audio capabilities might experience the event in a large virtual concert hall with deep bass response and complex reverberation. Meanwhile, a user in a space with limited audio capabilities might experience the same concert in a smaller virtual venue with modified acoustics, but the system maintains synchronization of key musical elements and social interactions between users. In some examples, a user wearing a device with headphones or in-ear earbuds may choose among multiple available experiences.
[0247] The system may implement a layered content modification architecture to manage asymmetric experiences. A core experience layer defines fundamental elements that must remain consistent across all users, such as temporal synchronization of events or relative positions of key virtual objects. Additional layers may define experience elements that can be modified more freely, such as room acoustics, ambient effects, or secondary sound sources.
[0248] In some implementations, the system includes coordination mechanisms to maintain narrative and experiential consistency despite technical disparities. These mechanisms may analyze the intended emotional impact or key experience points of virtual content and ensure these elements are preserved through different technical implementations appropriate for each user's environment.
[0249] The system may employ machine learning algorithms to develop and refine mappings between different experience versions. These algorithms can analyze relationships between high-capability and limited-capability implementations of virtual experiences to optimize how content is adapted while maintaining essential experiential elements.
[0250] In some implementations, the system provides awareness mechanisms that help users understand and adapt to asymmetric capabilities. For example, visual cues might indicate when another user's environment enables different acoustic features, or the system might adjust social interaction mechanics to account for different audio capabilities.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0251] Additional details regarding mirrored queues, shared playback sessions, shared moodscapes, and / or other shared experiences or contents can be found in the following patents and applications, each of which is hereby incorporated by reference in its entirety for all purposes: (1) U.S. Patent No. 10,587,693, issued March 10, 2020, titled “Mirrored Queues””; (2) U.S. Patent No. 11,204,737, issued May 13, 2021, titled “Playback Queues for Shared Experiences,” and (3) International Patent Application No. PCT / US2024 / 026459, filed April 26, 2024, titled “Providing Moodscapes and Other Media Experiences.”g. Real-World Characteristics Tokenization
[0252] In some implementations, the system implements comprehensive tokenization of real-world characteristics, enabling precise representation and control of physical system capabilities through blockchain mechanisms. This tokenization extends beyond basic asset management to create programmable representations of physical system capabilities and states.
[0253] The system may create different classes of tokens representing various aspects of real-world audio systems. Capability tokens might represent specific technical abilities of playback devices, such as frequency response ranges, power handling, or spatial audio processing features. State tokens could represent current operational parameters like volume levels, equalization settings, or bonding configurations. Configuration tokens might define relationships between multiple devices or acoustic treatment elements.
[0254] In some implementations, smart contracts govern how tokenized characteristics can be modified or combined. For example, a smart contract might enforce rules about valid speaker configurations based on the capabilities represented by different tokens. Another smart contract might manage how7environmental tokens can be updated based on sensor data from environmental analysis components.
[0255] The system may implement a token hierarchy that reflects relationships between different physical system characteristics. Parent tokens might represent overall system capabilities, while child tokens represent specific feature implementations. This hierarchy can help manage complex systems with interdependent characteristics and ensure consistent system behavior.
[0256] In some implementations, the system includes mechanisms for validating and verifying tokenized characteristics. These mechanisms might require consensus from multiple sensors or verification through acoustic measurements before updating token states. The system may maintain an auditable history of changes to tokenized characteristics, enabling analysis of system modifications over time.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0257] The tokenization architecture may include interfaces that enable virtual reality engines and other system components to query and understand available capabilities. These interfaces can provide standardized methods for discovering system features, checking operational states, and requesting modifications to system configurations.
[0258] Smart contracts may implement access control and modification rights for tokenized characteristics. Different system components or user roles might have varying levels of authority to read or modify specific tokens. The system may include delegation mechanisms that allow temporary transfer of control rights under specific conditions.
[0259] In some implementations, the system can generate derived tokens based on combinations or analysis of primary tokenized characteristics. For example, the system might create acoustic performance tokens that represent the calculated capabilities of multiple devices working together. These derived tokens can simplify system management while maintaining precise representation of capabilities.
[0260] In addition to the benefits described above, the tokenization of real-world characteristics can also be used to verify data provenance as originating from human source(s) / activity as opposed to synthetic data generated de novo via a generative media module. There are situations in which the disambiguation of human sourced data versus synthetic data may be desirable, or even critical. Consider, for instance, so-called “deep fakes” in which Al-generated content (e.g., images, videos, audio) purports to represent a person in a particular context, or doing or saying something that, in reality, never happened. In another example, data generated by a person (or people) based on particular circumstances, activities, contexts, etc. can be monetized via a data marketplace. However, potential data buyers on the marketplace may prefer to confirm that the data being purchased resulted from human activity'.
[0261] The tokenization of real-world characteristics can address these problems by providing a token tied to a physical, real-world presence or interaction, akin to “a physical fingerprint.” In some examples, the combination of tokens may provide a verification of human activity and / or presence. For instance, a validation approach that combines real-world characteristic tokens (e.g., a parent token, one or more child tokens, and / or a mix thereof) in combination with a token associated with another physical object (e.g., a tag such as an NFC tag(s) embedded in a physical object in the room and / or on the user’s person) can provide an interested entity7proof that a particular data source was the result of human activity and / or created as the result of physical, real-world interactions, rather than synthetically generated.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO h. Character-Based HRTF Perspective Simulation
[0262] In some implementations, the system provides character perspective simulation by generating and applying character-specific Head-Related Transfer Functions (HRTFs) that can be mapped to a user's individual HRTF characteristics. This approach enables users to experience spatial audio from the perceptual perspective of different virtual characters or entities within the virtual environment.
[0263] The system may generate character HRTFs based on various anatomical and physical characteristics of virtual characters. For example, in some implementations, the system analyzes three-dimensional models of character head shapes, ear geometries, and torso characteristics to compute acoustic transfer functions that represent how the character would perceive spatial audio. This analysis may account for factors such as the size and shape of the character's head, the position and geometry of their ears, and characteristics of their ear canals.
[0264] In some implementations, the system employs computational acoustic modeling to simulate sound wave interactions with the character's anatomical features. This modeling may utilize techniques such as boundary element methods (BEM), finite element analysis (FEA), or other suitable computational approaches to calculate how sound waves would be modified by the character's physical features before reaching their auditory systems. The resulting transfer functions capture the character-specific acoustic transformations that would occur in different spatial directions.
[0265] The system may maintain a database of character HRTF profiles stored on distributed ledger systems. In some implementations, these profiles may be represented as non-fungible tokens (NFTs) that contain both the HRTF data and metadata describing the character's relevant physical characteristics. Smart contracts may govern how these HRTF profiles can be accessed and modified.
[0266] To enable perceptual translation between character and user perspectives, the system implements HRTF mapping techniques. In some implementations, the system first analyzes the user's personal HRTF characteristics, which may be obtained through acoustic measurements, photographic analysis, previously stored profiles, machine learning classification techniques, etc. In some examples, the user’s HRTF may comprise a composite or “average” HRTF based on user biological characteristics (e.g., height, weight, gender, etc.). In some instances, the biological characteristics may be obtained via a measurement scan from a user device (e.g.. the XR device 202 of Figure 2, a tablet or smartphone such as the control device 130a of Figure 1H, and / or another suitable device). In certain examples, thePATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO measurement characteristics may be obtained via a holographic projection scanning process performed by the XR device 202 or another suitable device. In some instances, system may obtain user biological characteristics via a third-party entity such as, for example, an online retailer that obtains a measurement scan from a user for determining clothing size and fit. As those of ordinary' skill in the art will appreciate, scan data can be used to obtain subject anthropomorphic dimensions (e.g., head size and shape, pinna size and shape, torso dimensions, shoulder and chest width) useful in determining HRTFs. Moreover, in scenarios with multiple users, listeners, participants etc. in the same location (i.e., the same room or adjacent acoustic spaces), the system may generate a composite HRTF of the particular users using one or more techniques described above, or perhaps based on a general population HRTF and / or psychoacoustic model.
[0267] The system then develops a transformation function that maps between the user's HRTF and the character's HRTF while preserving key spatial cues. The mapping process may account for different types of spatial audio cues. The system can preserve interaural time differences between ears while accounting for different head sizes. Additionally, the system may adapt amplitude variations based on relative head shadow effects. The process can transform frequency-dependent modifications caused by outer ear geometries, while also accounting for different shoulder reflection and shadowing characteristics.
[0268] In some implementations, the system employs machine learning techniques to optimize the HRTF mapping process. These techniques may analyze relationships between different HRTF characteristics and develop models for translating spatial cues between different anatomical configurations while maintaining perceptual coherence.
[0269] The system may implement different mapping strategies depending on the degree of anatomical difference between the user and character. For characters with roughly humanoid proportions, the system may use direct feature mapping approaches. For characters with radically different anatomical configurations, the system may employ more abstract mappings that preserve key spatial relationships while adapting to human perceptual capabilities.
[0270] In some implementations, the system includes transition management mechanisms for switching between different HRTF perspectives. These mechanisms may implement smooth interpolation between HRTF characteristics when changing character perspectives to avoid abrupt perceptual shifts. The transition process may consider gradual morphing between HRTF frequency responses while maintaining spatial coherence during transitions. The systemPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO works to preserve ongoing audio events and adapt dynamic audio characteristics throughout the transition.
[0271] The system may provide visualization overlays that help users understand the character's acoustic perspective. In some implementations, these visualizations may highlight acoustic features such as areas of enhanced or diminished sensitivity, directional awareness capabilities, or frequency-specific spatial perceptions that differ from human norms.
[0272] Real-time adaptation capabilities may be implemented to account for character movement and interactions. In some implementations, the system dynamically updates HRTF processing based on character head orientation and movement. The system also considers environmental interactions that may temporarily modify acoustic properties, as well as changes in character state or form that affect acoustic perception. The processing additionally accounts for relative positions of sound sources and other characters.
[0273] The system may implement biometric monitoring to optimize the character perspective simulation. In some implementations, sensors track user responses such as head movement patterns, stress indicators, or other physiological markers. This data may be used to refine HRTF mappings and transition behaviors to improve comfort and perceptual effectiveness.
[0274] In some implementations, the system supports shared experiences where multiple users can experience the same character perspectives while accounting for their individual HRTF characteristics. The system maintains synchronized character state tracking across users while performing individual HRTF mapping optimizations for each participant. Temporal and spatial synchronization is preserved throughout the experience, and users maintain shared awareness of perspective-specific acoustic features.
[0275] The system may include calibration processes to optimize character perspective simulation for individual users. The system can perform initial HRTF measurement or profile selection, followed by refinement of mapping parameters through user feedback. Transition timing and interpolation characteristics may be adjusted based on user response. The system can additionally optimize visualization and feedback mechanisms to improve the user experience.
[0276] In some implementations, the system maintains usage history and preference data for different character perspectives via distributed ledgers. This data may inform future optimizations and enable personalized adjustments to HRTF mapping strategies. SmartPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO contracts may govern how this historical data is stored, accessed, and applied to subsequent experiences.
[0277] The system may implement error detection and correction mechanisms to maintain perceptual quality. The system monitors for spatial audio artifacts or discontinuities while checking for mapping inconsistencies between user and character perspectives. Transition anomalies or uncomfortable perceptual effects can be detected and addressed. The system additionally monitors for synchronization issues in multi-user scenarios.
[0278] The system may provide developer tools for creating and testing character HRTF profdes. These tools enable HRTF measurement and modeling while supporting mapping strategy- development and testing. Developers can configure transition behaviors and optimize multi-user experiences. The tools also provide performance monitoring and analysis capabilities to ensure high-quality implementations.i. Virtual Concert Hall
[0279] In some implementations, the system may provide a virtual concert hall experience built on distributed ledger technology. The virtual concert hall itself may be represented as a non-fungible token containing highly accurate impulse response data and acoustic signatures crafted by acoustic engineers. These acoustic signatures may capture the unique reverberations, reflections, and frequency responses that characterize the virtual venue's sound.
[0280] The system may enable performers to create performance tokens containing both musical content and performance data. This performance data may include information about musician positions, instrument characteristics, and intended acoustic behaviors. The system may analyze this data in conjunction with the venue's acoustic signature to generate appropriate modifications for different playback environments.
[0281] When users attend concerts in the virtual venue, the system may analyze their local acoustic environments and playback capabilities. The content generation components may then adjust the virtual venue's characteristics to maintain the intended acoustic experience within the constraints of each user's space. For example, in environments with limited low-frequency reproduction capabilities, the system may employ psychoacoustic techniques to preserve the sense of scale and power in orchestral performances.
[0282] Smart contracts implemented on the distributed ledger may manage various aspects of the concert experience. These contracts may handle ticket sales and access rights, automatically distribute royalties to performers, and govern how acoustic data can be used across different implementations. The contracts may additionally manage synchronizationPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO between multiple attendees to maintain coherent shared experiences despite varying local acoustic conditions.j. Interactive Audio Museum
[0283] In some implementations, the system may enable an interactive audio museum experience where historical artifacts are represented through both visual and acoustic elements. Each artifact may be associated with a token containing visual models, acoustic signatures, and various audio recordings related to the artifact. These recordings may include oral histories, reconstructed operation sounds, and expert commentary about the artifact's significance.
[0284] The system may analyze the physical space where users experience the virtual museum and modify acoustic presentations accordingly. For environments with different reverberation characteristics than traditional museum spaces, the content generation components may adjust the acoustic properties of exhibition spaces while maintaining the clarity and intelligibility of audio content. The system may identify opportunities to utilize auxiliary playback zones to enhance the sense of space and movement through the virtual galleries.
[0285] In implementations supporting multiple simultaneous visitors, the system may coordinate acoustic experiences between users while accounting for their relative positions and local acoustic conditions. As users approach exhibits, the system may generate spatialized audio that creates appropriate acoustic transitions and maintains natural sound localization. The system may additionally modify crowd noise simulation and ambient effects to maintain appropriate acoustic energy while avoiding interference with exhibit-specific audio content.
[0286] Museum curators or historians may create guided tour experiences represented as tokens on the distributed ledger. These tours may include spatial audio narratives, ambient soundscapes, and interactive acoustic elements. The system may analyze each user's environment and preferences to deliver personalized versions of these tours while maintaining the intended educational and emotional impact of the content. Smart contracts may manage access rights and ensure appropriate compensation for content creators.k. Decentralized Sound Effects Marketplace
[0287] In some implementations, the system may provide a decentralized marketplace for spatial audio assets, sound effects, and / or real-world characteristics. Each audio asset may be represented as a token containing the sound data, acoustic behavior specifications, metadata about its origin and intended use, and licensing terms. In the case of real-world characteristics, an entity associated with a particular location, venue, arena, facility, or otherwise may tokenizePATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO associated characteristics or sound effects (e.g., room acoustic characteristics based on a static configuration or dynamic configuration, historical or real-time sound effects) For instance, the marketplace may enable a user in the virtual concert hall example above to obtain a tokenized set of characteristics of their favorite venue and not only recreate its acoustics (and other characteristics) virtually, but, in some cases, parameters that the real-world audio system can used in realistically simulating the venue. The system may analyze these assets to determine how they can be optimized for different acoustic environments and playback configurations.
[0288] In some implementations, the system may dynamically generate and tokenize new audio assets during extended real ity sessions for potential marketplace distribution. When users engage in generative XR experiences, the system may identify novel audio elements, modifications, or combinations that emerge during the session and automatically capture these as discrete audio assets. For example, when anotable user such as a celebrity or content creator experiences a unique generative audio sequence, the system may record and tokenize that experience, including relevant contextual metadata, performance data, and environmental parameters that contributed to its creation. These newly created assets may then be made available through decentralized marketplaces where other users can discover and incorporate them into their own XR experiences. Smart contracts may govern how these dynamically generated assets can be accessed, modified, and reused while maintaining appropriate attribution and compensation mechanisms between the original experience creator and subsequent users. The system may additionally record relationships between derived experiences that incorporate these assets, enabling tracking of creative lineage and value distribution across multiple generations of content creation.
[0289] The content generation components may automatically adjust purchased audio assets based on the characteristics of a user's playback environment. For example, when a sound effect designed for a full-range speaker system is used in an environment with limited frequency response, the system may generate modified versions that preserve the essential character and emotional impact of the sound while working within available capabilities. The system may store these modifications as linked tokens that maintain relationships with the original assets.
[0290] Smart contracts may govern the licensing and usage of audio assets, automatically handling royalty payments and ensuring compliance with usage terms. These contracts may include provisions for how assets can be modified or incorporated into derivative works. The system may implement transaction recording and usage tracking through the distributed ledger to maintain transparent accounting of asset utilization.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0291] The marketplace may incorporate decentralized governance mechanisms that enable community curation of audio assets. Participants may evaluate and categorize assets, with these assessments recorded on the distributed ledger. The system may use this community input along with acoustic analysis to help users identify assets best suited to their specific implementation requirements and playback capabilities.l. Personalized Audio Training Simulator
[0292] In some implementations, the system may provide training simulations for high-stress occupations where acoustic awareness and response are essential to performance. Each training scenario may be represented as a token containing environmental models, interactive elements, and dynamic soundscape specifications. The system may analyze the acoustic characteristics of training spaces to optimize scenario presentation for different implementation environments.
[0293] The content generation components may adjust scenario acoustics based on both environmental capabilities and trainee performance metrics. In implementations where the training space has limited low-frequency reproduction capabilities, the system may employ frequency shifting and psychoacoustic techniques to maintain the impact of scenario events. For example, when simulating structural collapse in firefighter training, the system may generate modified sound profiles that create appropriate psychological pressure while working within available acoustic capabilities.
[0294] The system may incorporate biometric monitoring to enable dynamic scenario adaptation. Sensor data including heart rate, respiration, and other physiological markers may inform real-time modifications to acoustic intensity and complexify. In some implementations, the system may adjust ambient sound levels, introduce or modify distraction elements, and alter acoustic spatialization to create appropriate stress conditions.
[0295] Performance data and scenario modifications may be recorded on distributed ledgers to enable progress tracking and scenario refinement. Smart contracts may govern how scenarios adapt to different skill levels and control access to progressive training modules. The system may analyze accumulated performance data to identify' acoustic patterns that correlate with successful outcomes and adjust scenario presentations accordingly.m. Immersive Audiobook Experience
[0296] In some implementations, the system may provide enhanced audiobook experiences combining narrative audio with responsive virtual environments. Each audiobook implementation may be represented as a token containing narration, sound effects,PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO environmental specifications, and interaction parameters. The system may analyze playback environments to determine appropriate acoustic modifications while maintaining narrative clarity and emotional impact.
[0297] The content generation components may generate dynamic soundscapes that adapt to story progression and user movement through virtual spaces. When users navigate virtual environments corresponding to narrative locations, the system may adjust acoustic properties to match architectural and material characteristics described in the story. The system may identify opportunities to utilize auxiliary playback zones to enhance environmental immersion while ensuring narrative audio remains clearly intelligible.
[0298] In implementations supporting interactive storytelling, the system may modify both narrative and environmental audio based on user choices. The content generation components may analyze story branch characteristics and generate appropriate acoustic transitions that maintain narrative continuity. Smart contracts may govern how interactive elements can modify core narrative content while preserving artistic integrity and appropriate attribution.
[0299] The system may incorporate user preference data and behavioral analysis to personalize acoustic presentations. Reading history, genre preferences, and interaction patterns stored on distributed ledgers may inform how the system balances narrative and environmental audio elements. The system may adjust factors such as voice processing, ambient sound levels, and spatialization based on individual user preferences while maintaining core narrative experiences.n. Virtual Reality Meditation Retreat
[0300] In some implementations, the sy stem may provide meditation experiences with adaptive acoustic environments. The content generation components may generate calming soundscapes comprising synthesized and recorded natural sounds, adjusting acoustic characteristics based on environmental capabilities and user biometric responses. The system may analyze meditation spaces to identify acoustic properties that could enhance or detract from relaxation goals.
[0301] Biometric data including heart rate, breathing patterns, and movement may inform real-time adjustments to soundscape elements. The system may modify tempo, frequency content, and spatial distribution of acoustic elements to encourage physiological calming responses. In implementations where environmental noise could impact meditation effectiveness, the system may adjust acoustic masking techniques while maintaining peaceful ambient characteristics.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0302] The system may enable group meditation experiences with acoustics optimized for multiple participants. Content generation components may analyze relative positions and biometric data from multiple users to generate shared soundscapes that support group coherence while adapting to individual needs. The system may adjust spatialization and acoustic energy distribution to maintain appropriate social awareness without disrupting individual practice.
[0303] User preferences and meditation histories may be recorded on distributed ledgers to enable personalized experience optimization. Smart contracts may govern how different acoustic elements can be combined and modified while maintaining therapeutic effectiveness. The system may analyze accumulated meditation session data to identify acoustic patterns associated with successful outcomes and refine soundscape generation accordingly.o. Adaptive Audio for Fitness
[0304] In some implementations, the system may provide workout experiences with responsive audio environments. The content generation components may generate music and sound effects that synchronize with exercise activities while adapting to user performance and environmental capabilities. The system may analyze workout spaces to optimize acoustic energy distribution and avoid potential interference patterns.
[0305] Biometric monitoring may enable real-time adaptation of audio characteristics including tempo, intensity, and motivational elements. The system may adjust these parameters based on factors such as heart rate, movement patterns, and detected exertion levels. In implementations where environmental acoustics could impact exercise effectiveness, the system may modify audio spatialization and frequency content to maintain motivational impact while avoiding fatigue.
[0306] The system may support group workout scenarios where multiple users receive synchronized but individually optimized audio. Content generation components may analyze relative positions and performance metrics of multiple participants to generate coherent shared experiences while adapting to individual capabilities and preferences. The system may adjust acoustic elements to maintain group energy while providing personalized motivation.
[0307] Workout preferences, performance histories, and biometric response patterns may be recorded on distributed ledgers to enable experience optimization. Smart contracts may govern how' different audio elements can be combined and modified while maintaining exercise effectiveness. The system may analyze accumulated workout data to identify acoustic patterns associated with optimal performance and adjust generation parameters accordingly.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0308] Additional details regarding the use of a wearable (or non-wearable) device to assess user activity, mood, or other parameters can be found in International Patent Application No. PCT / US2021 / 071260, filed August 24, 2021, titled '‘Mood Detection and / or Influence Via Audio Playback Devices,” which is hereby incorporated by reference in its entirety for all purposes.p. Collaborative Content Creation
[0309] In some implementations, the system may enable group content (e.g., music, video, text, multimedia content, etc.) creation in virtual studio environments. The content generation components may analyze acoustic characteristics of each participant's physical space to optimize monitoring and playback capabilities. The system may adjust individual mix elements to compensate for frequency response limitations while maintaining consistent artistic intent across different playback environments.
[0310] The system may implement real-time synchronization of musical elements contributed by different participants. Content generation components may analyze network conditions and environmental acoustics to maintain temporal alignment while accounting for latency. In implementations where participants have significantly different playback capabilities, the system may generate optimized monitor mixes that preserve creative decisionmaking ability despite hardware variations.
[0311] Smart contracts may manage rights and contributions for collaboratively created content. The system may record individual contributions, modifications, and creative decisions on distributed ledgers to enable appropriate attribution and compensation. Generated content may be represented as tokens containing both the final compositions and detailed contribution metadata.
[0312] The system may support asymmetric collaboration scenarios where participants have varying roles and acoustic requirements. Content generation components may optimize monitoring for different instrumental and vocal contributions while maintaining coherent shared experiences. The system may adjust acoustic properties of virtual studio spaces to provide appropriate creative environments while working within physical space limitations.
[0313] In some implementations, the system may enable collaborative creation across various domains by providing shared virtual workspaces adapted to different types of creative and professional activities. While the system may support collaborative music creation in virtual recording studios, it may additionally facilitate other forms of artistic and professional collaboration such as virtual film production, visual art creation, literary composition, orPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO engineering design. The content modification components may analyze the acoustic requirements of different creative activities and adjust virtual workspace characteristics accordingly. For example, in a virtual film editing suite, the system may optimize audio monitoring for dialogue clarity and sound effect evaluation, while in a virtual engineering workspace, the system may enhance acoustic feedback for mechanical simulations and design reviews. For visual art creation, the system may generate appropriate ambient soundscapes that enhance creative focus while enabling clear communication between collaborators. The system may dynamically adjust these workspace characteristics based on the specific tools and activities being used, such as providing specialized acoustic environments for reviewing architectural acoustics in building designs or evaluating sound design in game development. Smart contracts may manage rights and contributions across different types of creative output, maintaining appropriate attribution and compensation mechanisms regardless of the creative domain. The system may record collaborative sessions and generated content on distributed ledgers, enabling transparent tracking of contributions and modifications across complex creative projects involving multiple participants and disciplines.q. User Museum Tour
[0314] In some implementations, the system may enable shared museum experiences with coordinated spatial audio. The content generation components may analyze acoustic characteristics of each participant's environment to optimize exhibit audio presentation. The system may adjust spatial audio rendering to maintain appropriate acoustic relationships between participants while accommodating different playback capabilities.
[0315] When participants approach virtual exhibits, the system may generate coordinated audio experiences that maintain spatial and temporal coherence across different physical spaces. The content generation components may adjust acoustic properties of virtual gallery spaces to provide appropriate environmental context while ensuring exhibit audio remains clearly intelligible. In implementations where participants have varying movement capabilities, the system may modify acoustic transitions to maintain synchronized experiences despite different navigation patterns.
[0316] The system may implement dynamic soundscape generation responsive to participant interactions and positions. Content generation components may analyze group dynamics and acoustic conditions to generate appropriate ambient effects and crowd simulation. The system may adjust acoustic energy distribution to maintain social presence while avoiding interference with exhibit-specific audio content.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0317] Smart contracts may manage access rights and guide content presentation for different types of tours and visitor groups. The system may record tour participation data and interaction patterns on distributed ledgers to enable experience optimization. Generated soundscapes and acoustic modifications may be represented as tokens that maintain relationships with original exhibit content while enabling personalized presentation.r. Interactive Storytelling
[0318] In some implementations, the system may enable shared storytelling experiences with dynamic acoustic environments. The content generation components may analyze participant locations and environmental capabilities to generate appropriate soundscapes for narrative enhancement. The system may adjust acoustic properties of virtual gathering spaces to support story progression while maintaining social presence.
[0319] When participants contribute to evolving narratives, the system may generate acoustic transitions and environmental effects that maintain story coherence while adapting to different playback capabilities. The content generation components may analyze narrative emotional content and participant responses to adjust acoustic intensity and complexity appropriately. In implementations where participants have varying acoustic environments, the system may modify spatial audio rendering to preserve key narrative elements while optimizing for individual conditions.
[0320] The system may implement real-time generation of sound effects and ambient elements responsive to story development. Content generation components may analyze narrative patterns and participant engagement to produce appropriate acoustic enhancement. The system may adjust generated audio elements to maintain dramatic impact while working within environmental limitations.
[0321] Smart contracts may manage rights for collaboratively created stories and associated acoustic content. The system may record narrative contributions and acoustic modifications on distributed ledgers to enable appropriate attribution and compensation. Generated content may be represented as tokens containing both story elements and acoustic specifications while maintaining relationships between collaborative contributions.s. Augmented Reality Escape Room
[0322] In some implementations, the system may provide puzzle experiences combining phy sical and virtual acoustic elements. The content generation components may analyze room acoustics and available playback devices to optimize audio cue presentation. The system mayPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO adjust spatial audio rendering to maintain accurate sound localization while accounting for environmental interference patterns.
[0323] When participants interact with puzzle elements, the system may generate appropriate acoustic feedback while maintaining consistency with physical space capabilities. The content generation components may analyze room acoustics and participant positions to produce localized audio cues that guide discovery’ while preserving challenge appropriate to skill levels. In implementations where physical spaces have varying acoustic properties, the system may modify audio presentations to maintain puzzle solvability while optimizing for local conditions.
[0324] The system may implement dynamic difficulty adjustment based on participant performance and environmental capabilities. Content generation components may analyze puzzle progression and acoustic limitations to generate appropriate hints and feedback. The system may adjust audio cue complexify and spatial distribution to maintain engagement while ensuring critical information remains accessible.
[0325] Smart contracts may manage puzzle progression and achievement recording across different physical implementations. The system may record completion data and acoustic modification patterns on distributed ledgers to enable experience optimization. Generated audio content may be represented as tokens that maintain relationships with puzzle elements while enabling adaptation to different environments.t. Virtual Sports Event
[0326] In some implementations, the system may enable shared viewing of sporting events with dynamic crowd simulation. The content generation components may analyze acoustic characteristics of viewing spaces to optimize atmosphere generation. The system may adjust crowd noise synthesis and spatialization to create appropriate energy levels while working within playback limitations.
[0327] When multiple viewers participate from different locations, the system may generate synchronized but individually optimized acoustic experiences. The content generation components may analyze viewer reactions and environmental capabilities to produce appropriate crowd responses while maintaining temporal alignment. In implementations where viewing spaces have varying acoustic properties, the system may modify’ crowd simulation to preserve excitement while optimizing for local conditions.
[0328] The system may implement real-time generation of crowd reactions based on event progression and viewer engagement. Content generation components may analyze game eventsPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO and viewer responses to produce appropriate acoustic enhancement. The system may adjust generated crowd elements to maintain emotional impact while avoiding interference with critical event audio.
[0329] Smart contracts may manage viewing rights and social interaction capabilities across distributed viewing locations. The system may record viewing patterns and acoustic preferences on distributed ledgers to enable experience optimization. Generated crowd content may be represented as tokens that maintain relationships with event broadcasts while enabling personalized presentation.V. Conclusion
[0330] The above discussions relating to playback devices, controller devices, playback zone configurations, and media content sources provide only some examples of operating environments within which funchons and methods described below may be implemented. Other operating environments and configurations of media playback systems, playback devices, and network devices not explicitly described herein may also be applicable and suitable for implementation of the functions and methods.
[0331] The description above discloses, among other things, various example systems, methods, apparatus, and articles of manufacture including, among other components, firmware and / or software executed on hardware. It is understood that such examples are merely illustrative and should not be considered as limiting. For example, it is contemplated that any or all of the firmware, hardware, and / or software aspects or components can be embodied exclusively in hardware, exclusively in software, exclusively in firmware, or in any combination of hardware, softw are, and / or firmw are. Accordingly, the examples provided are not the only ways) to implement such systems, methods, apparatus, and / or articles of manufacture.
[0332] Additionally, references herein to ‘’example’’ means that a particular feature, structure, or characteristic described in connection with the example can be included in at least one example or embodiment of an invention. The appearances of this phrase in various places in the specification are not necessarily all referring to the same example, nor are separate or alternative examples mutually exclusive of other examples. As such, the examples described herein, explicitly and implicitly understood by one skilled in the art, can be combined with other examples.
[0333] The specification is presented largely in terms of illustrative environments, systems, procedures, steps, logic blocks, processing, and other symbolic representations that directly orPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO indirectly resemble the operations of data processing devices coupled to networks. These process descriptions and representations are typically used by those skilled in the art to most effectively convey the substance of their work to others skilled in the art. Numerous specific details are set forth to provide a thorough understanding of the present disclosure. However, it is understood to those skilled in the art that certain examples of the present technology can be practiced without certain, specific details. In other instances, well known methods, procedures, components, and circuitry have not been described in detail to avoid unnecessarily obscuring aspects of the examples. Accordingly, the scope of the present disclosure is defined by the appended claims rather than the foregoing description of examples.
[0334] When any of the appended claims are read to cover a purely software and / or firmware implementation, at least one of the elements in at least one example is hereby expressly defined to include a tangible, non-transitory medium such as a memory, DVD, CD, Blu-ray, and so on, storing the software and / or firmware.
[0335] The disclosed technology is illustrated, for example, according to various examples described below. Various examples of examples of the disclosed technology are described as numbered examples (1, 2, 3, etc.) for convenience. These are provided as examples and do not limit the disclosed technology. It is noted that any of the dependent examples may be combined in any combination, and placed into a respective independent example. The other examples can be presented in a similar manner.
[0336] Example 1. A method comprising: determining one or more acoustic characteristics associated with a real-world environment; receiving, via a network interface, virtual environment data defining acoustic properties of a virtual environment; modifying the virtual environment data based on the determined acoustic characteristics of the real-world environment to generate a modified virtual environment having adjusted acoustic properties that account for acoustic capabilities and limitations of the real-world environment; and providing the modified virtual environment to an extended reality (XR) display device for playback.
[0337] Example 2. The method of Example 1. wherein obtaining the acoustic characteristics comprises analyzing sound reflections, reverberation time, and / or frequency response characteristics of the real-world environment.
[0338] Example 3. The method of any one of Examples 1-2, w herein modifying the virtual environment data comprises adjusting sizes, shapes, and / or materials of virtual objects to createPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO sound reflections and reverberations that match acoustic capabilities of the real-world environment.
[0339] Example 4. The method of any one of Examples 1-3, wherein modifying the virtual environment data comprises adjusting positions of virtual sound sources based on locations of audio playback devices in the real-world environment.
[0340] Example 5. The method any one of Examples 1-4, further comprising: receiving biometric data from sensors associated with a user; and adjusting the acoustic properties of the modified virtual environment based on the biometric data.
[0341] Example 6. The method of any one of Examples 1-5, further comprising analyzing audio equipment specifications in the real-world environment to determine frequency response limitations and spatial audio reproduction capabilities.
[0342] Example 7. The method of Example 6, wherein modifying the virtual environment data comprises adjusting acoustic properties to avoid frequencies that cannot be accurately reproduced by the audio equipment.
[0343] Example 8. The method of any one of Examples 1-7. further comprising: determining that the real-world environment includes an auxiliary zone suitable for additional audio playback; and modifying the virtual environment data to utilize the auxiliary zone for enhanced audio reproduction.
[0344] Example 9. The method of any one of Examples 1-8. further comprising: obtaining head-related transfer function (HRTF) data for a user; and personalizing spatial audio rendering in the modified virtual environment based on the EIRTF data.
[0345] Example 10. The method of any one of Examples 1-9, further comprising: analyzing emotional characteristics of the virtual environment; and adjusting acoustic properties of the modified virtual environment to enhance emotional impact while remaining within acoustic capabilities of the real-world environment.
[0346] Example 11. The method of any one of Examples 1-10, wherein modifying the virtual environment data comprises substituting certain audio frequencies with psychoacoustic alternatives that create similar perceived effects within limitations of the real-world environment.
[0347] Example 12. The method of any one of Examples 1-11, further comprising: monitoring real-time changes in the acoustic characteristics of the real-world environment; and dynamically updating the modified virtual environment based on the real-time changes.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0348] Example 13. A method comprising: obtaining first acoustic data from a first device in a first real-world environment and second acoustic data from a second device in a second real-world environment; analyzing the first and second acoustic data to determine respective acoustic capabilities of the first and second real-world environments; generating first modified virtual environment data for the first device and second modified virtual environment data for the second device, wherein the first and second modified virtual environment data are customized based on the respective acoustic capabilities while maintaining a shared virtual experience between users of the first and second devices; and providing the first and second modified virtual environment data to the respective devices for playback.
[0349] Example 14. The method of Example 13, wherein generating the first and second modified virtual environment data comprises creating different audio mixes optimized for respective acoustic capabilities of the first and second real-world environments while maintaining synchronized timing between the environments.
[0350] Example 15. The method of any one of Examples 13-14, further comprising: receiving first user preference data associated with a first user of the first device and second user preference data associated with a second user of the second device; and customizing the first and second modified virtual environment data based on the respective user preference data.
[0351] Example 16. The method of any one of Examples 13-15, wherein the shared virtual experience comprises a virtual concert at a venue, and wherein generating the first and second modified virtual environment data comprises adjusting virtual acoustics of the concert venue differently for each device based on respective acoustic capabilities of the first and second real-world environments.
[0352] Example 17. The method of any one of Examples 13-16, further comprising: analyzing audio equipment specifications for the first and second devices; and generating the first and second modified virtual environment data to optimize audio reproduction for the respective audio equipment specifications.
[0353] Example 18. The method of any one of Examples 13-17, further comprising: monitoring relative positions of users in their respective real-world environments; and adjusting spatial audio rendering in the first and second modified virtual environment data based on the relative positions.
[0354] Example 19. The method of any one of Examples 13-18, further comprising: detecting acoustic interference in one of the real -world environments; and modifyingPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO corresponding virtual environment data to minimize impact of the interference while maintaining the shared virtual experience.
[0355] Example 20. The method of any one of Examples 13-19, wherein generating the first and second modified virtual environment data comprises creating different virtual room sizes or material properties for each device while maintaining consistent relative spatial relationships between virtual objects.
[0356] Example 21. The method of any one of Examples 13-20, further comprising generating spatialized audio content separately for each device based on respective acoustic capabilities while maintaining synchronized timing of audio events.
[0357] Example 22. The method of any one of Examples 13-21, further comprising: monitoring network latency between the devices; and adjusting audio synchronization in the modified virtual environment data to compensate for the network latency.
[0358] Example 23. A method comprising: determining real-world environment data characterizing acoustic properties of a physical space; storing the real-world environment data and associated user preference data via a blockchain; receiving a request to generate a virtual environment for an extended reality (XR) device located in the physical space; retrieving the stored real -world environment data and user preference data via the blockchain; generating, via a generative media module, customized virtual environment data based on the retrieved real-world environment data and user preference data; and providing the customized virtual environment data to the XR device for playback.
[0359] Example 24. The method of Example 23, wherein storing via the blockchain comprises: creating a non-fungible token (NFT) containing the real-world environment data and user preference data; and recording ownership and access rights for the NFT on the blockchain.
[0360] Example 25. The method of any one of Examples 23-24, further comprising: receiving biometric data from sensors associated with a user; storing the biometric data via the blockchain; and using the generative media module to adjust the customized virtual environment data based on the biometric data.
[0361] Example 26. The method of any one of Examples 23-25, wherein generating the customized virtual environment data comprises: using machine learning to analyze historical user interaction data stored on the blockchain; and adjusting acoustic properties based on patterns identified in the historical user interaction data.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO
[0362] Example 27. The method of any one of Examples 23-26, further comprising: storing virtual sound source specifications as NFTs on the blockchain; and incorporating the virtual sound sources into the customized virtual environment data according to smart contract terms.
[0363] Example 28. The method of any one of Examples 23-27, wherein generating the customized virtual environment data comprises using the generative media module to create personalized audio content based on emotional characteristics identified in the user preference data.
[0364] Example 29. The method of any one of Examples 23-28, further comprising: storing acoustic signatures of different physical spaces on the blockchain; and using the acoustic signatures to optimize the customized virtual environment data for different locations.
[0365] Example 30. The method of any one of Examples 23-29, further comprising implementing smart contracts to manage access rights and usage terms for virtual audio assets incorporated in the customized virtual environment data.
[0366] Example 31. The method of any one of Examples 23-30, wherein generating the customized virtual environment data comprises using the generative media module to create dynamic soundscapes that adapt to user movement patterns stored on the blockchain.
[0367] Example 32. The method of any one of Examples 23-31, further comprising: storing collaborative audio experiences as NFTs on the blockchain; and incorporating elements of the collaborative audio experiences into the customized virtual environment data based on user preferences.
[0368] Example 33. One or more computer-readable media storing instructions that, when executed by one or more processors of a computing device or system, cause the computing device or system to perform operations comprising the method of any one of Examples 1-32.
[0369] Example 34. An extended reality XR system comprising: one or more processors; and the one or more computer-readable media of Example 33.
[0370] Example 35. An extended reality (XR) device comprising: one or more processors; and the one or more computer-readable media of Example 33.
Claims
PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO CLAIMS1. A method comprising:determining one or more acoustic characteristics associated with a real-world environment;receiving, via a network interface, virtual environment data defining acoustic properties of a virtual environment;modifying the virtual environment data based on the determined acoustic characteristics of the real-world environment to generate a modified virtual environment having adjusted acoustic properties that account for acoustic capabilities and limitations of the real-world environment; and providing the modified virtual environment to an extended reality (XR) display device for playback.
2. The method of claim 1. wherein obtaining the acoustic characteristics comprises analyzing sound reflections, reverberation time, and / or frequency response characteristics of the real-world environment.
3. The method of any one of claims 1-2, wherein modifying the virtual environment data comprises adjusting sizes, shapes, and / or materials of virtual objects to create sound reflections and reverberations that match acoustic capabilities of the real-world environment.
4. The method of any one of claims 1-3, wherein modifying the virtual environment data comprises adjusting positions of virtual sound sources based on locations of audio playback devices in the real-world environment.
5. The method any one of claims 1-4, further comprising:receiving biometric data from sensors associated with a user; andadjusting the acoustic properties of the modified virtual environment based on the biometric data.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO 6. The method of any one of claims 1-5, further comprising analyzing audio equipment specifications in the real-world environment to determine frequency response limitations and spatial audio reproduction capabilities.
7. The method of claim 6, wherein modifying the virtual environment data comprises adjusting acoustic properties to avoid frequencies that cannot be accurately reproduced by the audio equipment.
8. The method of any one of claims 1-7, further comprising:determining that the real-world environment includes an auxiliary zone suitable for additional audio playback; andmodifying the virtual environment data to utilize the auxiliary’ zone for enhanced audio reproduction.
9. The method of any one of claims 1-8, further comprising:obtaining head-related transfer function (HRTF) data for a user; and personalizing spatial audio rendering in the modified virtual environment based on the HRTF data.
10. The method of any one of claims 1-9, further comprising:analyzing emotional characteristics of the virtual environment; andadjusting acoustic properties of the modified virtual environment to enhance emotional impact while remaining within acoustic capabilities of the real- world environment.
11. The method of any one of claims 1-10, wherein modifying the virtual environment data comprises substituting certain audio frequencies with psychoacoustic alternatives that create similar perceived effects within limitations of the real-world environment.
12. The method of any one of claims 1-11, further comprising:monitoring real-time changes in the acoustic characteristics of the real-world environment; andPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO dynamically updating the modified virtual environment based on the real-time changes.
13. A method comprising:obtaining first acoustic data from a first device in a first real-world environment and second acoustic data from a second device in a second real-world environment;analyzing the first and second acoustic data to determine respective acoustic capabilities of the first and second real-world environments;generating first modified virtual environment data for the first device and second modified virtual environment data for the second device, wherein the first and second modified virtual environment data are customized based on the respective acoustic capabilities while maintaining a shared virtual experience between users of the first and second devices; andproviding the first and second modified virtual environment data to the respective devices for playback.
14. The method of claim 13, wherein generating the first and second modified virtual environment data comprises creating different audio mixes optimized for respective acoustic capabilities of the first and second real-world environments while maintaining synchronized timing between the environments.
15. The method of any one of claims 13-14. further comprising:receiving first user preference data associated with a first user of the first device and second user preference data associated with a second user of the second device; andcustomizing the first and second modified virtual environment data based on the respective user preference data.
16. The method of any one of claims 13-15, wherein the shared virtual experience comprises a virtual concert at a venue, and wherein generating the first and second modified virtual environment data comprises adjusting virtual acoustics of the concert venuePATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO differently for each device based on respective acoustic capabilities of the first and second real-world environments.
17. The method of any one of claims 13-16, further comprising:analyzing audio equipment specifications for the first and second devices; and generating the first and second modified virtual environment data to optimize audio reproduction for the respective audio equipment specifications.
18. The method of any one of claims 13-17, further comprising: monitoring relative positions of users in their respective real-world environments; and adjusting spatial audio rendering in the first and second modified virtual environment data based on the relative positions.
19. The method of any one of claims 13-18. further comprising:detecting acoustic interference in one of the real-world environments; and modifying corresponding virtual environment data to minimize impact of the interference while maintaining the shared virtual experience.
20. The method of any one of claims 13-19, wherein generating the first and second modified virtual environment data comprises creating different virtual room sizes or material properties for each device while maintaining consistent relative spatial relationships between virtual objects.
21. The method of any one of claims 13-20, further comprising generating spatialized audio content separately for each device based on respective acoustic capabilities while maintaining synchronized timing of audio events.
22. The method of any one of claims 13-21, further comprising: monitoring network latency between the devices; andadjusting audio synchronization in the modified virtual environment data to compensate for the network latency.
23. A method comprising:PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO determining real-world environment data characterizing acoustic properties of a physical space;storing the real-world environment data and associated user preference data via a blockchain;receiving a request to generate a virtual environment for an extended reali ty (XR) device located in the physical space;retrieving the stored real-world environment data and user preference data via the blockchain;generating, via a generative media module, customized virtual environment data based on the retrieved real-world environment data and user preference data; andproviding the customized virtual environment data to the XR device for playback.
24. The method of claim 23, wherein storing via the blockchain comprises: creating a non-fungible token (NFT) containing the real-world environment data and user preference data; and recording ownership and access rights for the NFT on the blockchain.
25. The method of any one of claims 23-24, further comprising:receiving biometric data from sensors associated with a user;storing the biometric data via the blockchain; andusing the generative media module to adjust the customized virtual environment data based on the biometric data.
26. The method of any one of claims 23-25, wherein generating the customized virtual environment data comprises:using machine learning to analyze historical user interaction data stored on the blockchain; andadjusting acoustic properties based on patterns identified in the historical user interaction data.
27. The method of any one of claims 23-26, further comprising:storing virtual sound source specifications as NFTs on the blockchain; andPATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO incorporating the virtual sound sources into the customized virtual environment data according to smart contract terms.
28. The method of any one of claims 23-27, wherein generating the customized virtual environment data comprises using the generative media module to create personalized audio content based on emotional characteristics identified in the user preference data.
29. The method of any one of claims 23-28, further comprising:storing acoustic signatures of different physical spaces on the blockchain; and using the acoustic signatures to optimize the customized virtual environment data for different locations.
30. The method of any one of claims 23-29, further comprising implementing smart contracts to manage access rights and usage terms for virtual audio assets incorporated in the customized virtual environment data.
31. The method of any one of claims 23-30, wherein generating the customized virtual environment data comprises using the generative media module to create dynamic soundscapes that adapt to user movement patterns stored on the blockchain.
32. The method of any one of claims 23-31, further comprising:storing collaborative audio experiences as NFTs on the blockchain; and incorporating elements of the collaborative audio experiences into the customized virtual environment data based on user preferences.
33. One or more computer-readable media storing instructions that, when executed by one or more processors of a computing device or system, cause the computing device or system to perform operations comprising the method of any one of claims 1-32.
34. An extended reality XR system comprising:one or more processors; andthe one or more computer-readable media of claim 33.PATENT Attorney Docket No. 24-1202-PCT Fortem Ref. No. SNS.163WO 35. An extended reality (XR) device comprising:one or more processors; andthe one or more computer-readable media of claim 33.