Generative audio playback via wearable playback device
By realizing the playback of generative media content and context data-driven intelligent control on wearable playback devices, the needs of synchronous playback of music in multi-device and multi-room environments are solved, and a personalized and dynamic audio experience is realized, enhancing the user's listening comfort and immersion.
Patent Information
- Application Number
- CN202380070004.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-30
- Filing Date
- 2023-09-22
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has difficulty providing an easy-to-use approach in consumer products to enhance the listening experience of digital media, especially the need for synchronous playback of music in multi-room and multi-device environments, etc. is not adequately addressed.
By configuring a wearable playback device to play back the generated media content, and intelligently controlling the playback based on context data, such as detecting the user to automatically start the soundscape playback after wearing the headset, and dynamically adjust the audio content using sensor data and environmental information.
It realizes a personalized and dynamic audio experience in multi-device and multi-room environments, enhancing user listening comfort and immersion, while simplifying the operation process.
Smart Images

Figure CN119998781A_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Patent Application No. 63 / 377,776, filed on September 30, 2022, the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present technology relates to consumer products and, more particularly, to methods, systems, products, features, services, and other elements related to media playback systems or some aspect thereof. Background Art
[0003] Options for accessing and listening to digital audio from an external speaker setup were limited until 2003, when SONOS filed one of its first patent applications, entitled "Method for Synchronizing Audio Playback between Multiple Networked Devices," and began selling a media playback system in 2005. The SONOS wireless HiFi system enables people to experience music from many sources through one or more networked playback devices. Through a software control application installed on a smartphone, tablet, or computer, a person is able to play the content he or she desires in any room with a networked playback device. In addition, by using a controller, for example, different songs can be streamed to each room with a playback device, rooms can be grouped together for synchronized playback, or the same song can be listened to synchronously in all rooms.
[0004] Given the growing interest in digital media, there remains a need to develop technology that is easy for consumers to use to further enhance the listening experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] The features, aspects and advantages of the disclosed techniques may be better understood with reference to the following description, appended claims and accompanying drawings.
[0006] Figure 1A is a partial cross-sectional diagram of an environment having a media playback system configured according to aspects of the disclosed technology.
[0007] Figure 1B yes Figure 1A A schematic diagram of a media playback system and one or more networks.
[0008] Figure 1C yes Figure 1A and Figure 1B Schematic diagram of a media playback system comprising a wearable playback device for playing back generative audio content.
[0009] Figure 2A is a functional block diagram of an example playback device.
[0010] Figure 2B yes Figure 2A An isometric diagram of an example housing of a playback device.
[0011] Figure 2C yes Figure 2A FIG. 1 is a diagram of another example housing of a playback device.
[0012] Figure 2D yes Figure 2A FIG. 1 is a diagram of another example housing of a playback device.
[0013] Figure 2E yes Figure 2A FIG. 1 is a diagram of another example housing of a playback device.
[0014] FIG. 3A to FIG. 3E is a diagram illustrating an example playback device configuration according to aspects of the present disclosure.
[0015] Figure 4A is a functional block diagram of an example controller device according to aspects of the present disclosure.
[0016] Figure 4B and Figure 4C is a controller interface according to aspects of the present disclosure.
[0017] FIG. 5A to FIG. 5D An example method for generative audio playback via a wearable audio playback device according to aspects of the present disclosure is shown.
[0018] Figure 6 is a schematic diagram of a system for generating and playing back generative media content according to aspects of the present disclosure.
[0019] Figure 7 is a schematic diagram of another system for generating and playing back generative media content according to aspects of the present disclosure.
[0020] FIG. 8A to FIG. 8D An example method for playback of generative audio based on position according to aspects of the present disclosure is shown.
[0021] Fig. 9 An example scenario involving playback of generative audio via a wearable playback device within a home environment is shown in accordance with aspects of the present disclosure.
[0022] Fig. 10A An example method for exchanging playback of generated audio between a wearable playback device and an external playback device according to aspects of the present disclosure is shown.
[0023] Fig. 10BAn example scenario involving the exchange of playback of generated audio between a wearable playback device and an external playback device in accordance with aspects of the present disclosure is shown.
[0024] Fig. 10C is a schematic diagram of a system for generating and playing back generative media content according to aspects of the present disclosure.
[0025] Fig. 10D An example method of playing back generative audio via both a wearable playback device and an outboard playback device according to aspects of the present disclosure is shown.
[0026] Fig.11 An example rules engine for limiting playback of generative audio within an environment is shown in accordance with aspects of the present disclosure.
[0027] The drawings are for the purpose of illustrating example embodiments, but it is understood that the invention is not limited to the arrangements and instrumentalities shown in the drawings. In the drawings, the same reference numerals identify at least substantially similar elements. To facilitate discussion of any particular element, the most significant digit or digits in any reference numeral refer to the drawing in which the element is first introduced. For example, element 103a is referred to in reference numeral 103b. Figure 1A Introduced and discussed for the first time. DETAILED DESCRIPTION
[0028] 1. Overview
[0029] Generative media content is content that is dynamically synthesized, created, and / or modified based on an algorithm, whether implemented in software or a physical model. Generative media content can be changed over time based on an algorithm alone or in combination with contextual data (e.g., user sensor data, environmental sensor data, occurrence data). In various examples, such generative media content can include generative audio (e.g., music, ambient soundscapes, etc.), generative visual images (e.g., abstract visual designs that dynamically change shapes, colors, etc.), or any other suitable media content or a combination thereof. As described elsewhere herein, generative audio can be created at least in part via an algorithm and / or a non-human system that utilizes rule-based computing to produce novel audio content.
[0030] Because generative media content can be modified dynamically in real time, it can achieve a unique user experience that cannot be obtained by conventional media playback using pre-recorded content. For example, generative audio can be an infinite and / or dynamic audio that changes as the input of the algorithm (e.g., input parameters associated with user input, sensor data, media source data, or any other suitable input data) changes. In some examples, generative audio can be used to guide the user's mood to a desired emotional state, wherein one or more characteristics of the generative audio change in response to real-time measurements that reflect the user's emotional state. As used in examples of the present technology, the system can provide generative audio based on the user's current and / or desired emotional state, based on the user's activity level, based on the number of users present in the environment, or any other suitable input parameters.
[0031] Listening to audio content (whether generated audio or pre-existing audio) via a wearable playback device (e.g., headphones, earbuds) typically requires the wearable playback device to be bound or connected to another device, such as a smartphone, tablet, laptop, etc. Certain wearable playback devices (e.g., WiFi-enabled devices) may not require binding to another local device, but may still require the user to actively initiate audio playback (e.g., via voice commands). In some cases, it would be desirable for the listener to use a wearable device that is not bound to another device, where playback can be initiated automatically, and optionally, in response to contextual data (e.g., data related to the listener's environment).
[0032] The present technology relates to a wearable playback device configured to play back generative media content and intelligently control such playback based on contextual data. For example, the wearable playback device may detect engagement of at least one of a user's ears (e.g., detecting an earphone on one ear, or an earphone placed in at least one of the user's ears, etc.), and automatically initiate playback of a soundscape (e.g., generative audio, media content, other sounds such as sounds related to the user's environment). The generative audio soundscape may include a generative musical composition based at least in part on one or more media content stems and / or audio cues derived from contextual data. The contextual data may include information related to the user's environment (e.g., location, place, or orientation relative to the environment, time of day, temperature, circadian rhythm, humidity, number of people nearby, ambient light level) or other types of indicators (e.g., doorbells, alarms, events). In some embodiments, certain other indicators (e.g., doorbells, alarms, etc.) may be selectively communicated to the wearable playback device to notify the user while the user remains within the immersive audio experience. Furthermore, in general, audio captured in this environment can serve as input to a generative media algorithm, and in some cases the output of the generative media algorithm can itself serve as input (i.e., resampled) to future iterations of generative media content.
[0033] In various examples, contextual data may be obtained via onboard sensors carried by the wearable playback device, sensors associated with other playback devices within the environment, or any other suitable sensor data source. In some examples, the wearable playback device may recognize the user when the wearable playback device is placed on the user's head, and the wearable playback device may further customize the generative media content based on the user's profile, current or desired emotional state, and / or other biometric data (e.g., brainwave activity, heart rate, breathing rate, skin moisture content, ear shape, head direction or orientation). In addition, in some cases, playback of the generative audio content may be dynamically swapped or switched between playback via the wearable playback device and playback via one or more external playback devices within the listening environment, either alternately or simultaneously.
[0034] Although some embodiments described herein may involve functions being performed by given actors (e.g., "users" and / or other entities), it should be understood that this description is for purposes of explanation only. The claims should not be interpreted as requiring any such example actors to perform an action unless the language of the claim itself expressly requires it.
[0035] 2. Example operating environment
[0036] Figures 1A to 1CAn example configuration of a media playback system 100 (or "MPS 100") capable of implementing one or more embodiments disclosed herein is shown. Figure 1A , the MPS 100 shown is associated with an example home environment having multiple rooms and spaces, which may be collectively referred to as a "home environment," "smart home," or "environment 101." Environment 101 includes a home having several rooms, spaces, and / or playback zones, including a master bathroom 101a, a master bedroom 101b (referred to herein as "Nick's room"), a second bedroom 101c, a family room or study 101d, an office 101e, a living room 101f, a dining room 101g, a kitchen 101h, and an outdoor patio 101i. Although certain embodiments and examples are described below in the context of a home environment, the techniques described herein may be implemented in other types of environments. In some embodiments, for example, the MPS 100 may be in one or more commercial environments (e.g., restaurants, malls, airports, hotels, retail stores, or other stores), one or more vehicles (e.g., sports utility vehicles, buses, cars, ships, boats, airplanes), multiple environments (e.g., a combination of a home environment and a vehicle environment), and / or another suitable environment where multi-zone audio may be desired.
[0037] Within these rooms and spaces, MPS 100 includes one or more computing devices. Figures 1A to 1C , such computing devices may include playback devices 102 (individually identified as playback devices 102a through 102p), network microphone devices 103 (individually identified as "NMDs" 103a through 102i), and controller devices 104a and 104b (collectively referred to as "controller devices 104"). Figure 1B , the home environment may include additional and / or other computing devices, including local network devices, such as one or more smart lighting devices 108 ( Figure 1B ), smart thermostat 110 and local computing device 105 ( Figure 1A In the following embodiments, one or more of the various playback devices 102 may be configured as portable playback devices, while other playback devices may be configured as fixed playback devices. For example, the headset 102o ( Figure 1B) is a portable playback device, while playback device 102d on the bookcase may be a fixed device. As another example, playback device 102c on the patio may be a battery-powered device that may allow it to be transported to various areas within environment 101 as well as outside of environment 101 when it is not plugged into a wall outlet, etc. In addition, one or more of the various playback devices 102 may be configured as a wearable playback device (e.g., a playback device configured to be worn on, by, or around a user, such as headphones, earbuds, smart glasses with integrated audio transducers, etc.). Playback devices 102o and 102p ( Figure 1C ) is an example of a wearable playback device. In contrast to a wearable playback device, one or more of the various playback devices 102 may be configured as an out-of-body playback device (e.g., a playback device configured to output audio content to a listener and / or multiple users at a distance from the playback device, rather than a private listening experience associated with a wearable playback device).
[0038] refer to Figure 1B , the various playback devices, network microphones, and controller devices 102 to 104 and / or other network devices of the MPS 100 may be coupled to each other via a network 111 that may include a network router 109, via point-to-point connections that may be wired and / or wireless, and / or through other connections. For example, the study room 101d ( Figure 1A ) can have a point-to-point connection with the playback device 102a, which is also located in the study 101d and can be designated as the "right" device. In a related embodiment, the left playback device 102j can communicate with other network devices (e.g., playback device 102b), which can be designated as the "front" device, via a local network 111 via a point-to-point connection and / or other connections. The local network 111 can, for example, be a network that interconnects one or more devices within a limited area (e.g., a residence, an office building, a car, a personal workspace, etc.). The local network 111 can, for example, include: one or more local area networks (LANs), such as a wireless local area network (WLAN) (e.g., a WI-FI network, a Z-WAVE network, etc.); and / or one or more personal area networks (PANs), such as a Bluetooth network, a wireless USB network, a ZIGBEE network, and an IRDA network.
[0039] like Figure 1BAs further shown, MPS 100 can be coupled to one or more remote computing devices 106 via a wide area network ("WAN") 107. In some embodiments, each remote computing device 106 can be in the form of one or more cloud servers. Remote computing devices 106 can be configured to interact with computing devices in environment 101 in various ways. For example, remote computing devices 106 can be configured to facilitate streaming and / or controlling playback of media content, such as audio, in home environment 101.
[0040] In some implementations, the various playback devices, NMDs, and / or controller devices 102-104 may be communicatively coupled to at least one remote computing device associated with a voice assistant service (“VAS”) and at least one remote computing device associated with a media content service (“MCS”). For example, in Figure 1B In the example shown, remote computing device 106a is associated with VAS 190 and remote computing device 106b is associated with MCS 192. Although for clarity, the remote computing device 106a is associated with VAS 190 and MCS 192 is associated with MCS 192. Figure 1B In the example of FIG. 1 , only a single VAS 190 and a single MCS 192 are shown, but the MPS 100 can be coupled to multiple, different VASs and / or MCSs. In some embodiments, the VAS can be operated by one or more of AMAZON, GOOGLE, APPLE, MICROSOFT, NUANCE, SONOS, or other voice assistant providers. In some embodiments, the MCS can be operated by one or more of SPOTIFY, PANDORA, AMAZON MUSIC, or other media content services.
[0041] like Figure 1B As further shown in FIG. 1 , the remote computing device 106 also includes a remote computing device 106c configured to perform certain operations (e.g., remotely facilitating media playback functions, managing device and system status information, directing communications between devices of the MPS 100 and one or more VASs and / or MCSs, etc.). In one example, the remote computing device 106c provides a cloud server for one or more SONOS Wireless HiFi systems.
[0042] In various embodiments, one or more playback devices 102 may take the form of or include an onboard (e.g., integrated) network microphone device. For example, playback devices 102a to 102e include or are otherwise equipped with corresponding NMDs 103a to 103e, respectively. Unless otherwise specified in this description, playback devices including or equipped with NMDs are interchangeably referred to herein as playback devices or NMDs. In some cases, one or more NMDs 103 may be stand-alone devices. For example, NMD 103f and NMD 103g may be stand-alone devices. Stand-alone NMDs may omit components and / or functions typically included in playback devices, such as speakers or related electronics. For example, in this case, the stand-alone NMD may not produce audio output or may produce limited audio output (e.g., relatively low-quality audio output).
[0043] The various playback devices of MPS 100 and network microphone devices 102 and 103 may each be associated with a unique name that may be assigned to the respective device by a user, such as during setup of one or more of these devices. Figure 1B As shown in the illustrated example, a user may assign the name "bookcase" to playback device 102d because it is physically located on a bookcase. Similarly, NMD 103f may be assigned the name "island" because it is physically located on an island countertop in kitchen 101h ( Figure 1A ). Some playback devices may be assigned names based on zones or rooms, such as playback devices 102e, 102l, 102m, and 102n named "bedroom," "dining room," "living room," and "office," respectively. In addition, some playback devices may have functionally descriptive names. For example, playback devices 102a and 102b are assigned the names "right" and "front," respectively, because these two devices are configured to provide specific audio channels during media playback in the zone of study 101d ( Figure 1A ). The playback device 102c in the patio may be named portable because it is battery powered and / or easily transported to different areas of the environment 101. Other naming conventions are also possible.
[0044] As described above, the NMD can detect and process sounds from its environment, such as sounds including background noise mixed with speech spoken by people near the NMD. For example, when the NMD detects sound in the environment, the NMD can process the detected sound to determine whether the sound includes speech that includes speech input intended for the NMD and ultimately a particular VAS. For example, the NMD can identify whether the speech includes a wake-up word associated with a particular VAS.
[0045] exist Figure 1B In the example shown, the NMD 103 is configured to interact with the VAS 190 via the local network 111 and / or the router 109. For example, when the NMD identifies a potential wake-up word in the detected sound, interaction with the VAS 190 may be initiated. The identification causes a wake-up word event, which in turn causes the NMD to begin sending the detected sound data to the VAS 190. In some embodiments, the various local network devices 102 to 105 ( Figure 1A ) and / or the remote computing device 106 c can exchange various feedback, information, instructions and / or related data with the remote computing device associated with the selected VAS. This exchange can be related to or independent of the sent message containing the voice input. In some embodiments, the remote computing device and the media playback system 100 can exchange data via a communication path as described herein and / or using metadata exchanges described in U.S. Patent Publication No. 2017-0242653, issued on August 24, 2017 and entitled “Voice Control of a Media Playback System”, the entire contents of which are incorporated herein by reference.
[0046] After receiving the sound data stream, VAS 190 determines whether there is voice input in the streaming data from the NMD, and if so, VAS 190 will also determine the potential intent in the voice input. VAS 190 can then send a response back to MPS 100, which can include sending a response directly to the NMD that caused the wake-up word event. The response is usually based on the intent determined by VAS 190 to be present in the voice input. As an example, in response to VAS 190 receiving a voice input with the utterance "Play Hey Jude by The Beatles", VAS 190 can determine that the potential intent of the voice input is to initiate playback and further determine that the intended voice input is to play a specific song "Hey Jude". After these determinations, VAS 190 can send a command to a specific MCS 192 to retrieve content (i.e., the song "Hey Judy"), and MCS 192 in turn directly or indirectly provides (e.g., streams) the content to MPS 100 via VAS 190. In some implementations, VAS 190 may send a command to MPS 100 that causes MPS 100 itself to retrieve the content from MCS 192 .
[0047] In some implementations, when speech input is recognized among speech detected by two or more NMDs located in proximity to each other, the NMDs may facilitate arbitration between each other. For example, environment 101 ( Figure 1A) is relatively close to the living room playback device 102m equipped with an NMD, and both devices 102d and 102m may detect the same sound at least sometimes. In this case, this may require arbitration as to which device is ultimately responsible for providing the detected sound data to the remote VAS. Examples of arbitration between NMDs can be found, for example, in previously cited U.S. Patent Publication No. 2017-0242653.
[0048] In some implementations, an NMD may be assigned to or associated with a designated or default playback device that may not include an NMD. For example, kitchen 101h ( Figure 1A ) can be assigned to a restaurant playback device 1021 that is relatively close to the island NMD 103f. In practice, the NMD can instruct the assigned playback device to play audio in response to the remote VAS receiving a voice input from the NMD to play audio, which the NMD may have sent to the VAS in response to a user speaking a command to play a certain song, album, playlist, etc. Additional details about assigning NMDs and playback devices as designated or default devices can be found, for example, in previously referenced U.S. Patent Publication No. 2017-0242653.
[0049] refer to Figure 1C , the media playback system 100 can be configured to generate and play back generative media content via one or more wearable playback devices 102o, 102p and / or via one or more external playback devices 102. In the example shown, the wearable playback devices 102o, 102p each take the form of headphones, although any suitable wearable playback device may be used. The external playback device 102 may include any suitable device configured to output audio content for external listening in an environment. Some external playback devices include integrated audio transducers (e.g., sound bars, subwoofers), while other external playback devices may include amplifiers that are configured to provide output signals to be played back via other devices (e.g., hub devices, set-top boxes, etc.).
[0050] Figure 1CThe various communication links shown can be wired or wireless network connections that can be facilitated at least in part via a router 109. The wireless connection can include WiFi, Bluetooth, or any other suitable communication protocol. As shown, one or more local sources 150 can be connected to the wearable playback device 102o and the external playback device 102 via a network (e.g., via a router 109). The local sources 150 can include any suitable media and / or audio sources, such as display devices (e.g., televisions, projectors, etc.), microphones, analog playback devices (e.g., turntables), portable data storage devices (e.g., USB sticks), computer storage devices (e.g., hard drives of laptop computers), etc. These local sources 150 can optionally provide audio content to be played back via the wearable playback device 102o, 102p and / or the external playback device 102. In some examples, the media content obtained from the local sources 150 can provide input to the generative media module so that the resulting generative media content is based on and / or combines features of the media content from the local sources 150.
[0051] The first wearable playback device 102o can be communicatively coupled to a control device 104 (e.g., a smartphone, tablet, laptop, etc.), for example, via WiFi, Bluetooth, or other suitable wireless connection. The control device 104 can be used to select content and / or otherwise control the playback of audio via the wearable playback device 102o and / or any additional playback devices 102.
[0052] The first wearable playback device 102o is also optionally connected to the local area network via the router 109 (e.g., via a WiFi connection). The second wearable playback device 102p can be connected to the first wearable playback device 102o via a direct wireless connection (e.g., Bluetooth) and / or a wireless network (e.g., via a WiFi connection of the router 109). The wearable playback devices 102o, 102p can communicate with each other to transmit audio content, timing information, sensor data, generative media content parameters (e.g., content models, algorithms, etc.), or any other suitable information to facilitate the generation, selection and / or playback of audio content. The second wearable playback device 102p can also optionally connect to one or more external playback devices via a wireless connection (e.g., a Bluetooth or WiFi connection to a hub device).
[0053] The media playback system 100 can also communicate with one or more remote computing devices 154, which are associated with media content providers and / or generative audio sources. These remote computing devices 154 can provide media content, optionally including generative media content. As described in more detail elsewhere herein, in some cases, generative media content can be produced via one or more generative media modules, which can be instantiated via remote computing devices 154, via one or more in the playback device 102, and / or its certain combination.
[0054] The generative media module may generate generative media based at least in part on input parameters, which may include sensor data (e.g., received from sensor data source 152) and / or other suitable input parameters. With respect to sensor input parameters, sensor data source 152 may include data from any suitable sensor, regardless of where the sensor is located relative to various playback devices and any values measured thereby. Examples of suitable sensor data include physiological sensor data, such as data obtained from biometric sensors, wearable sensors, and the like. Such data may include physiological parameters, such as heart rate, respiratory rate, blood pressure, brain waves, activity level, movement, body temperature, and the like.
[0055] Suitable sensors include wearable sensors configured to be worn or carried by a user, such as a wearable playback device, headphones, a watch, a mobile device, a brain-computer interface, a microphone, or other similar device. In some examples, the sensor may be a non-wearable sensor or fixed to a fixed structure. The sensor may provide sensor data, which may include data corresponding to, for example, brain activity, a user's mood or emotional state, speech, position, motion, heart rate, pulse, body temperature, and / or perspiration.
[0056] In some examples, sensor data source 152 includes data obtained from networked device sensor data (e.g., Internet of Things (IoT) sensors, such as networked lights, cameras, temperature sensors, thermostats, presence detectors, microphones, etc.). Additionally or alternatively, sensor data source 152 may include environmental sensors (e.g., measuring or indicating weather, temperature, time / day / week / month) as well as user calendar events, proximity to other electronic devices, historical usage patterns, blockchain, or other distributed data sources, etc.
[0057] In one example, a user may wear a biometric device that can measure various biometric parameters of the user, such as heart rate or blood pressure. The generative media module (whether residing at a remote computing device 154 or at one or more local playback devices 102) can use these parameters to further adjust the generative audio, such as by increasing the tempo of the music in response to detecting a high heart rate (because this can indicate that the user is engaged in a high-intensity activity) or decreasing the tempo of the music in response to detecting high blood pressure (because this can indicate that the user is stressed and can benefit from calming music). In yet another example, one or more microphones of the playback device can detect the user's voice. The captured voice data can then be processed to determine, for example, the user's mood, age, or gender (to identify a particular user from among several users within a household), or any other such input parameter. Other examples are also possible.
[0058] Additional details regarding the generation and playback of generative media content may be found in commonly owned International Patent Application Publication No. WO 2022 / 109556, entitled “Playback of Generative Media Content,” the entire contents of which are incorporated herein by reference.
[0059] Other aspects regarding the different components of the example MPS 100 and how the different components may interact to provide a media experience to a user may be found in the following sections. Although the discussion herein may generally relate to the example MPS 100, the techniques described herein are not limited to application within the above-described home environment, etc. For example, the techniques described herein may be useful in other home environment configurations that include more or less of any playback device, network microphone, and / or controller device 102 to 104. For example, the techniques herein may be used in an environment with a single playback device 102 and / or a single NMD 103. In some examples of such situations, the local network 111 ( Figure 1B ), and a single playback device 102 and / or a single NMD 103 can communicate directly with remote computing devices 106a to 106d. In some embodiments, a telecommunications network (e.g., an LTE network, a 5G network, etc.) can communicate with various playback devices, network microphones, and / or controller devices 102 to 104 independently of a local network 111.
[0060] Although the above has been about Figures 1A to 1C Specific implementations of the MPS are described, but there are many configurations of the MPS, including but not limited to configurations that do not interact with remote services, systems that do not include a controller, and / or any other configuration suitable for the requirements of a given application.
[0061] a. Example playback device and network microphone device
[0062] Figure 2A It shows Figures 1A to 1C 1 is a functional block diagram of certain aspects of one of the playback devices 102 of the MPS 100. As shown, the playback device 102 includes various components, each of which is discussed in more detail below, and the various components of the playback device 102 may be operably coupled to each other via a system bus, a communication network, or some other connection mechanism. Figure 2A In the example shown, playback device 102 may be referred to as an "NMD-equipped" playback device because it includes components that support the functionality of an NMD, such as Figure 1A One of the NMDs 103 is shown.
[0063] As shown, the playback device 102 includes at least one processor 212, which may be a clock-driven computing component configured to process input data according to instructions stored in a memory 213. The memory 213 may be a tangible, non-transitory computer-readable medium configured to store instructions executable by the processor 212. For example, the memory 213 may be a data storage device that may be loaded with software code 214 that may be executed by the processor 212 to implement certain functions.
[0064] In one example, these functions may involve playback device 102 retrieving audio data from an audio source that may be another playback device. In another example, the function may involve playback device 102 sending audio data, detected voice data (e.g., corresponding to voice input) and / or other information to another device on the network via at least one network interface 224. In another example again, the function may involve playback device 102 causing one or more other playback devices to play back audio synchronously with playback device 102. In another example again, the function may involve playback device 102 facilitating pairing with one or more other playback devices or otherwise binding with one or more other playback devices to create a multi-channel audio environment. Many other example functions are possible, some of which are discussed below.
[0065] As just mentioned, certain functions may include synchronizing playback of audio content by playback device 102 with one or more other playback devices. During synchronized playback, a listener may not perceive time delay differences between the playback of audio content by the synchronized playback devices. U.S. Patent No. 8,234,395, entitled "System and method for synchronizing operations among a plurality of independently clocked digital data processing devices," filed on April 4, 2004, provides some examples of audio playback synchronization between playback devices in more detail, and the entire contents of the U.S. Patent No. 8,234,395, which is incorporated herein by reference in its entirety, is hereby incorporated by reference.
[0066] To facilitate audio playback, the playback device 102 includes an audio processing component 216 that is generally configured to process the audio before the playback device 102 presents the audio. In this regard, the audio processing component 216 may include one or more digital-to-analog converters ("DACs"), one or more audio pre-processing components, one or more audio enhancement components, one or more digital signal processors ("DSPs"), etc. In some implementations, the one or more audio processing components 216 may be subcomponents of the processor 212. In operation, the audio processing component 216 receives analog audio and / or digital audio and processes and / or otherwise intentionally alters the audio to produce an audio signal for playback.
[0067] The generated audio signal may then be provided to one or more audio amplifiers 217 for amplification and playback through one or more speakers 218 operatively coupled to the amplifiers 217. The audio amplifiers 217 may include components configured to amplify the audio signal to a level for driving the one or more speakers 218.
[0068] Each speaker 218 may include a separate transducer (e.g., a "driver"), or the speaker 218 may include a complete speaker system including a housing having one or more drivers. For example, specific drivers for the speakers 218 may include, for example, a woofer (e.g., for low frequencies), a mid-band driver (e.g., for mid-range frequencies), and / or a tweeter (e.g., for high frequencies). In some cases, the transducers may be driven by separate corresponding audio amplifiers of the audio amplifier 217. In some embodiments, the playback device may not include the speaker 218, but may include a speaker interface for connecting the playback device to external speakers. In some embodiments, the playback device may include neither the speaker 218 nor the audio amplifier 217, but may include an audio interface (not shown) for connecting the playback device to an external audio amplifier or audio-visual receiver.
[0069] In addition to generating audio signals for playback by the playback device 102, the audio processing component 216 may also be configured to process audio to be sent to one or more other playback devices for playback via the network interface 224. In an example scenario, audio content to be processed and / or played back by the playback device 102 may be received from an external source, for example, via an audio line-in interface (e.g., an auto-detecting 3.5 mm audio line-in connection) of the playback device 102 (not shown) or via the network interface 224, as described below.
[0070] As shown, at least one network interface 224 may be in the form of one or more wireless interfaces 225 and / or one or more wired interfaces 226. The wireless interface may provide the playback device 102 with network interface functionality to communicate wirelessly with other devices (e.g., other playback devices, NMDs, and / or controller devices) according to a communication protocol (e.g., any wireless standard, including IEEE 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.11ad, 802.11af, 802.11ah, 802.11ai, 802.11aj, 802.11aq, 802.11ax, 802.11ay, 802.15, Bluetooth, 4G mobile communication standards, 5G mobile communication standards, etc.). The wired interface may provide the playback device 102 with network interface functionality to communicate with other devices via a wired connection according to a communication protocol (e.g., IEEE 802.3). Although Figure 2A The illustrated network interface 224 includes both a wired interface and a wireless interface, but the playback device 102 may include only a wireless interface or only a wired interface in some implementations.
[0071] Typically, the network interface 224 facilitates the data flow between the playback device 102 and one or more other devices on the data network. For example, the playback device 102 can be configured to receive audio content from an audio content source on one or more other playback devices, a network device within a LAN, and / or a WAN (e.g., the Internet) via a data network. In one example, the audio content and other signals sent and received by the playback device 102 can be sent in the form of digital packet data including a source address based on an Internet Protocol (IP) and a destination address based on IP. In this case, the network interface 224 can be configured to parse the digital packet data so that the data going to the playback device 102 is correctly received and processed by the playback device 102.
[0072] like Figure 2A As shown, the playback device 102 also includes a voice processing component 220 that is operably coupled to one or more microphones 222. The microphones 222 are configured to detect sounds (i.e., sound waves) in the environment of the playback device 102, and then provide the sounds to the voice processing component 220. More specifically, each microphone 222 is configured to detect sounds and convert the sounds into digital signals or analog signals representing the detected sounds, which can then enable the voice processing component 220 to perform various functions based on the detected sounds, as described in more detail below. In one embodiment, the microphones 222 are arranged as a microphone array (e.g., an array of six microphones). In some embodiments, the playback device 102 includes more than six microphones (e.g., eight microphones or twelve microphones) or less than six microphones (e.g., four microphones, two microphones, or a single microphone).
[0073] In operation, the speech processing component 220 is generally configured to detect and process sounds received via the microphone 222, identify potential speech input in the detected sounds, and extract the detected sound data for use by the VAS (e.g., the VAS 190 ( Figure 1B)) is capable of processing voice input identified in the detected sound data. In addition to other example voice processing components, the voice processing component 220 may also include one or more analog-to-digital converters, an acoustic echo canceller ("AEC"), a spatial processor (e.g., one or more multi-channel Wiener filters, one or more other filters, and / or one or more beamformer components), one or more buffers (e.g., one or more circular buffers), one or more wake-up word engines, one or more voice extractors, and / or one or more voice processing components (e.g., a component configured to recognize the voice of a specific user or a specific set of users associated with a home). In an example embodiment, the voice processing component 220 may include one or more DSPs, or one or more modules of a DSP, or otherwise take the form of one or more DSPs, or one or more modules of a DSP. In this regard, certain voice processing components 220 may be configured with specific parameters (e.g., gain and / or spectral parameters) that can be modified or otherwise tuned to achieve specific functions. In some embodiments, one or more voice processing components 220 may be subcomponents of the processor 212.
[0074] In some embodiments, the voice processing component 220 can detect and store a voice profile of the user, which can be associated with the user account of the MPS 100. For example, the voice profile can be stored as a set of command information or variables stored in a data table and / or compared to a set of command information or variables stored in a data table. The voice profile can include aspects of the pitch or frequency of the user's voice and / or other unique aspects of the user's voice, such as those described in previously referenced U.S. Patent Publication No. 2017 / 0242653.
[0075] like Figure 2A As further shown in FIG. 1 , the playback device 102 also includes a power supply assembly 227. The power supply assembly 227 may include at least an external power supply interface 228, which may be coupled to a power supply (not shown) via a cable or the like that physically connects the playback device 102 to an electrical outlet or some other external power source. Other power supply assemblies may include, for example, transformers, converters, and similar assemblies configured to format electrical power.
[0076] In some embodiments, the power supply component 227 of the playback device 102 may additionally include an internal power supply 229 (e.g., one or more batteries) that is configured to power the playback device 102 without a physical connection to an external power source. When equipped with an internal power supply 229, the playback device 102 can operate independently of the external power supply. In some such embodiments, the external power interface 228 can be configured to facilitate charging the internal power supply 229. As discussed above, playback devices including internal power supplies may be referred to as "portable playback devices" herein. Those portable playback devices that weigh no more than fifty ounces (e.g., between three ounces and fifty ounces, between five ounces and fifty ounces, between ten ounces and fifty ounces, between ten ounces and twenty-five ounces, etc.) may be referred to as "ultra-portable playback devices" herein. Playback devices that operate using an external power supply rather than an internal power supply may be referred to as "fixed playback devices" herein, although such devices can actually be moved around a home or other environment.
[0077] The playback device 102 may also include a user interface 231 that may facilitate user interaction independent of or in conjunction with user interaction facilitated by one or more of the controller devices 104. In various embodiments, the user interface 231 includes one or more physical buttons and / or supports a graphical interface provided on a touch-sensitive screen and / or surface, etc., for a user to directly provide input. The user interface 231 may also include one or more lights (e.g., LEDs) and speakers to provide visual and / or audio feedback to the user.
[0078] like Figure 2A As shown, playback device 102 may also include one or more sensors 209. Sensor 209 includes any suitable sensor, regardless of what value it measures. Examples of suitable sensors include user engagement sensors for determining whether a user is wearing or touching a wearable playback device, microphones or other audio capture devices, cameras or other imaging devices, accelerometers, gyroscopes, or other motion or activity sensors, physiological sensors for measuring physiological parameters such as heart rate, breathing rate, blood pressure, brain waves, activity level, movement, body temperature, etc. In some cases, sensor 209 may be wearable (e.g., playback device 102 itself is wearable, or sensor 209 is separate from playback device 102 but remains communicatively coupled to playback device 102).
[0079] The playback device 102 may also optionally include a generative media module 211 that is configured to generate media content alone or in combination with other devices (eg, other local playback devices and non-playback devices, remote computing devices 154 ( Figure 1C) etc.) to generate generative media content). As previously mentioned, generative media content can include any media content (e.g., audio, video, audio-visual output, tactile output or any other media content) dynamically created, synthesized and / or modified by a non-human rule-based process (such as an algorithm or model), even if the process involves manual input. Although such processes can be rule-based, they do not need to be completely deterministic, but can be combined with a certain randomness aspect or other random behavior. Such creation or modification can occur for real-time or near real-time playback. Additionally or alternatively, generative media content can be asynchronously generated or modified (e.g., in advance before requesting playback), and then a specific item of generative media content can be selected for later playback. As used herein, a "generative media module" includes any system that can generate generative media content based on one or more inputs, whether implemented in software, physical models or a combination thereof. In some examples, such generative media content includes new media content, which can be created as brand-new media content or can be created by mixing, combining, manipulating or otherwise modifying one or more pre-existing media content segments. As used herein, a "generative media content model" includes any algorithm, pattern, or rule set that can be used to generate new types of generative media content using one or more inputs (e.g., sensor data, audio captured by an onboard microphone, parameters provided by an artist, media segments such as audio clips or samples, etc.). In an example, a generative media module can use a variety of different generative media content models to produce different generative media content. In some cases, artists or other collaborators can interact with, create, and / or update the generative media content model to produce specific generative media content. Although several examples throughout this discussion involve audio content, the principles disclosed herein can be applied to other types of media content, such as video, audio-visual, tactile, or other media content in some examples.
[0080] In some examples, the generative media module 211 can utilize one or more input parameters in the form of playback device capabilities (e.g., number and type of transducers, output power, other system architectures), device location (e.g., location relative to other playback devices, location relative to one or more users), e.g., data from sensor 209, from other sensors. Additional inputs can include device states of one or more devices within the group, such as thermal state (e.g., if a particular device is in danger of overheating, the generative content can be modified to reduce the temperature), battery power (e.g., bass output can be reduced in a portable playback device with low battery power), and binding state (e.g., whether a particular playback device is configured as part of a stereo pair, bound to a subwoofer, or as part of a home theater arrangement, etc.). Any other suitable device characteristics or states can similarly be used as inputs for generating generative media content.
[0081] In the case of a wearable playback device, the position and / or orientation of the playback device relative to the environment can serve as an input to the generative media module 211. For example, a spatial soundscape can be provided so that as the user moves around the environment, the corresponding audio produced changes (e.g., the sound of a waterfall in one corner of the room, birds singing in another corner). Additionally or alternatively, as the user's orientation changes (facing one direction, tilting her head up or down, etc.), the corresponding audio output can be modified accordingly.
[0082] The identity of the user (as determined via sensor 209 or otherwise) may also serve as an input to the generative media module 211, so that the specific generative media produced by the generative media module 211 is customized to the individual user. Another example input parameter includes user presence - for example, when a new user enters a space where the generative audio is being played back, the presence of the user may be detected (e.g., via a proximity sensor, beacon, etc.), and the generative audio may be modified in response. Such modification may be based on the number of users (e.g., providing ambient, meditation audio for 1 user, relaxing music for 2 to 4 users, and party or dance music for more than 4 users). Modifications may also be based on the identification of the users present (e.g., a user profile based on user characteristics, listening history, or other such indicia).
[0083] As an illustrative example, Figure 2B An example housing 230 of the playback device 102 is shown, which includes a user interface in the form of a control area 232 at the top 234 of the housing 230. The control area 232 includes buttons 236a to 236c for controlling audio playback, volume level, and other functions. The control area 232 also includes a button 236d for switching the microphone 222 to an on state or an off state.
[0084] like Figure 2B As further shown in FIG. 1 , the control area 232 is at least partially surrounded by an aperture formed in the top 234 of the housing 230, the microphone 222 (in Figure 2B The microphone 222 may be arranged at various locations along and / or within the top 234 or other area of the housing 230 to detect sound from one or more directions relative to the playback device 102.
[0085] As described above, playback device 102 may be configured as a portable playback device (eg, an ultra-portable playback device) that includes an internal power source. Figure 2C An example housing 240 of such a portable playback device 102 is shown. As shown, the housing 240 of the portable playback device includes a user interface in the form of a control area 242 at the top 244 of the housing 240. The control area 242 may include a capacitive touch sensor for controlling audio playback, volume level, and other functions. The housing 240 of the portable playback device may be configured to engage with a charging dock 246 connected to an external power source via a cable 248. The charging dock 246 may be configured to provide power to the portable playback device to recharge the internal battery. In some embodiments, the charging dock 246 may include a collection of one or more conductive contacts (not shown) located at the top of the charging dock 246, which engage with conductive contacts (not shown) on the bottom of the housing 240. In other embodiments, the charging dock 246 may provide power to the portable playback device from the cable 248 without using conductive contacts. For example, the charging dock 246 may wirelessly charge the portable playback device via one or more inductive coils integrated into each of the charging dock 246 and the portable playback device.
[0086] In some embodiments, playback device 102 may be in the form of wired and / or wireless headphones (e.g., ear-hook headphones, on-ear headphones, or in-ear headphones). Figure 2D An example housing 250 for such an implementation of the playback device 102 is shown. As shown, the housing 250 includes a headband 252 that couples a first earphone 254a to a second earphone 254b. Each of the earphones 254a and 254b can house any portion of the electronic components in the playback device, such as one or more speakers. In addition, one or more of the earphones 254a and 254b can include a control area 258 for controlling audio playback, volume levels, and other functions. The control area 258 can include any combination of the following: capacitive touch sensors, buttons, switches, and dials. Figure 2DAs shown, the housing 250 may also include ear pads 256a and 256b coupled to the earphones 254a and 254b, respectively. The ear pads 256a and 256b may provide a soft barrier between the user's head and the earphones 254a and 254b, respectively, to improve user comfort and / or provide acoustic isolation from the surrounding environment (e.g., passive noise reduction (PNR)). As described above, the playback device 102 may include one or more sensors configured to detect various parameters, optionally including on-ear detection that indicates when the user is wearing and not wearing the headphone device. In some embodiments, the wired and / or wireless headphones may be ultra-portable playback devices that are powered by an internal energy source and weigh less than 50 ounces.
[0087] In some embodiments, playback device 102 may take the form of an in-ear headset or a hearing aid. Figure 2E An example housing 260 for such an implementation of a playback device 102 is shown. As shown, the housing 260 includes: an in-ear portion 262, configured to be arranged in or adjacent to a user's ear; and an ear-hanging portion 264, configured to extend above and behind the user's ear. The housing 260 can accommodate any part of the electronic components in the playback device, such as one or more audio transducers, microphones, and audio processing components. Multiple control areas 266 can facilitate user input for controlling audio playback, volume levels, noise cancellation, pairing with other devices, and other functions. The control area 258 can include any combination of the following items: one or more buttons, switches, dials, capacitive touch sensors, etc. As described above, the playback device 102 may include one or more sensors configured to detect various parameters, optionally including in-ear detection indicating when the user is wearing and not wearing an in-ear headphone device.
[0088] It should be understood that the playback device 102 can take the form of other wearable devices that are separate and spaced apart from the headphones. Wearable devices can include those devices that are configured to be worn around a portion of a subject (e.g., the head, neck, torso, arm, wrist, finger, leg, ankle, etc.). For example, the playback device 102 can take the form of a pair of glasses that includes a front portion of a frame (e.g., configured to hold one or more lenses), a first temple rotatably coupled to the front portion of the frame, and a second temple rotatably coupled to the front portion of the frame. In this example, the pair of glasses can include one or more transducers that are integrated into at least one of the first temple and the second temple and are configured to project sound toward the subject's ear.
[0089] Although the above has been about FIG. 2A to FIG. 2ESpecific implementations of playback devices and network microphone devices are described, but there are multiple configurations of devices, including but not limited to devices without UIs, microphones at different locations, multiple microphone arrays positioned in different arrangements, and / or any other configuration suitable for the requirements of a given application. For example, the UI and / or microphone array can be implemented in other playback devices and / or computing devices other than the playback devices and / or computing devices described herein. In addition, although specific examples of playback devices 102 are described with reference to MPS 100, those skilled in the art will recognize that the playback devices described herein can be used in a variety of different environments, including (but not limited to) environments with more and / or fewer elements, without departing from the present invention. Similarly, the MPS described herein can be used with a variety of different playback devices.
[0090] For example, SONOS currently offers (or has offered) certain playback devices for sale, including "SONOS ONE", "FIVE", "PLAYBAR", "AMP", "CONNECT:AMP", "PLAYBASE", "BEAM", "ARC", "CONNECT", "MOVE", "ROAM" and "SUB", which can implement certain embodiments disclosed herein. Any other past, present and / or future playback devices may additionally or alternatively be used to implement the playback devices of the example embodiments disclosed herein. In addition, it should be understood that the playback devices are not limited to FIG. 2A to FIG. 2D Examples or SONOS product offerings shown. For example, the playback device can be an integral part of another device or assembly such as a television, a lighting fixture, or some other device used indoors or outdoors.
[0091] b. Example playback device configuration
[0092] FIG. 3A to FIG. 3E An example configuration of a playback device is shown. Figure 3A In some example instances, a single playback device may belong to a zone. For example, playback device 102c ( Figure 1A ) may belong to zone A. In some embodiments described below, multiple playback devices may be "bound" to form a "bound pair" that together form a single zone. For example, Figure 3A The playback device 102f named "Bed 1" ( Figure 1A ) can be bound to Figure 3A The playback device 102g named "Bed 2" ( Figure 1A) to form Zone B. The bound playback devices may have different playback responsibilities (e.g., channel responsibilities). In another implementation described below, multiple playback devices may be merged to form a single zone. For example, a playback device 102d named "Bookcase" may be merged with a playback device 102m named "Living Room" to form a single Zone C. The merged playback devices 102d and 102m may not be specifically assigned different playback responsibilities. That is, in addition to synchronously playing audio content, the merged playback devices 102d and 102m may also play audio content individually as if they were not merged.
[0093] For control purposes, each zone in MPS 100 may be represented as a single user interface (“UI”) entity. For example, as displayed by controller device 104, zone A may be provided as a single entity named “Portable”, zone B may be provided as a single entity named “Stereo”, and zone C may be provided as a single entity named “Living Room”.
[0094] In various embodiments, a zone may take the name of one of the playback devices belonging to the zone. For example, zone C may take the name of living room device 102m (as shown). In another example, zone C may alternatively take the name of bookcase device 102d. In yet another example, zone C may take a name that is some combination of bookcase device 102d and living room device 102m. The selected name may be selected by a user via input at controller device 104. In some embodiments, a zone may be given a different name than the devices belonging to the zone. For example, Figure 3A Zone B in the example is named "Stereo", but none of the devices in Zone B have that name. In one aspect, Zone B is a single UI entity representing a single device named "Stereo", which is composed of constituent devices "Bed 1" and "Bed 2". In one embodiment, the Bed 1 device may be the master bedroom 101h ( Figure 1A ) in the playback device 102f, and the Bed 2 device may be also in the master bedroom 101h ( Figure 1A ) in the playback device 102g.
[0095] As mentioned above, bound playback devices can have different playback responsibilities, such as playback responsibilities for certain audio channels. Figure 3B As shown, bed 1 device 102f and bed 2 device 102g can be bound to produce or enhance a stereo effect of audio content. In this example, bed 1 playback device 102f can be configured to play the left channel audio component, while bed 2 playback device 102g can be configured to play the right channel audio component. In some embodiments, this stereo binding can be referred to as "pairing."
[0096] Furthermore, playback devices configured to be bound may have additional and / or different corresponding speaker drivers. Figure 3C As shown, a playback device 102b named "front" can be bound to a playback device 102k named "subwoofer". The front device 102b can render a mid-range to high-range frequency, while the subwoofer device 102k, for example as a woofer, can render a low-range frequency. When unbound, the front device 102b can be configured to render a full range of frequencies. As another example, Figure 3D The front device 102b and the subwoofer device 102k are shown further bound to the right playback device 102a and the left playback device 102j, respectively. In some embodiments, the right device 102a and the left device 102j can form the surround or "satellite" channels of the home theater system. The bound playback devices 102a, 102b, 102j, and 102k can form a single D zone ( Figure 3A ).
[0097] In some implementations, playback devices may also be "merged." In contrast to certain bound playback devices, merged playback devices may have no assigned playback responsibilities, but may each present the full range of audio content that each respective playback device is capable of providing. However, the merged device may be represented as a single UI entity (i.e., a zone, as described above). For example, Figure 3E Playback devices 102d and 102m are shown merged in a living room, which would result in these devices being represented by a single UI entity in section C. In one embodiment, playback devices 102d and 102m may play back audio in sync, during which each outputs the full range of audio content that each respective playback device 102d and 102m is capable of presenting.
[0098] In some embodiments, an independent NMD may be alone in a zone. For example, Figure 1A The NMD 103h in the Figure 3A An NMD may also be bound or merged with another device to form a zone. For example, NMD device 103f named "Island" may be bound with playback device 102i, which together form zone F, also referred to as "Kitchen". Additional details regarding the assignment of NMDs and playback devices as designated or default devices may be found, for example, in previously referenced U.S. Patent Publication No. 2017-0242653. In some embodiments, an independent NMD may not be assigned to a zone.
[0099] Zones of separate, bonded, and / or merged devices may be arranged to form a group of playback devices that synchronously play back audio. Such a group of playback devices may be referred to as a "group," "zone group," "sync group," or "playback group." In response to input provided via controller device 104, playback devices may be dynamically grouped and ungrouped to form new or different groups of synchronously playing back audio content. For example, referring to Figure 3A , zone A can be grouped with zone B to form a zone group of playback devices that includes both zones. As another example, zone A can be grouped with one or more other zones C to I. Zones A to I can be grouped and ungrouped in a variety of ways. For example, three, four, five or more (e.g., all) of zones A to I can be grouped together. Zones of individual playback devices and / or bound playback devices can play back audio in sync with each other when grouped together, as described in previously cited U.S. Patent No. 8,234,395. Grouped devices and bound devices are example types of associations between portable playback devices and fixed playback devices that can be caused in response to a trigger event, as discussed above and described in more detail below.
[0100] In various embodiments, zones in an environment may be assigned specific names, which may be default names for zones within a zone group or combinations of names for zones within a zone group (e.g., "restaurant+kitchen"), such as Figure 3A In some embodiments, the zone group can be given a unique name selected by the user, such as "Nick's Room", as shown in FIG. Figure 3A The name "Nick's Room" may be a name selected by the user instead of a previous name for the zone group, such as the room name "Master Bedroom".
[0101] Return to reference Figure 2A , certain data may be stored in memory 213 as one or more state variables that are periodically updated and used to describe the state of a playback zone, playback device, and / or zone group associated therewith. Memory 213 may also include data associated with the state of other devices of media playback system 100, which may be shared between devices from time to time so that one or more devices have the latest data associated with the system.
[0102] In some embodiments, the memory 213 of the playback device 102 can store instances of various variable types associated with the state. The variable instances can be stored with identifiers (e.g., tags) corresponding to the types. For example, some identifiers can be a first type "a1" for identifying playback devices of a zone, a second type "b1" for identifying playback devices that can be bound to a zone, and a third type "c1" for identifying a zone group to which the zone can belong. As a related example, in Figure 1A, an identifier associated with a patio may indicate that the patio is the only playback device for a particular zone and is not in a zone group. An identifier associated with a living room may indicate that the living room is not grouped with other zones, but includes bound playback devices 102a, 102b, 102j, and 102k. An identifier associated with a restaurant may indicate that the restaurant is part of a restaurant+kitchen group and that devices 103f and 102i are bound. Since the kitchen is part of a restaurant+kitchen zone group, an identifier associated with the kitchen may indicate the same or similar information. Other example zone variables and identifiers are described below.
[0103] In yet another example, the MPS 100 may include other associated variables or identifiers representing zones and zone groups, such as identifiers associated with areas, such as Figure 3A A region can refer to a cluster of blocks and / or blocks that are not in a block group. For example, Figure 3A A first zone named "first zone" and a second zone named "second zone" are shown. The first zone includes zones and zone groups of a terrace, a study, a dining room, a kitchen, and a bathroom. The second zone includes zones and zone groups of a bathroom, Nick's room, a bedroom, and a living room. In one aspect, a zone can be used to call a group of zones and / or a cluster of zones that share one or more zones and / or groups of zones of another cluster. In this regard, such a zone is different from a zone group, and a zone group does not share a zone with another zone group. Additional examples of techniques for implementing zones can be found, for example, in U.S. Patent Publication No. 2018-0107446, issued on April 19, 2018 and entitled "Room Association Based on Name" and U.S. Patent No. 8,483,853, filed on September 11, 2007 and entitled "Controlling and manipulating groupings in a multi-zone media system", the entire contents of each of which are incorporated herein by reference. In some embodiments, MPS 100 may not implement regions, in which case the system may not store variables associated with regions.
[0104] The memory 213 may also be configured to store other data. Such data may be about audio sources accessible by the playback device 102 or a playback queue with which the playback device (or some other playback device) may be associated. In the embodiment described below, the memory 213 is configured to store a command data set for selecting a particular VAS when processing speech input.
[0105] During operation, Figure 1AOne or more playback zones in an environment may each play different audio content. For example, a user may be grilling in a patio zone and listening to hip-hop music being played by playback device 102c, while another user may be preparing food in a kitchen zone and listening to classical music being played by playback device 102i. In another example, a playback zone may play back the same audio content in sync with another playback zone. For example, a user may be in an office zone, where playback device 102n is playing the same hip-hop music being played by playback device 102c in the patio zone. In this case, playback devices 102c and 102n may play hip-hop music in sync, so that the user may seamlessly (or at least substantially seamlessly) enjoy the audio content being played in an external speaker mode while moving between different playback zones. Synchronization between playback zones may be achieved in a manner similar to synchronization between playback devices, as described in previously cited U.S. Patent No. 8,234,395.
[0106] As suggested above, the zone configuration of the MPS 100 can be modified dynamically. Thus, the MPS 100 can support multiple configurations. For example, if a user physically moves one or more playback devices to or from a zone, the MPS 100 can be reconfigured to accommodate the changes. For example, if a user physically moves playback device 102c from a patio zone to an office zone, the office zone can now include playback devices 102c and 102n. In some cases, the user can use, for example, one of the controller devices 104 and / or voice input to pair or group the moved playback device 102c with the office zone and / or rename the players in the office zone. As another example, if one or more playback devices 102 are moved to a specific space in a home environment that is not yet a playback zone, the moved playback device can be renamed or associated with the playback zone of the specific space.
[0107] In addition, different playback zones of the MPS 100 can be dynamically combined into zone groups or divided into separate playback zones. For example, the dining room zone and the kitchen zone can be combined into a zone group for a banquet so that playback devices 102i and 102l can present audio content synchronously. As another example, the bound playback devices in the study zone can be divided into (i) a TV zone and (ii) a separate listening zone. The TV zone can include the front playback device 102b. The listening zone can include the right playback device 102a, the left playback device 102j, and the subwoofer playback device 102k, which can be grouped, paired, or merged as described above. Dividing the study zone in this manner can allow one user to listen to music in a listening zone in one area of the living room space while another user watches TV in another area of the living room space. In a related example, a user can utilize NMD 103a or 103b ( Figure 1B) to control the study area before the study area is separated into the TV area and the listening area. Once the listening area is separated, it can be controlled by, for example, a user near NMD 103a, while the TV area can be controlled by, for example, a user near NMD 103b. However, as described above, any of NMDs 103 can be configured to control various playback and other devices of MPS 100.
[0108] c. Example Controller Device
[0109] Figure 4A It shows Figure 1A 104 of the MPS 100. The controller devices according to several embodiments of the present invention may be used in various systems, such as (but not limited to) Figure 1A Such a controller device may also be referred to herein as a "control device" or a "controller". Figure 4A The controller device shown may include components that are generally similar to certain components of the network devices described above, such as a processor 412, a memory 413 storing program software 414, at least one network interface 424, and one or more microphones 422. In one example, the controller device may be a dedicated controller for MPS 100. In another example, the controller device may be a network device, such as an iPhone™, iPad™, or any other smart phone, tablet computer, or network device (e.g., a networked computer such as a PC or Mac™), on which media playback system controller application software may be installed.
[0110] The memory 413 of the controller device 104 may be configured to store controller application software and other data associated with the MPS 100 and / or users of the system 100. The memory 413 may be loaded with instructions in the software 414, which may be executed by the processor 412 to implement certain functions, such as facilitating user access, control, and / or configuration of the MPS 100. The controller device 104 may be configured to communicate with other network devices via a network interface 424, which may take the form of a wireless interface, as described above.
[0111] In one example, system information (e.g., such as state variables) may be communicated between controller device 104 and other devices via network interface 424. For example, controller device 104 may receive playback zone and zone group configurations in MPS 100 from a playback device, an NMD, or another network device. Likewise, controller device 104 may send such system information to a playback device or another network device via network interface 424. In some cases, the other network device may be another controller device.
[0112] Controller device 104 may also transmit playback device control commands, such as volume control and audio playback control, to the playback device via network interface 424. As suggested above, changes to the configuration of MPS 100 may also be performed by a user using controller device 104. Such configuration changes may include: adding or removing one or more playback devices to or from a zone, adding or removing one or more zones to or from a zone group, forming a bound or merged player, isolating one or more playback devices from a bound or merged player, etc.
[0113] like Figure 4A As shown, the controller device 104 may also include a user interface 440, which is generally configured to facilitate user access to and control of the MPS 100. The user interface 440 may include a touch screen display or other physical interface configured to provide various graphical controller interfaces, such as Figure 4B and Figure 4C Controller interfaces 440a and 440b are shown. Figure 4B and Figure 4C , controller interfaces 440a and 440b include playback control area 442, playback zone area 443, playback status area 444, playback queue area 446, and source area 448. The user interface shown is only a user interface that can be used on a network device (e.g., Figure 4A 1 and can be accessed by a user to control a media playback system (e.g., MPS 100). Alternatively, other user interfaces of varying formats, styles, and interaction sequences can be implemented on one or more network devices to provide similar control access to the media playback system.
[0114] Playback control area 442 ( Figure 4B ) may include selectable (e.g., by touch or by using a cursor) icons that cause the playback device in the selected playback zone or zone group to play or pause, fast forward, rewind, skip to next, skip to previous, enter / exit shuffle mode, enter / exit repeat mode, enter / exit crossfade mode, etc. The playback control area 442 may also include selectable icons that, when selected, modify equalization settings and / or playback volume, etc.
[0115] Playback area area 443 ( Figure 4C ) may include representations of playback zones within MPS 100. Playback zones area 443 may also include representations of zone groups (e.g., restaurant + kitchen zone groups), as shown. In some embodiments, the graphical representations of playback zones may be selectable to bring up additional selectable icons to manage or configure playback zones in MPS 100, e.g., create bound zones, create zone groups, detach zone groups, rename zone groups, etc.
[0116] For example, as shown, a "grouping" icon may be provided within each graphical representation of a playback zone. The "grouping" icon provided within the graphical representation of a particular zone may be selectable to bring up an option for selecting one or more other zones in the MPS 100 to be grouped with the particular zone. Once grouped, playback devices in the zone that have been grouped with the particular zone will be configured to play audio content in sync with the playback devices in the particular zone. Similarly, a "grouping" icon may be provided within the graphical representation of a zone group. In this case, the "grouping" icon may be selectable to bring up an option for deselecting one or more zones in the zone group to be removed from the zone group. Other interactions and implementations for grouping and degrouping zones via a user interface are also possible. When a playback zone or zone group configuration is modified, the playback zone may be dynamically updated in the playback zone area 443 ( Figure 4C ) in the representation.
[0117] Playback status area 444 ( Figure 4B ) may include a graphical representation of the audio content currently playing, previously played, or scheduled to play next in the selected playback zone or zone group. The selected playback zone or zone group may be visually distinguished on the controller interface, for example, within the playback zone area 443 and / or the playback status area 444. The graphical representation may include the track title, artist name, album name, album year, track length, and other relevant information that would be useful for the user to know when controlling the MPS 100 via the controller interface 400.
[0118] Playback queue area 446 can comprise the graphical representation of the audio content in the playback queue associated with selected playback zone or zone group.In certain embodiments, each playback zone or zone group can be associated with playback queue, and this playback queue comprises the corresponding information with zero or more audio items played back by this playback zone or zone group.For example, each audio item in the playback queue can comprise uniform resource identifier (URI), uniform resource locator (URL) or some other identifiers, it can be used for searching and / or retrieving audio items from local audio content source or networked audio content source by the playback device in playback zone or zone group, and this audio item can be played back by playback device then.
[0119] In one example, a playlist can be added to a playback queue, in which case information corresponding to each audio item in the playlist can be added to the playback queue. In another example, the audio items in the playback queue can be saved as a playlist. In another example, when a playback zone or zone group is continuously playing streaming media audio content (e.g., internet radio, which can be continuously played until stopped), rather than discrete audio items with playback duration, the playback queue can be empty or filled but "unused". In an alternative embodiment, the playback queue can include internet radio and / or other streaming media audio content items, and is "in use" when a playback zone or zone group is playing these items. Other examples are also possible.
[0120] When playback zones or zone groups are "grouped" or "ungrouped," the playback queues associated with the affected playback zones or zone groups may be cleared, or reassociated. For example, if a first playback zone including a first playback queue is grouped with a second playback zone including a second playback queue, the established zone group may have an associated playback queue that is initially empty, contains audio items from the first playback queue (e.g., if the second playback zone is added to the first playback zone), or contains audio items from the second playback queue (e.g., if the first playback zone is added to the second playback zone), or contains a combination of audio items from both the first playback queue and the second playback queue. Subsequently, if the established zone group is ungrouped, the resulting first playback zone may be reassociated with the previous first playback queue, or may be associated with a new playback queue that is empty or contains audio items from a playback queue associated with a zone group that was established before the established zone group was ungrouped. Similarly, the resulting second playback zone may be reassociated with the previous second playback queue, or may be associated with a new playback queue that is empty or contains audio items from a playback queue associated with a zone group established before the established zone group was ungrouped. Other examples are also possible.
[0121] Still reference Figure 4B and Figure 4C , the audio content is in the playback queue area 446 ( Figure 4B) can comprise track title, artist name, track length and / or other relevant information that is associated with the audio content in the playback queue.In one example, the graphical representation of audio content can be selectable, to call out additional selectable icons to manage and / or manipulate the audio content represented in the playback queue and / or the playback queue.For example, represented audio content can be removed from the playback queue, represented audio content is moved to different positions in the playback queue, or represented audio content is selected to play immediately, or to play after any currently played audio content etc.The playback queue that is associated with playback zone or zone group can be stored in one or more playback devices in this playback zone or zone group, not in the playback device in this playback zone or zone group and / or in some other specified devices.The playback of this playback queue can relate to one or more playback devices and may play the media items in the sequential or random order playback queue.
[0122] Source area 448 may include graphical representations of selectable audio content sources and / or selectable voice assistants associated with corresponding VASs. VASs may be selectively assigned. In some examples, multiple VASs (e.g., AMAZON's Alexa, MICROSOFT's Cortana, etc.) may be invoked by the same NMD. In some embodiments, a user may specifically assign a VAS to one or more NMDs. For example, a user may assign a first VAS to Figure 1A One or both of the NMDs 102a and 102b in the living room are shown, and a second VAS is assigned to the NMD 103f in the kitchen. Other examples are possible.
[0123] d. Sample audio content source
[0124] The audio source in source area 448 can be an audio content source, and selected playback zone or zone group can retrieve and play audio content from this audio content source.One or more playback devices in zone or zone group can be configured to obtain playback audio content (for example, according to corresponding URI or URL of audio content) from various available audio content sources.In one example, playback device can (for example, connect via line input) directly retrieve audio content from corresponding audio content source.In another example, can be on the network, audio content is provided to playback device by one or more other playback devices or network equipment.As described in more detail below, in some embodiments, audio content can be provided by one or more media content services.
[0125] Example audio content sources may include: a media playback system (e.g., Figures 1A to 1CThe media playback system may include, for example, a memory of one or more playback devices in the MPS 100 of the present invention, a local music library on one or more network devices (e.g., a controller device, a network-enabled personal computer, or a network attached storage (“NAS”)), a streaming media audio service that provides audio content over the Internet (e.g., a cloud-based music service), or an audio source connected to the media playback system via a line-in connection on a playback device or network device.
[0126] In certain embodiments, audio content sources can be added in a media playback system such as media playback system 100, or audio content sources can be removed therefrom. In one example, whenever one or more audio content sources are added, removed or updated, audio item indexing can be performed. Audio item indexing can include: scanning the identifiable audio items in all folders / directories shared on the network accessible by the playback device in the media playback system, and generating or updating an audio content database comprising metadata (for example, title, artist, album, track length, etc.) and other associated information (for example, URI or URL of each identifiable audio item found). Other examples for managing and maintaining the audio content source are also possible.
[0127] 3. Example of Generative Audio Playback via Wearable Playback Device
[0128] Generative audio content enables the creation and delivery of audio content customized to a specific user, a specific environment, and / or a specific time. In the case of a wearable playback device, the generative audio content can be further customized based on parameters detected via the wearable playback device, and playback can be controlled based on contextual data and distributed between various playback devices in the environment. For example, a wearable playback device can detect that a user is wearing it, and automatically initiate playback of generative audio content, which may include generative music works based at least in part on one or more media content trunks and / or audio prompts derived from contextual data. Contextual data may include information related to the user's environment (e.g., time of day, temperature, circadian rhythm, humidity, number of people nearby, ambient light level) or other types of indicators (e.g., doorbells, alarms, events). Such contextual data may be obtained via onboard sensors carried by a wearable playback device, sensors associated with other playback devices in the environment, or any other suitable sensor data source. In some examples, when the wearable playback device is placed on the user's head, the wearable playback device can identify the specific user, and the wearable playback device can further customize the generative media content according to the user's profile, current or desired emotional state, and / or other biometric data. In addition, in some cases, the playback of the generative audio content can be dynamically exchanged or switched between playback only via the wearable playback device and playback alternately or simultaneously via one or more external playback devices within the listening environment. In some examples, the wearable playback device provides one or more contextual inputs to the generative soundscape that is played back via one or more external playback devices and optionally via the wearable playback device. For example, consider the following scenario: one or more listeners want to monitor the biometric data (e.g., heart rate, breathing rate, body temperature, blood sugar level, blood oxygen saturation percentage) of the wearer of the wearable playback device because the wearer may have a disease, be an elderly person, a child, or otherwise have a capability difference with one or more external listeners. One or more generative soundscapes can be generated and played back using one or more biometric parameters of the wearer to allow the external listener to monitor the parameters via the soundscape.
[0129] a. Produce and play back generative audio content via a wearable playback device
[0130] FIG. 5A to FIG. 5D An example method for generative audio playback via a wearable audio playback device is shown. Figure 5A, method 500 begins at block 502, where on-ear detection is performed. For example, on-board sensors may determine when the wearable playback device is placed over or against a user's head (e.g., ear cups are over the user's ears, or headphones are placed in the user's ears, etc.). The on-board sensors may take any suitable form, such as optical proximity sensors, capacitive or inductive touch sensors, gyroscopes, accelerometers, or other motion sensors, etc.
[0131] Once on-ear detection has occurred, method 500 proceeds to box 504, which involves automatically starting playback of generative audio content via a wearable playback device. As previously described, generative media content (such as generative audio content) can be generated via an onboard generative media module resident on a wearable playback device, via a generative media module resident on other local playback devices or computing devices (e.g., accessible via a local area network (e.g., WiFi) or a direct wireless connection (e.g., Bluetooth)). Additionally or alternatively, the generative media content can be generated by one or more remote computing devices, optionally using one or more input parameters provided by the wearable playback device and / or a media playback system including the wearable playback device. In some cases, the generative media content can be generated by a combination of any of these devices.
[0132] At block 506, the on-ear position is no longer detected (e.g., via one or more on-board sensors), and in block 508, playback of the generative audio content is stopped. Optionally, playback of the generative media content may be automatically switched or swapped for playback via one or more other playback devices within the environment (e.g., one or more external playback devices). FIG. 10A to FIG. 10D Let's describe this playback control including switching in more detail.
[0133] exist Figure 5A In the illustrated process, the wearable playback device is configured to automatically start playback of the generated audio content when being worn by the user, and automatically terminate playback of the generated audio content once removed by the user. Optionally, the process can be performed without the wearable playback device being bound or otherwise connected to another device.
[0134] In various examples, the generative audio content can be a soundscape customized to the user and / or the user's environment. For example, the generative audio can be responsive to the user's current emotional state (e.g., determined via one or more sensors of the wearable playback device) and / or contextual data (e.g., information about the user's environment or home, external data such as events outside the user's environment, time of day, etc.).
[0135] Figure 5BAnother example method 510 is shown, which begins at box 512, where generative audio is played back via a wearable playback device. At box 514, a Bluetooth connection to an audio source (e.g., a mobile phone, tablet computer, etc.) or a WiFi audio source is detected. After the detection, the method 510 continues to box 516, where the playback of the generative audio is switched to the playback of the source audio. Optionally, the switch can only occur when the playback of the source audio begins, thereby ensuring that the user will not be left with unwanted silence.
[0136] In some examples, the transition may include a crossfade over a predetermined amount of time (e.g., 5 seconds, 10 seconds, 20 seconds, 30 seconds) so that the transition between the generated audio content and the source audio is gradual. The transition may also be based on other factors, such as transients or "drops" in the content, or caused by a frequency band and the energy detected within that frequency band. In some cases, the source audio type (e.g., music, spoken word, phone call) may also be detected, and the crossfade time may be adjusted accordingly. For example, for music, the crossfade may be a longer time (e.g., 10 seconds), for spoken word, the crossfade may be a medium time (e.g., 5 seconds), and for phone calls, the crossfade may be a shorter time (e.g., 0 seconds, 1 second, 2 seconds). Optionally, the transition may involve initiating active noise control (ANC) if such function has not been used during playback, or terminating active noise control (ANC) if such function was previously used.
[0137] Figure 5C An example method 520 is shown. At block 522, a Bluetooth, WiFi, or other wirelessly connected audio source is no longer detected at the wearable playback device. In block 524, the wearable playback device automatically switches to playing back the generative audio content. The switch may also include cross-fading and / or switching ANC functionality as described above.
[0138] Optionally, as shown in block 526, previously played back audio from a Bluetooth or WiFi source (e.g., Figure 5B The generated audio content may be used as an input parameter (e.g., a stem or seed) to the generative media module such that when playback of the generated audio content is resumed in block 524, the generated audio content includes certain characteristics or is otherwise based at least in part on the previous source audio. This may facilitate a sense of continuity between different media content being played back.
[0139] Figure 5DAn example method 530 and accompanying scenario for playing back generative audio content via multiple playback devices is shown. The method 530 begins at block 532, where a group generative audio playback session is initiated. For example, a wearable playback device may be grouped with one or more additional playback devices for synchronized playback, which may include one or more external playback devices. In this configuration, the various playback devices may synchronize playback of generative media content. Figure 5D In the example shown, wearable device 540 is grouped with a second wearable playback device 542 and external playback devices 544 , 546 , and 548 .
[0140] In block 534, the first playback device generates generative audio content. In various examples, the first playback device can be a wearable playback device (e.g., playback device 540), an external playback device (e.g., playback device 546), or a combination of playback devices that work together to generate generative audio content. As described elsewhere, the generative audio content can be based at least in part on sensor data and / or contextual data, which can be derived at least in part from the playback device itself.
[0141] The method 530 continues to block 536 where the generated audio content is sent to additional playback devices within the environment. The various playback devices can then play back the generated audio content in sync with each other.
[0142] In some examples, wearable playback device 542 may be selected as the group coordinator for the sync group, particularly if wearable playback device 542 is already playing back generative audio content. However, in other examples, playback device 548 may be selected as the group coordinator, e.g., because it has the highest computing power, it is plugged into a power source, has better network hardware or network connectivity, etc.
[0143] In some examples, an external playback device 548 (e.g., a subwoofer) is tied to an external playback device 546 (e.g., a soundbar), and low-frequency audio content may be played back via the external playback device 548 without playing back any audio via the external playback device 546. For example, a listener of the wearable playback device 540 and / or the wearable playback device 542 may wish to utilize the low-frequency capabilities provided by the subwoofer (e.g., the external playback device 548) when listening to the generative audio content via the wearable playback devices 540, 542. In some implementations, the subwoofer may play back low-frequency content that is intended to be heard with multiple but different high-frequency content items (e.g., a single subwoofer and subwoofer content channel may provide bass content for both headphone and external listeners, even though the high-frequency content for each may be different).
[0144] Figure 66 is a schematic diagram of an example distributed generative media playback system 600. As shown, an artist 602 can provide a plurality of media segments 604 and one or more generative content models 606 to a generative media module 211 stored via one or more remote computing devices. The media segments can correspond to, for example, specific audio segments or seeds (e.g., individual notes or chords, short tracks of n bars, non-musical content, etc.). In some examples, the generative content model 606 can also be provided by the artist 602. This can include providing the entire model, or the artist 602 can provide input to the model 606, for example, by changing or adjusting certain aspects (e.g., rhythm, melodic constraints, harmonic complexity parameters, chord change density parameters, etc.). In addition, the process can involve seamless looping, which can use segments of longer stems and recombined infinitely.
[0145] Generative media module 211 may receive both media segments 604 and one or more input parameters 603 (as described elsewhere herein). Based on these inputs, generative media module 211 may output generative media. Figure 6 As shown, the artist 602 may optionally audition the generative media module 211, for example, by receiving exemplary outputs (e.g., media segments 604 and / or generative content models 606) based on inputs provided by the artist 602. In some cases, the audition may play back variations of the generative media content to the artist 602 depending on various different input parameters (e.g., one version corresponding to a high energy level intended to produce an exciting or uplifting effect, another version corresponding to a low energy level intended to produce a calming effect, etc.). Based on the output via this audition step, the artist 602 may dynamically update the settings of the media segments 604 and / or the generative content model 606 until the desired output is achieved.
[0146] In the example shown, at block 608, there may be an iteration every n hours (or minutes, days, etc.) where the generative media module 211 may produce multiple different versions of generative media content. In the example shown, there are three versions: version A in block 610, version B in block 612, and version C in block 614. These outputs are then stored (e.g., via a remote computing device) as generative media content 616. A particular version of these versions (in this example, version C in block 618) may be sent (e.g., streamed) to the local playback device 102a for playback.
[0147] Although three versions are shown here by way of example, there may actually be many more versions of the generative media content generated via the remote computing device. These versions may vary along a number of different dimensions, such as being suitable for different energy levels, suitable for different intended tasks or activities (e.g., studying vs. dancing), suitable for different times of the day, or any other appropriate variation.
[0148] In the example shown, playback device 102a can periodically request the generated media content of a specific version from a remote computing device. This request can be based on, for example, user input (e.g., selection via the user of a controller device), sensor data (e.g., the number of people present in the room, background noise level, etc.) or other suitable input parameters. As shown, input parameter 603 can optionally be provided to playback device 102a (or detected by playback device 102a). Additionally or alternatively, input parameter 603 can be provided to remote computing device 106 (or detected by remote computing device 106). In some examples, playback device 102a sends input parameter to remote computing device 106, and this remote computing device 106 then provides suitable version to playback device 102a, without playback device 102a specifically requesting specific version.
[0149] The external playback device 102a may send the generated media content 616 to the wearable playback device 102o via WiFi, Bluetooth, or other suitable wireless connection. Optionally, the wearable playback device 102o may also provide one or more input parameters (e.g., sensor data from the wearable playback device 102o) that can be used as input to the generative media module 211. Additionally or alternatively, the wearable playback device 102o may receive one or more input parameters 603 that may be used to modify the playback of the generated audio content via the wearable playback device 102o.
[0150] Figure 7 is a schematic diagram of an example system 700 for generating and playing back generative media content. The system 700 may be similar to the system described above with respect to FIG. 1 , except that the generative media content 616 (eg, version C 618) may be sent directly from a remote computing device to the wearable playback device 102o. Figure 6 Described system 600. As previously described, the generated media content can then be further distributed from the wearable playback device 102o to other playback devices within the environment, including other wearable playback devices and / or external playback devices.
[0151] b. Positioning and customization of generative media playback
[0152] In various examples, the production, distribution, and playback of generative media content may be controlled and / or modulated based on location information, user identity, or other such parameters. FIG. 8A to FIG. 8D An example method for playing back generative audio based on position is shown. Fig. 8A , method 800 begins at box 802, where in-ear detection is performed. Optionally, after on-ear detection, method 800 proceeds to box 804, which involves identifying a user. User identification can be performed using any suitable technology (e.g., voice recognition, fingerprint touch sensor, entering a user-specific code via user input, facial recognition via an imaging device, or any other suitable technology). In some cases, the user identity can be inferred based on the connected Bluetooth device (e.g., if the wearable playback device is connected to John Smith's mobile phone, John Smith is identified as the user after on-ear detection). Optionally, the user identity can be a pseudo-identity, where preferences for different profiles are stored but are not matched to a single user account or real-world identity.
[0153] At block 806, the wearable playback device plays back specific generated audio content based on the user's location and optionally the user's identity. The location can be determined using any suitable positioning technology (e.g., exchanging positioning signals (e.g., UWB positioning, WiFi RSSI evaluation, acoustic positioning signals, etc.) with other devices within the environment. In some examples, the wearable playback device may include a GPS component configured to provide absolute location data. As shown, process 800 may branch to the following regarding Figure 8B , Figure 8C and Fig.8D Described methods 820, 830, and / or 840. Additional examples regarding position determination may be found in the attached appendix.
[0154] In some examples, in addition to location information, the generative audio content generated and played back at block 806 is based on certain real-time contextual information (e.g., temperature, light, humidity, time of day, emotional state of the user). If the identity of the user has been determined, the generative audio content may be generated based on one or more user preferences or other input parameters specific to the user (e.g., the user's listening history, favorite artists, etc.). In some cases, the wearable playback device may initiate playback of the generative audio content only when the identified user is authorized to use or otherwise associated with the wearable device.
[0155] At decision box 808, if the location has not been changed, the method returns to box 806 and continues to play back the generated audio content based on the location and the optional user identity. If at decision box 808, the location has been changed (e.g., determined via sensor data), method 800 continues to box 810, where a new location is identified, and in box 812, specific generated audio content is played back based on the new location. For example, moving from the kitchen to the home office within a home environment may cause the generated audio content to automatically convert to a calmer, more focused audio content, while moving to the home gym may cause the generated audio content to convert to a more cheerful, high-energy audio content. The way this conversion occurs may have a significant impact on the user experience. In most cases, a slow / smooth conversion will be preferred, but the conversion may also be more complex (e.g., increasing the sense of space so that the sound feels contained by a given room, or even partially blended with the audio of a nearby user when nearby).
[0156] In some examples, the generative audio content can be generated based at least in part on a theme, mood, and / or comfort landscape associated with a particular location. For example, a room can have a theme (e.g., a forest room, a beach room, a tropical jungle room, a gym), and the generative audio content that is generated and played back is based on the soundscape associated with the room. The generative audio content played back via the wearable playback device can be the same, substantially similar, or otherwise related to the audio content that is typically played back by one or more external devices in the room. In some examples, context-based audio cues can be layered onto the generative audio content associated with the room.
[0157] Figure 8B An example method 820 is shown, wherein, at box 822, audio and / or sound within a current location is detected. For example, a wearable playback device may include one or more microphones (and / or other microphones within the environment may capture sound data and send it to the wearable playback device). At box 824, the detected audio and / or sound may be used as input for generative audio content. External audio played back via a nearby playback device, ambient noise and sound within a user's environment, and / or audio played back near another user via another wearable playback device may be used as input to a generative media module to produce generative audio content that is responsive to, consistent with, and / or otherwise at least partially affected by the audio detected within the environment.
[0158] Figure 8CAn example method 830 for controlling light sources in a manner corresponding to generative audio content is shown. Method 830 begins at box 832, which involves determining available light sources based on the identified location. These can be, for example, internet-connected light bulbs, televisions, and / or other displays that can be remotely controlled. The light sources can be distributed in specific rooms or other locations within the environment. In box 834, the current generative audio content characteristics are mapped to the available light sources. And in box 836, the light sources adjust their corresponding lighting parameters based on the mapped characteristics.
[0159] Characteristics of generative audio content that can be mapped to available light sources may include, for example, mood, tempo, individual notes, chords, etc. Based on the mapped generative audio content characteristics, individual light sources may have corresponding parameters (e.g., color, brightness / intensity, color temperature) adjusted accordingly.
[0160] Fig.8D An example method 840 for detecting a location and determining the room acoustics of a particular listening environment (e.g., an environment in which a listener of a wearable playback device is located) is shown. The method 840 begins at box 842, where a location identifier is detected, and in box 844, the method 840 involves identifying the location based on the detected identifier. This can include using a suitable sensor modality or combination of modalities (e.g., UWB positioning, acoustic positioning, BLE positioning, ultrasonic positioning, motion sensor data (e.g., IMU), wireless RSSI positioning, etc.) to determine the location where the user is currently located. In some examples, the user location can be mapped to predefined rooms or spaces, which can be stored via the media playback system for the purpose of managing synchronized audio playback across user environments.
[0161] In decision block 846, if the room acoustics are to be determined, the process may proceed to Fig. 8AMethod 800 is shown. If the room acoustics are to be determined, method 840 proceeds to decision box 848. If the device is in a room with calibration data (e.g., the cell has performed a calibration process in the room and the calibration data is available to the media playback system), the method proceeds to box 850 to obtain spatial calibration data from a neighboring device in the room (or from another component of the media playback system). If in box 848, the device is not in a room with pre-existing calibration data, then in box 852, the wearable playback device can directly determine the calibration coefficients, for example by performing a calibration process. Examples of suitable calibration processes can be found in commonly owned U.S. Patent No. 9,7906,323, entitled “Playback Device Calibration” and U.S. Patent No. 9,763,018, entitled “Calibration of Audio Playback Devices”, each of which is incorporated herein by reference in its entirety. In various examples, the wearable playback device may access calibration data obtained by another playback device within a particular room or location, or the wearable playback device itself may perform a calibration process to determine calibration coefficients.
[0162] If the wearable device itself performs or is otherwise involved in determining room acoustics, the external playback device can be configured to emit calibration audio (e.g., a tone, sweep, chirp, or other suitable audio output) that can be detected and recorded via one or more microphones of the wearable playback device. Calibration can then be performed using the calibration audio data (e.g., via the wearable playback device, the external playback device, a control device, a remote computing device, or other suitable device or combination of devices) to determine a room acoustic profile that can be applied to the output of the wearable playback device. Optionally, the calibration determination can also be used to calibrate the audio output of other devices in the room (e.g., the external playback device used to emit calibration audio).
[0163] At block 854, based on the device type and the calibration data, the method may determine wearable playback device audio parameters to simulate the acoustics of the identified location. For example, using calibration data from blocks 850 or 852, playback via the wearable playback device may be modified to simulate the acoustics of external playback at the location. Once the audio parameters are determined, method 840 may proceed to Fig. 8A Method 800 is shown.
[0164] Fig. 9An example scenario involving playback of generative audio content via a wearable playback device within a home environment is shown. As shown, a user wearing a wearable playback device can move to various locations 901 to 905 within the home environment. Optionally, the wearable playback device can be wirelessly connected (e.g., via Bluetooth, WiFi, or other means) to a nearby playback device to relay data (e.g., audio content, playback control, timing and / or sensor data, etc.). In the illustrated scenario, at location 901, a user can exercise in a den and listen to a sports soundscape with a brisk rhythm. When the user moves to location 902 in the kitchen, the wearable playback device can automatically switch to playback more relaxed generative audio content suitable for cooking, eating, or socializing. When the user moves to location 903 in the home office, the wearable playback device can automatically switch to playback generative audio content suitable for work and concentration. Next, at location 904 in the TV viewing area, the wearable playback device can switch to receiving audio data from a home theater main playback device to play back audio accompanying video content played back via the TV. Finally, when the user enters the bedroom at location 905, the wearable playback device can automatically switch from playing back TV audio to outputting relaxing generative audio content to facilitate the user's relaxation in preparation for sleep. In these and other examples, the wearable playback device can automatically switch playback by switching from generative audio content to other content sources, or by modulating the generative audio content itself based on the user's location and / or other contextual data.
[0165] c. Playback control between wearable playback device and external playback device
[0166] Fig. 10A An example method 1000 for exchanging playback of generative audio content between a wearable playback device and an external playback device is shown, and Fig. 10B An example scenario involving exchanging playback of generative audio content between a wearable playback device and an external playback device is shown. Method 1000 begins at block 1002, where an exchange is initiated. An exchange may involve moving playback of generative audio output from one device (e.g., a wearable playback device) to one or more nearby target devices (e.g., an external playback device within substantially audible range of the wearable device).
[0167] In block 1006, in an example where the generated audio is exchanged to two or more target devices, a speaker target group coordinator device may be selected among the target devices. The group coordinator may be selected based on parameters such as network connectivity, processing power, available memory, remaining battery power, location, etc. In some examples, the speaker group coordinator is selected based on a predicted direction of a user's path within the environment. For example, in Fig. 10BIn the example of FIG. 5 , if the user is moving from the master bedroom (lower right corner) through the living room toward the kitchen, the playback device in the kitchen can be selected as the group coordinator.
[0168] At block 1008 , the group coordinator may then determine the playback responsibilities of the target device, which may depend on the device location, audio output capabilities (eg, subwoofer versus ultra-portable device), and other suitable parameters.
[0169] Block 1010 involves mapping the generative audio content channel to the playback responsibility determined at block 1008, and in block 1012, playback is switched from the wearable playback device to the target external playback device. In some cases, the switch may involve gradually decreasing the playback volume via the wearable playback device and gradually increasing the playback volume via the external playback device. Alternatively, the switch may involve abruptly terminating playback via the wearable playback device and simultaneously initiating playback via the external playback device.
[0170] Group coordinator selection may involve certain rules, such as no portable devices and / or no subwoofer devices. However, in some examples, it may make sense to use a subwoofer or similar device as a group coordinator because it is likely to remain in the soundscape as the user moves around the home because its low frequency output will be used regardless of the user's location.
[0171] In some examples, the group coordinator is not a target device, but rather another local device (e.g., a local hub device) or a remote computing device (such as a cloud server). In certain examples, the wearable device continues to produce generative audio content and coordinates audio between the exchanged target devices.
[0172] Fig. 10C is a schematic diagram of a system 1020 for generating and playing back generative media content according to aspects of the present disclosure. The system 1020 may be similar to the system 1020 described above with respect to Figure 6 and Figure 7 The described systems 600 and 700, and Fig. 10C Some components omitted from the description may be included in various implementations. Some aspects of the generation of generative media content are omitted, and only the generative media module 211 and the resulting generative media content 616 are shown here. However, in various embodiments, any method or technique described elsewhere herein or known to those skilled in the art may be incorporated into the generation of the generative media content 616.
[0173] In various examples, the generated media content 616 may include multi-channel media content. The generated media content 616 may then be sent to the wearable playback device 102o and / or the group coordinator 1022 for playback via the external playback devices 102a, 102b, and 102c. In the case of a swap action, the generated media content 616 played back via the wearable playback device 102o may be swapped to the group coordinator 1022, so that playback via the wearable playback device 102o is stopped and playback via the external playback devices 102a to 102c is started. A reverse swap may also be performed, where the external generated audio content stops being played back via the external playback devices 102a to 102c and starts being played back via the wearable playback device 102o.
[0174] In some cases, the wearable playback device 102o may include its own onboard generative media module 211 , and switching playback from the wearable playback device 102o to the external playback device involves sending generative audio content produced via the wearable playback device 102o from the wearable playback device 102o to the group coordinator 1022 .
[0175] In some examples, swapping to a group of targeted external playback devices includes sending details or parameters of the generative media module to the target group coordinator 1022. In other examples, the wearable device 102i receives generative audio from the cloud-based generative media module, and swapping the audio from the wearable device 102o to the target devices 102a-102c involves the target group coordinator 1022 requesting the generative media content from the cloud-based generative media module 211.
[0176] Some or all playback devices 102 may be configured to receive one or more input parameters 603. As previously described, input parameters 603 may include any suitable input, such as user input (e.g., user selection via a controller device), sensor data (e.g., the number of people present in a room, background noise levels, time of day, weather data, etc.), or other suitable input parameters. In various examples, input parameters 603 may optionally be provided to playback device 102, and / or may be detected or determined by playback device 102 itself.
[0177] In some examples, to determine specific playback responsibilities and to coordinate synchronized playback between various devices, the group coordinator 1022 can send timing information and / or playback responsibility information to the playback devices 102. Additionally or alternatively, the playback devices 102 themselves can determine their respective playback responsibilities based on the received multi-channel media content together with the input parameters 603.
[0178] Fig. 10DAn example method 1030 for playing back generated audio content via a wearable playback device and an external playback device is shown. At box 1032, method 1030 involves: while continuing to play back the generated audio via the wearable playback device, playing back the generated audio via the external playback device at a first volume level, which can be lower than a final second volume level that has not yet been reached. In box 1034, the wearable playback device can enable a transparent mode, and at box 1036, gradually reduce the wearable playback device volume from a fourth volume level to a third volume level that is less than the fourth volume level. For example, the fourth volume level can be the current playback volume, and the third volume level can be an intermediate volume level between the fourth volume level and silence.
[0179] Block 1038 involves gradually increasing the playback volume of the external playback device from a first volume level to a second volume level greater than the first volume. And in block 1040, playback of the generated audio content via the wearable playback device stops completely. In various examples, gradually decreasing the playback volume of the wearable playback device can occur simultaneously with gradually increasing the playback volume of the external playback device.
[0180] In some examples, when in transparency mode (block 1034), audio played back via the wearable playback device can be modified to reduce the amount of audio played back in a frequency band associated with human speech. This can enable the user to hear speech even when the audio is being played back via the wearable playback device. Additionally or alternatively, any active noise cancellation functionality can be reduced or completely turned off when in transparency mode.
[0181] d. Rules engine for constraining generative media playback
[0182] In some cases, it may be desirable to provide user-level and / or location-level restrictions on the playback of soundscapes or other audio content. Fig.11 An example rule engine for limiting the playback of generative audio content in an environment is shown. As shown, there may be location restrictions that vary from user to user. In this example, user 1 is allowed to play back audio content in some locations (terrace, entrance study, kitchen / dining room, living room, bedroom 1, and bathroom 1), but is restricted to playing back audio content in other locations (bedroom 2, bathroom 2). Similarly, user 2 is allowed to play back audio content in some locations (terrace, bedroom 2, entrance study, kitchen / dining room, living room), but is restricted to playing back audio content in other locations (bathroom 2, bedroom 1, bathroom 1). This restriction may be based on user preferences, device capabilities, user characteristics (e.g., age), and may prohibit users from playing back soundscapes or other audio in certain areas or rooms within an environment or home.
[0183] IV. Conclusion
[0184] The above description discloses, among other things, various example systems, methods, devices, and articles of manufacture, including, among other things, firmware and / or software executed on hardware. It should be understood that these examples are merely illustrative and should not be considered restrictive. For example, it is contemplated that any or all of these firmware, hardware, and / or software aspects or components may be implemented exclusively in hardware, exclusively in software, exclusively in firmware, or in any combination of hardware, software, and / or firmware. Therefore, the examples provided are not the only way to implement these systems, methods, devices, and / or articles of manufacture.
[0185] It should be understood that the sending of information to a specific component, device, and / or system mentioned herein should be understood to include sending information (e.g., messages, requests, responses) to the specific component, device, and / or system indirectly or directly. Therefore, the information sent to the specific component, device, and / or system may pass through any number of intermediate components, devices, and / or systems before reaching its destination. For example, the control device may send the information to the playback device by first sending the information to the computing system (which in turn sends the information to the playback device). In addition, the intermediate components, devices, and / or systems may modify the information. For example, the intermediate components, devices, and / or systems may modify a portion of the information, reformat the information, and / or merge additional information.
[0186] Similarly, the information received from a particular component, device, and / or system mentioned herein should be understood to include receiving information (e.g., messages, requests, responses) indirectly or directly from the particular component, device, and / or system. Therefore, the information received from the particular component, device, and / or system may pass through any number of intermediate components, devices, and / or systems before being received. For example, a control device may receive information indirectly from a playback device by receiving information from a cloud server originating from the playback device. In addition, the intermediate components, devices, and / or systems may modify the information. For example, the intermediate components, devices, and / or systems may modify a portion of the information, reformat the information, and / or merge additional information.
[0187] This specification is presented primarily in terms of illustrative environments, systems, processes, steps, logic blocks, processing, and other symbolic representations that are directly or indirectly similar to the operations of a data processing device coupled to a network. Those skilled in the art typically use these processing descriptions and representations to communicate their work content to other technicians of the field. Various specific details are set forth to provide a thorough understanding of the present disclosure. However, those skilled in the art should understand that the present disclosure can be practiced without specific, specific details. In other instances, well-known methods, processes, components, and circuits are not described to avoid unnecessarily obscuring aspects of the embodiments. Therefore, the scope of the present disclosure is defined by the appended claims, rather than the description of the above embodiments.
[0188] The various methods disclosed herein can be implemented by any device described herein or any other device known now or developed later. The various methods described herein may include one or more operations, functions or actions shown by the blocks. Although each block is shown in a continuous order, these blocks may also be performed in parallel, and / or in a sequence different from the sequence disclosed and described herein. In addition, based on the desired implementation, each block can be combined into fewer blocks, divided into more blocks, and / or removed.
[0189] In addition, for the methods disclosed herein, the flowchart shows some examples of possible implementation functions and operations. In this regard, each box can represent a module, segment or part of a program code, and the program code includes one or more instructions executable by one or more processors for implementing a specific logical function or step in the process. The program code can be stored on any type of computer-readable medium, for example, a storage device including a disk or hard drive. The computer-readable medium may include a non-transitory computer-readable medium, for example, a tangible non-transitory computer-readable medium for storing data for a short time, such as a register memory, a processor cache, and a random access memory (RAM). The computer-readable medium may also include a non-transitory medium, for example, an auxiliary storage or a persistent long-term storage device, such as a read-only memory (ROM), an optical disk or disk, a compact disk read-only memory (CD-ROM), etc. The computer-readable medium may also be any other volatile or non-volatile storage system. The computer-readable medium may be considered to be a computer-readable storage medium, such as a tangible storage device. In addition, for these methods and other processes and methods disclosed herein, each box in the accompanying drawings may represent a circuit connected to perform a specific logical function in the process.
[0190] When any of the following claims is understood to cover a pure software and / or firmware implementation, at least one element in at least one example is expressly defined herein to include a non-transitory tangible medium storing the software and / or firmware, such as a memory, DVD, CD, Blu-ray, etc.
[0191] 5. Examples
[0192] For example, the disclosed technology is illustrated according to the various examples described below. For convenience, various examples of embodiments of the disclosed technology are described as numbered examples. These examples are provided as examples and do not limit the disclosed technology. Note that any one of the dependent examples can be combined in any combination and placed in the corresponding independent examples. Other examples can be presented in a similar manner.
[0193] Example 1: A wearable playback device comprising: one or more audio transducers; a network interface; one or more processors; and one or more tangible, non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the wearable playback device to perform operations comprising: detecting whether a user is wearing the wearable playback device; obtaining one or more input parameters via a network interface of the wearable playback device; after detecting that the user is wearing the wearable playback device, generating generative media content based at least in part on the one or more input parameters; playing back the generative media content via the one or more audio transducers; detecting that the user is no longer wearing the wearable playback device; and after detecting that the user is no longer wearing the wearable playback device, ceasing to play back the generative media content via the one or more audio transducers.
[0194] Example 2. A wearable playback device according to any of the examples in the present invention, wherein one or more input parameters include position information, and when the position information changes over time, the generated media content changes at least in part based on the changing position information.
[0195] Example 3. A wearable playback device according to any of the examples described herein, wherein the location information includes an identification of a room or space within the user's environment.
[0196] Example 4. A wearable playback device according to any of the examples in this document, wherein one or more input parameters include user identity information.
[0197] Example 5. A wearable playback device according to any of the examples described in the present invention, wherein creating the generative media content includes: (i) sending at least one of the one or more input parameters to a second playback device via a network interface of the wearable playback device, and (ii) receiving the generative media content generated by the second playback device via the network interface of the wearable playback device.
[0198] Example 6. A wearable playback device according to any of the examples described herein, wherein the one or more input parameters include one or more first input parameters, and wherein creating the generative media content includes: receiving at least one second input parameter from a network interface of a second playback device via a network interface of the wearable playback device.
[0199] Example 4. A wearable playback device according to any of the examples herein, wherein the one or more input parameters include sound data obtained via a microphone of the wearable playback device.
[0200] Example 8. A wearable playback device according to any of the examples herein, wherein the sound data comprises a seed for a generative media content engine.
[0201] Example 9. A wearable playback device according to any of the examples described herein, wherein the operations also include: while playing back the generated media content, adjusting (turning on / off, changing color, etc.) one or more light sources in the user's environment based at least in part on one or more input parameters.
[0202] Example 10. A wearable playback device according to any of the examples described herein, wherein one or more input parameters include spatial calibration information of the user's position, and playing back the generated media content includes playing back audio via the wearable playback device using the spatial calibration information.
[0203] Example 11. A wearable playback device according to any of the examples described herein, wherein the operations further include: when playing back the generated media content via the wearable playback device, initiating synchronized playback of at least a portion of the generated media content via a second audio playback device.
[0204] Example 12. A wearable playback device according to any of the examples herein, wherein the operations further include: after initiating synchronized playback of the generative media content via the second audio playback device, reducing the playback volume of the wearable playback device.
[0205] Example 13. A wearable playback device as described in any of the examples herein, wherein initiating synchronized playback of the generative media content via the second audio playback device is based at least in part on a user instruction.
[0206] Example 14. A wearable playback device as described in any of the examples herein, wherein initiating synchronized playback of the generative media content via the second audio playback device is based at least in part on a proximity determination between the wearable playback device and the second audio playback device.
[0207] Example 15. A wearable playback device according to any of the examples described in the present invention, wherein the operations also include: detecting a wireless audio source connection when playing back the generated media content via the wearable playback device; and after detecting the wireless audio source connection, stopping playback of the generated media content and initiating playback of audio data from the wireless audio source.
[0208] Example 16. A wearable playback device according to any of the examples described herein, wherein the operations further include: determining that the wireless audio source is disconnected; and after determining that the wireless audio source is disconnected, initiating playback of the generated media content.
[0209] Example 17. A wearable playback device as described in any of the examples herein, wherein the generated media content is based at least in part on audio data from a wireless audio source.
[0210] Example 18. A wearable playback device according to any of the examples described in the present invention, wherein the one or more input parameters include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, temperature, breathing rate, brain waves)); networked device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); wearable playback device capability data (e.g., number and type of sensors, output power); wearable playback device state (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is bound to another playback device); or user data (e.g., user identity, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, breathing rate, brain activity, voice speech characteristics), user mood data).
[0211] Example 19. A method comprising: detecting whether a user is wearing a wearable playback device; obtaining one or more input parameters via a network interface of the wearable playback device; after detecting that the user is wearing the wearable playback device, generating generative media content based at least in part on the one or more input parameters; playing back the generative media content via the wearable playback device; detecting that the user is no longer wearing the wearable playback device; and after detecting that the user is no longer wearing the wearable playback device, ceasing to play back the generative media content via the wearable playback device.
[0212] Example 20. A method according to any of the examples in the present invention, wherein the one or more input parameters include position information, and when the position information changes over time, the generated media content changes at least in part based on the changing position information.
[0213] Example 21. A method according to any of the examples in the present document, wherein the location information includes an identification of a room or space within the user's environment.
[0214] Example 22. A method according to any of the examples in the present document, wherein one or more input parameters include user identity information.
[0215] Example 23. A method according to any of the examples in the present invention, wherein creating generative media content includes: (i) sending at least one of the one or more input parameters to a second playback device via a network interface of the wearable playback device, and (ii) receiving the generative media content generated by the second playback device via the network interface of the wearable playback device.
[0216] Example 24. A method according to any of the examples in the present document, wherein the one or more input parameters include one or more first input parameters, and wherein creating the generative media content includes: receiving at least one second input parameter from a network interface of a second playback device via a network interface of the wearable playback device.
[0217] Example 25. A method according to any of the examples in the present document, wherein one or more input parameters include sound data obtained via a microphone of the wearable playback device.
[0218] Example 26. A method according to any of the examples in the present document, wherein the sound data comprises a seed for a generative media content engine.
[0219] Example 27. The method according to any of the examples in the present document also includes: while playing back the generated media content, adjusting (turning on / off, changing color, etc.) one or more light sources in the user's environment based at least in part on one or more input parameters.
[0220] Example 28. A method according to any of the examples in the present document, wherein one or more input parameters include spatial calibration information of the user's position, and playing back the generated media content includes playing back audio via a wearable playback device using the spatial calibration information.
[0221] Example 29. The method of any of the examples herein further comprises: while the generative media content is being played back via the wearable playback device, initiating synchronized playback of at least a portion of the generative media content via a second audio playback device.
[0222] Example 30. The method of any of the examples herein further includes: after initiating synchronized playback of the generative media content via the second audio playback device, reducing the playback volume of the wearable playback device.
[0223] Example 31. A method as in any of the examples herein, wherein initiating synchronized playback of the generative media content via the second audio playback device is based at least in part on a user instruction.
[0224] Example 32. A method as in any of the examples herein, wherein initiating synchronized playback of the generative media content via the second audio playback device is based at least in part on a proximity determination between the wearable playback device and the second audio playback device.
[0225] Example 33. The method described in any of the examples in the present invention further includes: detecting a wireless audio source connection when playing back the generated media content via a wearable playback device; and after detecting the wireless audio source connection, stopping playback of the generated media content and initiating playback of audio data from the wireless audio source.
[0226] Example 34. The method of any of the examples described herein further includes: determining that the wireless audio source is disconnected; and after determining that the wireless audio source is disconnected, initiating playback of the generated media content.
[0227] Example 35. A method as described in any of the examples herein, wherein the generated media content is based at least in part on audio data from a wireless audio source.
[0228] Example 36. A method according to any of the examples in the examples herein, wherein one or more input parameters include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, temperature, breathing rate, brain waves)); networked device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); wearable playback device capability data (e.g., number and type of sensors, output power); wearable playback device status (e.g., device temperature, battery charge, current audio playback, playback device location, whether the playback device is bound to another playback device); or user data (e.g., user identity, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, breathing rate, brain activity, voice speech characteristics), user mood data).
[0229] Example 37. One or more tangible, non-transitory computer-readable media storing instructions that, when executed by one or more processors of a wearable playback device, cause the wearable playback device to perform operations comprising: detecting whether a user is wearing the wearable playback device; obtaining one or more input parameters via a network interface of the wearable playback device; after detecting that the user is wearing the wearable playback device, generating generative media content based at least in part on the one or more input parameters; playing back the generative media content via one or more audio transducers of the wearable playback device; detecting that the user is no longer wearing the wearable playback device; and after detecting that the user is no longer wearing the wearable playback device, ceasing to play back the generative media content via the one or more audio transducers.
[0230] Example 38. One or more computer-readable media according to any of the examples described herein, wherein one or more input parameters include position information, and when the position information changes over time, the generated media content changes at least in part based on the changing position information.
[0231] Example 39. One or more computer-readable media according to any of the examples herein, wherein the location information includes an identification of a room or space within the user's environment.
[0232] Example 40. One or more computer-readable media according to any of the examples described herein, wherein the one or more input parameters include user identity information.
[0233] Example 41. One or more computer-readable media according to any of the examples described herein, wherein creating generative media content includes: (i) sending at least one of the one or more input parameters to a second playback device via a network interface of the wearable playback device, and (ii) receiving the generative media content generated by the second playback device via the network interface of the wearable playback device.
[0234] Example 42. One or more computer-readable media according to any of the examples described herein, wherein the one or more input parameters include one or more first input parameters, and wherein creating the generative media content includes: receiving at least one second input parameter from a network interface of a second playback device via a network interface of the wearable playback device.
[0235] Example 43. One or more computer-readable media according to any of the examples herein, wherein the one or more input parameters include sound data obtained via a microphone of the wearable playback device.
[0236] Example 44. One or more computer-readable media as described in any of the examples herein, wherein the sound data comprises a seed for a generative media content engine.
[0237] Example 45. One or more computer-readable media according to any of the examples described herein, wherein the operations also include: while playing back the generated media content, adjusting (turning on / off, changing color, etc.) one or more light sources in the user's environment based at least in part on one or more input parameters.
[0238] Example 46. One or more computer-readable media according to any of the examples described herein, wherein one or more input parameters include spatial calibration information of the user's position, and playing back the generated media content includes playing back audio via a wearable playback device using the spatial calibration information.
[0239] Example 47. One or more computer-readable media according to any of the examples herein, wherein the operations further include: when playing back the generative media content via the wearable playback device, initiating synchronized playback of at least a portion of the generative media content via a second audio playback device.
[0240] Example 48. One or more computer-readable media according to any of the examples described herein, wherein the operations further include: after initiating synchronized playback of the generative media content via the second audio playback device, reducing the playback volume of the wearable playback device.
[0241] Example 49. One or more computer-readable media as described in any of the examples herein, wherein initiating synchronized playback of the generative media content via the second audio playback device is based at least in part on a user instruction.
[0242] Example 50. One or more computer-readable media according to any of the examples herein, wherein initiating synchronized playback of the generative media content via the second audio playback device is determined at least in part based on proximity between the wearable playback device and the second audio playback device.
[0243] Example 51. One or more computer-readable media according to any of the examples described herein, wherein the operations further include: detecting a wireless audio source connection when playing back the generated media content via a wearable playback device; and after detecting the wireless audio source connection, stopping playback of the generated media content and initiating playback of audio data from the wireless audio source.
[0244] Example 52. One or more computer-readable media according to any of the examples herein, wherein the operations further include: determining that the wireless audio source is disconnected; and after determining that the wireless audio source is disconnected, initiating playback of the generated media content.
[0245] Example 53. One or more computer-readable media as described in any of the examples herein, wherein the generated media content is based at least in part on audio data from a wireless audio source.
[0246] Example 54. One or more computer-readable media according to any of the examples in the present invention, wherein the one or more input parameters include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, temperature, breathing rate, brain waves)); networked device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); wearable playback device capability data (e.g., number and type of sensors, output power); wearable playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is bound to another playback device); or user data (e.g., user identity, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, breathing rate, brain activity, voice speech characteristics), user mood data).
Claims
1. A method comprising: obtaining one or more input parameters via a network interface of the wearable playback device; causing generation of generative media content based at least in part on the one or more input parameters; After detecting that the user is wearing the wearable playback device, playing back the generated media content via the wearable playback device; as well as When the user no longer wears the wearable playback device, playback of the generated media content via the wearable playback device is stopped.
2. The method according to claim 1, wherein: The one or more input parameters include position information, and wherein, when the position information changes over time, the generative media content changes based at least in part on the changing position information.
3. The method according to claim 2, wherein: The location information includes an identification of a room or space within the user's environment.
4. A method according to any preceding claim, wherein: The one or more input parameters include user identity information.
5. A method according to any preceding claim, wherein: Such that creating the generative media content comprises: (i) sending at least one of the one or more input parameters to a second playback device via the network interface of the wearable playback device, and (ii) receiving, via the network interface of the wearable playback device, the generative media content generated via the second playback device.
6. A method according to any preceding claim, wherein: The one or more input parameters include one or more first input parameters, and wherein causing creation of the generative media content includes receiving at least one second input parameter from a network interface of a second playback device via the network interface of the wearable playback device.
7. A method according to any preceding claim, wherein: The one or more input parameters include sound data obtained via a microphone of the wearable playback device.
8. The method according to claim 7, wherein: The sound data comprises a seed for a generative media content engine.
9. The method according to any preceding claim, further comprising: While playing back the generative media content, one or more light sources in the user's environment are caused to be adjusted (turned on / off, changed color, etc.) based at least in part on the one or more input parameters.
10. A method according to any preceding claim, wherein: The one or more input parameters include spatial calibration information of the user's position, and wherein playing back the generative media content includes playing back audio via the wearable playback device using the spatial calibration information.
11. The method according to any preceding claim, further comprising: While the generative media content is being played back via the wearable playback device, synchronized playback of at least a portion of the generative media content via a second audio playback device is initiated.
12. The method according to claim 11, further comprising: After initiating synchronized playback of the generative media content via the second audio playback device, a playback volume of the wearable playback device is reduced.
13. The method according to claim 11 or 12, wherein: Initiating synchronized playback of the generative media content via the second audio playback device is based at least in part on a user instruction.
14. The method according to any one of claims 11 to 13, wherein: Initiating synchronized playback of the generative media content via the second audio playback device is based at least in part on a proximity determination between the wearable playback device and the second audio playback device.
15. The method according to any preceding claim, further comprising: detecting a wireless audio source connection while playing back the generative media content via the wearable playback device; as well as After detecting the wireless audio source connection, playback of the generative media content is stopped and playback of audio data from the wireless audio source is initiated.
16. The method according to claim 15, further comprising: determining that the wireless audio source is disconnected; as well as After determining that the wireless audio source is disconnected, playback of the generative media content is initiated.
17. The method according to claim 15 or 16, wherein: The generative media content is based at least in part on audio data from the wireless audio source.
18. A method according to any preceding claim, wherein: The one or more input parameters include one or more of the following: Physiological sensor data; Sensor data from connected devices; Environmental data; Wearable playback device capability data; Wearable playback device status; or User data.
19. One or more tangible non-transitory computer-readable media storing instructions that, when executed by one or more processors of a wearable playback device, cause the wearable playback device to perform a method according to any preceding claim.
20. A wearable playback device comprising: one or more audio transducers; Network interface; one or more processors; as well as One or more tangible, non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the wearable playback device to perform the method of one of claims 1 to 19.
Citation Information
Patent Citations
Voice Control of a Media Playback System
US20170242653A1
Room Association Based on Name
US20180107446A1
System and method for synchronizing operations among a plurality of independently clocked digital data processing devices
US8234395B2
Controlling and manipulating groupings in a multi-zone media system
US8483853B1
Calibration of audio playback devices
US9763018B1