Generating digital media based on blockchain data
By combining generative media content systems with blockchain data, the limitations of existing technologies in multi-device synchronous playback and dynamic audio experience are overcome, enabling personalized and seamless audio synchronization playback.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SONOS INC
- Filing Date
- 2023-05-09
- Publication Date
- 2026-05-29
Smart Images

Figure CN120418861B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. Patent Application No. 63 / 364,931, filed May 18, 2022, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to consumer products, and more specifically to methods, systems, products, features, services and other elements relating to media playback or some aspect thereof. Background Technology
[0004] Options for accessing and listening to digital audio from external speakers were limited until Sonos began developing a new playback system in 2002. Sonos then filed one of its first patent applications in 2003 entitled "Method for Synchronizing Audio Playback between Multiple Networked Devices," and began selling its first media playback system in 2005. Sonos' wireless home audio system allows people to experience music from many sources via one or more networked playback devices. Through a software control application installed on a controller (e.g., a smartphone, tablet, computer, voice input device), people can play whatever they want in any room with a networked playback device. Media content (e.g., songs, podcasts, video audio) can be streamed to the playback device, allowing each room with a playback device to play different media content. Furthermore, rooms can be grouped together to synchronously play the same media content, and / or the same media content can be heard synchronously in all rooms. Attached Figure Description
[0005] The features, aspects, and advantages of the disclosed technology can be better understood by referring to the following specification, appended claims, and drawings, as listed below. Those skilled in the art will understand that the features shown in the drawings are for illustrative purposes, and variations in the inclusion of different and / or additional features and their arrangement are possible.
[0006] Figure 1A It is a partial cross-sectional view of an environment having a media playback system configured according to aspects of the disclosed technology.
[0007] Figure 1B yes Figure 1A A schematic diagram of a media playback system and one or more networks.
[0008] Figure 1CThis is a block diagram of a playback device.
[0009] Figure 1D This is a block diagram of a playback device.
[0010] Figure 1E This is a block diagram of the binding playback device.
[0011] Figure 1F This is a block diagram of a network microphone device.
[0012] Figure 1G This is a block diagram of a playback device.
[0013] Figure 1H This is a partial schematic diagram of the control equipment.
[0014] Figures 1I to 1L A schematic diagram of the corresponding media playback system area is shown.
[0015] Figure 1M A schematic diagram of the media playback system area is shown.
[0016] Figure 2 This is a functional block diagram of a system for playing back generative media content, based on an example of this technology.
[0017] Figure 3 This is a functional block diagram of generative media modules based on various aspects of this technology.
[0018] Figure 4 This is an example architecture for storing and retrieving generative media content based on various aspects of this technology.
[0019] Figure 5 This is a functional block diagram illustrating data exchange in a system for playing back generated media content according to various aspects of the present technology.
[0020] Figure 6 This is a schematic diagram of a distributed generative media playback system based on various aspects of this technology.
[0021] Figure 7 This is a schematic diagram of another example of a distributed generative media playback system based on various aspects of this technology.
[0022] Figures 8 to 13 This is a flowchart of a method for playing back generated media content based on various aspects of this technology.
[0023] The accompanying drawings are for illustrative purposes only, but those skilled in the art will understand that the techniques disclosed herein are not limited to the arrangements and / or means shown in the drawings. Detailed Implementation
[0024] I. Overview
[0025] Generative media content is content that is dynamically synthesized, created, and / or modified based on algorithms, whether implemented in software or a physical model. Generative media content can change over time based solely on algorithms or in combination with contextual data (e.g., user sensor data, environmental sensor data, event data). In various examples, such generative media content can include generative audio (e.g., music, ambient soundscapes, etc.), generative visual images (e.g., lighting, abstract visual designs that dynamically change shape, color, etc.), generative odors, generative haptic outputs (e.g., vibration, haptic outputs, etc.), or any other suitable media content or combinations thereof. As illustrated elsewhere in this document, generative media can be created at least in part via algorithms and / or non-human systems that utilize rule-based computation to generate novel media content.
[0026] Because generative media content can be modified dynamically in real time, it can achieve unique user experiences that cannot be obtained using conventional media playback with pre-recorded content. For example, generative audio can be infinitely and / or dynamically changing as the algorithm's input changes (e.g., input parameters associated with user input, sensor data, media source data, or any other suitable input data). In some examples, generative audio can be used to guide a user's emotions to a desired emotional state, where one or more characteristics of the generative audio change in response to real-time measurements reflecting the user's emotional state. As used in examples of this technology, the system can provide generative audio based on the user's current and / or desired emotional state, based on the user's activity level, based on the number of users present in the environment, or any other suitable input parameter.
[0027] As another example, generative audio can be created and / or modified based on one or more inputs, such as a user's location or activity, the number of users present in the room, the time of day, or any other input (e.g., determined by one or more sensors or by user input). For instance, a media playback system can automatically generate generative audio content suitable for focused study or work when a single user is sitting calmly at her desk, while the same system can automatically generate generative audio suitable for a social gathering or dance party when multiple users in the room are active and excited. In various examples, audio characteristics that can be dynamically modified to generate generative audio can include the selection of audio samples or clips, tempo, bass / treble / midrange volume, spatial filtering of the audio output, or any other suitable audio characteristics. Audio characteristics can be altered by using different pitches or sounds, the timing of pitches or sounds, and / or audio samples that may have the desired quality. In some cases, characteristics can also be altered by filtering or modulating the playback of content, such as equalization, phase, or reverb / delay. During the listening experience, the audio characteristics of generative music can be altered based on various inputs (such as time of day, geographic location, weather) or various user inputs (such as inferred emotions, level of collective activity) or physiological inputs (such as heart rate).
[0028] In some cases, generative media content, such as soundscapes, can be associated with data stored in one or more blockchain databases and / or layers. For example, non-fungible tokens (NFTs) typically consist of data files stored on a blockchain that are associated with digital or tangible assets, such as songs or albums, visual artworks, or literary works. While an NFT can refer to a digital artwork, it typically does not include the digital artwork itself because the required file size may be intractable (or too costly) to store on a blockchain layer. Instead, an NFT may include metadata (e.g., a URL or other locator) pointing to the location of the digital artwork and / or how to access it. In various examples, NFTs or other data stored via blockchain or other distributed ledger technologies can be used for media content creation. For example, in some examples, NFTs serve as input to generative media engines, resulting in generative media content with properties that depend at least in part on specific NFTs or other blockchain data.
[0029] While some of the examples described herein may relate to functions performed by a given actor such as a “user,” a “listener,” and / or other entity, it should be understood that this is for illustrative purposes only. Unless the language of the claims themselves explicitly requires it, the claims should not be construed as requiring any such example actor to perform an action.
[0030] In the accompanying drawings, the same reference numerals generally identify similar and / or identical elements. To facilitate discussion of any particular element, one or more of the most significant bits in the reference numerals refer to the drawing in which that element was first introduced. For example, element 110a in reference to... Figure 1A This is the first time it has been introduced and discussed. Many details, dimensions, angles, and other features shown in the accompanying drawings are merely illustrative of specific examples of the disclosed techniques. Therefore, other examples may have different details, dimensions, angles, and features without departing from the spirit or scope of this disclosure. Furthermore, those skilled in the art will understand that other examples of various disclosed techniques can be practiced without the several details described below.
[0031] II. Suitable operating environment
[0032] Figure 1A This is a partial cross-sectional view of a media playback system 100 distributed in an environment 101 (e.g., a house). The media playback system 100 includes one or more playback devices 110 (each identified as playback devices 110a to 110n), one or more network microphone devices (“NMD”), 120 (each identified as NMD 120a to NMD 120c), and one or more control devices 130 (each identified as control devices 130a and 130b).
[0033] As used herein, the term "playback device" can generally refer to a network device configured to receive, process, and output data to a media playback system. For example, a playback device can be a network device that receives and processes audio content. In some examples, a playback device includes one or more transducers or speakers powered by one or more amplifiers. However, in other examples, a playback device includes one of a speaker and an amplifier (or neither). For example, a playback device may include one or more amplifiers configured to drive one or more speakers external to the playback device via corresponding wiring or cables.
[0034] Furthermore, as used herein, the term NMD (i.e., "network microphone device") can generally refer to a network device configured for audio detection. In some examples, the NMD is a standalone device configured primarily for audio detection. In other examples, the NMD is incorporated into a playback device (or vice versa).
[0035] The term "control device" can generally refer to a network device configured to perform functions related to facilitating user access, control, and / or configuration of the media playback system 100.
[0036] Each of the playback devices 110 is configured to receive audio signals or data from one or more media sources (e.g., one or more remote servers, one or more local devices) and play back the received audio signals or data as sound. One or more NMDs 120 are configured to receive spoken commands, and one or more control devices 130 are configured to receive user input. In response to received spoken commands and / or user input, the media playback system 100 can play back audio via one or more playback devices 110. In some examples, the playback devices 110 are configured to begin playback of media content in response to a trigger. For example, one or more playback devices 110 may be configured to play back a morning playlist when an associated trigger condition is detected (e.g., the presence of a user in the kitchen, detection of coffee machine operation). In some examples, for instance, the media playback system 100 is configured to synchronously play back audio from a first playback device (e.g., playback device 110a) with a second playback device (e.g., playback device 100b). References below... Figures 1B to 1H The interaction between the playback device 110, NMD 120 and / or control device 130 of the media playback system 100 configured according to various examples of this disclosure is described in more detail.
[0037] exist Figure 1A In the illustrated example, environment 101 includes a home with multiple rooms, spaces, and / or playback zones, comprising (clockwise from the top left corner) a master bathroom 101a, a master bedroom 101b, a secondary bedroom 101c, a family room or study 101d, an office 101e, a living room 101f, a dining room 101g, a kitchen 101h, and an outdoor terrace 101i. While certain embodiments and examples are described below in the context of a home environment, the techniques described herein can be implemented in other types of environments. In some examples, for instance, media playback system 100 may be implemented in one or more commercial environments (e.g., restaurants, shopping malls, airports, hotels, retail stores, or other shops), one or more vehicles (e.g., SUVs, buses, cars, ships, airplanes), multiple environments (e.g., a combination of a home environment and a vehicle environment), and / or another suitable environment that may require multi-zone audio.
[0038] The media playback system 100 may include one or more playback zones, some of which may correspond to rooms in environment 101. The media playback system 100 may have one or more playback zones initially, and additional zones may be added or removed to form, for example... Figure 1AThe configuration is shown. Each zone can be named according to different rooms or spaces (e.g., office 101e, master bathroom 101a, master bedroom 101b, secondary bedroom 101c, kitchen 101h, dining room 101g, living room 101f, and / or outdoor terrace 101i). In some aspects, a single playback zone may include multiple rooms or spaces. In other aspects, a single room or space may include multiple playback zones.
[0039] exist Figure 1A In the example shown, the master bathroom 101a, secondary bedroom 101c, office 101e, living room 101f, dining room 101g, kitchen 101h, and outdoor terrace 101i each include a playback device 110, and the master bedroom 101b and study 101d include multiple playback devices 110. In the master bedroom 101b, playback devices 110l and 110m can be configured, for example, as individual playback devices among playback devices 110, as a bundled playback zone, as a merged playback device, and / or any combination thereof, to synchronously play back audio content. Similarly, in the study 101d, playback devices 110h to 110j can be configured, for example, as individual playback devices among playback devices 110, as one or more bundled playback devices, and / or as one or more merged playback devices to synchronously play back audio content. See below for reference. Figure 1B and Figure 1E Describe additional details regarding bundled playback devices and merged playback devices.
[0040] In some respects, one or more playback zones in environment 101 may be playing different audio content. For example, a user may be barbecuing on terrace 101i and listening to hip-hop music being played by playback device 110c, while another user is preparing food in kitchen 101h and listening to classical music played by playback device 110b. In another example, a playback zone may be playing the same audio content synchronously with another playback zone. For example, a user may be in office 101e listening to playback device 110f playing the same hip-hop music that playback device 110c is playing on terrace 101i. In some respects, the synchronous playback of hip-hop music by playback devices 110c and 110f makes the user perceive that the audio content is playing seamlessly (or at least substantially seamlessly) as it moves between different playback zones. Additional details regarding audio playback synchronization between playback devices and / or zones can be found, for example, in U.S. Patent No. 8,234,395 entitled “System and method for synchronizing operations among aplurality of independently clocked digital data processing devices,” the entire contents of which are incorporated herein by reference.
[0041] a. A suitable media playback system
[0042] Figure 1B This is a schematic diagram of the media playback system 100 and the cloud network 102. For ease of explanation, from... Figure 1B Some devices of the media playback system 100 and the cloud network 102 are omitted. One or more communication links 103 (hereinafter referred to as "link 103") communicatively couple the media playback system 100 and the cloud network 102.
[0043] Link 103 may include, for example, one or more wired networks, one or more wireless networks, one or more wide area networks (WANs), one or more local area networks (LANs), one or more personal area networks (PANs), one or more telecommunications networks (e.g., one or more Global System for Mobile Communications (GSM) networks, Code Division Multiple Access (CDMA) networks, Long Term Evolution (LTE) networks, 5G communication networks, and / or other suitable data transmission protocol networks). Cloud network 102 is configured to deliver media content (e.g., audio content, video content, photos, social media content) to media playback system 100 in response to a request sent from media playback system 100 via link 103. In some examples, cloud network 102 is also configured to receive data (e.g., voice input data) from media playback system 100 and accordingly send commands and / or media content to media playback system 100.
[0044] Cloud network 102 includes computing devices 106 (identified as first computing device 106a, second computing device 106b, and third computing device 106c, respectively). Computing devices 106 may include various computers or servers, such as media streaming service servers storing audio and / or other media content, voice service servers, social media servers, media playback system control servers, etc. In some examples, one or more computing devices 106 include modules of a single computer or server. In some examples, one or more computing devices 106 include one or more modules, computers, and / or servers. Furthermore, although cloud network 102 is described in the context of a single cloud network, in some examples, cloud network 102 includes multiple cloud networks that include communication-coupled computing devices. Additionally, although cloud network 102 is described in... Figure 1B The cloud network 102 is shown as having three computing devices 106, but in some examples, the cloud network 102 includes fewer (or more) three computing devices 106.
[0045] Media playback system 100 is configured to receive media content from network 102 via link 103. The received media content may include, for example, a Uniform Resource Identifier (URI) and / or a Uniform Resource Locator (URL). For example, in some examples, media playback system 100 may stream, download, or otherwise obtain data from a URI or URL corresponding to the received media content. Network 104 communicatively couples link 103 to at least a portion of the devices of media playback system 100 (e.g., one or more of playback device 110, NMD 120, and / or control device 130). Network 104 may include, for example, a wireless network (e.g., a WiFi network, Bluetooth, Z-Wave network, ZigBee, and / or other suitable wireless communication protocol network) and / or a wired network (e.g., including Ethernet, Universal Serial Bus (USB), and / or other suitable wired communication networks). As will be understood by those skilled in the art, as used herein, “WiFi” can refer to several different communication protocols that transmit at 2.4 GHz, 5 GHz and / or other suitable frequencies, including, for example, IEEE 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.11ac, 802.11ad, 802.11af, 802.11ah, 802.11ai, 802.11aj, 802.11aq, 802.11ax, 802.11ay, 802.15, etc.
[0046] In some examples, network 104 includes a dedicated communication network used by media playback system 100 to send messages between devices and / or send and receive media content from media content sources (e.g., one or more computing devices 106). In some examples, network 104 is configured to be accessible only by devices within media playback system 100, thereby reducing interference and competition with other home appliances. However, in other examples, network 104 includes an existing home communication network (e.g., a home Wi-Fi network). In some examples, link 103 and network 104 include one or more of the same networks. In some aspects, for example, link 103 and network 104 include telecommunications networks (e.g., LTE networks, 5G networks). Furthermore, in some examples, media playback system 100 is implemented without network 104, and devices including media playback system 100 can communicate with each other, for example, via one or more direct connections, PANs, telecommunications networks, and / or other suitable communication links.
[0047] In some examples, audio content sources may be added or removed periodically in media playback system 100. In some examples, for instance, media playback system 100 performs an indexing of media items when one or more media content sources are updated, added to, and / or removed from media playback system 100. Media playback system 100 may scan some or all folders and / or directories accessible by playback device 110 for identifiable media items and generate or update a media content database including metadata (e.g., title, artist, album, track length) and other associated information (e.g., URI, URL) for each identifiable media item found. In some examples, for instance, the media content database is stored on one or more of playback device 110, NMD 120, and / or control device 130.
[0048] exist Figure 1B In the illustrated example, playback devices 110l and 110m comprise group 107a. Playback devices 110l and 110m may be located in different rooms of a home and are grouped together in group 107a on a temporary or permanent basis based on user input received at control device 130a and / or another control device 130 in media playback system 100. When arranged in group 107a, playback devices 110l and 110m can be configured to synchronously play back the same or similar audio content from one or more audio content sources. In some examples, for instance, group 107a includes a binding area where playback devices 110l and 110m each include the left and right audio channels of multi-channel audio content, respectively, thereby producing or enhancing the stereo effect of the audio content. In some examples, group 107a includes an additional playback device 110. However, in other examples, the media playback system 100 omits other grouping arrangements of group 107a and / or playback device 110.
[0049] The media playback system 100 includes NMD 120a and NMD 120d, each NMD including one or more microphones configured to receive voice output from a user. Figure 1BIn the illustrated example, NMD 120a is a standalone device, while NMD 120d is integrated into playback device 110n. For example, NMD 120a is configured to receive voice input 121 from user 123. In some examples, NMD 120a sends data associated with the received voice input 121 to a Voice Assistant Service (VAS) configured to (i) process the received voice input data and (ii) send the corresponding command to media playback system 100. In some aspects, for example, computing device 106c includes one or more modules and / or servers of a VAS (e.g., a VAS operated by one or more of SONOS®, AMAZON®, GOOGLE®, APPLE®, MICROSOFT®). Computing device 106c can receive voice input data from NMD 120a via network 104 and link 103. In response to receiving voice input data, computing device 106c processes the voice input data (i.e., "play Hey Jude by the Beatles") and determines that the processed voice input includes a command to play the song (e.g., "Hey Jude"). Computing device 106c then sends a command to media playback system 100 to play "Hey Jude" by the Beatles on one or more playback devices 110 from a suitable media service (e.g., via one or more computing devices 106).
[0050] b. Suitable playback equipment
[0051] Figure 1CThis is a block diagram of a playback device 110a including input / output 111. Input / output 111 may include analog I / O 111a (e.g., one or more cabling, cables, and / or other suitable communication links configured to carry analog signals) and / or digital I / O 111b (e.g., one or more cabling, cables, or other suitable communication links configured to carry digital signals). In some examples, analog I / O 111a is an audio line input connection, including, for example, an automatically detected 3.5mm audio line input connection. In some examples, digital I / O 111b includes a Sony / Philips Digital Interface Format (S / PDIF) communication interface and / or cable and / or Toshiba Link (TOSLINK) cable. In some examples, digital I / O 111b includes a High Definition Multimedia Interface (HDMI) interface and / or cable. In some examples, digital I / O 111b includes one or more wireless communication links, including, for example, radio frequency (RF), infrared, Wi-Fi, Bluetooth, or another suitable communication protocol. In some examples, analog I / O 111a and digital I / O 111b include interfaces (e.g., ports, plugs, jacks) that are configured to receive connectors for cables transmitting analog and digital signals, respectively, without necessarily including the cables.
[0052] Playback device 110a may receive media content (e.g., audio content including music and / or other sounds) from local audio source 105 via input / output 111 (e.g., cable, cabling, PAN, Bluetooth connection, self-organizing wired or wireless communication network and / or other suitable communication link). Local audio source 105 may include, for example, a mobile device (e.g., smartphone, tablet, laptop) or another suitable audio component (e.g., television, desktop computer, amplifier, phonograph, Blu-ray player, storage for storing digital media files). In some aspects, local audio source 105 includes a local music library on a smartphone, computer, network attached storage (NAS), and / or another suitable device configured to store media files. In some examples, one or more of playback device 110, NMD 120, and / or control device 130 include local audio source 105. However, in other examples, the media playback system omits local audio source 105 entirely. In some examples, playback device 110a does not include input / output 111 and receives all audio content via network 104.
[0053] The playback device 110a also includes electronics 112, a user interface 113 (e.g., one or more buttons, knobs, dials, touch-sensitive surfaces, displays, touchscreens), and one or more transducers 114 (hereinafter referred to as "transducers 114"). Electronics 112 is configured to receive audio from an audio source (e.g., a local audio source 105) via input / output 111, and via network 104 (…). Figure 1B The playback device 110a receives audio from one or more computing devices 106a to 106c, amplifies the received audio, and outputs the amplified audio via one or more transducers 114 for playback. In some examples, the playback device 110a may optionally include one or more microphones 115 (e.g., a single microphone, multiple microphones, a microphone array) (hereinafter referred to as "microphone 115"). In some examples, for example, the playback device 110a with one or more optional microphones 115 can be used as an NMD configured to receive voice input from a user and perform one or more operations accordingly based on the received voice input.
[0054] exist Figure 1C In the illustrated example, electronic device 112 includes one or more processors 112a (hereinafter referred to as "processor 112a"), memory 112b, software component 112c, network interface 112d, one or more audio processing components 112g (hereinafter referred to as "audio component 112g"), one or more audio amplifiers 112h (hereinafter referred to as "amplifier 112h"), and power supply 112i (e.g., one or more power supplies, power cords, power outlets, batteries, induction coils, Power over Ethernet (PoE) interfaces, and / or other suitable power sources). In some examples, electronic device 112 may optionally include one or more other components 112j (e.g., one or more sensors, video displays, touchscreens, battery charging docks).
[0055] Processor 112a may include clock-driven computing components configured to process data, and memory 112b may include a computer-readable medium (e.g., a tangible non-transitory computer-readable medium, a data store loaded with one or more software components 112c) configured to store instructions for performing various operations and / or functions. Processor 112a is configured to execute instructions stored in memory 112b to perform one or more operations. Such operations may include, for example, causing playback device 110a to access an audio source (e.g., computing devices 106a to 106c). Figure 1BOne or more of the playback devices 110 and / or another playback device in playback device 110 retrieve audio data. In some examples, the operation also includes causing playback device 110a to send audio data to another playback device in playback device 110a and / or another device (e.g., one of NMD 120). Some examples include operations that pair playback device 110a with another playback device in one or more playback devices 110 to implement a multi-channel audio environment (e.g., stereo pair, bonded area).
[0056] Processor 112a can also be configured to perform operations that synchronize the playback of audio content between playback device 110a and another playback device among one or more playback devices 110. As will be understood by those skilled in the art, during synchronized playback of audio content on multiple playback devices, the listener will preferably not perceive any time delay difference between playback device 110a and the playback of audio content by one or more other playback devices 110. Additional details regarding audio playback synchronization between playback devices can be found, for example, in U.S. Patent No. 8,234,395, which is incorporated above by reference.
[0057] In some examples, memory 112b is also configured to store data associated with playback device 110a, such as one or more zones and / or zones that playback device 110a is a member of, audio sources accessible to playback device 110a, and / or playback queues that playback device 110a (and / or another playback device among one or more playback devices) can be associated with. The stored data may include one or more state variables that are periodically updated and used to describe the state of playback device 110a. Memory 112b may also include data associated with the state of one or more devices in other devices of media playback system 100 (e.g., playback device 110, NMD 120, control device 130). In some aspects, for example, state data is shared among at least a portion of the devices in media playback system 100 during predetermined time intervals (e.g., every 5 seconds, every 10 seconds, every 60 seconds), such that one or more devices have the most up-to-date data associated with media playback system 100.
[0058] Network interface 112d is configured to facilitate playback device 110a with a data network (e.g., link 103 and / or network 104). Figure 1BData transfer between one or more other devices on the network interface 112d. The network interface 112d is configured to send and receive data corresponding to media content (e.g., audio content, video content, text, photos) and other signals including digital packet data (e.g., non-transient signals) containing an Internet Protocol (IP)-based source address and / or an IP-based destination address. The network interface 112d can parse the digital packet data, enabling the electronic device 112 to correctly receive and process data destined for the playback device 110a.
[0059] exist Figure 1C In the example shown, network interface 112d includes one or more wireless interfaces 112e (hereinafter referred to as "wireless interface 112e"). Wireless interface 112e (e.g., a suitable interface including one or more antennas) can be configured to wirelessly communicate with one or more other devices (e.g., one or more of other playback devices 110, NMD 120, and / or control devices 130) according to a suitable wireless communication protocol (e.g., WiFi, Bluetooth, LTE), which are communicatively coupled to network 104. Figure 1B In some examples, network interface 112d may optionally include wired interface 112f (e.g., an interface or receptacle of a network cable configured to receive network cables such as Ethernet, USB-A, USB-C, and / or Thunderbolt cables), which is configured to communicate with other devices via a wired connection according to a suitable wired communication protocol. In some examples, network interface 112d includes wired interface 112f and does not include wireless interface 112e. In some examples, electronic device 112 completely excludes network interface 112d and sends and receives media content and / or other data via another communication path (e.g., input / output 111).
[0060] Audio component 112g is configured to process and / or filter data including media content received by electronic device 112 (e.g., via input / output 111 and / or network interface 112d) to generate an output audio signal. In some examples, audio processing component 112g includes, for example, one or more digital-to-analog converters (DACs), audio preprocessing components, audio enhancement components, digital signal processors (DSPs), and / or other suitable audio processing components, modules, circuitry, etc. In some examples, one or more of audio processing components 112g may include one or more sub-components of processor 112a. In some examples, electronic device 112 omits audio processing component 112g. In some aspects, for example, processor 112a executes instructions stored in memory 112b to perform audio processing operations to generate an output audio signal.
[0061] Amplifier 112h is configured to receive and amplify an audio output signal generated by audio processing component 112g and / or processor 112a. Amplifier 112h may include electronic devices and / or components configured to amplify the audio signal to a level sufficient to drive one or more transducers 114. In some examples, for instance, amplifier 112h includes one or more switching or Class D power amplifiers. However, in other examples, the amplifier includes one or more other types of power amplifiers (e.g., linear gain power amplifiers, Class A amplifiers, Class B amplifiers, Class AB amplifiers, Class C amplifiers, Class D amplifiers, Class E amplifiers, Class F amplifiers, Class G and / or Class H amplifiers, and / or other suitable types of power amplifiers). In some examples, amplifier 112h includes a suitable combination of two or more of the aforementioned types of power amplifiers. Furthermore, in some examples, the respective amplifiers in amplifier 112h correspond to the respective transducers in transducers 114. However, in other examples, electronic device 112 includes a single amplifier in amplifier 112h configured to output the amplified audio signal to multiple transducers 114. In some other examples, the amplifier 112h is omitted from the electronic device 112.
[0062] Transducer 114 (e.g., one or more loudspeakers and / or loudspeaker drivers) receives amplified audio signals from amplifier 112h and presents or outputs the amplified audio signals as sound (e.g., audible sound waves with frequencies between about 20 Hz and about 20 kHz). In some examples, transducer 114 may include a single transducer. However, in other examples, transducer 114 includes multiple audio transducers. In some examples, transducer 114 includes more than one type of transducer. For example, transducer 114 may include one or more low-frequency transducers (e.g., subwoofers, woofers), mid-frequency transducers (e.g., mid-frequency transducers, mid-frequency woofers), and one or more high-frequency transducers (e.g., one or more tweeters). As used herein, “low frequency” can generally refer to audible frequencies below about 500 Hz, “mid frequency” can generally refer to audible frequencies between about 500 Hz and about 2 kHz, and “high frequency” can generally refer to audible frequencies above 2 kHz. However, in some examples, one or more transducers 114 include transducers that do not adhere to the aforementioned frequency range. For example, one of the transducers 114 may include a mid-frequency bass transducer configured to output sound at frequencies between about 200 Hz and about 5 kHz.
[0063] For example, SONOS currently offers (or has offered) certain playback devices, including, for example, "SONOS ONE," "PLAY:1," "PLAY:3," "PLAY:5," "PLAYBAR," "PLAYBASE," "CONNECT:AMP," "CONNECT," and "SUB." Other suitable playback devices may be used additionally or alternatively to implement the playback devices of the examples disclosed herein. Furthermore, those skilled in the art will understand that the playback devices are not limited to the examples described herein or SONOS's product offerings. In some examples, for instance, one or more playback devices 110 include wired or wireless headphones (e.g., over-ear headphones, on-ear headphones, in-ear headphones). In other examples, one or more of playback devices 110 include a docking station and / or an interface configured to interact with a docking station of a personal mobile media playback device. In some examples, the playback device may be part of another device or component, such as a television, lighting equipment, or some other device used indoors or outdoors. In some examples, the playback device omits a user interface and / or one or more transducers. For example, Figure 1D It is a block diagram of a playback device 110p that includes input / output 111 and electronic devices 112 but does not have a user interface 113 or a transducer 114.
[0064] Figure 1E This is a block diagram of a bound playback device 110q, which includes components associated with playback device 110i (e.g., a subwoofer). Figure 1A Acoustically bonded playback device 110a ( Figure 1C In the illustrated example, playback devices 110a and 110i are separate playback devices 110 housed in a single housing. However, in some examples, a bundled playback device 110q comprises a single housing housing both playback devices 110a and 110i. The bundled playback device 110q can be configured to function differently from unbundled playback devices (e.g., Figure 1C Playback device 110a) and / or paired or bound playback devices (e.g., Figure 1B The playback devices 110a and 110i process and reproduce sound in a manner consistent with the playback devices 110a and 110i. In some examples, for instance, playback device 110a is a full-range playback device configured to present low-frequency, mid-frequency, and high-frequency audio content, and playback device 110i is a subwoofer configured to present low-frequency audio content. In some aspects, playback device 110a, when paired with a first playback device, is configured to present only the mid-frequency and high-frequency components of a specific audio content, while playback device 110i presents the low-frequency components of the specific audio content. In some examples, the paired playback device 110q includes an additional playback device and / or another paired playback device.
[0065] c. Suitable network microphone equipment (NMD)
[0066] Figure 1F It is NMD 120a ( Figure 1A and Figure 1B The NMD 120a includes one or more speech processing components 124 (hereinafter referred to as "speech components 124") and a reference playback device 110a. Figure 1C The description includes several components including processor 112a, memory 112b, and microphone 115. NMD 120a optionally includes components also included in playback device 110a. Figure 1C Other components in the NMD 120a include, for example, the user interface 113 and / or the transducer 114. In some examples, the NMD 120a is configured as a media playback device (e.g., one or more playback devices 110) and also includes, for example, one or more audio components 112g. Figure 1C The NMD 120a includes an amplifier 114 and / or other playback device components. In some examples, the NMD 120a includes Internet of Things (IoT) devices such as thermostats, alarm panels, fire and / or smoke detectors, etc. In some examples, the NMD 120a includes a microphone 115, a voice processing unit 124, and the components mentioned above. Figure 1B The components of the described electronic device 112 are only a portion. In some aspects, for example, the NMD 120a includes a processor 112a and a memory 112b. Figure 1B The NMD 120a omits one or more other components of the electronic device 112. In some examples, the NMD 120a includes additional components (e.g., one or more sensors, cameras, thermometers, barometers, hygrometers).
[0067] In some examples, NMD can be integrated into playback devices. Figure 1G This is a block diagram of a playback device 110r including the NMD 120d. The playback device 110r may include many or all of the components of the playback device 110a and also includes a microphone 115 and a voice processing unit 124. Figure 1F The playback device 110r may optionally include an integrated control device 130c. The control device 130c may, for example, include a user interface (e.g., Figure 1B The user interface 113 is configured to receive user input (e.g., touch input, voice input) without a separate control device. However, in other examples, the playback device 110r receives input from another control device (e.g., ...). Figure 1B The control device 130a) receives commands.
[0068] Refer again Figure 1F The microphone 115 is configured to receive signals from the environment (e.g., Figure 1A The environment 101) and / or the room where the NMD 120a is located acquire, capture, and / or receive sound. The received sound may include, for example, spoken words, audio played back by the NMD 120a and / or another playback device, background speech, ambient sounds, etc. The microphone 115 converts the received sound into electrical signals to generate microphone data. The voice processing 124 receives and analyzes the microphone data to determine if voice input is present in the microphone data. For example, voice input may include an activation word followed by a utterance that includes a user request. As will be understood by those skilled in the art, an activation word is a word or other audio cue that indicates user voice input. For example, when querying AMAZON® VAS, a user might say the activation word “Alexa.” Other examples include “Ok, Google” for invoking GOOGLE® VAS and “Hey, Siri” for invoking APPLE® VAS.
[0069] After detecting the activation word, the voice processing unit monitors microphone data accompanying the user request in the voice input. The user request may include, for example, a command to control a third-party device, such as a thermostat (e.g., a NEST® thermostat), a lighting device (e.g., a PHILIPSHUE® lighting device), or a media playback device (e.g., a Sonos® playback device). For example, a user might say the activation word “Alexa” and then say the phrase “set the thermostat to 68 degrees” to set up their home (e.g.,...). Figure 1A The temperature in environment 101). The user might say the same activation word and then say "light up the living room" to turn on the lighting in the living room area of the home. The user could similarly say the activation word and then request that a specific song, album, or music playlist be played on the playback devices in the home.
[0070] d. Suitable control equipment
[0071] Figure 1H It is control equipment 130a ( Figure 1A and Figure 1BA partial schematic diagram of the media playback system 100 is shown. As used herein, the term "control device" may be used interchangeably with "controller" or "control system." Among other features, the control device 130a is configured to receive user input related to the media playback system 100 and, in response, cause one or more devices in the media playback system 100 to perform actions or operations corresponding to the user input. In the illustrated example, the control device 130a includes a smartphone (e.g., an iPhone™, an Android phone) on which media playback system controller application software is installed. In some examples, the control device 130a includes, for example, a tablet computer (e.g., an iPad™), a computer (e.g., a laptop computer, a desktop computer), and / or other suitable devices (e.g., a television, a car audio head unit, an IoT device). In some examples, the control device 130a includes a dedicated controller for the media playback system 100. In other examples, as described above regarding... Figure 1G As described, control device 130a is integrated into another device in media playback system 100 (e.g., playback device 110, NMD 120 and / or other suitable devices configured to communicate over a network).
[0072] Control device 130a includes electronics 132, a user interface 133, one or more speakers 134, and one or more microphones 135. Electronics 132 includes one or more processors 132a (hereinafter referred to as "processor 132a"), memory 132b, software components 132c, and a network interface 132d. Processor 132a may be configured to perform functions related to facilitating user access to, control of, and configuration of media playback system 100. Memory 132b may include a data storage device that may be loaded with one or more software components executable by processor 132a to perform those functions. Software components 132c may include applications and / or other executable software configured to facilitate control of media playback system 100. Memory 132b may be configured to store, for example, software components 132c, media playback system controller application software, and / or other data associated with media playback system 100 and the user.
[0073] Network interface 132d is configured to facilitate network communication between control device 130a and one or more other devices and / or one or more remote devices in media playback system 100. In some examples, network interface 132d is configured to operate according to one or more suitable communication industry standards (e.g., infrared, radio, wired standards including IEEE 802.3, and wireless standards including IEEE 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, 4G, LTE). For example, network interface 132d may be configured to communicate with playback device 110, NMD 120, other devices in control device 130, Figure 1B The network interface 132d sends and / or receives data from one of the computing devices 106, including one or more other media playback systems. The sent and / or received data may include, for example, playback device control commands, status variables, playback zone and / or zone group configurations. For example, based on user input received at user interface 133, network interface 132d can send playback device control commands (e.g., volume control, audio playback control, audio content selection) from control device 130 to one or more playback devices 110. Network interface 132d can also send and / or receive configuration changes, such as adding or removing one or more playback devices 110 from a zone, adding or removing one or more zones from a zone group, forming a bound or merged player, detaching one or more playback devices from a bound or merged player, etc. See below for further details. Figures 1I to 1M Additional descriptions for districts and groups were found.
[0074] User interface 133 is configured to receive user input and facilitate control of media playback system 100. User interface 133 includes media content art 133a (e.g., album art, lyrics, video), playback status indicators 133b (e.g., elapsed time and / or remaining time indicators), media content information area 133c, playback control area 133d, and zone indicators 133e. Media content information area 133c may include a display of relevant information (e.g., title, artist, album, genre, release year) about the currently playing media content and / or media content in the queue or playlist. Playback control area 133d may include optional icons (e.g., via touch input and / or via a cursor or another suitable selector) to cause one or more playback devices in the selected playback zone or group of zones to perform playback actions, such as play or pause, fast forward, rewind, skip to the next, skip to the previous, enter / exit shuffle mode, enter / exit repeat mode, enter / exit crossfade mode, etc. The playback control area 133d may also include selectable icons for modifying equalization settings, playback volume, and / or other appropriate playback actions. In the illustrated example, the user interface 133 includes a display presented on a touchscreen interface of a smartphone (e.g., an iPhone™, an Android phone). However, in some examples, alternative user interfaces with varying formats, styles, and sequences of interactions may be implemented on one or more network devices to provide similar control access to the media playback system.
[0075] One or more speakers 134 (e.g., one or more transducers) may be configured to output sound to a user of control device 130a. In some examples, the one or more speakers include individual transducers configured to output low, mid, and / or high frequencies accordingly. In some aspects, for example, control device 130a is configured as a playback device (e.g., one of playback devices 110). Similarly, in some examples, control device 130a is configured as an NMD (e.g., one of NMD 120) that receives voice commands and other sounds via one or more microphones 135.
[0076] One or more microphones 135 may include, for example, one or more condenser microphones, electret condenser microphones, dynamic microphones, and / or other suitable types of microphones or transducers. In some examples, two or more microphones 135 are arranged to capture location information of an audio source (e.g., speech, audible sound) and / or configured to facilitate the filtering of background noise. Furthermore, in some examples, control device 130a is configured to function as a playback device and NMD. However, in other examples, control device 130a omits one or more speakers 134 and / or one or more microphones 135. For example, control device 130a may include a device (e.g., a thermostat, IoT device, network device) that includes a portion of electronics 132 and a user interface 133 (e.g., a touchscreen) without any speakers or microphones.
[0077] Suitable playback equipment configuration
[0078] Figures 1I to 1M Example configurations of playback devices in zones and zones are shown. First refer to... Figure 1M In one example, a single playback device can belong to one zone. For example, secondary bedroom 101c ( Figure 1A Playback device 110g in the diagram can belong to zone C. In some implementations described below, multiple playback devices can be "bonded" to form a "bonded pair" that together form a single zone. For example, playback device 110l (e.g., a left playback device) can be bonded to playback device 110l (e.g., a left playback device) to form zone A. The bonded playback devices can have different playback responsibilities (e.g., channel responsibilities). In another implementation described below, multiple playback devices can be merged to form a single zone. For example, playback device 110h (e.g., a front playback device) can be merged with playback device 110i (e.g., a subwoofer) and playback devices 110j and 110k (e.g., left and right surround speakers, respectively) to form a single zone D. In another example, playback devices 110g and 110h can be merged to form a merged group or zone group 108b. The merged playback devices 110g and 110h may not be specifically assigned different playback responsibilities. That is, in addition to synchronously playing back audio content, the merged playback devices 110h and 110i can each play back audio content as if they were not merged.
[0079] Each zone in the media playback system 100 can be provided for control as a single user interface (UI) entity. For example, zone A can be provided as a single entity named "Master Bathroom". Zone B can be provided as a single entity named "Master Bedroom". Zone C can be provided as a single entity named "Secondary Bedroom".
[0080] The bound playback device can have different playback responsibilities, such as the responsibility of certain audio channels. For example, such as... Figure 1I As shown, playback devices 110l and 110m can be paired to produce or enhance the stereo effect of audio content. In this example, playback device 110l can be configured to play back the left channel audio component, while playback device 110k can be configured to play back the right channel audio component. In some implementations, this stereo pairing may be referred to as "pairing".
[0081] Furthermore, the bundled playback device may have additional and / or different corresponding speaker drivers. For example... Figure 1J As shown, a playback device 110h, named "front," can be paired with a playback device 110i, named "subwoofer." The front device 110h can be configured to reproduce the mid-to-high frequency range, while the subwoofer device 110i can be configured to reproduce the low frequencies. However, when not paired, the front device 110h can be configured to reproduce the entire frequency range. As another example, Figure 1K The diagram also shows a front device 110h and a subwoofer device 110i, which are respectively bound to the left playback device 110j and the right playback device 110k. In some implementations, the left device 110j and the right device 110k can be configured to form surround or "satellite" channels for a home theater system. The bound playback devices 110h, 110i, 110j, and 110k can form a single D-zone ( Figure 1M ).
[0082] The merged playback devices may not have assigned playback responsibilities, but each can present the full range of audio content that the respective playback device can provide. However, the merged devices can be represented as a single UI entity (i.e., a zone, as described above). For example, playback devices 110a and 110n in the main bathroom have a single UI entity for zone A. In one example, playback devices 110a and 110n can both synchronously output the full frequency range of audio content that each of the respective playback devices 110a and 110n can handle.
[0083] In some examples, the NMD is paired with or merged with another device to form a zone. For example, the NMD 120b may be paired with the playback device 110e, and together they form zone F, called the living room. In other examples, a standalone network microphone device may be in a zone on its own. However, in other examples, a standalone network microphone device may not be associated with a zone. Additional details regarding associating a network microphone device and a playback device as a designated or default device can be found, for example, in U.S. Patent Application No. 15 / 438,749, which was previously cited.
[0084] Individual, bundled, and / or merged devices can be grouped together to form partitions. For example, see reference. Figure 1MArea A can be grouped with Area B to form area group 108a, which includes two areas. Similarly, Area G can be grouped with Area H to form area group 108b. As another example, Area A can be grouped with one or more other areas C through I. Areas A through I can be grouped and ungrouped in various ways. For example, three, four, five, or more (e.g., all) of areas A through I can be grouped together. When areas of individual playback devices and / or bundled playback devices are grouped together, they can play back audio synchronously with each other, as described in previously cited U.S. Patent No. 8,234,395. Playback devices can dynamically group and ungroup to form new or different groups of synchronously played back audio content.
[0085] In various implementations, a region within an environment can be the default name of a region within a group or a combination of region names within a region group. For example, region group 108b could have been assigned a name such as "Restaurant + Kitchen". Figure 1M As shown. In some examples, block groups can be assigned unique names chosen by the user.
[0086] Some data can be stored as one or more state variables in the memory of the playback device (e.g., Figure 1C In the memory 112b), these state variables are periodically updated and used to describe the state of the playback area, playback device, and / or associated groups of areas. The memory may also include data associated with the state of other devices in the media system and is shared between devices from time to time, so that one or more devices have the latest data associated with the system.
[0087] In some examples, the memory can store instances of various variable types associated with a state. Variable instances can be stored along with identifiers (e.g., labels) corresponding to the type. For example, some identifiers could be a first type "a1" for identifying playback devices in a zone, a second type "b1" for identifying playback devices that can be bound to a zone, and a third type "c1" for identifying the zone group to which the zone can belong. As a related example, the identifier associated with secondary bedroom 101c could indicate that the playback device is the only playback device in zone C and is not in a zone group. The identifier associated with the study could indicate that the study is not grouped with other zones, but includes bound playback devices 110h to 110k. The identifier associated with the dining room could indicate that the dining room is part of the dining room + kitchen zone group 108b and that devices 110b and 110d are grouped together (…). Figure 1L Since the kitchen is part of the dining room + kitchen area group 108b, identifiers associated with the kitchen can indicate the same or similar information. Other example area variables and identifiers are described below.
[0088] In yet another example, the media playback system 100 may represent other associated variables or identifiers for zones and zone groups, such as identifiers associated with zones, like... Figure 1M As shown. A region can involve blocks and / or clusters of blocks that are not within a single block. For example, Figure 1M An upper region 109a comprising zones A through D and a lower region 109b comprising zones E through I are shown. In one aspect, a region can be used to invoke a group of zones and / or a cluster of zones that share another cluster. In another aspect, this differs from a group of zones that does not share a zone with another group. Further examples of techniques for implementing regions can be found, for example, U.S. Application No. 15 / 682,506, filed August 21, 2017, entitled “Room Association Based on Name,” and U.S. Patent No. 8,483,853, filed September 11, 2007, entitled “Controlling and manipulating groupings in a multi-zone media system.” Each of these applications is incorporated herein by reference in its entirety. In some examples, the media playback system 100 may not implement regions, in which case the system may not store variables associated with regions.
[0089] III. Multi-device playback of generative media content
[0090] Figure 2This is a functional block diagram of a system 200 for playing back generative media content. As previously described, generative media content can include any media content (e.g., audio, video, audiovisual output, haptic output, or any other media content) dynamically created, synthesized, and / or modified by non-human rule-based processes (such as algorithms or models). Such creation or modification can occur for real-time or near-real-time playback. Additionally or alternatively, generative media content can be generated or modified asynchronously (e.g., in advance before playback is requested), and specific items of the generative media content can then be selected for later playback. As used herein, a “generative media module” includes any system, whether implemented in software, a physical model, or a combination thereof, that can generate generative media content based on one or more inputs. In some examples, such generative media content includes novel media content that can be created as entirely new media content or created by mixing, combining, manipulating, or otherwise modifying one or more pre-existing segments of media content. As used herein, a “generative media content model” includes any algorithm, pattern, or set of rules that can be used to generate novel generative media content using one or more inputs, such as sensor data, parameters provided by artists, media segments such as audio clips or samples. In the examples, the generative media module can use a variety of different generative media content models to produce different generative media content. In some cases, artists or other collaborators can interact with, create, and / or update the generative media content model to produce specific generative media content. Although several examples throughout this discussion involve audio content, the principles disclosed herein can be applied to other types of media content, such as video, audiovisual, haptic, or other media content, in some examples.
[0091] like Figure 2 As shown, system 200 includes a generative media group coordinator 210 that communicates with generative media group members 250a and 250b, as well as sensor data source 218, media content source 220, and control device 130. This communication can be performed via network 102, which, as described above, can include any suitable wired or wireless network connection or a combination thereof (e.g., WiFi network, Bluetooth, Z-Wave network, ZigBee, Ethernet connection, Universal Serial Bus (USB) connection, etc.).
[0092] One or more remote computing devices 106 may also communicate with group coordinator 210 and / or group members 250a and 250b via network 102. In various examples, remote computing device 106 may be a cloud-based server associated with a device manufacturer, media content provider, voice assistant service, or other suitable entity. Figure 2As shown, remote computing device 106 may include generative media module 214. As described in more detail elsewhere herein, remote computing device 106 can generate generative media content remotely from local devices (e.g., coordinator 210 and members 250a and 250b). The generative media content can then be sent to one or more of these local devices for playback. Additionally or alternatively, the generative media content may be generated wholly or partially via local devices (e.g., group coordinator 210 and / or group members 250a and 250b). In some examples, group coordinator 210 itself may be a remote computing device, such that it is communicatively coupled to group members 250a and 250b via a wide area network, and these devices do not need to be located in the same environment (e.g., home, business location, etc.).
[0093] a. Example of generative media group operations
[0094] In the illustrated example, the generative media group includes a generative media group coordinator 210 (also referred to herein as "coordinator device 210") and a first generative media group member 250a and a second generative media group member 250b (also referred to herein as "first member device" 250a" and "second member device" 250b" and collectively as "member device 250"). Optionally, one or more remote computing devices 106 may also form part of the generative media group. In operation, these devices may communicate with each other and / or with other components (e.g., sensor data source 218, control device 130, media content source 220, or any other suitable data source or component) to facilitate the generation and playback of generative media content.
[0095] In various examples, some or all of devices 210 and / or 250 may be located together in the same environment (e.g., in the same home, store, etc.). In some examples, at least some of devices 210 and / or 250 may be located far apart from each other, for example, in different homes, different cities, etc.
[0096] Coordinator device 210 and / or member device 250 may include the above-mentioned... Figures 1A to 1H Some or all of the components of the described playback device 110 or network microphone device 120. For example, the coordinator device 210 and / or member device 250 may optionally include playback components 212 (e.g., transducers, amplifiers, audio processing components, etc.), or such components may be omitted in some cases.
[0097] In some examples, the coordinator device 210 is itself a playback device and can therefore also be used as a member device 250. In other examples, the coordinator device 210 may be connected to one or more member devices 250 (e.g., via a direct wired connection or via network 102), but the coordinator device 210 itself does not play back generative media content. In various examples, the coordinator device 210 may be implemented on a bridge-like device on a local network, on a playback device that is not itself part of a generative media group (i.e., the playback device itself does not play back generative media content), and / or on a remote computing device (e.g., a cloud server).
[0098] In various examples, one or more devices may include a generative media module 214 thereon. This generative media module 214 can, for example, use a suitable generative media content model based on one or more inputs to generate novel synthetic media content. Figure 2 As shown, in some examples, the coordinator device 210 may include a generative media module 214 for generating generative media content, which can then be sent to member devices 250a and 250b for simultaneous and / or synchronized playback. Additionally or alternatively, member devices 250 (e.g., such as...) Figure 2 Some or all of the member devices 250b shown may include a generative media module 214, which can be used by the member device 250 to locally generate generative media content based on one or more inputs. In various examples, one or more input parameters received from the local device may optionally be used to generate generative media content via the remote computing device 106. The generative media content may then be sent to one or more local devices for coordination and / or playback.
[0099] In some examples, at least some of the member devices 250 do not include the generative media module 214. Alternatively, in some cases, each member device 250 may include the generative media module 214 and may be configured to generate generative media content locally. In at least some examples, none of the member devices 250 includes the generative media module 214. In this case, the generative media content may be generated by the coordinator device 210. This generative media content can then be sent to the member devices 250 for simultaneous and / or synchronized playback.
[0100] exist Figure 2In the example shown, coordinator device 210 additionally includes coordination component 216. As described in more detail herein, in some cases, coordinator device 210 can facilitate the playback of generative media content via multiple different playback devices, which may or may not include coordinator device 210 itself. In operation, coordination component 216 is configured to facilitate the synchronization of both generative media creation (e.g., using one or more generative media modules 214 that may be distributed across various devices) and generative media playback. For example, coordinator device 210 may send timing data to member device 250 to facilitate synchronized playback. Additionally or alternatively, coordinator device 210 may send inputs, generative media model parameters, or other data related to generative media module 214 to one or more member devices 250, such that member devices 250 can generate generative media locally (e.g., using locally stored generative media module 214), and / or allow member devices 250 to update or modify generative media module 214 based on inputs received from coordinator device 210.
[0101] As described in more detail elsewhere in this document, generative media module 214 can be configured to generate generative media based on one or more inputs using a generative media content model. These inputs may include sensor data (e.g., provided by sensor data source 218), user input (e.g., received from control device 130 or via direct user interaction with coordinator device 210 or member device 250), and / or media content source 220. For example, generative media module 214 can generate and continuously modify generative audio by adjusting various characteristics of the generative audio based on one or more input parameters (e.g., sensor data associated with one or more users of devices 210, 250).
[0102] b. Example media content source
[0103] In various examples, media content source 220 may include one or more local and / or remote media content sources. For example, media content source 220 may include one or more local audio sources 105 as described above (e.g., audio received via an input / output connection such as from a mobile device (e.g., a smartphone, tablet, laptop) or another suitable audio component (e.g., a television, desktop computer, amplifier, phonograph, Blu-ray player, memory storing digital media files)). Additionally or alternatively, media content source 220 may include one or more remote computing devices accessible via a network interface (e.g., via communications on network 102). Such remote computing devices may include separate computers or servers, such as media streaming service servers storing audio and / or other media content.
[0104] In various examples, the media available via media content source 220 may include a complete sound, a song, a portion of a song (e.g., a sample), or any audio component (e.g., pre-recorded audio of a particular instrument, synthesized beats or other audio segments, pre-recorded audio segments in the form of non-musical audio (such as spoken or natural sounds, etc.). In operation, generative media module 214 may utilize such media to generate generative media content, for example, by combining, mixing, overlaying, manipulating, or otherwise modifying retrieved media content to produce novel generative media content for playback via one or more devices. In some examples, generative media content may take the form of a combination of pre-recorded audio segments (e.g., pre-recorded songs, spoken recordings, etc.) and novel synthesized audio being created and overlaid with the pre-recorded audio. As used herein, "generative media content" or "generated media content" may include any such combination.
[0105] c. Example Generative Media Module
[0106] As described above, generative media module 214 may include any system that can generate generative media content based on one or more inputs, whether instantiated as software, a physical model, or a combination thereof. In various examples, generative media module 214 may utilize a generative media content model, which may include one or more algorithms or mathematical models that determine how to generate media content based on relevant input parameters. In some cases, such as based on instructions received from one or more remote computing devices (e.g., a cloud server associated with a music service or other entity), or based on input received from other group member devices in the same or different environments, or any other suitable input, these algorithms and / or mathematical models themselves may be updated over time. In some examples, various devices within the group may have different generative media modules 214 on them—for example, a first member device may have a different generative media module 214 than a second member device. In other cases, each device within the group having a generative media module 214 may include substantially the same model or algorithm.
[0107] Generative media content can be generated using any suitable algorithm or combination of algorithms. Examples of such algorithms include algorithms using machine learning techniques (e.g., generative adversarial networks, neural networks, etc.), formal grammars, Markov models, finite state automata, and / or any algorithm implemented within currently available products (such as OpenAI's JukeBox, Amazon's AWS DeepComposer, Google's Magenta, Amper Music's AmperAI, etc.). In various examples, generative media module 214 can leverage any suitable generative algorithm that exists now or is being developed in the future.
[0108] Based on the above discussion, generating generative media content (e.g., audio content) can involve altering various characteristics of the media content in real time and / or generating novel media content through algorithms in real time or near real time. In the context of audio content, this can be achieved by storing multiple audio samples in a database (e.g., within media content source 220), which can be remotely located and accessible via network 102 by coordinator device 210 and / or member device 250, or alternatively, the audio samples can be maintained locally on devices 210, 250 themselves. Audio samples can be associated with one or more metadata tags corresponding to one or more audio characteristics of the sample. For example, a given sample can be associated with metadata tags indicating that the sample includes audio of a specific frequency or frequency range (e.g., bass / mid / treble) or a specific instrument, genre, rhythm, tonality, release date, geographic region, timbre, reverberation, distortion, sound texture, or any other audio characteristic that will be apparent.
[0109] In operation, the generative media module 214 (e.g., of the coordinator device 210 and / or the second member device 250b) can retrieve certain audio samples based on their associated tags and blend these audio samples together to create generative audio. As the generative media module 214 retrieves audio samples with different tags and / or different audio samples with the same or similar tags, the generative audio can evolve in real time. The audio samples retrieved by the generative media module 214 can depend on one or more inputs (such as sensor data, time of day, geographic location, weather) or various user inputs (such as mood selection) or physiological inputs (such as heart rate). In this way, as the input changes, the generative audio also changes. For example, if the user selects a calm or relaxed mood input, the generative media module 214 can retrieve audio samples and blend them with tags corresponding to audio content that the user can find calm or relaxed. Examples of such audio samples can include audio samples labeled as low tempo or low harmonic complexity, or audio samples that have been pre-defined as calm or relaxed and have been so labeled. In some examples, audio samples can be identified as calm or relaxed based on an automated process of analyzing the temporal and spectral content of the signal. Other examples are also possible. In any of the examples in this paper, the generative media module 214 can adjust the characteristics of the generative audio by retrieving and mixing audio samples associated with different metadata tags or other suitable identifiers.
[0110] Modifying the characteristics of generated audio can include: manipulating volume, balance, or one or more of these; removing certain instruments or tones; and altering the audio's tempo, gain, reverb, spectral equalization, timbre, or sound texture. In some examples, generated audio can be played back in different ways on different devices, such as emphasizing certain characteristics of the generated audio on a specific playback device closest to the user. For instance, the nearest playback device could emphasize certain instruments, beats, tones, or other characteristics, while the other playback devices could act as background audio sources.
[0111] As described elsewhere in this document, media content module 214 can be configured to generate media designed to guide a user's emotional and / or physiological state in a desired direction. In some examples, the user's current state (e.g., mood, emotional state, activity level, etc.) is continuously and / or iteratively monitored or measured (e.g., at predetermined intervals) to ensure that the user's current state is transitioning toward the desired state or at least not in the opposite direction. In such an example, the generative audio content can be modified to steer the user's current state toward the desired final state.
[0112] In any of the examples herein, the generative media module may use hysteresis to avoid rapid adjustments to the generative audio that could negatively impact the listening experience. For example, if the generative media module modifies the media based on the user's position relative to the playback device, the playback device may rapidly change the generative audio in any of the ways described herein when the user moves rapidly closer to or further away. Such rapid adjustments can be unpleasant for the user. To reduce these rapid adjustments, the generative media module 214 may be configured to employ hysteresis by delaying the adjustment of the generative audio for a predetermined period of time when the user's movement or other activity triggers the adjustment. For example, if the playback device detects that the user has moved within a threshold distance of the playback device, instead of immediately performing one of the aforementioned adjustments, the playback device may wait a predetermined amount of time (e.g., a few seconds) before making the adjustment. If, after the predetermined amount of time, the user remains within the threshold distance, the playback device may continue to adjust the generative audio. However, if, after the predetermined amount of time, the user does not remain within the threshold distance, the generative media module 214 may avoid adjusting the generative audio. Generative media module 214 can similarly apply lag to other generative media adjustments described herein.
[0113] Figure 3 A flowchart is shown for a process 300 for generating generative audio content using various input parameters. In various examples, one or more of these input parameters can be modified based on user input. For example, an artist can select... Figure 2 The various parameters, constraints, or available audio segments shown are used, and these choices can subsequently at least partially determine the final output of the generative audio content. As previously described, such generative media modules can be stored and operated on one or more playback devices for local playback (e.g., via the same playback device and / or via other playback devices coupled via local area network communication). Additionally or alternatively, such generative media modules can be stored and operated on one or more remote computing devices, wherein the resulting output is transmitted via a wide area network to one or more remote devices for playback.
[0114] As shown in the figure, the process begins at box 302 and proceeds to the clock / metronome in box 304, where it receives inputs of tempo 306 and time stamp 308. Tempo 306 and time stamp 308 can be selected by the artist or automatically determined or generated using a model. The process continues to box 310, where chord changes can be triggered and chord change frequency parameter 312 is received as input. Artists can choose to have higher chord change frequencies in music designed to evoke a higher energy experience (e.g., dance music, uplifting ambient music, etc.). Conversely, lower chord change frequencies can be associated with lower energy output (e.g., calming music).
[0115] At box 314, select a chord from the available chord segments 316. Multiple chord information parameters 318, 320, and 322 can also be provided as input to chord segment 316. These inputs can be used to determine the specific chord to be played next and output as box 324. In some examples, the artist can provide information for each chord, such as weight, frequency of use of the specific chord, etc.
[0116] Next, in box 326, chord changes are selected based at least in part on the harmonic complexity parameter used as input. The harmonic complexity parameter 328 can be tuned or selected by the artist or can be determined automatically. Typically, higher harmonic complexity parameters can be associated with higher-energy audio outputs, and lower harmonic complexity parameters can be associated with lower-energy audio outputs. In some cases, the harmonic complexity parameter may include inputs such as chord inversions, voices, and harmonic density.
[0117] In box 330, the process obtains the root note of the chord, and in box 332, the bass band to be played is selected from the available bass bands 334. These bass bands then pass through bus processing 336, where equalization, filtering, timing, and other processing can be performed.
[0118] Returning to the chord changes in box 326, the process continues separately to box 338 to play the harmony selected from the available harmonic segments 340. This harmonic segment then passes through bus processing 342. Similar to the bass bus processing, harmonic segment bus processing 342 can involve equalization, filtering, timing, and may perform other processing.
[0119] Returning to the selected chord 324, the process continues separately to filter melodic notes in box 344, which utilizes the input of melodic constraint 346. The output in box 348 is the available melodic notes to be played. Melodic constraint 346 can be provided by the artist and can, for example, specify which notes to play or not play, limit the melodic range, or provide other such constraints, depending on the specific selected chord 324.
[0120] In box 350, the process determines which melody note (from the available melody notes 348) to play. This determination can be made automatically based on model values, artist-provided input, randomization effects, or any other suitable input. In the example shown, one input comes from the trigger melody note box 352, which is then based on the melody density parameter 354. The artist can provide the melody density parameter 354, which partially determines the complexity and / or high energy level of the audio output. Based on this parameter, the melody notes can be triggered more or less frequently and at specific times, using box 352, which is input to box 350, to determine which melody note to play. In various examples, the output of box 350 can be provided as input to box 350 in the form of a feedback loop, such that the next melody note selected in box 350 depends at least in part on the last melody note selected in box 350. A melody segment is then selected from the available melody segments 358 in box 356, and then the melody segment is processed through bus 360.
[0121] Returning to the beginning in box 302, the process proceeds separately to box 362 to play non-musical content. This could be, for example, natural sounds, spoken audio, or other such non-musical content. Various non-musical segments 364 can be stored and made available for playback. These non-musical content segments can also be processed via a bus in box 366.
[0122] The outputs of these various paths (e.g., selected bass segments, harmonic segments, melody segments, and / or non-musical segments) can each be processed separately via a bus before being combined at box 368 via mixing and main processing. Here, the combination level can be set, various filters can be applied, relative timing can be established, and any other suitable processing steps can be performed before the generative audio content is output in box 370. In various examples, some of these paths can be omitted entirely. For example, the generative media module can omit the option to play back non-musical content along with the generative musical content. Figure 3 The process 300 shown is merely exemplary, and those skilled in the art will understand that suitable modifications can be made to the process 300 shown herein, and additionally, there are many suitable alternative processes that can be used to generate generative media content.
[0123] Figure 4 This is an example architecture for storing and retrieving generative media content. In this example, the generative media content includes various discrete tracks (each track has multiple variations associated with an energy level or another parameter), which can be selected and played back in various orders and groups depending on specific input parameters.
[0124] As shown in the figure, generative media content 404 can be stored as one or more audio files associated with global generative media content metadata 402. Such metadata may include, for example, global tempo (e.g., beats per minute), global trigger frequency (e.g., the frequency at which changes in input parameters are checked), and / or global crossfade duration (e.g., the duration of fade-in / fade-out between different selected energies).
[0125] Generative media content 404 contains multiple different tracks 406, 408, and 410. In operation, these tracks can be selected and played back in various arrangements (e.g., random groups with some overlays, or playback according to a predetermined order). In some examples, the generative media content 404, including tracks 406, 408, and 410, can be stored locally via one or more playback devices, while one or more remote computing devices can periodically send updated versions of these tracks, the generative media content, and / or global generative media content metadata. In some examples, the playback devices can periodically poll or query the remote computing devices, and in response to the query or polling, the remote computing devices can provide updates to the generative media modules stored on the local playback devices.
[0126] For each track, there may exist corresponding subsets of that track corresponding to different energy levels. For example, the first energy level (EL) of track 1 is at 412, the second energy level of track 1 is at 414, and the nth energy level of track 1 is at 416. Each of these tracks may include both metadata (e.g., metadata 418, 420, 422) and specific media files corresponding to a particular energy level (e.g., media files 424, 426, 428). In some examples, each track may include multiple media files arranged in a particular manner (e.g., media file 424), and the corresponding metadata (e.g., metadata 418) may expect the arrangement and combination of these multiple media files. The media files may, for example, be any suitable format that can be played back and / or streamed to a playback device for playback. In some examples, one or more of media files 424, 426, 428 may be... Figure 3 The output of the described generative model. Metadata may include, for example, tempo (if different from the global tempo), trigger frequency (if different from the global trigger frequency), sequence information (e.g., whether a particular file is played sequentially, randomly, or with percentage weights), crossfade duration (if different from the global crossfade), spatial information (e.g., for rendering audio content in space using multiple transducers), polyphonic information (e.g., allowing multiple audio files to play immediately in the segment), and / or level (e.g., level adjustment in dB, or random level within a predefined range).
[0127] In operation, one or more input parameters (e.g., the number of people in the room, time of day, etc.) can be used to determine a target energy level. This determination can be made using a playback device and / or one or more remote computing devices. Based on this determination, specific media files corresponding to the determined energy level can be selected. The generative media module can then arrange and play back these selected tracks according to a generative content model. This can involve playing the selected tracks in a specific predefined order, playing them in a random or pseudo-random order, or any other suitable method. In some examples, tracks can be played back in a manner that is at least partially overlapping. It may be useful to vary the amount of overlap between tracks so that casual listeners do not hear repetitive loops of audio content but perceive the generative audio as a non-repeating, endless stream of audio.
[0128] although Figure 4 The examples shown use energy levels as a parameter to distinguish different generative audio content, but in various examples, specific variations or arrangements of generative audio content can vary along other dimensions (e.g., genre, time of day, associated user task, etc.).
[0129] d. Example sensor data source and other input parameters
[0130] As previously described, the generative media module 214 can generate generative media at least in part based on input parameters that may include (e.g., received from sensor data source 218) sensor data and / or other suitable input parameters. Regarding sensor input parameters, sensor data source 218 may include data from any suitable sensor, regardless of the sensor's location relative to the generative media assembly and any values measured therefrom. Examples of suitable sensor data include physiological sensor data, such as data obtained from biometric sensors, wearable sensors, etc. This data may include physiological parameters such as heart rate, respiratory rate, blood pressure, brain waves, activity level, exercise, body temperature, etc.
[0131] Suitable sensors include wearable sensors configured to be worn or carried by a user, such as headsets, watches, mobile devices, brain-computer interfaces (e.g., neural links), headphones, microphones, or other similar devices. In some examples, the sensor may be a non-wearable sensor or fixed to a fixed structure. The sensor can provide sensor data, which may include data corresponding to, for example, brain activity, speech, location, movement, heart rate, pulse, body temperature, and / or perspiration. In some examples, the sensor may correspond to multiple sensors. For example, as illustrated elsewhere in this document, the sensor may correspond to a first sensor worn by a first user, a second sensor worn by a second user, and a third sensor not worn by the user (e.g., fixed to a fixed structure). In such an example, the sensor data may correspond to multiple signals received from each of the first, second, and third sensors.
[0132] Sensors can be configured to acquire or generate information that typically corresponds to a user's mood or emotional state. In one example, the sensor is a wearable brain-sensing headband, one of many examples of sensors described herein. Such a headband could, for example, include an electroencephalogram (EEG) headband with multiple sensors. In some examples, the headband could correspond to any Muse. TM Headband (InteraXon; Toronto, Canada). Sensors may be located at different locations around the inner surface of the headband, for example, to correspond to different brain anatomy structures of the user (e.g., frontal, parietal, temporal, and sphenoid bones). Thus, each of these sensors can receive different data from the user. Each of these sensors may correspond to a specific channel that can be streamed from the headband to system devices 210 and / or 250. Such sensor data can be used to detect the user's emotions, for example, by classifying the frequencies and intensities of various brain waves or by performing other analyses. Additional details regarding the use of brain-sensing headbands for generating audio content can be found in co-owned U.S. Application No. 62 / 706,544, filed August 24, 2020, entitled "MOOD DETECTION AND / OR INFLUENCE VIA AUDIO PLAYBACK DEVICES," the entire contents of which are incorporated herein by reference.
[0133] In some examples, sensor data source 218 includes data obtained from sensor data from connected devices (e.g., Internet of Things (IoT) sensors, such as connected lights, cameras, temperature sensors, thermostats, presence detectors, microphones, etc.). Additionally or alternatively, sensor data source 218 may include environmental sensors (e.g., measuring or indicating weather, temperature, time / day / week / month, etc.).
[0134] In some examples, the generative media module 214 may use inputs in the form of playback device capabilities (e.g., the number and type of transducers, output power, other system architecture), device location (e.g., location relative to other playback devices, location relative to one or more users). Additional examples of creating and modifying generative audio based on user and device location are described in more detail in co-owned U.S. Application No. 62 / 956,771, filed January 3, 2020, entitled “GENERATIVE MUSICBASED ON USER LOCATION,” the entire contents of which are incorporated herein by reference. Additional inputs may include device status of one or more devices within the group, such as thermal status (e.g., if a particular device is at risk of overheating, the generative content can be modified to reduce the temperature), battery level (e.g., in a portable playback device with low battery power, bass output can be reduced), and tethering status (e.g., whether a particular playback device is configured as part of a stereo pair, tethered to a subwoofer, or as part of a home theater setup, etc.). Any other suitable device characteristics or states can be similarly used as input for generating generative media content.
[0135] Another example input parameter includes user presence—for example, when a new user enters a space where the generative audio is playing back, the user's presence can be detected (e.g., via a proximity sensor, beacon, etc.), and the generative audio can be modified in response. This modification can be based on the number of users (e.g., ambient or meditation audio for one user, relaxing music for two to four users, and party or dance music for more than four users). Modification can also be based on the identifiers of the present users (e.g., user profiles based on user characteristics, listening history, or other such markers).
[0136] In one example, a user can wear a biometric device that measures various biometric parameters of the user (such as heart rate or blood pressure) and reports these parameters to devices 210 and / or 250. Generative media modules 214 of these devices 210 and / or 250 can use these parameters to further adapt the generative audio, such as increasing the tempo of the music in response to the detection of a high heart rate (as this can indicate that the user is engaged in high-intensity activity) or decreasing the tempo of the music in response to the detection of high blood pressure (as this can indicate that the user is stressed and can benefit from calming music).
[0137] In yet another example, one or more microphones of the playback device (e.g., Figure 1FThe microphone 115 can detect the user's voice. The captured voice data can then be processed to determine, for example, the user's mood, age, or gender (to identify a specific user from several users within a household) or any other such input parameter. Other examples are also possible.
[0138] e. Example coordination among group members
[0139] Figure 5 This is a functional block diagram illustrating data exchange in a system for playing back generated media content. For illustrative purposes, Figure 5 The illustrated system 500 includes interaction between coordinator device 210 and member devices 250b. However, the interactions and procedures described herein can be applied to interactions involving multiple additional coordinator devices 210 and / or member devices 250. Figure 5 As shown, the coordinator device 210 includes a generative media module 214a that receives input parameters 502 (e.g., sensor data, media content, model parameters for the generative media module 214a, or other such inputs) and clock and / or timing data 504. In various examples, the clock and / or timing data 504 may include synchronization signals for synchronizing playback and / or synchronizing generative media generated by various devices within the group. In some examples, the clock and / or timing data 504 may be provided by an internal clock, a processor, or other such components incorporated within the coordinator device 210 itself. In some examples, the clock and / or timing data 504 may be received from a remote computing device via a network interface.
[0140] Based on these inputs, generative media module 214a can output generative media content 404a. Optionally, the output generative media content 404a itself can be used as input to generative media module 214a in the form of a feedback loop. For example, generative media module 214a can use a model or algorithm that depends at least in part on the previously generated content to produce subsequent content (e.g., audio frames).
[0141] In the illustrated example, member device 250b also includes a generative media module 214b, which may be substantially the same as the generative media module 214a of coordinator device 210, or may differ from it in one or more respects. Generative media module 214b may also receive input parameters 502 and clock and / or timing data 504. These inputs may be received from coordinator device 210, from other member devices, from other devices on the local network (e.g., a locally networked smart thermostat providing temperature data), and / or from one or more remote computing devices (e.g., a cloud server providing clock and / or timing data 504, or weather data, or any other such input). Based on these inputs, generative media module 214b may output generative media content 404b. This generated generative media content 404b may optionally be fed back to generative media module 214b as part of a feedback loop. In some examples, the generated media content 404b may include, or consist of, the generated media content 404a that has already been sent over the network to member device 250b (generated via coordinator device 210). In other cases, the generated media content 404b may be generated independently of and solely with the generated media content 404a generated via coordinator device 210.
[0142] Generative media content 404a and 404b can then be played back via devices 210 and 250b themselves, and / or by other devices within the group. In various examples, generative media content 404a and 404b can be configured for simultaneous and / or synchronous playback. In some cases, generative media content 404a and 404b can be substantially the same or similar to each other, with each generative media module 214 using the same or similar algorithms and the same or similar inputs. In other cases, generative media content 404a and 404b can be different from each other, but are still configured for synchronous or simultaneous playback.
[0143] f. Example generative media using a distributed architecture
[0144] As previously mentioned, media content generation can be computationally intensive, and in some cases, performing it entirely on the local playback device alone may be impractical. In some examples, the generative media module of the local playback device may request generative media content from generative media modules stored on one or more remote computing devices (e.g., cloud servers). This request may include or be based on specific input parameters (e.g., sensor data, user input, context information, etc.). In response to this request, the remote generative media module may stream the specific generative media content to the local device for playback. The specific generative media content provided to the local playback device may change over time, depending on specific input parameters, the configuration of the generative media module, or other such parameters. Additionally or alternatively, the playback device may store discrete tracks for playback (e.g., tracks with different variations associated with different energy levels, such as...). Figure 4 (As depicted). The remote computing device can then periodically provide the local playback device with new files for updating tracks for playback, or alternatively, it can provide the generative media module with updates to determine when and how to play specific files stored locally on the playback device.
[0145] In this way, the tasks required to generate and play back generative audio are distributed among one or more remote computing devices and one or more local playback devices. Overall efficiency can be improved by sending at least some of the computationally intensive tasks associated with generating novel media content to the remote computing devices, and optionally by reducing the need for real-time computing. By generating a discrete number of alternative tracks or track variations based on a specific media content model via the remote computing devices before playback, the local playback devices can request and receive specific variations based on real-time or near-real-time input parameters (e.g., sensor data). For example, the remote computing devices can generate different versions of media content, and the playback devices can request a specific version in real-time based on input parameters. The result is the playback of appropriate generative media content based on real-time or near-real-time input parameters (e.g., sensor data) without the need for real-time regeneration of such media content.
[0146] Figure 6This is a schematic diagram of an example distributed generative media playback system 600. As shown, an artist 602 can provide multiple media segments 604 and one or more generative content models 606 to a generative media module 214 stored via one or more remote computing devices. Media segments may correspond to, for example, specific audio segments or seeds (e.g., individual notes or chords, short tracks of n bars, non-musical content, etc.). In some examples, the generative content model 606 may also be provided by the artist 602. This can include providing the entire model, or the artist 602 can provide input to the model 606, for example, by changing or tuning certain aspects (e.g., rhythm, melodic constraints, harmonic complexity parameters, chord change density parameters, etc.).
[0147] Generative media module 214 can receive both media segment 604 and one or more input parameters 502 (as described elsewhere in this document). Based on these inputs, generative media module 214 can output generative media. Figure 6 As shown, artist 602 can optionally audition the generative media module 214, for example, by receiving exemplary outputs based on inputs provided by artist 602 (e.g., media segment 604 and / or generative content model 606). In some cases, the audition may depend on various different input parameters playing back variations of the generative media content to artist 602 (e.g., one version corresponds to a high energy level intended to produce an exciting or invigorating effect, another version corresponds to a low energy level intended to produce a calming effect, etc.). Based on the output via this audition step, artist 602 can dynamically update the settings of media segment 604 and / or generative content model 606 until the desired output is achieved.
[0148] In the example shown, at box 608, an iteration can occur every n hours (or minutes, days, etc.), during which the generative media module 214 can produce multiple different versions of generative media content. In the example shown, there are three versions: version A in box 610, version B in box 612, and version C in box 614. These outputs are then (e.g., via a remote computing device) stored as generative media content 616. A particular version of these versions (version C in box 618 in this example) can be sent (e.g., streamed) to a local playback device 250 for playback. In some examples, a particular version may correspond to... Figure 4 Tracks 406, 408 and 410 shown.
[0149] Although three versions are shown here as examples, there may actually be many more versions of generative media content generated via remote computing devices. These versions can vary along several different dimensions, such as to suit different energy levels, to suit different intended tasks or activities (e.g., learning versus dancing), to suit different times of day, or any other appropriate variations.
[0150] In the illustrated example, playback device 250 may periodically request a specific version of generated media content from a remote computing device. This request may be based on, for example, user input (e.g., user selection via a controller device), sensor data (e.g., the number of people present in the room, background noise levels, etc.) or other suitable input parameters. As shown, input parameter 502 may optionally be provided to (or detected by) playback device 250. Additionally or alternatively, input parameter 502 may be provided to (or detected by) remote computing device 106. In some examples, playback device 250 sends the input parameter to remote computing device 106, which in turn provides a suitable version to playback device 250 without playback device 250 specifically requesting a particular version.
[0151] g. Example methods for generating digital content based on blockchain data
[0152] As mentioned earlier, in some cases, systems used to generate and play back generative media content can interact with blockchain data (or data stored via other distributed ledger technologies). For example, such as Figure 6As shown, the blockchain layer 620 can be used to provide data as input to other components of the generative media playback system 600, such as playback device 250, input parameters 502, media segment 604, generative content model 606, and / or generative media module 214. In various cases, the blockchain layer 620 can store data that can be used as one or more input parameters 502, including or usable data for obtaining or influencing a specific media segment 604, including or usable data for obtaining or influencing a specific generative content model 606, and / or including or usable data for obtaining or influencing a specific generative media module 214. Furthermore, some or all of these components can communicate with the blockchain layer 620 to write data to the blockchain, record transactions, or otherwise interact with the blockchain layer 620. For example, playback device 250 can record transactions reflecting the playback of a specific track on the blockchain layer 620. Data stored via the blockchain layer 620 may include specific input parameters 502, or data stored via the blockchain layer 620 may be used to generate suitable input parameters 502. Similarly, specific media segments 604, generative content models 606, generative media modules 214, input parameters 502, or other suitable data can be written to the blockchain layer 620 to create an immutable record of such content, transactions, or other data. Additional details regarding the use of blockchain technology (or other suitable distributed ledger technology) in creating and playing back generative media content are described in more detail below.
[0153] Figure 7 This is a schematic diagram of another example of a distributed generative media playback system 700. As shown, system 700 includes or communicates with a blockchain layer 620, for example, to obtain, generate, or store input parameters 502, generative content model 606, or other data or parameters used when creating and playing back generative media content.
[0154] Examples of such blockchain layer 620 include public distributed ledgers such as Ethereum, Solana, Avalanche, and Polygon. While blockchain layer 620 is shown, any suitable distributed ledger technology can be used in various examples, including private or semi-private blockchains and non-blockchain implementations such as directed acyclic graphs (DAGs) (e.g., Nano, IOTA, etc.). In various examples, participants using blockchain layer 620 can transact with each other in a peer-to-peer manner, and the operation of blockchain layer 620 can be distributed, so that no central entity controls the operation of the network. This distributed ledger can be used to track the creation, exchange, and redemption of certain real-world assets (e.g., currencies). Due to the practical immutability of the data stored on the blockchain, this approach enables robust auditing of asset transactions. Currency is just one of many assets that might be expected to be tracked on a distributed ledger. Other types of assets may differ from currencies in how they manage one or more actions related to asset creation, exchange, and / or redemption. Furthermore, different blockchain architectures may differ in their strategies and protocols, and may also differ in the tools used to program the behavior of assets.
[0155] Typically, a distributed ledger where each unit of an asset is represented by some form of digital token can be programmed to assign that token a set of behaviors appropriate to the asset it represents. For example, "fungibility" allows an asset to be exchanged for other assets of the same class. Each unit of a currency of a given denomination (e.g., the US dollar) is fungible because it has the same value as every other unit of the same denomination. In contrast, property ownership is "non-fungible" because its value depends on the size, location, and other aspects of the property being owned. For each asset represented by a token, appropriate fungibility or non-fungibility behaviors are programmed to be used within that token class in the virtual ledger that tracks the asset.
[0156] According to some embodiments, tokens traded via blockchain layer 620 can be non-fungible. Such non-fungible tokens (NFTs) may be unique and not interchangeable with any other tokens. NFTs may include unique digital artwork and / or music, domain names, digital collectibles (e.g., CryptoKitties, memes), event tickets, parts of a virtual world, digital objects used in games, avatars or characters, items with utility (e.g., specific functions such as providing voting or governance rights), and / or associated with them. In various examples, the NFT itself may include associated data (e.g., the original audio data of a music NFT may be stored on-chain), or the NFT may include pointers (e.g., URLs or URIs) to data stored elsewhere (e.g., audio data stored on a server maintained by the music NFT issuer).
[0157] These tokens, whether fungible or nonfungible, can be stored by users via digital wallets. A digital wallet can be a device, physical medium, program, or service that stores the public and / or private keys for blockchain transactions. In some cases, a digital wallet can store multiple public / private key pairs for various different blockchains, allowing users to store assets associated with different blockchains in a single wallet. Examples include MetaMask, Phantom, Coinbase Wallet, Ledger Nano, and others. In operation, users can sign blockchain transactions via their wallets using the appropriate private key (or authorize the wallet to sign transactions using its private key). If the transaction signature is valid, the transaction is then confirmed and added to the corresponding block in the blockchain. In some examples, the wallet identifier itself can be used as an input parameter to the generative model, regardless of any tokens held in a particular wallet.
[0158] In various implementations, blockchain layer 620 can be configured to automatically execute transactions under one or more conditions. Such self-executing transactions can be referred to as "smart contracts." Smart contracts can include computer code stored on the blockchain, configured to execute only under specified conditions and in a specified manner. For example, a smart contract can be configured to execute a specific transaction when a threshold is exceeded, at a specific time, based on one or more other transactions, or any other suitable criterion. In some examples, generative media module 214 and / or generative content model 606 can be implemented as smart contracts such that interaction with the smart contract via blockchain layer 620 results in the smart contract outputting generative media content, a generative content model, or data or instructions capable of generating such generative media content or a generative content model.
[0159] A unique organizational structure in blockchain is the Decentralized Autonomous Organization (DAO). A DAO is typically a community-led entity without a central authority. Such a DAO can be fully autonomous and transparent, with smart contracts providing the underlying rules and executing agreed-upon decisions. Community voting can be conducted via on-chain transactions through token holders. Based on the results of a specific vote, smart contracts can execute certain transactions or other code to implement the decisions of the DAO members. Typically, a DAO issues tokens to users in exchange for monetary investments or donations, or without compensation (e.g., via an "airdrop"). Token holders then typically retain certain voting rights, which can be proportional to the number of tokens they hold. In some cases, token holders also receive financial rewards, such as a portion of the transaction fees collected by the DAO.
[0160] exist Figure 7In the example system 700 shown, the generative media module 214 can receive multiple different inputs and output one or more generative content versions 610 in response. These content versions 610 are stored in a generative media content storage device 616, from which a specific selected generative content version 618 can be selected for playback via playback device 250 or other output devices (e.g., the light component of the generative media content version 618 can cause lighting device 702 to output light, for example, with a specific hue, color temperature, brightness, on / off, or other modes, based on the generative media content version 618). (This is consistent with previous information regarding...) Figure 6 The method described is similar; the inputs to generative media module 214 include generative content model 606 and input parameters 502. Also as previously mentioned, in some cases, generative content model 606 may be stored via blockchain layer 620, or may be obtained from data stored via blockchain layer 620. Similarly, one or more of the input parameters 502 may include or be based on data stored via blockchain layer 620. For example, blockchain data may be used in generative media module 214 to produce suitable outputs, such as the “audioization” of a data stream (e.g., data may be converted into a corresponding audio output). In some cases, blockchain data may include data provided by one or more “oracles,” which are typically third-party services that provide external information to smart contracts (e.g., price feeds, weather data, election results, etc.). In some examples, generative media module 606 and / or generative media content 616 may be stored locally via playback device 250, in which case input parameters 502 may be streamed to playback device 250 to generate a new version 610 of generative media content 616.
[0161] Additionally or alternatively, the generative media module 214 may receive the outputs of one or more smart contracts 706 as input. In some cases, the generative media module 214 itself may take the form of a smart contract, such that the program code is stored on the blockchain and runs automatically under certain conditions (e.g., user 708 interacts with smart contract 706 and, in response, sends specific generative media content to a designated destination). In some examples, the generative media module 214 may run locally or via a remote server (rather than as a smart contract on the blockchain), but may communicate with smart contract 706 to receive input parameters 502 from it or to provide appropriate outputs to smart contract 706. For example, specific generative media content output by the generative media module 214 may be used to generate one or more NFTs via smart contract 706. In some implementations, each specific version of the generated content generated by the generative media module 214 may have a corresponding NFT, such that each NFT generated by smart contract 706 based on inputs from the generative media module 214 is unique. Figure 7 This is illustrated graphically, where multiple discrete NFTs 710a to NFT 710f, labeled NFT1 to NFTn, can be generated via smart contract 706. These NFTs 710 can additionally be provided as input to generative media module 214. For example, generative media module 214 can dynamically generate different content, at least in part, based on a specific NFT 710 with which it interacts. Additionally or alternatively, if user 708 holds the appropriate NFT 710 in her digital wallet, that user may only be able to access a specific generative media module 214.
[0162] In some examples, one or more of the NFTs 710 are existing NFTs owned by an individual third party or entity (e.g., a person or entity unrelated to user 708) and are temporarily accessible by the generative media module 214. In some examples, one or more of the NFTs 710 do not include audio data, but instead include other data types (e.g., video, images, or other data) that can be “sound-ified” or otherwise converted into a form usable for generative media content by the generative media module 214 (or another suitable component).
[0163] In the example shown, smart contract 706 can also output NFT 712, labeled NFT0, which can be held by user 708. Furthermore, NFT 712 can interact with or be generated by artist DAO 714, who in turn can communicate with one or more smart contracts 706. As previously mentioned, DAOs are typically community-led organizations where members hold tokens (e.g., NFT 712) that designate membership, provide voting rights or other governance rights, and optionally entitle token holders to economic benefits, such as proceeds from future DAO revenue. In some examples, artist DAO 714 may hold certain music royalties, which, when collected, can be partially distributed to the holder of the appropriate NFT or other token (e.g., user 708 can receive economic benefits from artist DAO 714 at least in part based on the user's ownership of NFT 712).
[0164] Optionally, the data corresponding to the NFT 712 can be stored or embedded via physical media. For example, such as Figure 7 As shown, the data corresponding to NFT 712 can be embedded in vinyl record 716 (e.g., via a unique QR code, a code embedded in a groove in vinyl record 716), or any other suitable technology for storing the data via physical media (e.g., NFC or other RF tags). Although vinyl record 716 is shown, in various examples, the physical substrate can take many forms, such as a playback device, a physical card or ticket, a poster, etc.
[0165] Including blockchain layer 620, smart contracts 706, DAO 714, and / or NFTs 710 and 712 can provide several benefits to generative media playback system 700. For example, by associating a specific NFT with generative media content (e.g., a soundscape), user 708 can obtain a personalized experience, which may also be transferred via the distributed peer-to-peer transaction mechanism of blockchain layer 620. However, one problem with this approach is that the data included in NFTs is often static, which contradicts the dynamic, context-aware nature of generative soundscapes or other generative media content. Another problem associated with NFTs is link decay, where data in an NFT becomes outdated because the locator in the NFT no longer references the artwork associated with it. One way to mitigate the possibility of link decay is to store the generative media content or generative media engine within blockchain layer 620. Furthermore, in some cases, the artwork itself can be embedded in the NFT (e.g., artwork data is stored on-chain rather than on a separate server).
[0166] In some cases, NFT 710 may include a specific seed used by generative media module 214. Examples of such a seed include... Figure 6 Media segment 604, Figure 4 The track, energy level or metadata, or Figure 3 Any of the various components shown. Optionally, the properties of a particular NFT (at least in relation to its use by the generative media module 14) may depend on its transaction history. For example, the specific seed associated with that NFT may dynamically change depending on when the NFT was last traded or the number of transactions that have taken place. Additionally or alternatively, different combinations of NFTs 710 connected to the generative media module 214 may produce different content versions 610, such that a particular generative media output will depend on which NFTs 710 are used as input.
[0167] In some examples, additional data associated with the generative media playback system 700 can be stored via blockchain layer 620. For example, user 708's listening history can be stored on blockchain layer 620 to provide an immutable listening history. This can include listening history for generative media content or non-generative content (e.g., standard pre-recorded tracks or other content). In some cases, "followers" of a particular user 708 can subscribe to that user's listening history by accessing data stored via blockchain layer 620. Because blockchains are generally permissionless and transparent, followers have free access to listening history (or other content data associated with a specific network address). In another example, followers can subscribe to generative media content for a particular user 708, so that while user 708's generative media content is dynamically created based on various inputs, other followers can enjoy the same media content.
[0168] For example, consider an artist who wants to use generative media module 214 to create a specific soundscape. The artist's fans can listen to the soundscape in real-time or near real-time via blockchain data. In some cases, followers can utilize their own local generative media module 214 (or a generative media module running on another device), which then uses inputs, pointers, or other data from the artist's generative media module 214 to generate corresponding generative media content. In at least some examples, such a local generative media module 214 can also utilize additional local inputs (e.g., specific playback device characteristics, local sensor data, etc.) to merge the artist's generative media content with content generated by the user's own local generative media module 214 to produce slightly altered generative media content that still reflects the artist's intent. Optionally, an artist's followers may need to hold specific NFTs or other tokens to access the artist's generative media content. In another example, specific featured playlists or radio stations may only be accessible to users holding specific NFTs or tokens.
[0169] h. Example methods for generating and playing back generative audio
[0170] Figures 8 to 13 This is a flowchart of an example method for playing back generative audio content via multiple discrete playback devices. Methods 800, 900, 1000, 1100, 1200, and 1300 can be implemented by any device or system described herein or any other device or system now known or hereafter developed.
[0171] Examples of methods 800, 900, 1000, 1100, 1200, and 1300 include one or more operations, functions, or actions represented by boxes. Although the boxes are shown in a sequential order, they may also be executed in parallel and / or in an order different from that disclosed and described herein. Furthermore, depending on the desired implementation, the boxes may be combined into fewer boxes, divided into more boxes, and / or removed.
[0172] Furthermore, for methods 800, 900, 1000, 1100, 1200, 1300, and other processes and methods disclosed herein, flowcharts illustrate possible implementations of functionality and operation for some examples. In this regard, each box may represent a module, segment, or portion of program code, which includes one or more instructions executable by one or more processors for implementing a specific logical function or step in the process. The program code may be stored on any type of computer-readable medium, such as storage devices including disks or hard disk drives. Computer-readable media may include non-transitory computer-readable media, such as tangible non-transitory computer-readable media for short-term data storage, such as register memory, processor cache, and random access memory (RAM). Computer-readable media may also include non-transitory media, such as secondary storage or persistent long-term storage, such as read-only memory (ROM), optical discs or disks, compact disk read-only memory (CD-ROM), etc. Computer-readable media may also be any other volatile or non-volatile storage system. Computer-readable media can be considered as computer-readable storage media, such as tangible storage devices. Furthermore, regarding these methods and other processes and methods disclosed herein, Figures 8 to 13 Each box in the diagram can represent a circuit that is connected to perform a specific logical function in the process.
[0173] refer to Figure 8 Method 800 begins at box 802, which relates to receiving a command to play back generated media content via a group or binding area of a playback device. Such a command may be received via, for example, control device 130 or other suitable user input.
[0174] At box 804, method 800 relates to a group coordinator device providing timing information to generative group member devices. The timing information may include contextual timing data based on a common clock (e.g., time data associated with sensor input or other user input), generative media playback timing data (e.g., timestamps and synchronization data for facilitating synchronized playback of generative media), and / or media content streaming timing data.
[0175] At box 806, the method optionally includes determining a generative media content model to be used to generate generative media. Such a model can be, for example, as described above regarding... Figures 2 to 6The media content module 214 described is implemented in this way. In some examples, each of these member devices may use the same or substantially the same generative media content model, while in others, some or all of these member devices may use generative media content models that differ from each other. For example, a first generative media content model may produce rhythmic beats, while a second generative media content model may produce ambient natural sounds. When the generative audio produced by these different generative media content models is played back simultaneously, it can provide a pleasant listening experience for the user. In some examples, the selection of a particular generative media content model may itself be based on one or more input parameters, such as device capabilities, device location, the number of users present, user sensor data, etc.
[0176] In box 808, method 800 includes the coordinator device and member devices receiving context data and / or other input data. For example, the input data may include sensor data, user input, context data, or any other relevant data that can be used as input to a generative media content model.
[0177] Method 800 continues in box 810, where the coordinator and member devices synchronously generate and play back generative media content.
[0178] Figure 9 Another method 900 for playing back generative audio content via multiple playback devices is illustrated. Method 900 begins at block 902, where one or more input parameters are received at a group coordinator device. As previously described, the input parameters may include sensor data, user input, context data, or any other input that can be used by the generative media module to generate generative audio for playback.
[0179] In box 904, the coordinator device sends input parameters to one or more discrete playback devices that have generative media modules on it. For example, the coordinator device may acquire sensor data and other input parameters and send them to multiple discrete playback devices within an environment or even distributed across multiple environments. In some examples, these input parameters may include characteristics of the generative content model itself, such as providing instructions to update the generative media modules stored locally by one or more of these discrete playback devices.
[0180] In block 906, the method involves sending timing data from a coordinator device to a playback device. The timing data may include, for example, clock data or other synchronization signals configured to facilitate the coordinated generation of generative media content and the synchronized playback of that generative media content via discrete playback devices.
[0181] Method 900 continues in block 908, at least in part, by simultaneously replaying generative media content via playback devices based on input parameters. As previously described, various playback devices may play back the same generative audio, or each playback device may play back different generative audio, which, when played back synchronously, produces the desired psychoacoustic effect for the users present.
[0182] exist Figure 9 In the example, discrete playback devices can generate generative media content locally, and these discrete playback devices play back their own generative audio content in parallel with each other. Figure 10 In the alternative method 1000 shown, the generated media content is generated at the coordinator device, which then sends the generated media content along with timing data to the discrete playback device for synchronous playback.
[0183] In box 1002, method 1000 relates to receiving one or more input parameters at a group coordinator device. Examples of input parameters are described elsewhere in this document and include sensor data, user input, context data, or any other input that can be used by the generative media module to generate generative audio for playback.
[0184] In block 1004, the coordinator device generates a first and a second generative media stream based at least in part on input parameters, and in block 1006, sends the first and second media streams to a first and a second discrete playback device, respectively. For example, the coordinator device may generate two streams forming different channels of generative audio, such as two streams having a left channel to be played back by the first playback device and a corresponding right channel to be played back by the second playback device. Additionally or alternatively, the two streams may be different tracks, but can still be played back synchronously, such as the rhythm and beat in one stream and the ambient natural sounds in the other. Many other variations are possible. Although this example describes two streams for two playback devices, in various other examples, there may be one or more streams that can be provided to any number of playback devices for synchronous playback. In at least some examples, one or more playback devices may be located in different environments that are far apart from each other (e.g., in different homes, different cities, etc.).
[0185] In box 1008, a first playback device plays back a first generated media stream, and a second playback device simultaneously plays back a second generated media stream. In some examples, this simultaneous playback can be facilitated by using timing data received from a coordinator device.
[0186] Figure 11Another example method 1100 for generating and playing back generative media content is shown. As described above, it may be beneficial to use one or more remote computing devices (e.g., cloud-based servers) to perform at least a portion of the processing required to generate generative media content in order to reduce the computational demands placed on a local playback device and / or to use components of the local playback device to perform operations that would be infeasible. Method 1100 begins at block 1102, where one or more input parameters are received at the playback device. As previously described, the input parameters may include sensor data, user input, context data, or any other input that the generative media module may use to generate generative audio for playback.
[0187] In box 1104, method 1100 involves accessing a library comprising multiple pre-existing media segments. For example, multiple discrete media segments (e.g., tracks) may be stored at a playback device and may be arranged and / or mixed according to a generative content model for playback. Additionally or alternatively, the library may be stored on one or more remote computing devices, wherein the individual media segments are sent from the remote computing devices to the playback device for playback.
[0188] Method 1100 continues in box 1106, wherein media content is generated by arranging a selection of pre-existing media segments in the library for playback, based on a generative media content model and at least in part on the input parameters. As described elsewhere in this document, the generative media content model may receive one or more input parameters as input. Based on this input, and using the generative media content model, specific generative media content can be output. In the example, the generative media content may include an arrangement of pre-existing media segments, such as arranging them in a specific order, with or without overlap between specific media segments, and / or performing additional processing or mixing steps to produce the desired output.
[0189] In box 1108, the playback device plays back the generated media content. In various examples, this playback can be performed simultaneously and / or synchronously with an attached playback device.
[0190] Figure 12Another example method 1200 for generating and replaying generative media content is shown. As mentioned above, it can be beneficial to combine or rely on blockchain data to generate generative media content. Method 1200 begins at box 1202, where blockchain data stored via a distributed ledger is accessed via a playback device. The distributed ledger can be a public blockchain, such as Ethereum, Solana, etc., or it can be a private or semi-private blockchain or a non-blockchain ledger. The blockchain data can include one or more pre-existing media segments or other seeds to be used to generate generative media content. In some examples, such data is stored directly on the blockchain itself, while in other cases, the blockchain may store pointers (e.g., URLs or URIs) to the storage location of the media segments or other seed data. Optionally, the blockchain data may take the form of one or more non-fungible tokens (NFTs).
[0191] Method 1200 continues in box 1204, wherein media content is generated, at least in part, based on blockchain data via a playback device. In some examples, this generation may include accessing a library of pre-existing media segments stored on the playback device or other suitable storage location (e.g., a remote server, other devices on a local network, etc.). In some cases, these media segments can be retrieved from the blockchain or other remote location and stored via the playback device. The playback device can then arrange the selection of pre-existing media segments from the library for playback according to a generative media content model. This selection may be at least in part based on blockchain data. For example, as described elsewhere in this document, the generative media content model may communicate with smart contracts or decentralized autonomous organizations (DAOs) in a way that influences specific generative media content output by the model. In some cases, specific NFTs or other tokens may influence the output of the generative media content model. In some cases, such NFTs or other tokens may be combined such that a specific combination of NFTs or other tokens produces a unique output via the generative media content model. In at least some cases, two or more blockchains can be utilized simultaneously in this manner (e.g., a generative media content model can modify the output based on a user holding a first NFT on the Solana network and a second NFT on the Ethereum network). In box 1206, method 1200 relates to replaying the generated media content via a playback device.
[0192] Figure 13Another example method 1300 for generating and replaying generative media content is shown. As described above, smart contracts or other self-executing code can be used to generate, store, and replay generative media content. Method 1300 begins at box 1302, where data associated with a first token is sent over a network to a network address of a distributed ledger via a playback device. This address can be associated with a generative media smart contract configured to generate a model of generative media content. In some examples, the first token can be an NFT, and optionally, multiple such tokens can be sent to the smart contract address.
[0193] In box 1304, method 1300 involves receiving a generative media content model from a network address associated with a generative media smart contract via a playback device. For example, when the smart contract is executed, it can generate a specific generative media content model based at least in part on data associated with a first token. This generative media content model can then be provided to a user to generate novel media content customized based on the first token data. In addition to the first token data, the smart contract can also generate different outputs based on other input parameters described elsewhere (e.g., sensor data, playback device characteristic data, playback device status, user listening history data, etc.), provided that this data is provided to the smart contract address. Additionally or alternatively, the generative media content model provided to the user can output different media content based on one or more of these other input parameters.
[0194] Next, in box 1306, the playback device generates media content based at least in part on a generative media content model. In some examples, this generation may include accessing a library of pre-existing media segments stored on the playback device or other suitable storage location (e.g., a remote server, another device on a local network, etc.). The playback device can then arrange the selection of pre-existing media segments from that library for playback according to the generative media content model. In box 1308, method 1300 relates to playing back the generated media content via the playback device.
[0195] This article describes various examples of generative media playback. Those skilled in the art will understand that a wide variety of different generative media modules, algorithms, inputs, sensor data, and playback device configurations can be conceived and used based on this technology.
[0196] IV. Conclusion
[0197] The above discussion of playback devices, controller devices, playback zone configurations, and media content sources provides only examples of operating environments in which the functions and methods described below can be implemented. Configurations and other operating environments for media playback systems, playback devices, and network devices not explicitly described in this document are also applicable and suitable for implementing the functions and methods.
[0198] The above description discloses, in particular, various example systems, methods, apparatuses, and articles of art, including, especially, firmware and / or software executed on hardware. It should be understood that these examples are illustrative only and should not be considered limiting. For example, it is conceivable that any or all of these firmware, hardware, and / or software aspects or components may be implemented specifically in hardware, specifically in software, specifically in firmware, or in any combination of hardware, software, and / or firmware. Therefore, the examples provided are not the only ways to implement these systems, methods, apparatuses, and / or articles of art.
[0199] Furthermore, references to "example" herein mean that a particular feature, structure, or characteristic described in connection with an example may be included in at least one example or embodiment of the invention. The appearance of this phrase throughout the specification does not necessarily refer to the same example, nor is it a separate or alternative example mutually exclusive with other examples. Therefore, those skilled in the art should understand, explicitly and implicitly, that the examples described herein can be combined with other examples.
[0200] This specification is set forth primarily in terms of illustrative environments, systems, processes, steps, logical blocks, handling, and other symbolic representations that are directly or indirectly similar to the operation of data processing devices coupled to a network. These processing descriptions and representations are commonly used by those skilled in the art to communicate their work to others skilled in the art. Various specific details are set forth to provide a thorough understanding of this disclosure. However, those skilled in the art will understand that certain examples of this technology can be practiced without specific, concrete details. In other instances, well-known methods, processes, components, and circuits have not been described to avoid unnecessarily obscuring aspects of the examples. Therefore, the scope of this disclosure is defined by the appended claims, and not by the description of the foregoing examples.
[0201] When any of the appended claims is understood to cover pure software and / or firmware implementations, at least one element in at least one example is expressly defined herein to include non-transitory tangible media for storing software and / or firmware, such as memory, DVD, CD, Blu-ray, etc.
[0202] For example, the disclosed techniques are illustrated by various examples described below. For convenience, various examples of the disclosed techniques are described as numbered examples (1, 2, 3, etc.). These examples are provided as examples and do not limit the disclosed techniques. Note that any dependent example can be combined in any combination and placed within its respective independent example. Other examples may be presented in a similar manner.
[0203] Example 1: A method comprising: receiving input parameters at a coordinator device; sending the input parameters from the coordinator device to a plurality of playback devices, each playback device having a generative media module therein; and sending timing data from the coordinator device to the plurality of playback devices such that the playback devices simultaneously play back generative media content based at least in part on the input parameters.
[0204] Example 2: The method described in any of the examples in this paper, wherein both the first playback device and the second playback device play back different generative audio content based at least in part on the input parameters.
[0205] Example 3: The method described according to any example in the examples in this document, wherein the input parameters include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, temperature, respiratory rate, EEG)); networked device sensor data (e.g., camera, light, temperature sensor, thermostat, presence detector, microphone); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is paired with another playback device); or user data (e.g., user identifier, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiratory rate, brain activity), user emotion data).
[0206] Example 4: The method described according to any of the examples in this article, wherein the timing data includes at least one of the following: clock data or one or more synchronization signals.
[0207] Example 5: The method described according to any of the examples in this document further includes sending a signal from the coordinator device to at least one of a plurality of playback devices, the signal causing the generative media module of the playback device to be modified.
[0208] Example 6: The method described according to any of the examples in this article, wherein the generative media content includes at least one of the following: generative audio content or generative visual content.
[0209] Example 7: The method described according to any of the examples in this paper, wherein the generative media module includes an algorithm for automatically generating novel media outputs based on inputs that include at least input parameters.
[0210] Example 8: A device includes: a network interface; one or more processors; and a tangible non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the device to perform operations including: receiving input parameters via the network interface; sending the input parameters to a plurality of playback devices via the network interface, each playback device having a generative media module therein; and sending timing data to the plurality of playback devices via the network interface, such that the playback devices simultaneously play back generative media content at least in part based on the input parameters.
[0211] Example 9: A device according to any of the examples in this document, wherein both the first playback device and the second playback device play back different generative audio content based at least in part on input parameters.
[0212] Example 10: A device according to any of the examples in this document, wherein the input parameters include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, temperature, respiratory rate, brain waves)); networked device sensor data (e.g., camera, light, temperature sensor, thermostat, presence detector, microphone); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is paired with another playback device); or user data (e.g., user identifier, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiratory rate, brain activity), user emotion data).
[0213] Example 11: A device according to any of the examples in this document, wherein the timing data includes at least one of the following: clock data or one or more synchronization signals.
[0214] Example 12: A device according to any of the examples in this document, wherein the operation further includes: sending a signal from a coordinator device to at least one of a plurality of playback devices via a network interface, the signal causing the generative media module of the playback device to be modified.
[0215] Example 13: A device according to any of the examples in this document, wherein the generative media content includes at least one of the following: generative audio content or generative visual content.
[0216] Example 14: A device according to any of the examples in this paper, wherein the generative media module includes an algorithm for automatically generating novel media outputs based on inputs that include at least input parameters.
[0217] Example 15: A tangible, non-transitory, computer-readable medium storing instructions that, when executed by one or more processors of a device, cause the device to perform operations including: receiving input parameters at a coordinator device; sending the input parameters from the coordinator device to a plurality of playback devices, each playback device having a generative media module therein; and sending timing data from the coordinator device to the plurality of playback devices such that the playback devices simultaneously play back generative media content, at least in part based on the input parameters.
[0218] Example 16: A computer-readable medium according to any of the examples in this document, wherein both the first playback device and the second playback device play back different generative audio content based at least in part on input parameters.
[0219] Example 17: A computer-readable medium according to any of the examples described herein, wherein the input parameters include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, temperature, respiratory rate, brain waves)); networked device sensor data (e.g., camera, light, temperature sensor, thermostat, presence detector, microphone); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is paired with another playback device); or user data (e.g., user identifier, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiratory rate, brain activity), user emotion data).
[0220] Example 18: A computer-readable medium according to any of the examples in this document, wherein the timing data includes at least one of the following: clock data or one or more synchronization signals.
[0221] Example 19: A computer-readable medium according to any of the examples herein further includes sending a signal from a coordinator device to at least one of a plurality of playback devices, the signal causing the generative media module of the playback device to be modified.
[0222] Example 20: A computer-readable medium according to any of the examples in this document, wherein generative media content includes at least one of the following: generative audio content or generative visual content.
[0223] Example 21: A computer-readable medium according to any of the examples in this document, wherein the generative media module includes an algorithm for automatically generating novel media outputs based on inputs that include at least input parameters.
[0224] Example 22: A method comprising: receiving input parameters at a coordinator device; generating a first media content stream and a second media content stream via a generative media module of the coordinator device; sending the first media content stream to a first playback device via the coordinator device; and sending the second media content stream to a second playback device via the coordinator device such that the first media content stream and the second media content stream are simultaneously played back via the first playback device and the second playback device.
[0225] Example 23: The method described according to any of the examples in this document further includes sending timing data from the coordinator device to each of the first playback device and the second playback device.
[0226] Example 24: The method described according to any of the examples in this article, wherein the timing data includes at least one of the following: clock data or one or more synchronization signals.
[0227] Example 25: The method described according to any of the examples in this article, wherein the first media content stream and the second media content stream are different.
[0228] Example 26: The method described according to any example in the examples herein, wherein the input parameters include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, temperature, respiratory rate, EEG)); networked device sensor data (e.g., camera, light, temperature sensor, thermostat, presence detector); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is paired with another playback device); or user data (e.g., user identifier, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiratory rate, brain activity), user emotion data).
[0229] Example 27: The method described according to any of the examples in this paper further includes modifying the generative media module of the coordinator device.
[0230] Example 28: The method described according to any example in the examples of this article, wherein each of the first generative media content stream and the second generative media content stream includes at least one of the following: generative audio content or generative visual content.
[0231] Example 29: The method described according to any of the examples in this paper, wherein the generative media module includes an algorithm for automatically generating novel media outputs based on inputs that include at least input parameters.
[0232] Example 30: A device includes: a network interface; a generative media module; one or more processors; and a tangible non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the device to perform operations including: receiving input parameters via the network interface; generating a first media content stream and a second media content stream via the generative media module; sending the first media content stream to a first playback device via the network interface; and sending the second media content stream to a second playback device via the network interface, such that the first media content stream and the second media content stream are simultaneously played back via the first playback device and the second playback device.
[0233] Example 31: The device according to any of the examples in this document, wherein the operation further includes: sending timing data to each of the first playback device and the second playback device via a network interface.
[0234] Example 32: A device according to any of the examples in this document, wherein the timing data includes at least one of the following: clock data or one or more synchronization signals.
[0235] Example 33: A device according to any of the examples in this article, wherein the first media content stream and the second media content stream are different.
[0236] Example 34: A device according to any of the examples in this document, wherein the input parameters include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, temperature, respiratory rate, brain waves)); networked device sensor data (e.g., camera, light, temperature sensor, thermostat, presence detector); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is paired with another playback device); or user data (e.g., user identifier, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiratory rate, brain activity), user emotion data).
[0237] Example 35: A device according to any of the examples in this document, wherein the operation further includes modifying the generative media module.
[0238] Example 36: A device according to any of the examples in this document, wherein each of the first and second generative media content streams includes at least one of the following: generative audio content or generative visual content.
[0239] Example 37: A device according to any of the examples in this paper, wherein the generative media module includes an algorithm for automatically generating novel media outputs based on inputs that include at least input parameters.
[0240] Example 38: A tangible, non-transitory, computer-readable medium storing instructions that, when executed by one or more processors of a coordinator device, cause the coordinator device to perform operations including: receiving input parameters at the coordinator device; generating a first media content stream and a second media content stream via a generative media module of the coordinator device; sending the first media content stream to a first playback device via the coordinator device; and sending the second media content stream to a second playback device via the coordinator device such that the first media content stream and the second media content stream are simultaneously played back via the first playback device and the second playback device.
[0241] Example 39: The computer-readable medium according to any example herein also includes sending timing data from the coordinator device to each of the first and second playback devices.
[0242] Example 40: A computer-readable medium according to any of the examples in this document, wherein the timing data includes at least one of the following: clock data or one or more synchronization signals.
[0243] Example 41: A computer-readable medium according to any of the examples in this document, wherein the first media content stream and the second media content stream are different.
[0244] Example 42: A computer-readable medium according to any of the examples described herein, wherein the input parameters include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, temperature, respiratory rate, brain waves)); networked device sensor data (e.g., camera, light, temperature sensor, thermostat, presence detector); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is paired with another playback device); or user data (e.g., user identifier, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiratory rate, brain activity), user emotion data).
[0245] Example 43: A computer-readable medium according to any of the examples in this document, wherein the operation further includes modifying the generative media module of the coordinator device.
[0246] Example 44: A computer-readable medium according to any of the examples in this document, wherein each of the first and second generative media content streams includes at least one of the following: generative audio content or generative visual content.
[0247] Example 45: A computer-readable medium according to any of the examples in this document, wherein the generative media module includes an algorithm for automatically generating novel media outputs based on inputs that include at least input parameters.
[0248] Example 46: A playback device includes: one or more amplifiers configured to drive one or more audio transducers; one or more processors; and a data storage device having instructions that, when executed by the one or more processors, cause the playback device to perform an operation including: receiving one or more first input parameters at the playback device; generating first media content via the playback device based at least in part on the one or more first input parameters, the generation including: accessing a library stored on the playback device, the library comprising a plurality of pre-existing media segments; and arranging a first selection of the pre-existing media segments in the library for playback according to a generative media content model and at least in part based on the one or more input parameters; and playing back the generated first media content via the one or more amplifiers.
[0249] Example 47: A playback device according to any of the examples herein, wherein the operation further includes: receiving at the playback device one or more second input parameters different from the first input parameters; generating second media content, different from the first media content, via the playback device based at least in part on the one or more second input parameters, the generation including: accessing the library; and arranging a second selection of pre-existing media segments in the library for playback based on a generative media content model and at least in part on the one or more second input parameters; and playing back the generated second media content via one or more amplifiers.
[0250] Example 48: The playback device of claim 1, wherein a first selection of arranging the pre-existing media segments in the library for playback comprises arranging two or more media segments of the pre-existing media segments in a manner that is at least partially time-offset.
[0251] Example 49: A playback device according to any of the examples in this document, wherein a first option for arranging pre-existing media segments in the library for playback includes arranging two or more pre-existing media segments in a manner that at least partially overlaps in time.
[0252] Example 50: A playback device according to any of the examples in this document, wherein the first selection of arranging pre-existing media segments in the library for playback includes applying different equalization adjustments to the different pre-existing media segments.
[0253] Example 51: A playback device according to any of the examples in this document, wherein the first selection of the arrangement of pre-existing media segments in the library for playback includes applying varying gain levels to different pre-existing media segments over time.
[0254] Example 52: A playback device according to any of the examples in this document, wherein the first selection of the arrangement of the pre-existing media segments in the library for playback includes randomizing the starting point for playing back a particular pre-existing media segment.
[0255] Example 53: A playback device according to any of the examples in this document, wherein both the generated first media content and the generated second media content include novel media content.
[0256] Example 54: A playback device according to any of the examples in this document, wherein the generated first media content includes audio content, and a plurality of pre-existing media segments include a plurality of pre-existing audio segments.
[0257] Example 55: A playback device according to any of the examples in this document, wherein the generated first media content includes audiovisual content, and the plurality of pre-existing media segments include a plurality of pre-existing audio segments, pre-existing visual media segments, or pre-existing audiovisual media segments.
[0258] Example 56: The playback device according to any of the examples in this document further includes: receiving additional pre-existing media segments via a network interface; and updating the library to include at least the additional pre-existing media segments.
[0259] Example 57: A playback device according to any of the examples in this document, wherein the first input parameter and the second input parameter include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, temperature, respiratory rate, EEG)); networked device sensor data (e.g., camera, light, temperature sensor, thermostat, presence detector, microphone); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is paired with another playback device); or user data (e.g., user identifier, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiratory rate, brain activity, speech characteristics), user emotion data).
[0260] Example 58: A method comprising: receiving one or more first input parameters at a playback device; generating first media content via the playback device based at least in part on the one or more first input parameters, the generation comprising: accessing a library stored on the playback device, the library comprising a plurality of pre-existing media segments; and arranging a first selection of the pre-existing media segments in the library for playback based on a generative media content model and at least in part on the one or more input parameters; and playing back the generated first media content via the playback device.
[0261] Example 59: The method according to any of the examples in this document further includes: receiving, at a playback device, one or more second input parameters different from the first input parameters; generating, via the playback device, second media content different from the first media content based at least in part on the one or more second input parameters, the generation including: accessing the library; and arranging a second selection of pre-existing media segments in the library for playback based on a generative media content model and at least in part on the one or more second input parameters; and playing back the generated second media content via the playback device.
[0262] Example 60: The method according to any example in the examples of this article, wherein a first selection of arranging the pre-existing media segments in the library for playback includes arranging two or more media segments of the pre-existing media segments in a manner that is at least partially time-offset.
[0263] Example 61: The method according to any example in the examples herein, wherein a first option for arranging the pre-existing media segments in the library for playback includes arranging two or more media segments in a manner that at least partially overlaps in time.
[0264] Example 62: The method described according to any of the examples in this paper, wherein arranging the first selection of pre-existing media segments in the library for playback includes applying different equalization adjustments to the different pre-existing media segments.
[0265] Example 63: The method according to any of the examples in this paper, wherein the first selection of arranging the pre-existing media segments in the library for playback includes applying varying gain levels to different pre-existing media segments over time.
[0266] Example 64: The method described according to any of the examples in this paper, wherein the first selection of arranging the pre-existing media segments in the library for playback includes randomizing the starting point for playing back a particular pre-existing media segment.
[0267] Example 65: The method described according to any of the examples in this paper, wherein both the generated first media content and the generated second media content include novel media content.
[0268] Example 66: The method described according to any example in the examples of this article, wherein the generated first media content includes audio content, and the plurality of pre-existing media segments include a plurality of pre-existing audio segments.
[0269] Example 67: The method according to any example in the examples herein, wherein the generated first media content includes audiovisual content, and the plurality of pre-existing media segments include a plurality of pre-existing audio segments, pre-existing visual media segments, or pre-existing audiovisual media segments.
[0270] Example 68: The method according to any of the examples in this document further includes: receiving an additional pre-existing media segment via a network interface; and updating the library to include at least the additional pre-existing media segment.
[0271] Example 69: The method described according to any example in the examples herein, wherein the first input parameter and the second input parameter include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, temperature, respiratory rate, EEG)); networked device sensor data (e.g., camera, light, temperature sensor, thermostat, presence detector, microphone); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is paired with another playback device); or user data (e.g., user identifier, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiratory rate, brain activity, speech characteristics), user emotion data).
[0272] Example 70: A tangible, non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a playback device, cause the playback device to perform operations including: receiving one or more first input parameters at the playback device; generating first media content via the playback device based at least in part on the one or more first input parameters, the generation including: accessing a library stored on the playback device, the library comprising a plurality of pre-existing media segments; arranging a first selection of the pre-existing media segments in the library for playback according to a generative media content model and at least in part based on the one or more input parameters; and playing back the generated first media content via the playback device.
[0273] Example 71: A computer-readable medium according to any of the examples herein, wherein the operation further includes: receiving one or more second input parameters different from the first input parameters at a playback device; generating second media content, different from the first media content, via the playback device based at least in part on the one or more second input parameters, the generation including: accessing the library; arranging a second selection of pre-existing media segments in the library for playback according to a generative media content model and at least in part based on the one or more second input parameters; and playing back the generated second media content via one or more amplifiers.
[0274] Example 72: A computer-readable medium according to any of the examples in this document, wherein a first option for arranging pre-existing media segments in the library for playback includes arranging two or more pre-existing media segments in a manner that is at least partially time-offset.
[0275] Example 73: A computer-readable medium according to any of the examples in this document, wherein a first option for arranging pre-existing media segments in the library for playback includes arranging two or more pre-existing media segments in a manner that at least partially overlaps in time.
[0276] Example 74: A computer-readable medium according to any of the examples in this document, wherein a first selection of arranging pre-existing media segments for playback in the library includes applying different equalization adjustments to the different pre-existing media segments.
[0277] Example 75: A computer-readable medium according to any of the examples in this document, wherein a first selection of the arrangement of pre-existing media segments in the library for playback includes applying varying gain levels to different pre-existing media segments over time.
[0278] Example 76: A computer-readable medium according to any of the examples in this document, wherein a first selection of the arrangement of pre-existing media segments in the library for playback includes randomizing the starting point for playback of a particular pre-existing media segment.
[0279] Example 77: A computer-readable medium according to any of the examples in this document, wherein both the generated first media content and the generated second media content include novel media content.
[0280] Example 78: A computer-readable medium according to any of the examples in this document, wherein the generated first media content includes audio content, and a plurality of pre-existing media segments include a plurality of pre-existing audio segments.
[0281] Example 79: A computer-readable medium according to any of the examples herein, wherein the generated first media content includes audiovisual content, and the plurality of pre-existing media segments include a plurality of pre-existing audio segments, pre-existing visual media segments, or pre-existing audiovisual media segments.
[0282] Example 80: A computer-readable medium according to any of the examples in this document further includes: receiving an additional pre-existing media segment via a network interface; and updating a library to include at least the additional pre-existing media segment.
[0283] Example 81: A computer-readable medium according to any example in the examples herein, wherein the first input parameter and the second input parameter include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, temperature, respiratory rate, brain waves)); networked device sensor data (e.g., camera, light, temperature sensor, thermostat, presence detector, microphone); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is paired with another playback device); or user data (e.g., user identifier, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiratory rate, brain activity, speech characteristics), user emotion data).
[0284] Example 82: A system including a first playback device and a second playback device. The first playback device includes: a first network interface; one or more first processors; and a data storage device having instructions thereon that, when executed by the one or more processors, cause the first playback device to perform an operation including: receiving one or more input parameters; generating media content at least in part based on the one or more input parameters, the generated media content including a first portion and at least a second portion, the generation including: accessing a library stored on the playback device, the library including a plurality of pre-existing media segments; and arranging the selection of pre-existing media segments in the library for playback according to a generative media content model and at least in part based on the one or more input parameters; transmitting a signal via the first network interface, the signal including the second portion of the generated media content and corresponding timing information; and inducing playback of the first portion of the generated media content. The second playback device includes: a second network interface; one or more audio transducers; one or more second processors; and a data storage device having instructions that, when executed by the one or more second processors, cause the second playback device to perform operations including: receiving a transmission signal from the first playback device via the second network interface; and playing back a second portion of the generated media content substantially synchronously with the playback of a first portion of the generated media content via the one or more transducers, based on the timing information.
[0285] Example 83: A system according to any of the examples in this document further includes: a network device including: a third network interface; one or more processors; and a data storage device having instructions that, when executed by the one or more processors, cause the third playback device to perform an operation including: receiving a request from the first playback device via the third network interface on a data network; and, in response to receiving the request, sending an update library of a pre-existing media segment to the first playback device via the third network interface on the data network.
[0286] Example 84: A system according to any of the examples in this document, wherein the network device includes one or more of the following: a remote server, another playback device, a mobile computing device, a laptop computer, or a tablet computer.
[0287] Example 85: A system including a first playback device and a second playback device coupled via a local area network (LAN). The first playback device includes: one or more first processors; one or more first audio transducers; and a data storage device having instructions that, when executed by the one or more first processors, cause the first playback device to perform an operation including: receiving one or more input parameters; generating first media content at least in part based on the one or more input parameters, the generation including: accessing a first library stored on the first playback device, the first library including a plurality of pre-existing media segments; and arranging the selection of pre-existing media segments in the first library for playback according to a first generative media content model and at least in part based on the one or more input parameters; and playing back the first generative media content via the one or more first audio transducers. The second playback device includes: a second network interface;
[0288] One or more second audio transducers; one or more second processors; and a data storage device having instructions that, when executed by the one or more second processors, cause the second playback device to perform operations including: generating second media content at least in part based on one or more input parameters, the generated second media content being substantially identical to generated first media content, the generation including: accessing a second library stored on the second playback device, the second library comprising a plurality of pre-existing media segments; and arranging the selection of pre-existing media segments in the second library for playback according to a second generative media content model and at least in part based on one or more input parameters; and playing back the generated second media content synchronously with the playback of the generated first media content via the one or more second audio transducers.
[0289] Example 86: A system according to any of the examples in this paper, wherein the first generative media content model and the second generative media content model are substantially the same.
[0290] Example 87: A system based on any of the examples in this paper, wherein the first and second libraries are substantially the same.
[0291] Example 88: A method comprising: accessing blockchain data stored on a distributed ledger via a playback device; generating media content via the playback device, at least in part based on the blockchain data, the generation comprising: accessing a library stored on the playback device, the library comprising a plurality of pre-existing media segments; and arranging the selection of pre-existing media segments in the library for playback based on a generative media content model and at least in part based on the blockchain data; and playing back the generated media content via the playback device.
[0292] Example 89: The method according to any of the preceding examples, wherein the NFT data includes one or more pre-existing media segments, and wherein accessing the NFT data includes storing one or more pre-existing media segments in the library.
[0293] Example 90: The method according to any of the preceding examples, wherein the blockchain data includes first non-fungible token (NFT) data, wherein the distributed ledger is a first distributed ledger, the method further comprising: accessing data associated with a second NFT stored on a second distributed ledger via the playback device, and wherein the arrangement of the selection of pre-existing media segments in the library for playback according to the generative media content model is based at least in part on both the first NFT data and the second NFT data.
[0294] Example 91: The method according to any of the preceding examples, wherein the first distributed ledger is associated with a first blockchain layer, and wherein the second distributed ledger is associated with a second blockchain layer different from the first blockchain layer.
[0295] Example 92: The method described in any of the preceding examples, wherein blockchain data is associated with the playlist.
[0296] Example 93: The method according to any of the preceding examples, wherein the blockchain data depends at least in part on transactions involving non-fungible tokens (NFTs) recorded on the distributed ledger.
[0297] Example 94: The method according to any of the preceding examples, wherein the selection of pre-existing media segments in the library based on the generative media content model arrangement is also based at least in part on one or more input parameters.
[0298] Example 95: The method according to any of the preceding examples, wherein the input parameters include one or more of the following: physiological sensor data; networked device sensor data; environmental data; playback device characteristic data; playback device status; or user listening history data; oracle data stored via a distributed ledger; or user data.
[0299] Example 96: The method described according to any of the preceding examples, wherein the user listens to historical data via a distributed ledger.
[0300] Example 97: The method described in any of the preceding examples, wherein accessing blockchain data includes connecting to a user's wallet that holds non-fungible tokens (NFTs).
[0301] Example 98: The method according to any of the preceding examples, wherein accessing blockchain data includes accessing code associated with a physical media object (e.g., a QR code or other code printed on custom vinyl or other media) via a control device.
[0302] Example 99: The method according to any of the preceding examples, wherein arranging the pre-existing media segments in the library for playback includes arranging two or more of the pre-existing media segments in a manner that is at least partially time-offset.
[0303] Example 100: The method according to any of the preceding examples, wherein the selection of pre-existing media segments in the library for playback includes arranging two or more of the pre-existing media segments in a manner that at least partially overlaps in time.
[0304] Example 101: A method comprising: sending data associated with a first token via a network to a network address of a distributed ledger via a playback device, the address being associated with a generative media smart contract configured to generate a generative media content model; receiving the generative media content model from the network address associated with the generative media smart contract via the playback device; generating media content at least in part based on the generative media content model via the playback device, the generation comprising: accessing a library comprising a plurality of pre-existing media segments; and arranging the selection of pre-existing media segments in the library for playback according to the generative media content model; and playing back the generated media content via the playback device.
[0305] Example 102: According to any of the preceding examples, wherein the token data includes first non-fungible token (NFT) data, the method further includes: sending data associated with a second NFT stored on a distributed ledger to a network address associated with the generative media smart contract via the playback device; receiving a second generative media content model from the network address associated with the generative media smart contract via the playback device, the second generative media content model being different from the first generative media content model; generating second media content at least in part based on the second generative media content model via the playback device; and playing back the generated second media content via the playback device.
[0306] Example 103: The method described in any of the preceding examples, wherein token data is associated with a curated playlist.
[0307] Example 104: According to any of the examples above, the token data depends at least in part on transactions involving the token recorded on the distributed ledger.
[0308] Example 105: The method according to any of the preceding examples, wherein the selection of pre-existing media segments in the library based on the generative media content model arrangement is also based at least in part on one or more input parameters.
[0309] Example 106: The method according to any of the preceding examples, wherein the input parameters include one or more of the following: physiological sensor data; networked device sensor data; environmental data; playback device characteristic data; playback device status; or user listening history data; oracle data stored via a distributed ledger; or user data.
[0310] Example 107: The method described according to any of the preceding examples, wherein the user listens to historical data via a distributed ledger.
[0311] Example 108: The method according to any of the foregoing examples further includes: accessing the token data by connecting to the user wallet storing the token data before sending the token data.
[0312] Example 109: The method according to any of the foregoing examples further includes: accessing the token data via a control device, via a code associated with a physical media object (e.g., a QR code or other code printed on custom vinyl or other media), prior to sending the token data.
[0313] Example 110: The method according to any of the preceding examples, wherein arranging the pre-existing media segments in the library for playback includes arranging two or more of the pre-existing media segments in a manner that is at least partially time-offset.
[0314] Example 111: The method according to any of the preceding examples, wherein the selection of pre-existing media segments in the library for playback includes arranging two or more of the pre-existing media segments in a manner that at least partially overlaps in time.
[0315] Example 112: The method described according to any of the foregoing examples, wherein both the generated first media content and the generated second media content include novel media content.
[0316] Example 113: One or more tangible, non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause a media playback system or playback device to perform operations, including: the method according to any of the preceding examples.
[0317] Example 114: A media playback system comprising: one or more processors; and a computer-readable medium according to any of the preceding examples.
[0318] Example 115: A playback device comprising: one or more processors; and a computer-readable medium according to any of the preceding examples.
Claims
1. A computing system, comprising: Network interface; One or more processors: Memory, storing instructions that, when executed by the one or more processors, cause the computing system to perform operations including: Receive first data including one or more input parameters via the network interface; Based on the second data and the received first data, synthetic content is generated via one or more generative machine learning models, wherein the second data includes data corresponding to data retrieved from a blockchain-based distributed ledger, and wherein the one or more generative machine learning models include a first generative machine learning model configured to output a first type of media content and a second generative machine learning model configured to output a second type of media content, wherein the generated synthetic content includes a mixture of the outputs of the first generative machine learning model and the second generative machine learning model; and The generated composite content is sent to the network device via the network interface.
2. The computing system according to claim 1, wherein the first data corresponds to voice data.
3. The computing system of claim 1, wherein the first data corresponds to sensor data received via one or more sensors.
4. The computing system according to claim 3, wherein the sensor data includes at least one of the following: audio sensor data, image sensor data, or biometric sensor data.
5. The computing system of claim 1, wherein the first data includes pre-existing media content.
6. The computing system of claim 1, wherein the first data is received from the network device via at least one of: a local area network, a wide area network, or a cellular telecommunications network.
7. The computing system of claim 1, wherein the second data includes one or more seed parameters.
8. The computing system of claim 1, wherein the network device is a second network device, and wherein the first data is received via the first network device.
9. The computing system of claim 1, wherein the second data includes data associated with nonfungible tokens (NFTs).
10. The computing system of claim 9, wherein the second data includes data associated with two or more NFTs.
11. The computing system of claim 9, wherein the operation further includes receiving temporary access to the NFT via the network interface.
12. The computing system of claim 1, wherein the second data includes data associated with the output of the smart contract.
13. The computing system of claim 1, wherein the second data includes data associated with a decentralized autonomous organization (DAO).
14. The computing system of claim 1, wherein the one or more generative machine learning models are implemented via smart contracts.
15. The computing system of claim 14, wherein the first data includes data received via a blockchain oracle.
16. The computing system of claim 1, wherein the second data includes data associated with accessing the one or more generative machine learning models.
17. The computing system of claim 1, wherein the network device includes a server or other computing device storing a blockchain-based distributed ledger, and wherein the generated synthetic content includes the generated blockchain data.
18. The computing system of claim 17, wherein the generated blockchain data includes non-fungible tokens (NFTs).
19. The computing system of claim 17, wherein the generated blockchain data is produced via a smart contract.
20. The computing system of claim 1, wherein the first data includes media content of a first type, wherein the generated synthetic content includes media content of a second type, and wherein the media content of the second type is different from the media content of the first type.
21. The computing system of claim 1, wherein the first data includes pre-existing media content, and wherein the generated synthetic content includes a combination of the pre-existing media content and novel generative content.
22. The computing system of claim 1, wherein the network device includes a wearable playback device, and wherein the first data includes data corresponding to sensors carried by the wearable playback device.
23. The computing system of claim 22, wherein the second data corresponds to data associated with at least one of the following: the user of the wearable playback device, the characteristics of the wearable playback device, or the transaction history associated with the user or the wearable playback device.
24. The computing system of claim 1, wherein the operation further includes storing the data corresponding to the first data on a distributed ledger based on the blockchain.
25. The computing system of claim 1, wherein the first data includes data indicating the presence of a user or person.
26. The computing system of claim 1, wherein the second data corresponds to a transaction history associated with the network device.