Enhanced dialogue rendering

The system dynamically enhances dialogue in audio content by identifying and adjusting frequency ranges to improve comprehension, addressing the challenge of clear dialogue amidst competing sounds, thus enhancing listener experience.

WO2026085066A1PCT designated stage Publication Date: 2026-04-23SONOS INC
View PDF 12 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SONOS INC
Filing Date
2025-10-14
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Listeners often struggle to hear spoken dialogue clearly in audio content due to competing background sounds, and existing technologies apply dialogue enhancement techniques indiscriminately, affecting listener enjoyment.

Method used

The system dynamically enhances audio signals by identifying dialogue presence through metadata or frequency analysis, adjusting frequency ranges to improve dialogue comprehension while minimizing interference from non-dialogue sounds.

Benefits of technology

Enhances dialogue intelligibility without detracting from other audio content, providing a better listening experience by selectively applying dialogue enhancement techniques only when dialogue is present.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025050817_23042026_PF_FP_ABST
    Figure US2025050817_23042026_PF_FP_ABST
Patent Text Reader

Abstract

An example playback device may be configured to receive media content including first audio and second audio, determine a presence of dialogue in the media content, and generate, responsive to the presence of dialogue, adjusted second audio having reduced sound levels in a first frequency range including a set of frequencies associated with human speech. The playback device may cause playback of the first audio via a first transducer of the playback device and playback of the adjusted second audio via a second transducer of the playback device.
Need to check novelty before this filing date? Find Prior Art

Description

Atorney Docket No. SON00095WOU 1Client Docket No. 24-0404pENHANCED DIALOGUE RENDERINGRELATED APPLICATIONS

[0001] The present application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 707,535 entitled “Enhanced Dialogue Rendering” and filed October 15, 2024.FIELD OF THE DISCLOSURE

[0002] The present disclosure is related to consumer goods and, more particularly, to methods, systems, products, aspects, services, and other elements directed to media playback or some aspect thereof.BACKGROUND

[0003] Options for accessing and listening to digital audio in an out-loud setting were limited until in 2002, when Sonos, Inc. began development of a new type of playback system. Sonos then filed one of its first patent applications in 2003, entitled “Method for Synchronizing Audio Playback between Multiple Networked Devices,” and began offering its first media playback systems for sale in 2005. The SONOS Wireless Home Sound System enables people to experience music from many sources via one or more networked playback devices. Through a software control application installed on a controller (e.g., smartphone, tablet, computer, voice input device), one can play what she wants in any room having a networked playback device. Media content (e.g., songs, podcasts, video sound) can be streamed to playback devices such that each room with a playback device can play back corresponding different media content. In addition, rooms can be grouped together for synchronous playback of the same media content, and / or the same media content can be heard in all rooms synchronously.BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Aspects, and advantages of the presently disclosed technology may be better understood with regard to the following description, appended claims, and accompanying drawings, as listed below. A person skilled in the relevant art will understand that the elements shown in the drawings are for purposes of illustrations, and variations, including different and / or additional elements and arrangements thereof, are possible.

[0005] FIG. 1A is a partial cutaway view of an environment having an example media playback system configured in accordance with aspects of the disclosed technology.

[0006] FIG. IB is a schematic diagram of the media playback system of Figure 1A and one or more networks.Atorney Docket No. SON00095WOU 1 Client Docket No. 24-0404p

[0007] FIG. 1C is a block diagram of an example playback device.

[0008] FIG. ID is a block diagram of an example playback device.

[0009] FIG. IE is a block diagram of an example bonded playback device.

[0010] FIG. IF is a block diagram of an example network microphone device.

[0011] FIG. 1G is a block diagram of an example playback device.

[0012] FIG. 1H is a partial schematic diagram of an example control device.

[0013] FIG. 2 is a flow chart of an example method for enhancing dialogue rendering of an audio portion of media content.

[0014] FIG. 3 is a flow diagram of an example process for enhancing rendering of dialogue within audio content by adjusting both first audio signals containing the dialogue and second audio signals separate from the first audio signals.

[0015] FIG. 4 is a flow diagram of an example process for adjusting dialogue-containing audio using an extracted speech portion of the audio.

[0016] FIG. 5 is a block diagram of an example system and playback device configured for dynamic application of enhanced dialogue rendering.

[0017] The drawings are for the purpose of illustrating example embodiments, but those of ordinary skill in the art will understand that the technology disclosed herein is not limited to the arrangements and / or instrumentality shown in the drawings.DETAILED DESCRIPTIONI. Overview

[0018] Many listeners report being unable to hear spoken dialogue clearly while enjoying video content (e.g., television shows, movies, etc.) on home theater systems. Listeners complain that increasing the volume on their home theater system to better understand the spoken dialogue causes other sounds (e.g., music, explosions, crashes, special effects, etc.) in the video content to be too loud. Conversely, reducing the volume to a comfortable level for experiencing loud background sound elements of the video content can render the spoken dialogue difficult to hear and understand. The problem of understanding dialogue also exists with some types of audio-only content that includes substantial background sound (e.g., music and / or special effects), such as, in some examples, podcasts or audio blogs.

[0019] Sonos has developed technologies to enhance spoken dialogue in audio content to render the spoken dialogue in the audio content easier for listeners to hear and understand.Attorney Docket No. SON00095WOU 1Client Docket No. 24-0404p

[0020] For example, U.S. Patent Number 9,516,440, titled “Providing a multi-channel and a multi- zone audio environment” and issued December 6, 2016, discloses, among other features, playback devices configured to implement a “dialogue enhancement” mode of operation. When playing audio in the “dialogue enhancement” mode described in some embodiments of U.S. Patent Number 9,516,440, a playback device boosts the volume of the center channel speaker, rolls off the bass frequencies, lowers the volume of surround sound satellite speakers, and boosts the speech spectrum (e.g., 300 to 3400 Hz). Playback devices according to some embodiments disclosed in U.S. Patent Number 9,516,440 determine whether and when to activate the “dialogue enhancement” mode of operation based on, in some examples, whether the audio content is associated with video, which interface of which playback device received the audio content (e.g., via an interface connected to a television), and what application was used to initiate playback of the audio content. While operating in the “dialogue enhancement” mode described by Patent Number 9,516,440, the playback device plays all of the audio content according to the “dialogue enhancement” configuration, e.g., rolled off bass frequencies, lowered satellite volumes, and boosted speech spectrum (e.g., 300 to 3400 Hz), thereby helping to make the spoken dialogue in the audio content easier for listeners to hear and understand.

[0021] In another example, U.S. Patent Number 9,226,087, titled “Audio output balancing during synchronized playback” and issued on December 29, 2015, discloses, among other features, playback devices having limiters configured to attenuate audio content above a playback volume threshold such that the output of the playback device is capped at a certain volume level, thereby improving the acoustic output quality of the playback device. Playback devices according to some embodiments of U.S. Patent Number 9,226,087 may implement several different limiters based on, for example, acoustic limits associated with different playback devices configured to play audio content in a playback group and characteristics of the audio content to be played by the playback devices in the playback group. In some instances, use of limiters as described in U.S. Patent Number 9,226,087 can improve a listener's ability to hear and understand spoken dialogue in audio content by limiting the volume of very loud portions of the audio content that may otherwise cause the listener to reduce the volume of their playback device and / or cause distortion of the audio content during playback, which could make the spoken dialogue in the content more difficult to hear and understand absent application of the disclosed limiter features.Attorney Docket No. SON00095WOU 1Client Docket No. 24-0404p

[0022] In a further example, U.S. Patent Application Publication Number 20230195783, titled “Speech enhancement based on metadata associated with audio content” and published June 22, 2023 discusses media playback devices and computing systems designed to determine portions of audio content including speech dialogue using metadata associated with the audio content, and, for portions of the audio content including speech dialogue, identifying settings, such as, in some examples, equalization settings, surround sound settings, and volume settings, for application to the portions of the audio content including the speech dialogue. The dialogue enhancement parameters disclosed in U.S. Patent Application Publication Number 20230195783 include.

[0023] In view of the significant advances described above in making spoken dialogue (sometimes referred to herein simply as dialogue) in audio content easier for listeners to hear and understand, it is desirable to further improve the intelligibility of spoken dialogue in audio content in order to enhance the performance of playback devices and further increase listener enjoyment and satisfaction, particularly in the area of home theater. The embodiments disclosed herein continue to expand upon Sonos's history of innovation in audio technology in general, and home theater in particular, by further improving the intelligibility of dialog in audio content.

[0024] In contrast to systems that are designed to apply dialogue enhancement techniques regardless of whether audio content actually includes dialogue, in one aspect, the present disclosure relates to actively applying dialogue enhancement to specific timeframes of audio content determined to include human speech, while not applying the dialogue enhancement to other (“non-speech”) portions. This type of active dialogue enhancement enables listeners to better hear the dialogue within the audio content, while not detracting from listener enjoyment of other timeframes of the audio content without speech content that may have otherwise been negatively affected by the application of dialogue enhancement techniques.

[0025] In one aspect, the present disclosure relates to dynamically enhancing audio signals within media content responsive to identifying the presence of dialogue within a portion of the audio signals. A playback device, for example, may review streaming media content as it is received to identify the presence of dialogue within certain timeframes. Dynamic enhancement, for example, may include adjusting at least a portion of the audio signals of the media content. The adjustment may involve reducing intensity of competing sound within at least a portion of the frequency range of human speech, for example by attenuating signals within the frequency range with a predetermined gain reduction. Through applying a gain reduction to at least the frequencyAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p spectrum of the dialogue in a present timeframe of the media content, for example, non-dialogue sounds having the same frequenc(ies) will no longer be played at an intensity that could obscure comprehension of the dialogue. The non-dialogue audio signals may be attenuated, in a particular example, by applying a scoop filter to portions of the audio content not containing dialogue. The portions of the audio signals lacking dialogue, in illustration, may include right and / or left channel audio streams, where the center channel audio stream contains the dialogue. The adjustment may further involve increasing intensity of the portion of the audio signals including dialogue. For example, at least a portion of the frequency range of human speech may be amplified within a portion of the audio signals determined to contain dialogue. The portion of the audio signals having dialogue, for example, may include the center channel audio stream. The presence of dialogue may be identified, in some examples, from a metadata portion of the media content, through analysis of the frequency contents of audio signals, and / or through recognition of data provided within a stream of the media content dedicated to dialogue content.

[0026] In one aspect, the present disclosure relates to analyzing an audio portion of media content for the presence of dialogue and adjusting at least a portion of the audio signals of the media content to enhance comprehension of dialogue within the audio signals. Analyzing the audio portion, for example, may include, in some examples, performing a frequency spectrum analysis on the audio portion to identify signals within at least a subset of the frequency range of human speech. In another example, analyzing the audio portion may include applying one or more machine learning classifiers and / or artificial intelligence processes trained to determine a likelihood of dialogue content within the audio portion. In a further example, analyzing the audio portion may include comparing relative energy levels between portions of the audio signals received in relation to media content. Upon determining a sufficient likelihood of the presence of dialogue, for example, aspects of the audio signals may be adjusted to improve comprehension of the dialogue portion of the audio signals.

[0027] In one aspect, the present disclosure relates to extracting a dialogue portion of an audio signal stream for applying to enhancing comprehension of the dialogue through adjusting a frequency spectrum of the audio signal stream corresponding to the extracted dialogue portion. Extracting the dialogue portion may include identifying, from streaming media content, dialogue audio data through frame-by-frame analysis of an audio portion of the streaming media content. Identifying the dialogue audio data, in a first example, may include developing, from the dialogueAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p audio data, a speech mask. The speech mask may be used to adjust frequencies within the audio portion of the streaming audio content matching speech frequenc(ies) detected within a given frame of the audio portion. Identifying the dialogue audio data, in a second example, may include applying one or more machine learning processes to detect the speech audio data within the audio portion of the streaming media content. For example, certain systems and methods described herein employ at least one parametric machine learning model configured to dynamically differentiate over time between speech audio segments and non-speech audio segments within the audio data portion of the incoming stream of media content. The at least one parametric machine learning model, for example, may be configured to output a likelihood of speech content within a given frame of the audio portion of the streaming media content. The speech audio portion may be extracted responsive to the at least one parametric machine learning model indicating at least a threshold likelihood of speech content within the streaming media content.

[0028] While some examples described herein may refer to functions performed by given actors such as “users,” “listeners,” and / or other entities, it should be understood that such references are for purposes of explanation only. The claims should not be interpreted to require action by any such example actor unless explicitly required by the language of the claims themselves.

[0029] In the Figures, identical reference numbers identify generally similar, and / or identical, elements. To facilitate the discussion of any particular element, the most significant digit or digits of a reference number refers to the Figure in which that element is first introduced. For example, element 110a is first introduced and discussed with reference to FIG. 1A. Many of the details, dimensions, angles, and other features shown in the Figures are merely illustrative of particular embodiments of the disclosed technology. Accordingly, other embodiments can have other details, dimensions, angles, and features without departing from the spirit or scope of the disclosure. In addition, those of ordinary skill in the art will appreciate that further embodiments of the various disclosed technologies can be practiced without several of the details described below.II. Suitable Operating Environment

[0030] FIG. 1A is a partial cutaway view of a media playback system (MPS) 100 distributed in an environment 101 (e.g., a house). In the illustrated embodiment of FIG. 1A, the environment 101 includes a household having several rooms, spaces, and / or playback zones, including (clockwise from upper left) a master bathroom 101a, a master bedroom 101b, a second bedroom 101c, a family room or den 101 d, an office lOle, a living room 10 If, a dining room 101g, a kitchen lOlh, and anAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p outdoor patio lOli. While certain embodiments and examples are described below in the context of a home environment, the technologies described herein may be implemented in other types of environments. In some embodiments, for example, the media playback system 100 can be implemented in one or more commercial settings (e.g., a restaurant, mall, airport, hotel, a retail or other store), one or more vehicles (e.g., a sports utility vehicle, bus, car, a ship, a boat, an airplane, etc.), multiple environments (e.g., a combination of home and vehicle environments), and / or another suitable environment where multi-zone audio may be desirable.

[0031] Within the rooms and spaces of the environment 101, the MPS 100 includes one or more playback devices 110 (identified individually as playback devices HOa-n), one or more network microphone devices 120 (“NMDs”) (identified individually as NMDs 120a-c), and one or more control devices 130 (identified individually as control devices 130a and 130b).

[0032] As used herein the term “playback device” can generally refer to a network device configured to receive, process, and output data of a media playback system. For example, a playback device can be a network device that receives and processes audio content. In some embodiments, a playback device includes one or more transducers or speakers powered by one or more amplifiers. In other embodiments, however, a playback device includes one of (or neither of) the speaker and the amplifier. For instance, a playback device can have one or more amplifiers configured to drive one or more speakers external to the playback device via a corresponding wire or cable.

[0033] Moreover, as used herein the term “NMD” (i.e., a “network microphone device”) can generally refer to a network device that is configured for audio detection. In some embodiments, an NMD is a stand-alone device configured primarily for audio detection. A stand-alone NMD 120 may omit components and / or functionality that is typically included in a playback device 110, such as a speaker or related electronics. For instance, in such cases, a stand-alone NMD may not produce audio output or may produce limited audio output. In other embodiments, an NMD is incorporated into a playback device (or vice versa). A playback device 110 that includes components and functionality of an NMD 120 may be referred to as being “NMD-equipped.” Examples of playback devices 110 and NMDs 120 are described further below.

[0034] The term “control device” can generally refer to a network device configured to perform functions relevant to facilitating user access, control, and / or configuration of the media playback system 100. Examples of control devices are described further below.Attorney Docket No. SON00095WOU 1Client Docket No. 24-0404p

[0035] In some examples, one or more of the various playback devices 1 10 may be configured as portable playback devices, while others may be configured as stationary playback devices. For example, certain playback devices 110 may include an internal power source (e.g., a rechargeable battery) that allows the playback device to operate without being physically connected to a mains electrical outlet or the like. In this regard, such a playback device may be referred to herein as a “portable playback device.” On the other hand, playback devices that are configured to rely on power from a mains electrical outlet or the like may be referred to herein as “stationary playback devices,” although such devices may in fact be moved around a home or other environment. In practice, a person might often take a portable playback device to and from a home or other environment in which one or more stationary playback devices remain.

[0036] Each of the playback devices 110 is configured to receive audio signals or data from one or more media sources (e.g., one or more remote servers, one or more local devices, etc.) and play back the received audio signals or data as sound. The one or more NMDs 120 are configured to receive spoken word commands, and the one or more control devices 130 are configured to receive user input. In response to the received spoken word commands and / or user input, the media playback system 100 can play back audio via one or more of the playback devices 110. In certain embodiments, the playback devices 110 are configured to commence playback of media content in response to a trigger. For instance, one or more of the playback devices 110 can be configured to play back a morning playlist upon detection of an associated trigger condition (e.g., presence of a user in a kitchen, detection of a coffee machine operation, etc.). In some embodiments, for example, the media playback system 100 is configured to play back audio from a first playback device (e.g., the playback device 110a) in synchrony with a second playback device (e.g., the playback device 110b). Interactions between the playback devices 110, NMDs 120, and / or control devices 130 of the media playback system 100 configured in accordance with the various embodiments of the disclosure are described in greater detail below with respect to FIGS. 1B-1M.

[0037] The media playback system 100 can include one or more playback zones, some of which may correspond to the rooms in the environment 101. The media playback system 100 can be established with one or more playback zones, after which additional zones may be added, or removed, to form, for example, the configuration shown in FIG. 1A. Each zone may be given a name according to a different room or space such as the office lOle, master bathroom 101a, master bedroom 101b, the second bedroom 101c, kitchen lOlh, dining room 101g, living room 10 If,Attorney Docket No. SON00095WOU 1Client Docket No. 24-0404p and / or the balcony lOli. In some aspects, a single playback zone may include multiple rooms or spaces. In certain aspects, a single room or space may include multiple playback zones.

[0038] In the illustrated embodiment of FIG. 1A, the second bedroom 101c, the office 101 e, the living room 10 If, the dining room 101g, the kitchen lOlh, and the outdoor patio lOli each include one playback device 110, and the master bathroom 101a, the master bedroom 101b, and the den lOld include a collection of playback devices 110. In the master bedroom 101b, the playback devices 1101 and 110m may be configured, for example, to play back audio content in synchrony as individual ones of playback devices 110, as a bonded playback zone, as a consolidated playback device, and / or any combination thereof. Similarly, in the den 101 d, the playback devices 1 lOh-k can be configured, for instance, to play back audio content in synchrony as individual ones of playback devices 110, as one or more bonded playback devices, and / or as one or more consolidated playback devices. Additional details regarding bonded and consolidated playback devices are described below with respect to FIGS. IB, IE, and 1I-M.

[0039] In some aspects, one or more of the playback zones in the environment 101 may each be playing different audio content. For instance, a user may be grilling on the patio lOli and listening to hip hop music being played by the playback device 110c while another user is preparing food in the kitchen 101 h and listening to classical music played by the playback device 110b. In another example, a playback zone may play the same audio content in synchrony with another playback zone. For instance, the user may be in the office lOle listening to the playback device 1 lOf playing back the same hip hop music being played back by playback device 110c on the patio lOli. In some aspects, the playback devices 110c and 1 lOf play back the hip hop music in synchrony such that the user perceives that the audio content is being played seamlessly (or at least substantially seamlessly) while moving between different playback zones. Additional details regarding audio playback synchronization among playback devices and / or zones can be found, for example, in U.S. Patent No. 8,234,395 entitled, “System and method for synchronizing operations among a plurality of independently clocked digital data processing devices,” which is incorporated herein by reference in its entirety. a. Suitable Media Playback System

[0040] FIG. IB is a schematic diagram of the media playback system 100 and a cloud network 102. For ease of illustration, certain devices of the media playback system 100 and the cloud network 102 are omitted from FIG. IB. One or more communication links 103 (referred toAtorney Docket No. SON00095WOU 1Client Docket No. 24-0404p hereinafter as “the links 103”) communicatively couple the media playback system 100 and the cloud network 102.

[0041] The links 103 can include, for example, one or more wired networks, one or more wireless networks, one or more wide area networks (WAN), one or more local area networks (LAN), one or more personal area networks (PAN), one or more telecommunication networks (e.g., one or more Global System for Mobiles (GSM) networks, Code Division Multiple Access (CDMA) networks, Long-Term Evolution (LTE) networks, 5G communication networks, and / or other suitable data transmission protocol networks), etc. The cloud network 102 is configured to deliver media content (e.g., audio content, video content, photographs, social media content, etc.) to the media playback system 100 in response to a request transmitted from the media playback system 100 via the links 103. In some embodiments, the cloud network 102 is further configured to receive data (e.g., voice input data) from the media playback system 100 and correspondingly transmit commands and / or media content to the media playback system 100.

[0042] The cloud network 102 includes computing devices 106 (identified separately as a first computing device 106a, a second computing device 106b, and a third computing device 106c). The computing devices 106 can include individual computers or servers, such as, for example, a media streaming service server storing audio and / or other media content, a voice service server, a social media server, a media playback system control server, etc. In some embodiments, one or more of the computing devices 106 include modules of a single computer or server. In certain embodiments, one or more of the computing devices 106 include one or more modules, computers, and / or servers. Moreover, while the cloud network 102 is described above in the context of a single cloud network, in some embodiments the cloud network 102 includes a collection of cloud networks including communicatively coupled computing devices. Furthermore, while the cloud network 102 is shown in FIG. IB as having three of the computing devices 106, in some embodiments, the cloud network 102 has fewer (or more than) three computing devices 106.

[0043] The media playback system 100 is configured to receive media content from the networks102 via the links 103. The received media content can include, for example, a Uniform Resource Identifier (URI) and / or a Uniform Resource Locator (URL). For instance, in some examples, the media playback system 100 can stream, download, or otherwise obtain data from a URI or a URL corresponding to the received media content. A network 104 communicatively couples the links103 and at least a portion of the devices (e.g., one or more of the playback devices 110, NMDsAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p120, and / or control devices 130) of the media playback system 100. The network 104 can include, for example, a wireless network (e.g., a WI-FI network, a BLUETOOTH network, a Z-WAVE network, a ZIGBEE network, and / or other suitable wireless communication protocol network) and / or a wired network (e.g., a network such as Ethernet, Universal Serial Bus (USB), and / or another suitable wired communication). As those of ordinary skill in the art will appreciate, as used herein, “WI-FI” can refer to several different communication protocols including, for example, Institute of Electrical and Electronics Engineers (IEEE) 802.11a, 802.11b, 802.11g, 802. l ln, 802.11ac, 802.11ac, 802.11ad, 802.11af, 802.11ah, 802.11ai, 802.11aj, 802.11aq, 802.1 lax, 802.1 lay, 802.15, etc. transmitted at 2.4 Gigahertz (GHz), 5 GHz, and / or another suitable frequency.

[0044] In some embodiments, the network 104 includes a dedicated communication network that the media playback system 100 uses to transmit messages between individual devices and / or to transmit media content to and from media content sources (e.g., one or more of the computing devices 106). In certain embodiments, the network 104 is configured to be accessible only to devices in the media playback system 100, thereby reducing interference and competition with other household devices. In other embodiments, however, the network 104 includes an existing household or commercial facility communication network (e.g., a household or commercial facility WI-FI network). In some embodiments, the links 103 and the network 104 include one or more of the same networks. In some aspects, for example, the links 103 and the network 104 may include a telecommunication network (e.g., an LTE network, a 5G network, etc ). Moreover, in some embodiments, the media playback system 100 is implemented without the network 104, and devices including the media playback system 100 can communicate with each other, for example, via one or more direct connections, PANs, telecommunication networks, and / or other suitable communication links. The network 104 may be referred to herein as a “local communication network” to differentiate the network 104 from the cloud network 102 that couples the media playback system 100 to remote devices, such as cloud servers that host cloud services.

[0045] In some embodiments, audio content sources may be regularly added or removed from the media playback system 100. In some embodiments, for example, the media playback system 100 performs an indexing of media items when one or more media content sources are updated, added to, and / or removed from the media playback system 100. The media playback system 100 can scan identifiable media items in some or all folders and / or directories accessible to the playback devicesAtorney Docket No. SON00095WOU 1Client Docket No. 24-0404p110 and generate or update a media content database including metadata (e.g., title, artist, album, track length, etc.) and other associated information (e.g., URIs, URLs, etc.) for each identifiable media item found. In some embodiments, for example, the media content database is stored on one or more of the playback devices 110, network microphone devices 120, and / or control devices 130.

[0046] In the illustrated embodiment of FIG. IB, the playback devices 1101 and 110m form a group 107a. The playback devices 1101 and 110m can be positioned in different rooms and be grouped together in the group 107a on a temporary or permanent basis based on user input received at the control device 130a and / or another control device 130 in the media playback system 100. When arranged in the group 107a, the playback devices 1101 and 110m can be configured to play back the same or similar audio content in synchrony from one or more audio content sources. In certain embodiments, for example, the group 107a includes a bonded zone in which the playback devices 1101 and 110m have left audio and right audio channels, respectively, of multi-channel audio content, thereby producing or enhancing a stereo effect of the audio content. In some embodiments, the group 107a includes additional playback devices 110. In other embodiments, however, the media playback system 100 omits the group 107a and / or other grouped arrangements of the playback devices 110. Additional details regarding groups and other arrangements of playback devices are described in further detail below with respect to FIGS. II through IM.

[0047] The media playback system 100 includes the NMDs 120a and 120b, each including one or more microphones configured to receive voice utterances from a user. In the illustrated embodiment of FIG. IB, the NMD 120a is a standalone device and the NMD 120b is integrated into the playback device 1 lOn. The NMD 120a, for example, is configured to receive voice input 121 from a user 123. In some embodiments, the NMD 120a transmits data associated with the received voice input 121 to a voice assistant service (VAS) configured to (i) process the received voice input data and (ii) facilitate one or more operations on behalf of the media playback system 100.

[0048] In some aspects, for example, the computing device 106c includes one or more modules and / or servers of a VAS (e.g., a VAS operated by one or more of SONOS, AMAZON, GOOGLE, APPLE, MICROSOFT, etc.). The computing device 106c can receive the voice input data from the NMD 120a via the network 104 and the links 103.Attorney Docket No. SON00095WOU 1Client Docket No. 24-0404p

[0049] In response to receiving the voice input data, the computing device 106c processes the voice input data (i.e., “Play Hey Jude by The Beatles”), and determines that the processed voice input includes a command to play a song (e.g., “Hey Jude”). In some embodiments, after processing the voice input, the computing device 106c accordingly transmits commands to the media playback system 100 to play back “Hey Jude” by the Beatles from a suitable media service (e.g., via one or more of the computing devices 106) on one or more of the playback devices 110. In other embodiments, the computing device 106c may be configured to interface with media services on behalf of the media playback system 100. In such embodiments, after processing the voice input, instead of the computing device 106c transmitting commands to the media playback system 100 causing the media playback system 100 to retrieve the requested media from a suitable media service, the computing device 106c itself causes a suitable media service to provide the requested media to the media playback system 100 in accordance with the user’s voice utterance, b. Suitable Playback Devices

[0050] FIG. 1C is a block diagram of the playback device 110a including an input / output 111. The input / output 111 can include an analog I / O I l la (e.g., one or more wires, cables, and / or other suitable communication links configured to carry analog signals) and / or a digital I / O 111b (e.g., one or more wires, cables, or other suitable communication links configured to carry digital signals). In some embodiments, the analog VO 11 la is an audio line-in input connection including, for example, an auto-detecting 3.5mm audio line-in connection. In some embodiments, the digital I / O 111b includes a Sony / Philips Digital Interface Format (S / PDIF) communication interface and / or cable and / or a Toshiba Link (TOSLINK) cable. In some embodiments, the digital I / O 111b includes a High-Definition Multimedia Interface (HDMI) interface and / or cable. In some embodiments, the digital I / O 111b includes one or more wireless communication links such as, in some examples, a radio frequency (RF), infrared, WI-FI, BLUETOOTH, or another suitable communication link. In certain embodiments, the analog VO 11 la and the digital VO 11 lb includes interfaces (e.g., ports, plugs, jacks, etc.) configured to receive connectors of cables transmitting analog and digital signals, respectively, without necessarily including cables.

[0051] The playback device 110a, for example, can receive media content (e.g., audio content including music and / or other sounds) from a local audio source 105 via the input / output 111 (e.g., a cable, a wire, a PAN, a BLUETOOTH connection, an ad hoc wired or wireless communication network, and / or another suitable communication link). The local audio source 105 can be, in someAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p examples, a mobile device (e.g., a smartphone, a tablet, a laptop computer, etc.) or another suitable audio component (e.g., a television, a desktop computer, an amplifier, a phonograph (such as an LP turntable), a Blu-ray player, a memory storing digital media files, etc.). In some aspects, the local audio source 105 includes local music libraries on a smartphone, a computer, a networked- attached storage (NAS), and / or another suitable device configured to store media files. In certain embodiments, one or more of the playback devices 110, NMDs 120, and / or control devices 130 include the local audio source 105. In other embodiments, however, the media playback system omits the local audio source 105 altogether. In some embodiments, the playback device 110a does not include an input / output 111 and receives all audio content via the network 104.

[0052] In some embodiments, the playback device 110a further includes electronics 112, a user interface 113 (e.g., one or more buttons, knobs, dials, touch-sensitive surfaces, displays, touchscreens, etc.), and one or more transducers 114 (referred to hereinafter as “the transducers 114”). The electronics 112 are configured to receive audio from an audio source (e.g., the local audio source 105) via the input / output 111 or one or more of the computing devices 106a-c via the network 104 (FIG. IB), amplify the received audio, and output the amplified audio for playback via one or more of the transducers 114. In some embodiments, the playback device 110a optionally includes one or more microphones 115 (e.g., a single microphone, a collection of microphones, a microphone array) (hereinafter referred to as “the microphones 115”). In certain embodiments, for example, the playback device 110a having one or more of the optional microphones 115 can operate as an NMD configured to receive voice input from a user and correspondingly perform one or more operations based on the received voice input.

[0053] In the illustrated embodiment of FIG. 1C, the electronics 112 include one or more processors 112a (referred to hereinafter as “the processors 112a”), memory 112b, software components 112c, a network interface 112d, one or more audio processing components 112g (referred to hereinafter as “the audio components 112g”), one or more audio amplifiers 112h (referred to hereinafter as “the amplifiers 112h”), and power 112i (e.g., one or more power supplies, power cables, power receptacles, batteries, induction coils, Power-over Ethernet (POE) interfaces, and / or other suitable sources of electric power). In some embodiments, the electronics 112 optionally include one or more other components 112j (e.g., one or more sensors, video displays, touchscreens, battery charging bases, etc.).Attorney Docket No. SON00095WOU 1Client Docket No. 24-0404p

[0054] The processors 1 12a can include clock-driven computing component(s) configured to process data, and the memory 112b can include a computer-readable medium (e.g., a tangible, non-transitory computer-readable medium loaded with one or more of the software components 112c) configured to store instructions for performing various operations and / or functions. The processors 112a are configured to execute the instructions stored on the memory 112b to perform one or more of the operations. The operations can include, for example, causing the playback device 110a to retrieve audio data from an audio source (e.g., one or more of the computing devices 106a-c (FIG. IB)), and / or another one of the playback devices 110. In some embodiments, the operations further include causing the playback device 110a to send audio data to another one of the playback devices 110a and / or another device (e.g., one of the NMDs 120). Certain embodiments include operations causing the playback device 110a to pair with another of the one or more playback devices 110 to enable a multi-channel audio environment (e.g., a stereo pair, a bonded zone, etc.).

[0055] The processors 112a can be further configured to perform operations causing the playback device 110a to synchronize playback of audio content with another of the one or more playback devices 110. As those of ordinary skill in the art will appreciate, during synchronous playback of audio content on a collection of playback devices, a listener will preferably be unable to perceive time-delay differences between playback of the audio content by the playback device 110a and the other one or more other playback devices 110. Additional details regarding audio playback synchronization among playback devices can be found, for example, in U.S. Patent No. 8,234,395, which was incorporated by reference above.

[0056] In some embodiments, the memory 112b is further configured to store data associated with the playback device 110a, such as one or more zones and / or zone groups of which the playback device 110a is a member, audio sources accessible to the playback device 110a, and / or a playback queue that the playback device 110a (and / or another of the one or more playback devices) can be associated with. The stored data can include one or more state variables that are periodically updated and used to describe a state of the playback device 110a. The memory 112b can also include data associated with a state of one or more of the other devices (e.g., the playback devices 110, NMDs 120, control devices 130) of the media playback system 100. In some aspects, for example, the state data is shared during predetermined intervals of time (e.g., every 5 seconds, every 10 seconds, every 60 seconds, etc.) among at least a portion of the devices of the mediaAtorney Docket No. SON00095WOU 1Client Docket No. 24-0404p playback system 100, so that one or more of the devices have the most recent data associated with the media playback system 100.

[0057] The network interface 112d is configured to facilitate a transmission of data between the playback device 110a and one or more other devices on a data network such as, for example, the links 103 and / or the network 104 (FIG. IB). The network interface 112d is configured to transmit and receive data corresponding to media content (e.g., audio content, video content, text, photographs) and other signals (e.g., non-transitory signals) including digital packet data including an Internet Protocol (IP)-based source address and / or an IP -based destination address. The network interface 112d can parse the digital packet data such that the electronics 112 properly receive and process the data destined for the playback device 110a.

[0058] In the illustrated embodiment of FIG. 1C, the network interface 112d includes one or more wireless interfaces 112e (referred to hereinafter as “the wireless interface 112e”). The wireless interface 112e (e.g., a suitable interface having one or more antennae) can be configured to wirelessly communicate with one or more other devices (e.g., one or more of the other playback devices 110, NMDs 120, and / or control devices 130) that are communicatively coupled to the network 104 (FIG. IB) in accordance with a suitable wireless communication protocol (e.g., WIFI, BLUETOOTH, LTE, etc ). In some embodiments, the network interface 112d optionally includes a wired interface 112f (e.g., an interface or receptacle configured to receive a network cable such as an Ethernet, a USB-A, USB-C, and / or Thunderbolt cable) configured to communicate over a wired connection with other devices in accordance with a suitable wired communication protocol. In certain embodiments, the network interface 112d includes the wired interface 112f and excludes the wireless interface 112e. In some embodiments, the electronics 112 exclude the network interface 112d altogether and transmit and receive media content and / or other data via another communication path (e.g., the input / output 111).

[0059] The audio components 112g are configured to process and / or filter data including media content received by the electronics 112 (e.g., via the input / output 111 and / or the network interface 112d) to produce output audio signals. In some embodiments, the audio processing components 112g include, for example, one or more digital-to-analog converters (DACs), audio preprocessing components, audio enhancement components, digital signal processors (DSPs), and / or other suitable audio processing components, modules, circuits, etc. In certain embodiments, one or more of the audio processing components 112g can include one or more subcomponents of theAtorney Docket No. SON00095WOU 1Client Docket No. 24-0404p processors 112a. In some embodiments, the electronics 112 omit the audio processing components 112g. In some aspects, for example, the processors 112a execute instructions stored on the memory 112b to perform audio processing operations to produce the output audio signals.

[0060] The amplifiers 112h are configured to receive and amplify the audio output signals produced by the audio processing components 112g and / or the processors 112a. The amplifiers 112h can include electronic devices and / or components configured to amplify audio signals to levels sufficient for driving one or more of the transducers 114. In some embodiments, for example, the amplifiers 112h include one or more switching or class-D power amplifiers. In other embodiments, however, the amplifiers 112h include one or more other types of power amplifiers (e.g., linear gain power amplifiers, class-A amplifiers, class-B amplifiers, class-AB amplifiers, class-C amplifiers, class-D amplifiers, class-E amplifiers, class-F amplifiers, class-G amplifiers, class H amplifiers, and / or another suitable type of power amplifier). In certain embodiments, the amplifiers 112h include a suitable combination of two or more of the foregoing types of power amplifiers. Moreover, in some embodiments, individual ones of the amplifiers 112h correspond to individual ones of the transducers 114. In other embodiments, however, the electronics 112 include a single one of the amplifiers 112h configured to output amplified audio signals to the transducers 114. In some other embodiments, the electronics 112 omit the amplifiers 112h.

[0061] The transducers 114 (e.g., one or more speakers and / or speaker drivers) receive the amplified audio signals from the amplifier 112h and render or output the amplified audio signals as sound (e.g., audible sound waves having a frequency between about 20 Hertz (Hz) and 20 kilohertz (kHz)). In some embodiments, the transducers 114 represent a single transducer. In other embodiments, however, the transducers 114 include multiple audio transducers. In some embodiments, the transducers 114 include more than one type of transducer. For example, the transducers 114 can include one or more low frequency transducers (e g., subwoofers, woofers), mid-range frequency transducers (e.g., mid-range transducers, mid-woofers), and one or more high frequency transducers (e.g., one or more tweeters). As used herein, “low frequency” can generally refer to audible frequencies below about 500 Hz, “mid-range frequency” can generally refer to audible frequencies between about 500 Hz and about 2 kHz, and “high frequency” can generally refer to audible frequencies above 2 kHz. In certain embodiments, however, one or more of the transducers 114 include transducers that do not adhere to the foregoing frequency ranges. ForAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p example, one of the transducers 114 may include a mid-woofer transducer configured to output sound at frequencies between about 200 Hz and about 5 kHz.

[0062] By way of illustration, Sonos, Inc. presently offers (or has offered) for sale certain playback devices including, for example, a “SONOS ONE,” “PLAY:1,” “PLAY:3,” “PLAY:5,” “PLAYBAR ” “PLAYBASE,” “CONNECT: AMP,” “CONNECT,” “AMP,” “PORT,” and “SUB.” Other suitable playback devices may additionally or alternatively be used to implement the playback devices of example embodiments disclosed herein. Additionally, one of ordinary skill in the art will appreciate that a playback device is not limited to the examples described herein or to Sonos product offerings. In some embodiments, for example, one or more playback devices 110 include wired or wireless headphones (e.g., over-the-ear headphones, on-ear headphones, in-ear earphones, etc.). In other embodiments, one or more of the playback devices 110 include a docking station and / or an interface configured to interact with a docking station for personal mobile media playback devices. In certain embodiments, a playback device may be integral to another device or component such as a television, an LP turntable, a lighting fixture, or some other device for indoor or outdoor use. In some embodiments, a playback device omits a user interface and / or one or more transducers. For example, FIG. ID is a block diagram of a playback device I lOp including the input / output 111 and electronics 112 without the user interface 113 or transducers 114.

[0063] FIG. IE is a block diagram of a bonded playback device HOq including the playback device 110a (FIG. 1C) sonically bonded with the playback device HOi (e.g., a subwoofer) (FIG. 1A). In the illustrated embodiment, the playback devices 110a and HOi are separate ones of the playback devices 110 housed in separate enclosures. In some embodiments, however, the bonded playback device HOq includes a single enclosure housing both the playback devices 110a and HOi. The bonded playback device HOq can be configured to process and reproduce sound differently than an unbonded playback device (e.g., the playback device 110a of FIG. 1C) and / or paired or bonded playback devices (e.g., the playback devices 1101 and 110m of FIG. IB). In some embodiments, for example, the playback device 110a is a full-range playback device configured to render low frequency, mid-range frequency, and high frequency audio content, and the playback device 1 lOi is a subwoofer configured to render low frequency audio content. In some aspects, the playback device 110a, when bonded with the first playback device, is configured to render only the mid-range and high frequency components of a particular audio content, while the playback device HOi renders the low frequency component of the particular audio content. In someAtorney Docket No. SON00095WOU 1Client Docket No. 24-0404p embodiments, the bonded playback device l lOq includes additional playback devices and / or another bonded playback device. c. Suitable Network Microphone Devices (NMDs)

[0064] FIG. IF is a block diagram of the NMD 120a (FIGS. 1 A and IB). The NMD 120a includes one or more voice processing components 124 (hereinafter “the voice components 124”) and several components described with respect to the playback device 110a (FIG. 1C) including the processors 112a, the memory 112b, and the microphones 115. The NMD 120a optionally includes other components also included in the playback device 110a (FIG. 1C), such as the user interface 113 and / or the transducers 114. In some embodiments, the NMD 120a is configured as a media playback device (e.g., one or more of the playback devices 110), and further includes, for example, one or more of the audio components 112g (FIG. 1C), the amplifiers 112h, and / or other playback device components. In certain embodiments, the NMD 120a includes an Internet of Things (loT) device such as, for example, a thermostat, alarm panel, fire and / or smoke detector, etc. In some embodiments, the NMD 120a includes the microphones 115, the voice processing components 124, and only a portion of the components of the electronics 112 described above with respect to FIG. 1C. In some aspects, for example, the NMD 120a includes the processor 112a and the memory 112b (FIG. 1C), while omitting one or more other components of the electronics 112. In some embodiments, the NMD 120a includes additional components (e.g., one or more sensors, cameras, thermometers, barometers, hygrometers, etc ).

[0065] In some embodiments, an NMD can be integrated into a playback device. FIG. 1G is a block diagram of a playback device 1 lOr including an NMD 120d. The playback device 1 lOr can include many or all of the components of the playback device 110a and further include the microphones 115 and voice processing components 124 (FIG. IF). The playback device HOr optionally includes an integrated control device 130c. The control device 130c can include, for example, a user interface (e.g., the user interface 113 of FIG. 1C) configured to receive user input (e.g., touch input, voice input, etc.) without a separate control device. In other embodiments, however, the playback device HOr receives commands from another control device (e.g., the control device 130a of FIG. IB).

[0066] Referring again to FIG. IF, the microphones 115 are configured to acquire, capture, and / or receive sound from an environment (e.g., the environment 101 of FIG. 1 A) and / or a room in which the NMD 120a is positioned. The received sound can include, for example, vocal utterances, audioAtorney Docket No. SON00095WOU 1Client Docket No. 24-0404p played back by the NMD 120a and / or another playback device, background voices, ambient sounds, etc. The microphones 115 convert the received sound into electrical signals to produce microphone data. The voice processing components 124 receive and analyze the microphone data to determine whether a voice input is present in the microphone data. The voice input can include, for example, an activation word followed by an utterance including a user request. As those of ordinary skill in the art will appreciate, an activation word is a word or other audio cue signifying a user voice input. For instance, in querying the AMAZON VAS, a user might speak the activation word "Alexa." Other examples include "Ok, Google" for invoking the GOOGLE VAS and "Hey, Siri" for invoking the APPLE VAS.

[0067] After detecting the activation word, voice processing components 124 monitor the microphone data for an accompanying user request in the voice input. The user request may include, for example, a command to control a third-party device, such as a thermostat (e.g., NEST thermostat), an illumination device (e.g., a PHILIPS HUE lighting device), or a media playback device (e.g., a SONOS playback device). For example, a user might speak the activation word “Alexa” followed by the utterance “set the thermostat to 68 degrees” to set a temperature in a home (e.g., the environment 101 of FIG. 1A). The user might speak the same activation word followed by the utterance “turn on the living room” to turn on illumination devices in a living room area of the home. The user may similarly speak an activation word followed by a request to play a particular song, an album, or a playlist of music on a playback device in the home. d. Suitable Control Devices

[0068] FIG. 1H is a partial schematic diagram of the control device 130a (FIGS. 1A and IB). As used herein, the term “control device” can be used interchangeably with “controller” or “control system.” Among other aspects, the control device 130a is configured to receive user input related to the media playback system 100 and, in response, cause one or more devices in the media playback system 100 to perform an action(s) or operation(s) corresponding to the user input. In the illustrated embodiment, the control device 130a is a smartphone (e.g., an iPhone™ an Android phone, etc.) on which media playback system controller application software is installed. In some embodiments, the control device 130a may be, for example, a tablet (e.g., an iPad™), a computer (e.g., a laptop computer, a desktop computer, etc.), and / or another suitable device (e.g., a television, an automobile audio head unit, an loT device, etc.). In certain embodiments, the control device 130a is a dedicated controller for the media playback system 100. In other embodiments,Attorney Docket No. SON00095WOU 1Client Docket No. 24-0404p as described above with respect to FIG. 1G, the control device 130a is integrated into another device in the media playback system 100 (e.g., one more of the playback devices 110, NMDs 120, and / or other suitable devices configured to communicate over a network).

[0069] The control device 130a includes electronics 132, a user interface 133, one or more speakers 134, and one or more microphones 135. The electronics 132 include one or more processors 132a (referred to hereinafter as “the processors 132a”), a memory 132b, software components 132c, and a network interface 132d. The processor 132a can be configured to perform functions relevant to facilitating user access, control, and configuration of the media playback system 100. The memory 132b can include data storage that can be loaded with one or more of the software components executable by the processor 132a to perform those functions. The software components 132c can include applications and / or other executable software configured to facilitate control of the media playback system 100. The memory 132b can be configured to store, for example, the software components 132c, media playback system controller application software, and / or other data associated with the media playback system 100 and the user.

[0070] The network interface 132d is configured to facilitate network communications between the control device 130a and one or more other devices in the media playback system 100, and / or one or more remote devices. In some embodiments, the network interface 132d is configured to operate according to one or more suitable communication industry standards (e.g., infrared, radio, wired standards including IEEE 802.3, wireless standards including IEEE 802.11a, 802.11b, 802.11g, 802. l ln, 802.11ac, 802.15, 4G, LTE, etc.). The network interface 132d can be configured, for example, to transmit data to and / or receive data from the playback devices 110, the NMDs 120, other ones of the control devices 130, one of the computing devices 106 of FIG. IB, devices including one or more other media playback systems, etc. The transmitted and / or received data can include, for example, playback device control commands, state variables, playback zone and / or zone group configurations. For instance, based on user input received at the user interface 133, the network interface 132d can transmit a playback device control command (e.g., volume control, audio playback control, audio content selection, etc.) from the control device 130a to one or more of the playback devices 110. The network interface 132d can also transmit and / or receive configuration changes such as, for example, adding / removing one or more playback devices 110 to / from a zone, adding / removing one or more zones to / from a zone group, forming a bonded or consolidated player, separating one or more playback devices from a bonded or consolidatedAtorney Docket No. SON00095WOU 1Client Docket No. 24-0404p player, among others. Additional description of zones and groups can be found below with respect to FIGS. II through IM.

[0071] The user interface 133 is configured to receive user input and can facilitate control of the media playback system 100. The user interface 133 includes media content art 133a (e.g., album art, lyrics, videos, etc.), a playback status indicator 133b (e.g., an elapsed and / or remaining time indicator), media content information region 133c, a playback control region 133d, and a zone indicator 133e. The media content information region 133c can include a display of relevant information (e g., title, artist, album, genre, release year, etc.) about media content currently playing and / or media content in a queue or playlist. The playback control region 133d can include selectable (e.g., via touch input and / or via a cursor or another suitable selector) icons to cause one or more playback devices in a selected playback zone or zone group to perform playback actions such as, for example, play or pause, fast forward, rewind, skip to next, skip to previous, enter / exit shuffle mode, enter / exit repeat mode, enter / exit cross fade mode, etc. The playback control region 133d may also include selectable icons to modify equalization settings, playback volume, and / or other suitable playback actions. In the illustrated embodiment, the user interface 133 includes a display presented on a touch screen interface of a smartphone (e.g., an iPhone™ an Android phone, etc.). In some embodiments, however, user interfaces of varying formats, styles, and interactive sequences may alternatively be implemented on one or more network devices to provide comparable control access to a media playback system.

[0072] The one or more speakers 134 (e.g., one or more transducers) can be configured to output sound to the user of the control device 130a. In some embodiments, the one or more speakers include individual transducers configured to correspondingly output low frequencies, mid-range frequencies, and / or high frequencies. In some aspects, for example, the control device 130a is configured as a playback device (e.g., one of the playback devices 110). Similarly, in some embodiments the control device 130a is configured as an NMD (e.g., one of the NMDs 120), receiving voice commands and other sounds via the one or more microphones 135.

[0073] The one or more microphones 135 may include, for example, one or more condenser microphones, electret condenser microphones, dynamic microphones, and / or other suitable types of microphones or transducers. In some embodiments, two or more of the microphones 135 are arranged to capture location information of an audio source (e.g., voice, audible sound, etc.) and / or configured to facilitate filtering of background noise. Moreover, in certain embodiments, theAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p control device 130a is configured to operate as a playback device and an NMD. In other embodiments, however, the control device 130a omits the one or more speakers 134 and / or the one or more microphones 135. For instance, the control device 130a may include a device (e.g., a thermostat, an loT device, a network device, etc.) having a portion of the electronics 132 and the user interface 133 (e.g., a touch screen) without any speakers or microphones. e. Suitable Playback Device Configurations

[0074] FIGS. II through IM show example configurations of playback devices in zones and zone groups. Referring first to FIG. IM, in one example, a single playback device may belong to a zone. For example, the playback device 110g in the second bedroom 101c (FIG. 1A) may belong to Zone C. In some implementations described below, multiple playback devices may be “bonded” to form a “bonded pair” which together form a single zone. For example, the playback device 1101 (e.g., a left playback device) can be bonded to the playback device 110m (e.g., a right playback device) to form Zone B. Bonded playback devices may have different playback responsibilities (e.g., channel responsibilities). In another implementation described below, multiple playback devices may be merged to form a single zone. For example, the playback device I lOh (e.g., a front playback device) may be merged with the playback device 1 lOi (e.g., a subwoofer), and the playback devices HOj and 110k (e.g., left and right surround speakers, respectively) to form a single Zone D. In another example, the playback devices 110b and 1 lOd can be merged to form a merged group or a zone group 108b. The merged playback devices 110b and 1 lOd may not be specifically assigned different playback responsibilities. That is, the merged playback devices 110b and 1 lOd may, aside from playing audio content in synchrony, each play audio content as they would if they were not merged.

[0075] Each zone in the media playback system 100 may be provided for control as a single user interface (UI) entity. For example, Zone A may be provided as a single entity named Master Bathroom. Zone B may be provided as a single entity named Master Bedroom. Zone C may be provided as a single entity named Second Bedroom.

[0076] Playback devices that are bonded may have different playback responsibilities, such as responsibilities for certain audio channels. For example, as shown in FIG. II, the playback devices 1101 and 110m may be bonded so as to produce or enhance a stereo effect of audio content. In this example, the playback device 1101 may be configured to play a left channel audio component,Atorney Docket No. SON00095WOU 1Client Docket No. 24-0404p while the playback device 110m may be configured to play a right channel audio component. In some implementations, such stereo bonding may be referred to as “pairing.”

[0077] Additionally, bonded playback devices may have additional and / or different respective speaker drivers. As shown in FIG. 1 J, the playback device 1 lOh named Front may be bonded with the playback device 1 lOi named SUB. The Front device 1 lOh can be configured to render a range of mid to high frequencies and the SUB device 1 lOi can be configured to render low frequencies. When unbonded, however, the Front device 11 Oh can be configured to render a full range of frequencies. As another example, FIG. IK shows the Front and SUB devices 1 lOh and 1 lOi further bonded with Left and Right playback devices HOj and 110k, respectively. In some implementations, the Left and Right devices 1 lOj and 110k can be configured to form surround or “satellite” channels of a home theater system. The bonded playback devices 1 lOh, 1 lOi, 1 lOj, and 110k may form a single Zone D (FIG. IM).

[0078] Playback devices that are merged may not have assigned playback responsibilities and may each render the full range of audio content the respective playback device is capable of. Nevertheless, merged devices may be represented as a single UI entity (i.e., a zone, as discussed above). For instance, the playback devices 110a and 11 On in the master bathroom have the single UI entity of Zone A. In one embodiment, the playback devices 110a and 1 lOn may each output the full range of audio content each respective playback devices 110a and 1 lOn are capable of, in synchrony.

[0079] In some embodiments, an NMD is bonded or merged with another device so as to form a zone. For example, the NMD 120b may be bonded with the playback device 1 lOe, which together form Zone F, named Living Room. In other embodiments, a stand-alone network microphone device may be in a zone by itself. In other embodiments, however, a stand-alone network microphone device may not be associated with a zone. Additional details regarding associating network microphone devices and playback devices as designated or default devices may be found, for example, in subsequently referenced U.S. Patent No. 10,499,146.

[0080] Zones of individual, bonded, and / or merged devices may be grouped to form a zone group. For example, referring to FIG. IM, Zone A may be grouped with Zone B to form a zone group 108a that includes the two zones. Similarly, Zone G may be grouped with Zone H to form the zone group 108b. As another example, Zone A may be grouped with one or more other Zones C-I. The Zones A-I may be grouped and ungrouped in numerous ways. For example, three, four, five, orAtorney Docket No. SON00095WOU 1Client Docket No. 24-0404p more (e.g., all) of the Zones A-T may be grouped. When grouped, the zones of individual and / or bonded playback devices may play back audio in synchrony with one another, as described in previously referenced U.S. Patent No. 8,234,395. Playback devices may be dynamically grouped and ungrouped to form new or different groups that synchronously play back audio content.

[0081] In various implementations, the zones in an environment may be the default name of a zone within the group or a combination of the names of the zones within a zone group. For example, Zone Group 108b can be assigned a name such as “Dining + Kitchen”, as shown in FIG. IM. In some embodiments, a zone group may be given a unique name selected by a user.

[0082] Certain data may be stored in a memory of a playback device (e.g., the memory 112b of FIG. 1C) as one or more state variables that are periodically updated and used to describe the state of a playback zone, the playback device(s), and / or a zone group associated therewith. The memory may also include the data associated with the state of the other devices of the media system, and shared from time to time among the devices so that one or more of the devices have the most recent data associated with the system.

[0083] In some embodiments, the memory may store instances of various variable types associated with the states. Variable instances may be stored with identifiers (e.g., tags) corresponding to type. For example, certain identifiers may be a first type “al” to identify playback device(s) of a zone, a second type “bl” to identify playback device(s) that may be bonded in the zone, and a third type “cl” to identify a zone group to which the zone may belong. As a related example, identifiers associated with the second bedroom 101c may indicate that the playback device is the only playback device of the Zone C and not in a zone group. Identifiers associated with the Den may indicate that the Den is not grouped with other zones but includes bonded playback devices 11 Oh- 110k. Identifiers associated with the Dining Room may indicate that the Dining Room is part of the Dining + Kitchen zone group 108b and that devices 110b and HOd are grouped (FIG. IL). Identifiers associated with the Kitchen may indicate the same or similar information by virtue of the Kitchen being part of the Dining + Kitchen zone group 108b. Other example zone variables and identifiers are described below.

[0084] In yet another example, the memory may store variables or identifiers representing other associations of zones and zone groups, such as identifiers associated with Areas, as shown in FIG. IM. An area may involve a cluster of zone groups and / or zones not within a zone group. For instance, FIG. IM shows an Upper Area 109a including Zones A-D and I, and a Lower Area 109bAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p including Zones E-T. Tn one aspect, an Area may be used to invoke a cluster of zone groups and / or zones that share one or more zones and / or zone groups of another cluster. In another aspect, this differs from a zone group, which does not share a zone with another zone group. Further examples of techniques for implementing Areas may be found, for example, in U.S. Patent No. 10,712,997 filed August 21, 2017, and titled “Room Association Based on Name,” and U.S. Patent No. 8,483,853 filed September 11, 2007, and titled “Controlling and manipulating groupings in a multizone media system.” Each of these patents is incorporated herein by reference in its entirety. In some embodiments, the media playback system 100 may not implement Areas, in which case the system may not store variables associated with Areas.IV. Enhanced Dialogue Rendering

[0085] FIG. 5 is a functional block diagram showing a system 500 and playback device 502 configured to function, at least for a portion of its operation, in a dynamic dialogue enhancement mode, selectively adjusting audio input 580a to ensure clarity and comprehension of dialogue received from a playback source 520. The playback device 502, for example, may represent one or more of the playback devices 1 lOa-n of FIG. 1 A, playback device 1 lOp of FIG. ID and / or FIG. IE, and / or playback device HOr of FIG. 1G. The playback source 520, in some examples, may include an audio source connected to the playback device 502 via a network interface (e.g., a wired network connection, a wireless network connection), and / or a hardware input interface. Examples of the hardware input interface include the analog VO I l la, the digital VO 111b, a USB drive inserted into a USB connection, and a playback device or computer-readable medium drive connected via an auxiliary connection. Other example hardware input interfaces will be apparent in view of this disclosure. The playback device 502 includes voice capture components (“VCC”, or collectively “voice processor 560”) and at least one voice extractor 572, each of which is operably coupled to the voice processor 560. The voice processor 560 is configured to receive speech-containing audio signals from one of an array of microphones 522 (e.g., included in and / or external to the playback device 502) and process the speech containing audio to identify voice input commands from a user (e g., with the support of a natural language processing unit 576 of a voice services unit 510). The microphones 522, for example, provide detected sound 562, SD, from the environment of the playback device 502 to the voice processor 560. Each channel 562a- n of the detected sound 562 may correspond to a particular microphone 522. Since the playback device 502 contains human speech recognizing mechanisms in the voice processor 560 and voiceAtorney Docket No. SON00095WOU 1Client Docket No. 24-0404p services 510, in some embodiments, the preexisting human speech recognizing mechanisms may be used to recognize human speech within an audio portion of media content received for playback by the playback device. The system 500 further includes at least one network interface 524 which may connect the playback device 502, in some embodiments, to additional voice processing mechanisms (e.g., available via an edge server, a cloud computing service, another playback device or other local computing device, etc.).

[0086] The playback device 502 includes a playback digital signal processor (DSP) 530 (e.g., part of the audio processing components 112g of FIG. 1C). In further embodiments, the playback device 502 includes additional components, such as, in some examples, one or more audio amplifiers and / or media interfaces, which are not shown in FIG. 5 for purposes of clarity.

[0087] As further shown in FIG. 5, in some embodiments, the voice processor 560 includes a spatial processor 566, and one or more buffers 568 (e.g., at least one audio signal buffer). The spatial processor 566, in some embodiments, is configured to analyze the detected sound SD 562 and identify certain characteristics, such as, in some examples, a sound's amplitude (e.g., decibel level), frequency spectrum, and / or directionality. In one example, the spatial processor 566 may monitor metrics that distinguish speech from other sounds. Such metrics can include, in some examples, energy within the speech band relative to background noise and / or entropy within the speech band — a measure of spectral structure — which is typically lower in speech than in most common background noise. In some implementations, the spatial processor 566 determines a speech presence probability. The processed sound SDS 506 produced by the spatial processor 566 may be provided to the voice services unit 510 (e.g., directly or via one or more buffers 568) for voice command identification and interpretation. Similarly, a dynamic dialogue enhancement unit 550 of the playback DSP 530 may provide at least a portion of the audio input 580a, Sc512, for processing by the voice processor unit 560 in identifying the potential for dialogue within the received audio content.

[0088] In some embodiments, the playback device 502 includes speech processing components 576 (e.g., a speech processor or natural language processing (NLP) unit) configured to further facilitate voice processing. The speech processing components 576, for example, may perform voice recognition trained to recognize a particular user or a particular set of users associated with a household. Voice recognition software, for example, may implement voice-processing processes that are tuned to specific voice profile(s). The speech processing components 576, in someAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p embodiments, are configured to determine the intent of the words of a command or uttered in correspondence with a command (e.g., keywords within background speech). In another example, the natural language unit 576 may include one or more machine learning processes, neural networks, and / or artificial intelligence networks trained to recognize human vocalizations. The input SDS 506 derived from the microphones 522, for example, may be processed to determine an intent of the user. In relation to the portion of the audio input 580a forwarded by the dynamic dialogue enhancement unit 550 as audio signals Sc 512, the one or more machine learning processes, neural networks, and / or artificial intelligence networks may confirm presence of dialogue, quantify the dialogue content (e.g., a frequency spectrum thereof), and / or extract the dialogue portion from the audio content Sc 512.

[0089] In some implementations, at least one buffer 568 captures sound data Sc 512 utilizing a sliding window approach in which a given amount (e.g., a given window) of the most recently obtained audio input 580a (or portion thereof) is retained in the at least one buffer 568 while older sound data are overwritten when they fall outside of the window. For example, at least one buffer 568 may temporarily retain twenty frames of a sound specimen at a given time, discard the oldest frame after an expiration time, and then capture a new frame, which is added to the 19 prior frames of the sound specimen.

[0090] In some implementations, media content from a playback source 520 (e.g., local audio source 105 as described in relation to FIG. 1C) is received at one or more signal processors of the DSP 530. The DSP 530, for example, may include a collection of audio processing circuitry and / or software processes (e.g., the audio processing components 112g of FIG. 1C) for processing an audio portion (e.g., audio input) 580a of the media content. Further, the audio processing circuitry and / or software processes (e.g., programs, code, and / or instructions) may be arranged as separate signal processor units (e.g., signal processors), each signal processor unit configured to process a separate channel of a multi-channel audio input received from the playback source 520. The audio processing circuitry and / or software processes can include one or more computer processors and / or separate audio processing circuitry, such as analog electronic circuit elements or separate electronic elements configured to carry out particular audio processing operations. The audio processing circuitry and / or software processes of the DSP 530, in some embodiments, include one or more elements for enhancing the audio input 580a to improve the clarity of dialogue content responsive to the identification of dialogue within the audio input 580a (e.g., as recognized by theAtorney Docket No. SON00095WOU 1Client Docket No. 24-0404p voice processor unit 560 and / or voice services unit 510). The components of the playback DSP 530, in the illustrative example, include a decoder 532, an equalization / volume controller 534, an arraying processor 536, and a limiter 538. The playback DSP also includes the dynamic dialogue enhancement unit 550.

[0091] The input signal from the playback source 520 can be media content (e.g., audio content including music and / or other sounds) from a local or networked audio source. In one example, the audio input 580a may be a digital audio signal such as a packetized or non-packetized stream of audio from a music service or television, a digital audio file, an audio signal generated by the playback device 502 itself or a device connected to the playback device 502 (e.g., via a wired or wireless communication). For example, the packetized stream of audio may include 128 bits of audio data per packet. In another example, the audio signal from the playback source 520 may be an analog signal input from an auxiliary connection or a digital signal input from a USB connection. The audio input 580a may include frequency content that may range from 0 Hz to 22,050 Hz or some subset of this frequency range.

[0092] The decoder 232, in some embodiments, is configured to decode one or more audio formats such as, in some examples, Dolby and / or MP3. The equalizer / volume control 534 may include a user-adjusted volume control, user-adjusted treble and bass settings, and / or an equalizer. The array processor 536 may be configured to accommodate additional playback devices.

[0093] The limiter 538, in certain embodiments, can include various analog electrical circuit elements (e.g., capacitors, resistors, inductors) and / or digital filters that prevent the audio signal from exceeding a defined threshold. The limiter 538, for example, may be configured to attenuate an amplitude of the audio input 580a at one or more frequencies so that the playback device 502 continues to operate within its operational limit. The amount that the audio signal is reduced by the limiter 538 at any given moment is referred to herein as the “gain reduction” applied by the limiter 538. For example, if the limiter 538 received an audio signal 280a at 3 dB and output an audio signal 580b at 2 dB, then the gain reduction of the limiter 538 at that moment equals 1 dB. As audio signals are typically dynamic, the amount of gain reduction applied by the limiter 538 will generally vary over time.

[0094] In some implementations, the audio output 580b (e.g., a processed audio version of the audio signal 580a received from the playback source 520) is provided to the amplifier(s) 112h for amplification prior to broadcasting via the one or more transducers 114 (described in relation toAtorney Docket No. SON00095WOU 1Client Docket No. 24-0404pFIG. 1C). The playback device 502, for example, may incorporate at least a portion of the transducers 114. In another example, at least a portion of the transducers 114 may be external to the playback device 502.

[0095] In some embodiments, the playback DSP 530 includes the dynamic dialogue enhancement unit 550 configured to dynamically adjust aspects of the audio input 580a based on determining the presence of dialogue within a portion of the audio input 580a (e.g., through analysis performed by the voice processor 560 and / or the voice services unit 510). The dynamic dialogue enhancement unit 550, for example, may direct audio processing aspects of the playback DSP 530 to modify portions of the audio input 580a responsive to the indication that dialogue is contained therein.

[0096] Although illustrated and described in relation to a single voice processor unit 560 and voice services unit 510, along with individual processing aspects contained therein as discussed above, in some implementations, separate, dedicated audio processing circuitry and / or software processes, identical or similar in composition, may be provided in the playback device 502, where one set of audio processing circuitry and / or software processes may be configured to process the environmental sound captured by the microphones 522, while the other set of audio processing circuitry and / or software processes is configured to process the audio input 580a from the playback source 520. Conversely, in some embodiments, duplicate audio processing circuitry and / or software processes included in the voice processor unit 560 and / or the voice services unit 510 are allocated in an on-demand fashion to handle incoming environmental sound SD 562 and audio input 580a to support real-time or near real-time handling of incoming audio streams. Further, the playback DSP 530 may include certain additional resources, such as one or more additional equalization / volume controllers 534 and / or limiters 538configured to adjust aspects of the audio input 580a to increase clarity and comprehension of its dialogue contents. In an additional example, the playback device 502 may include one or more additional amplifiers 112h configured to selectively amplify portions of the audio input 580a to increase clarity and comprehension of the dialogue contents of the audio output 580b.

[0097] Turning to FIG. 2, a flow chart illustrates an example method 200 for enhancing rendering of dialogue by a playback device. The method 200, for example, may be performed by one of the playback devices 110 described in relation to FIG. 1A through FIG. IM. In another example, portions of the method 200 may be performed by a playback device 502 of FIG. 5. The dialogue,Atorney Docket No. SON00095WOU 1Client Docket No. 24-0404p in some examples, may be rendered by one or more of the transducer(s) 114 of the playback device 110a of FIG. 1C, at least one of the transducer(s) 114 of the network microphone device 120a of FIG. IF, and / or at least one of the transducer(s) 114 of the playback device 502 of FIG. 5.

[0098] In some implementations, the method 200 begins with obtaining, for playback via a transducer array, media content including first audio and second audio (202). The media content, in addition to the first audio and second audio, may include, in some examples, video content, metadata content, and / or additional audio content. The media content may have originated from a playback source, such as the local audio source 105 described in relation to FIG. 1C or the playback source 520 of FIG. 5. The media content may be obtained from a local audio source or a networked audio source. In one example, the first audio and the second audio may include audio signals such as those described in relation to the audio input 580a of FIG. 5. The media content may be obtained, for example, by the audio processing components 112g of the playback device 110a of FIG. 1C and / or the playback digital signal processor (DSP) 530 of FIG. 5.

[0099] The multiple portions of audio content, for example, may include multiple channels of audio signals or multiple aspects of audio content configured to be rendered according to settings or processes among a main (e.g., center) speaker as well as surround speakers. In another example, the multiple portions of audio content may include a dialogue audio track and at least one nondialogue audio track. The audio processing circuitry and / or software processes can include one or more computer processors and / or separate audio processing circuitry, such as analog electronic circuit elements or separate electronic elements configured to carry out particular audio processing operations. The audio content, for example, may be the audio input 580a of the system 500 of FIG. 5.

[0100] In some implementations, the media content is reviewed for presence of dialogue (204). The audio processing components 112g of FIG. 1C, for example, may review the first audio for presence of dialogue. The analysis may be performed in real-time or near-real time, for example as streaming media content is being received at a playback device. In a first example, the audio processing components 112g may analyze the first audio to identify one or more vocalizations. A voice processor (e.g., similar to the voice components 124 of the NMD 120a of FIG. IF), for example, may detect vocalizations within the first audio. The voice processor, for example, may be part of a voice services unit 510 of the playback device 502 of FIG. 5. The voice processor may monitor and analyze audio received within media content to determine if any humanAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p vocalizations are present. To identify a vocalization, in a second example, the first audio and / or the second audio may be filtered to determine whether the first audio likely contains dialogue. The filtering, for example, may extract a portion of the first audio for use in analyzing those signals within a frequency range of human speech (e.g., between 90 Hz and 255 Hz). Absence of dialogue, further to this example, may be indicative of lack of threshold sound within the frequency range of human speech in the first audio. In another example, the filtering may include filtering for energy characteristics of the first audio and / or second audio. In illustration, substantial low frequency energy below a threshold level of about 100 Hertz within the center channel audio (e.g., first audio) may be indicative of a lack of dialogue. For example, when significant low frequency energy is present, the scene is more likely to include an action sequence. Characteristics of energy within channels other than the center channel, conversely, may be used to infer the presence or absence of dialogue in the center channel audio signals. Identifying the one or more vocalizations, in a third example, may include analyzing a metadata portion of the media content to identify timings of dialogue content, such as sound metadata 508 of the audio input 580a of FIG. 5. The dialogue content, for example, may be flagged in part through subtitle (e.g., closed caption) metadata or other metadata including indications of speech content (e.g., time segments) that are time-aligned with the audio content. In a fourth example, machine learning analysis and / or an artificial intelligence network may analyze the first audio to recognize speech content. Although described as separate examples, in some embodiments, a combination of techniques may be used, such as performing analysis of metadata in addition to analyzing the energy characteristics of the first audio and / or second audio. In a fifth example, if a dialogue-dedicated track of audio is provided within the media content, the media content may be analyzed to determine whether sound content has been received in the dialogue audio track.

[0101] In some implementations, if dialogue is identified (206), adjusted second audio having reduced sound levels in a first frequency range including a set of frequencies associated with human speech is generated using the second audio (208). To increase speech perceptibility while adhering as much as possible to artistic intent within scenes of the media content, for example, sound within a range of speech frequencies may be de-emphasized in portions of the audio content (e.g., channels, surround speaker-directed audio portions of the media content, etc.) not containing dialogue. The second audio, for example, may include one or more audio portions (e.g., channels). When the second audio includes multiple portions, the portions of the second audio may be alteredAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p consistently (e.g., same for left and right channel) or strategically (e.g., differing modifications applied depending upon, in some examples, particular equipment, setup, and / or user preferences). Consistent application, for example, may be appropriate in circumstances where a segment of the audio signals of the media content are shared between the first audio and the second audio and / or between portions of the second audio. The second audio may be adjusted, for example, by the audio processing components 112g of the playback device 110a of FIG. 1C and / or components of the playback DSP 530, such as the equalizer 534 and / or the limiter 538.

[0102] A scoop filter, in some embodiments, may be applied to the second audio (e.g., left and / or right channel audio signals) to attenuate signals within a speech frequency range. The scoop filter, for example, may “scoop” the frequencies of the second audio conflicting with the detected dialogue, reducing and / or removing them from the content of the second audio. In some embodiments, different levels of attenuation may be applied, for example based on operating parameters (e.g., equipment used, surround sound settings, etc.), characteristics of the first audio portion and / or second audio portion (e.g., the energy characteristics discussed above), and / or user settings. The scoop filter, in some examples, may selectively attenuate at least a portion of the signals within the speech frequency range using around a 1, 2, 3, 5, 10, 20, or 30 decibel gain reduction. For example, if the scoop filter received a portion of the second audio at 3 dB and output the adjusted second audio having the portion reduced to 2 dB, then the gain reduction applied by the scoop filter at that moment equals 1 dB. In a particular example, the gain reduction may be within a range of about 1.8 dB to about 3.7 dB. One or more scoop filters may be provided in the audio processing components 112g of the playback device 110a of FIG. 1C. A scoop filter may be generated by a limiter, including various analog electrical circuit elements (e.g., capacitors, resistors, inductors) and / or digital filters that prevent the second from exceeding at least one defined threshold. The limiter, for example, may be configured to attenuate an amplitude of the second audio at one or more frequencies within the speech frequency range.

[0103] In some embodiments, de-emphasizing may include suppressing an operating setting of the playback device. In illustration, sharing signals designated as center channel sound (e.g., the first audio) across the center and surround speakers (e.g., both the first audio and the second audio), is commonly referred to as shouldering. Shouldering may be used for cinematic emphasis in surround sound applications. A shouldering setting may be suppressed during dialogue so that non-dialogue sounds within the range of human speech are not replicated in other channels, whichAtorney Docket No. SON00095WOU 1Client Docket No. 24-0404p would otherwise exacerbate any masking of the speech content. Tn another example, the shouldering setting may be switched to a non-speech shouldering setting, where only audio content outside the speech frequency range (or a portion thereof detected in the dialogue content) may be replicated across supporting transducers (e.g., surround speakers).

[0104] In some implementations, playback of the first audio is caused via a first transducer of the transducer array (212) and playback of the adjusted second audio is caused via a second transducer of the transducer array (210). For example, the first audio may be played via a first transducer of the transducers 114 of FIG. 1C and / or FIG. 5, and the second audio may be played via a second transducer of the transducers 114 of FIG. 1C and / or FIG. 5. One or more speakers, for example, may include and / or be fed by the transducer array. The first transducer may provide a center channel audio component (e.g., portion) of the media content, for example, while the second transducer may provide a left channel audio component of the media content and / or right channel audio component of the media content. The transducer array may be integrated into and / or connected with a surround sound speaker system.

[0105] In some implementations, the method 200 continues to repeat while additional media content is obtained (216).

[0106] Although described in relation to a particular set of operations, in other embodiments, the method 200 may include more or fewer operations. For example, if dialogue is identified (206), the first audio and / or second audio may be analyzed for evidence of competing sound, particularly within a frequency range of the dialogue (e.g., the frequency of human speech or a portion thereof). In illustration, enhancement may be deemed unnecessary where the dialogue content represents the majority of the sound within the media content for a given timeframe. In another example, upon identification of dialogue (206), a speech enhancement setting may be activated, leading to the generation of the adjusted second audio (208). Further to this example, when later timeframes of the media content are not determined to contain dialogue (206), the speech enhancement setting may be deactivated. In a third example, if the dialogue is identified (206) through a mechanism other than processing of metadata content, in some embodiments, a metadata file may be created or augmented to indicate portions of the media content containing dialogue. In this manner, when the media content is played at a different time, the metadata generated through reviewing the media content (204) during a prior playback session may be used to selectively adjust the second audio.Attorney Docket No. SON00095WOU 1Client Docket No. 24-0404p

[0107] Although the method 200 is described as being performed in a particular series of operations, in other embodiments, certain operations of the method 200 may be performed in a different order and / or concurrently. For example, although illustrated as a series of steps, playback of the first audio (212) would be performed concurrently with playback of either the second audio (214) (if dialogue was not identified) or the adjusted second audio (210) (if dialogue was identified) as indicated by the timing of the media content. In another example, a lookahead operation may be performed to flag segments of dialog within media content prior to playback (e.g., when obtained (202) but not actively playing due to the playback device controller being on “pause”, etc.). Other modifications of the method 200 are possible.

[0108] Turning to FIG. 3, a flow diagram of an example process 300 for enhancing rendering of dialogue within audio content by adjusting both first audio signals containing the dialogue and second audio signals separate from the first audio signals is illustrated. The process 300, for example, may be performed by one of the playback devices 110 described in relation to FIG. 1A through FIG. IM. The various processing units (e.g., engines) of the process 300, in some embodiments, are configured as software routines or processes (e.g., at least a portion of a software program) coded as instructions (e.g., software code) for executing on processing circuitry, such as one or more processors. Certain processing units or operations performed by certain processing units, in some embodiments, are configured as hardware logic (e.g., hardware-based operations) hard-coded or programmed into processing circuitry (e.g., logic circuitry), such as, in some examples, a programmable logic chip or other programmable logic device, an application-specific integrated circuit (ASIC), or a customized processor device. In an illustrative example, a software routine or process component of one of the processing units may provide data and / or variable values (e.g., within a non-transitory computer-readable medium) for use by hardware logic of that processing unit.

[0109] In some implementations, the process 300 begins with receiving, at a dialogue detection processing unit 304, first audio signals 302. The first audio signals 302, for example, may be included in media content, such as the media content described in relation to the operation 202 of the method 200 of FIG. 2. The dialogue detection processing unit 304, for example, may review the first audio signals 302 in a manner described in relation to operation 204 of the method 200 of FIG. 2. The dialogue detection processing unit 304, for example, may receive the first audio signals 302 from at least one media signal buffer configured to temporarily buffer an audio portionAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p of incoming media content for analysis and / or adjustment, such as one of the audio buffer(s) 582 of the playback DSP 530 of FIG. 5 (e.g., media signal buffer(s)). The first audio signals 302, for example, may be composed of frames or packets, each of which may include one or more sound samples. The frames may be streamed (e.g., read out) from one or more buffers for processing by the dialogue detection processing unit 304. The dialogue detection processing unit 304, for example, may be part of the voice services unit 510 of the playback device 502 of FIG. 5. For example, the dialogue detection processing unit 304 may use components of a voice extractor 572 to recognize elements of human speech within the first audio signals 302.

[0110] The dialogue detection processing unit 304, in some embodiments, includes a spatial processor, such as spatial processor 566 of FIG. 5, configured to analyze the first audio signals 302.[OHl] In some embodiments, the dialogue detection processing unit 304 includes one or more machine learning processes, neural networks, and / or artificial intelligence (Al) networks trained to recognize speech within the first audio signals. The dialogue detection processing unit 304, for example, may apply one or more machine learning techniques, neural network analysis techniques, and / or Al techniques to produce a speech mask and / or a set of speech metrics representing the spectral structure of vocalizations detected within the first audio signals 302. Creation of a speech mask, for example, is described in relation to Provisional Patent Application No. 63 / 700,280 entitled “Techniques for Speech Enhancement” and filed September 27, 2024.

[0112] The dialogue detection processing unit 304, in some embodiments, filters the first audio signals to extract a portion of the first audio signals including dialogue. Extraction of dialogue is described below in greater detail in relation to FIG. 4.

[0113] In some implementations, when the dialogue detection processing unit 304 determines that dialogue is present (306), a dynamic sound enhancement processing unit 308 determines one or more adjustment parameters 312 for adjusting the first audio signals 302 and / or second audio signals 316. The dynamic sound enhancement processing unit 308, for example, may be or include a dynamic dialog enhancement unit 550 of the playback DSP 530 of FIG. 5. The adjustment parameter(s) 312, for example, may be based in part on a characterization of the presence of dialogue generated by the dialogue detection processing unit 304. The characterization of the presence of dialogue, as discussed above, can include, in some examples, a speech mask, a set of speech metrics, and / or a speech presence probability. In a further example, the adjustmentAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p parameter(s) 312 may be based in part on a prior determination of the presence of dialogue (306) associated with previous first audio signals 302, such that the adjustment parameter(s) 312 match or are congruous with prior adjustment parameter(s) 312 to avoid a noticeable change in dialogue enhancement during a same dialogue-containing scene. The one or more adjustment parameters 312, in some examples, may include at least one attenuation parameter (e.g., a speech-enhancing filtering setting) for attenuating non-dialogue audio content, at least one emphasis parameter for emphasizing (e.g., a speech boosting setting) dialogue audio content, and / or at least one speech combining parameter (e.g., a speech summation setting) for adding extracted dialogue signals to the first audio signals 302. Further, the one or more adjustment parameters 312 may include, in relation to the attenuation parameter(s), the emphasis parameter(s), and / or the speech combining parameter(s), at least one level of adjustment selected from a set of adjustment levels corresponding to the type of adjustment parameter (e.g., an attenuation adjustment level, a boost adjustment level, or a summation adjustment level). In illustration, an attenuation parameter may be set to one of a low attenuation level, a mid-range attenuation level, or a high attenuation level.

[0114] The dynamic sound processing unit 308, in some embodiments, determines at least one of the one or more adjustment parameters 312 based at least in part on playback context 310. The playback context 310, in some examples, may relate to a media content context, a playback equipment context, and / or an audience context. Portions of the playback context 310, for example, may be collected and managed by a controller of the playback device, such as the control device 130a of FIG. 1A and FIG. IB. The memory 112b of the playback device 110a of FIG. 1C, for example, may be configured to store at least a portion of the playback context 310.

[0115] In some embodiments, at least one of the one or more adjustment parameters 312 is determined based at least in part on media content context. The media content context, for example, may include a type of media content (e.g., video, podcast, etc.), a genre of media content (e.g., action, drama, comedy, science fiction, thriller, musical, etc.), an audio stream quality type (e.g., high-definition (HD) audio, ultra HD audio, high-quality (HQ) audio, etc.), and / or an audio stream format type (e.g., sampling rate, packet size, etc.). Regarding the type of media content, speech enhancement may be applied to video more delicately (e.g., with less emphasis) than speech enhancement applied to podcasts, for example, to retain cinematic emphasis and meaning of the greater soundtrack corresponding to the video. Further, the type of video may weigh upon the adjustment parameters 312, where the background soundtrack of an action, science fiction, orAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p thriller movie, in particular in relation to scenes containing substantial non-dialogue audio content, may provide as strong or stronger context to the scene than the dialogue itself (e.g., which is often monosyllabic, expletives, or other speech that is of lesser value to the viewer’s comprehension of the scene). Thus, the adjustment parameters 312 may add greater dialogue emphasis in drama videos, generally, than in action videos. The audio stream quality type and / or format type, in other examples, may be used to designate, within the adjustment parameters 312, targeted adjustment (e.g., filtering, digital equalizing, etc.) options designed for the particular formatting of the audio content.

[0116] In some embodiments, at least one of the one or more adjustment parameters 312 is determined based at least in part on playback equipment context. The playback equipment context may include a type of equipment, such as a surround sound speaker system, a “sound base” (e.g., a device configured for pairing with a television or other display device to provide sound output for that device, including one or more transducers mounted with an interior volume of its enclosure and, optionally, other audio transducer(s) mounted on the exterior of the enclosure), or a “sound bar” (e.g., a single playback device including an array of transducers arranged to simulate or partially simulate a surround sound experience). In other examples, the playback equipment context may include a layout of playback equipment (e.g., relative positioning of speakers within a surround sound speaker system) and / or a target zone of the playback equipment, such as the zones described in relation to FIG. II through FIG. IM. In a further example, the playback equipment context may include whether a personal listening device, such as a style of headphones or a personal mobile media playback device, is in use. The type of equipment, such as the type of transducer, may be used to select adjustment parameters 312 appropriate to the style of audio output. For example, dialogue content of an audio portion destined for playing out a transducer of the playback equipment directed toward a listening area (e g., at the viewer(s)) may be handled in one manner according to the adjustment parameters 312 (e.g., emphasizing signals within at least a portion of the speech frequency range, such as a set of frequencies associated with human speech, at a particular level of emphasis), while another audio portion destined for playing out a transducer directed away from the listening area (e.g., away from the viewer(s)), may be handled in a different manner according to the adjustment parameters 312 (e.g., de-emphasizing signals within at least a portion of the speech frequency range at a particular level of attenuation). In further illustration, gain and / or attenuation levels (e.g., relative or absolute) as indicated in the adjustment parametersAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p312 may vary based on a type of transducer or speaker from which the audio signals will be output. In another illustration, audio content provided to a zone separate from a zone within the vicinity of the playback device may be treated differently than audio content provided in the zone having the playback device. In a multi-zone audio environment, for example, after leaving a room where video content is displayed, a listener may want to follow the story line (e.g., listen to broadcast news, track action within a sports game), for example using the zone grouping capability described above (see, e.g., FIG. IL and FIG. IM). In illustration, a listener may leave the den 101c of FIG. IK and enter the kitchen lOlh of FIG. IL to grab a snack, while continuing to monitor commentary related to a sporting event being viewed in the den 101c. Further to this illustration, according to the adjustment parameters 312, a stronger dialogue emphasis may be applied in the kitchen lOlh (e.g., to keep the listener apprised of the action without visual cues), while a more subtle dialogue emphasis may be applied in the den 101c (e.g., so the adjusted audio doesn’t detract from the viewer experience of being “inside the action”).

[0117] In some embodiments, at least one of the one or more adjustment parameters 312 is determined based at least in part on audience context. The audience context, in some examples, may include a number of listeners within a vicinity of the playback device and / or relative locations of listeners. In this manner, for example, the adjustment parameters 312 may consider directionality of transducers in relation to viewers in applying enhancement mechanisms. In a further example, the audience context may include identification of a particular user interacting with the playback device, such as a user identified, within user settings, as having hearing difficulties or a preference for a certain level of dialogue enhancement. The user settings, for example, may specify a level of enhancement of a set of available dialogue enhancement options (e.g., low / medium / high, a numeric value from 0 to N, etc.) designed to provide the user the option to specify a relative level of dialogue enhancement ranging from subtle (e.g., unlikely to detract from the overall programming experience) to strong (e.g., carrying a likelihood of substantially modifying artistic intent of the audio portion of the programming to significantly enhance the clarity of dialogue). The adjustment parameters 312, for example, may reflect such user settings. Further, the audience context may include an indication of ambient noise (e.g., conversations among listeners, a fan or other white noise, etc.).

[0118] The dynamic sound processing unit 308, in some embodiments, determines at least a portion of the adjustment parameters 312 based on a set of rules. Each rule of the set of rules, forAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p example, may match one or more items of playback context 310 to one or more adjustment parameters 312. A given rule, for example, may combine at least one element of the media content context with at least one element of the playback equipment context, while another rule may combine at least one element of the audience context with at least one element of the playback equipment context. In at least one example, the rules are stored as one or more logical implications. Any combination of types of context is possible. The set of rules, for example, may be stored to a memory location of the playback device, such as the memory 112b of FIG. 1C.

[0119] In some implementations, an audio adjustment processing unit 314 adjusts the first audio signals 302 and the second audio signals 316 according to the adjustment parameter(s) 312. The audio adjustment processing unit 314, for example, may perform the operation 208 of the method 200 of FIG. 2.

[0120] The audio adjustment processing unit 314, in some embodiments, implements at least one speech emphasizing mechanism to emphasize a dialogue portion of the first audio signals 302 in accordance with at least a portion of the adjustment parameters 312. The speech emphasizing mechanism, for example, may include emphasizing sound levels in at least a portion of the frequency range of human speech (e g., a set of frequencies associated with human speech). Emphasizing the sound levels in the portion of the frequency range of human speech, for example, may include boosting all signals within a designated set of frequencies by a voltage gain. The voltage gain, for example, may be identified via a setting or other parameter, selected from a set of voltage gains within a range from about 1 dB to about 4 dB, more particularly from about 1.8 dB to about 3.7 dB. Boost circuitry for selectively emphasizing the set of frequencies within the range of human speech may be provided in the audio processing components 112g of the playback device 110a of FIG. 1C. For example, the playback DSP 530 of FIG. 5 may include boost circuitry. In another example, signal processing software may selectively emphasize the set of frequencies, for example by performing equalization to boost the amplitude of the spectral content of the set of frequencies. The equalizer 534 of the Playback DSP 530 of FIG. 5, for example, may perform equalization to boost the amplitude of the spectral content of the set of frequencies. The boost circuitry and / or signal processing software, for example, may boost an amplitude of each frequency within the set of frequencies. The portion of the frequency range of human speech, for example, may include frequencies between about 90 Hz and about 255 Hz. The set of frequencies boosted, for example, may include a frequency spectrum (e g., a speech mask) of the dialogueAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p portion of the first audio signals 302, as detected by the dialogue detection processing unit 304. The emphasizing mechanism may include an audio summing mechanism (e.g., circuitry and / or software) to add the speech mask to the first audio signals. Implementing the at least one speech emphasizing mechanism, for example, results in adjusted first audio signals 318.

[0121] The audio adjustment processing unit 314, in some embodiments, implements at least one attenuation mechanism to reduce sound levels within at least a portion of a speech frequency range (e.g., a set of frequencies associated with human speech) in the second audio signals 316 in accordance with at least a portion of the adjustment parameters 312. Reducing the sound levels in the portion of the frequency range of human speech, for example, may include filtering all signals within a designated set of frequencies by a voltage gain reduction. The gain reduction, for example, may be identified via a setting or other parameter, selected from a set of voltage gain reductions within a range from about 1 dB to about 4 dB, more particularly from about 1.8 dB to about 3.7 dB. As described in relation to operation 208 of FIG. 2, for example, a scoop filter may be applied to the second audio signals 316 to attenuate signals within at least a portion of the speech frequency range. In a further example, the attenuation mechanism may include suppressing an operating setting of the playback device, such as a shouldering setting, as described in relation to operation 208 of FIG. 2. The playback DSP 530 of FIG. 5, for example, may include scoop filtering mechanisms for applying a gain reduction to the portion of the speech frequency range. In at least one example, implementing the at least one attenuation mechanism results in adjusted second audio signals 320.

[0122] In some implementations, a level of attenuation applied by the attenuation mechanism(s) is linked to a level of emphasis applied by the speech emphasizing mechanism(s). For example, the adjustment parameter(s) 312 may include a single set of parameters (e.g., either sound emphasizing parameters or sound attenuating parameters), and the audio adjustment processing unit 314 may determine the opposing sound emphasizing parameters or sound attenuating parameters. For example, a level of boost may be set to a same decibel level as the level of attenuation. In another example, the level of boost may be set to double the level of attenuation (or, conversely, the level of attenuation may be set to half the level of boost).

[0123] In some implementations, the adjusted first audio signals 318 and the adjusted second audio signals 320 are provided for playback by an audio playback processing unit 322. The first audio signals 318 and the second audio signals 320, for example, may be provided as the audio outputAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p580b of the playback DSP 530 of FIG. 5. The adjusted first audio signals 318 and the adjusted second audio signals 320, for example, may be temporarily stored to one or more audio buffers in communication with the audio playback processing unit 322. The audio playback processing unit 322, for example, may cause playback of the adjusted first audio signals 318 via a first transducer of a transducer array and cause playback of the adjusted second audio signals 320 via a second transducer of the transducer array (e.g., as described in relation to operation 210 of the method 200 of FIG. 2). Additionally, in the circumstance of a multi-zone audio environment, each of the adjusted first audio signals 318 and the adjusted second audio signals 320 may include two separate sets of audio signals, one prepared for each zone of the multi-zone environment. The audio playback processing unit 322, prior to causing playback of the adjusted first audio signals 318 and the adjusted second audio signals 320 (e.g., via a transducer array), may apply further processing to the adjusted first audio signals 318 and / or the adjusted second audio signals 320. For example, the audio playback processing unit 322 may include one or more audio processing components 112g and / or one or more audio amplifiers 112h for preparing the adjusted first audio signals 318 and / or the adjusted second audio signals 320 for playback. In example, the adjusted first audio signals 318 and / or the adjusted second audio signals 320 may undergo further processing by one or more components of the playback DSP 530 of FIG. 5.

[0124] In some implementations, where the dialogue detection processing unit 304 determines that dialog is not present (306), the first audio signals 302 and the second audio signals 316 are provided for playback by the audio playback processing unit 322. As described in the preceding paragraph, the audio playback processing unit 322 may apply further processing to the first audio signals 302 and / or the second audio signals 316 prior to playback. In other embodiments, the audio adjustment processing unit 314 may adjust the first audio signals 302 and / or the second audio signals 316 in a default (e g., non-speech-enhanced) manner. For example, the audio adjustment processing unit 314 may use default adjustment parameters (e.g., a default filtering setting, etc.) to prepare the first audio signals 302 and / or the second audio signals 316 for playback by the audio playback processing unit 322. The default adjustment parameters may be based on a portion of the playback context 310. Upon determining the dialogue is not present (306), for example, the dynamic sound enhancement processing unit 308 may revert the sound-enhancing filtering settings to the default settings, thereby maintaining the sound levels of the second audio signals 316 without adjustment within the speech frequency range.Attorney Docket No. SON00095WOU 1Client Docket No. 24-0404p

[0125] Although illustrated as a particular flow, in other embodiments, the process 300 may be performed in a different manner. For example, the dynamic sound enhancement processing unit 308 may determine, based on the playback context 310, that no adjustment of the first audio signals 302 and / or second audio signals 316 is needed. In illustration, if the ambient noise contains a threshold level of vocalizations, the dynamic sound enhancement processing unit 308 may determine that the playback is background noise within the environment that can proceed without adjustment to the first audio signals 302 and / or the second audio signals 316. In another example, the dynamic sound enhancement processing unit 308 may determine whether adjustment of the first audio signals 302 is desirable, or whether, based on the playback context 310, only the second audio signals 316 should be adjusted (e.g., in the manner described in relation to the method 200 of FIG. 2). Other modifications of the process 300 are possible.

[0126] Turning to FIG. 4, a flow diagram illustrates an example process 400 for adjusting audio content based on an extracted dialogue portion of the audio content. The example process 400, for example, may be performed at least in part by one of the playback devices 110 described in relation to FIG. 1A through FIG. IM and / or the playback device 502 of FIG. 5. The various processing units (e.g., engines) of the process 400, in some embodiments, are configured as software routines or processes (e.g., at least a portion of a software program) coded as instructions (e.g., software code) for executing on processing circuitry, such as one or more processors. Certain processing units or operations performed by certain processing units, in some embodiments, are configured as hardware logic (e.g., hardware-based operations) hard-coded or programmed into processing circuitry (e.g., logic circuitry), such as, in some examples, a programmable logic chip or other programmable logic device, an application-specific integrated circuit (ASIC), or a customized processor device. In an illustrative example, a software routine or process component of one of the processing units may provide data and / or variable values (e.g., within a non-transitory computer-readable medium) for use by hardware logic of that processing unit.

[0127] In some implementations, the process 400 begins with receiving the first audio signals 302 at a dialogue extraction processing unit 402. The dialogue extraction processing unit 402, for example, may include the voice extractor 572 of the voice services unit 510 of the playback device 502 of FIG. 5. In another example, the dialogue extraction processing unit 402 may be in network communication with a playback device, for example via a network interface 524 of the playback device 502 of FIG. 5.Attorney Docket No. SON00095WOU 1Client Docket No. 24-0404p

[0128] In some implementations, the dialogue extraction processing unit 402 separates a speech audio portion 302a from the first audio signals 302. Isolating the speech audio portion 302a, in some examples, may be performed by the voice processor unit 560 or the voice services unit 510 of the playback device 502 of FIG. 5. In another example, the speech audio portion 302a may be isolated by the playback DSP 530 (e.g., part of the dynamic dialogue enhancement unit 550) using audio signals temporarily stored by the buffer(s) 582. In certain embodiments, the dialogue extraction processing unit 402 separates first audio signals 302 into a non-speech audio portion 302b and the speech audio portion 302a. The first audio signals 302 may be divided, for example, through frequency analysis (e.g., separating sound within a speech frequency range from sound outside of a range of speech frequencies). The first audio signals 302, in another example, may be divided by using automatic speech recognition (ASR) analysis, confirming that the sounds within the speech frequency range correspond to recognizable verbalizations (e.g., according to natural language processing as performed, for example, by the natural language unit (NLU) 576 of FIG. 5). In another example, a metadata portion of the sound (e.g., sound metadata SM 508 of FIG. 5) may be used to find instances of speech and separate them from non-speech components of the first audio signals 302. In another example, one or more machine learning classifiers trained in recognizing speech patterns within audio content are applied to the first audio signals 302 to identify the speech audio portion 302a. The machine learning classifier(s), for example, may be applied on a frame-by-frame basis to the first audio signals 302, such that the dialogue extraction processing unit 402 may adaptively, in real-time, provide input to the audio adjustment processing unit 314 for dialogue enhancement purposes. The machine learning classified s), for example, may be included in the voice services unit 510 and / or provided as a networked (e.g., edge server, cloud, etc.) service accessible via the network interface 524 of the playback device 502 of FIG. 5. While illustrated as two separate signals, the non-speech audio portion 302b and the speech audio portion 302a may be included in the same digital output (e.g., including flags or markers differentiating the speech component from the non-speech component). In another example, the non-speech audio portion 302b may be created as a logical inversion of the speech audio portion 302a or vice- versa. The speech audio portion 302a and / or the non-speech audio portion 302b, for example, may be stored to a temporary buffer or high-speed memory region for future processing, such as the buffer(s) 582 of FIG. 5.Attorney Docket No. SON00095WOU 1Client Docket No. 24-0404p

[0129] In some embodiments, as described in greater detail in related U.S. Provisional Patent Application No. 63 / 700,280 entitled “Techniques for Speech Enhancement” and fded September 27, 2024, the dialogue extraction processing unit 402 operates in the frequency domain to separate speech content from non-speech content. For example, a short-time Fourier transform (STFT) may be applied to the first audio signals 302 to produce a corresponding input frequency spectrum. In the digital domain, the input frequency spectrum may be represented as a two-dimensional matrix of signal magnitude and frequency. The first audio signals 302 may be divided into a set of frequency bins (e.g., 256, 512, etc.) and the signal magnitude in each frequency bin may be recorded as a digital value. The dialogue extraction processing unit 402 may identify the speech audio portion 302a within the input frequency spectrum matrix by categorizing the frequency bins of the two-dimensional matrix as “speech” or “no speech,” in a binary fashion, for example producing a speech audio portion 302a represented by a matrix of signal magnitude, frequency, and speech / no-speech flag (e.g., a logical one or zero).

[0130] In some embodiments, dialogue extraction processing unit 402 applies machine learning techniques to identify speech content in the audio portion of the stream of media content. The machine learning techniques may be applied on the first audio signals 302 in its original format (e.g., as a time-domain signal) or in a converted format (e.g., as a frequency spectrum data stream). The machine learning techniques, for example, may be applied prior to and / or concurrently with at least a portion of the operations performed by the playback DSP 530 of FIG. 5 to automatically recognize speech signals within the audio input 580a while it is being prepared for broadcast as the audio output 580b. The machine learning techniques, for example, may be used to produce a speech mask for identifying a frequency spectrum corresponding to present dialogue content in the audio input 580a.

[0131] The machine learning techniques, in some embodiments, include one or more parametric machine learning processes (e.g., configured as at least one parameterized machine learning model) trained to identify speech in the audio portion of an incoming stream of media content by reducing the identification of speech within audio to a simplified function having a controlled set of coefficients (e.g., parameters). The parametric machine learning process(s), in some examples, may enable high speed analysis of the incoming stream of media content (e.g., in real time or near- real time) by reducing the complexity of the analysis. The parameters of the parametric machine learning process(s), for example, may yield a generalized function capable of predicting whetherAttorney Docket No. SON00095WOU 1Client Docket No. 24-0404p or not speech is likely present in a current frame of the input signal. The likelihood, in some examples, may be represented as a confidence level or metric (e.g., percentage or absolute value) of how likely the input represents audio including speech, or an uncertainty level or metric (e.g., percentage or absolute value) representing how likely the machine learning process is correct in its determination regarding whether or not the particular input (e.g., frame) contains speech.

[0132] In some embodiments, the parameterized machine learning model includes a neural network, such as a deep neural network (DNN) model or an artificial neural network (ANN) model. In further examples, the parameterized machine learning model may be a recurrent neural network (RNN), a convolutional neural network (CNN) model, a Gaussian mixture model (GMM), or a hidden Markov model (HMM). The various options for machine learning techniques are described in greater detail, for example, in relation to U.S. Provisional Patent Application No. 63 / 700,280 entitled “Techniques for Speech Enhancement” and filed September 27, 2024.

[0133] The audio adjustment processing unit 314, in some implementations, obtains the speech audio portion 302a and the non-speech audio portion 302b and analyzes the speech audio portion 302a in view of the non-speech audio portion 302b to determine whether to adjust at least one of the first audio signals 302 or the second audio signals 316 and / or how to apply the adjustment (e.g., at what level, using what type of technique, etc.). For example, the audio adjustment processing unit 314 may compare the speech audio portion 302a to the non-speech audio portion 302b to determine whether the speech audio portion 302a represents less than a threshold percentage of the energy level of the non-speech audio data 302b. The energy levels, for example, may be quantified at least in part by the spatial processor 566 of FIG. 5. In a similar example, the audio adjustment processing unit 314 may compare the speech audio portion 302a to the first audio signals 302 to determine whether the speech audio portion 302a represents less than a threshold percentage of the energy level of the first audio signals 302 in its entirety. For example, if the speech audio portion 302a already represents a majority of the energy of the first audio signals 302 (e.g., at least 50% of the first audio signals 302 or at least 100% of the second audio signals 316), dialogue enhancement may be deemed unnecessary due to the anticipated clarity of the dialogue based on its share of the energy content of the first audio signals 302. The threshold percentage, for example, may be within a range of about 50% to about 80% of the energy content of the first audio signals 302. The threshold, further, may depend at least in part on the type of extraction performed by the dialogue extraction processing unit 402, such as a relative confidence level inAtorney Docket No. SON00095WOU 1Client Docket No. 24-0404p the speech audio portion 302a representing dialogue content. Conversely, a low confidence level (e.g., below at least a 90% threshold, below at least a 95% threshold, etc.) may alone be sufficient for the audio adjustment processing unit 314 to determine that no audio adjustment is needed since it is unlikely the first audio signals 302 contain dialogue. The confidence level, for example, may be provided by the dialogue extraction processing unit 402 (e.g., as a metadata portion of the speech audio portion 302a).

[0134] In some implementations, the speech audio portion 302a and / or the non-speech audio portion 302b are made available to the audio adjustment processing unit 314 for producing adjusted first audio signals 404a and / or adjusted second audio signals 404b. The audio adjustment processing unit 314, for example, may add the speech audio portion 302a to the first audio signals 302 to produce adjusted first audio signals 404a that amplify the frequency spectrum of the speech audio portion 302a. In other words, gain achieved through adding the speech audio portion 302a to the first audio signals 404a results in a gain that matches the set of frequencies of the speech audio portion 302a. In another example, the audio adjustment processing unit 314 may filter the second audio signals 316 by masking a set of frequencies contained in the speech audio portion 302a, effectively reducing the volume of the non-speech audio portion 302b or otherwise adjusting the output of the non-speech audio portion 302b to avoid competition in the speech frequency range with the speech audio portion 302a.

[0135] In some implementations, the adjusted first audio signals 404a and / or the adjusted second audio signals 404b may be provided for output to a transducer array 406. For example, the adjusted first audio signals 404a and / or the adjusted second audio signals 404b may be buffered as the audio output 580b of FIG. 5 for output via the transducers 114. Conversely, if the audio adjustment processing unit 314 determined that one or both of the first audio signals 302 or the second audio signals 316 should remain without adjustment, the first audio signals 302 and / or the second audio signals 316 may be provided for output via the transducer array 406. For example, if the dialogue extraction processing unit 402, through either the process of attempting to extract the speech audio portion 302a itself or the level of confidence in the extraction of the speech audio portion 302a, determines that there is a low likelihood of dialogue within the first audio signals 302, the sound levels in the first audio signals 302 and the second audio signals 316 may remain unadjusted by the process 400 (e.g., maintaining the sound levels in the speech frequency range as-is in the first audio signals 302 and the second audio signals 316).Atorney Docket No. SON00095WOU 1Client Docket No. 24-0404p

[0136] Although illustrated as a particular flow, in other embodiments, the process 400 may be performed in a different manner. For example, as discussed above, rather than producing both a speech audio portion 302a and a non-speech audio portion 302b, speech-identifying data may be produced that includes information sufficient to identify both the speech audio portion 302a and the non-speech audio portion 302b, such as a speech mask. Further, as discussed above, the audio adjustment processing unit 314 may determine, based on the output of the dialogue extraction processing unit 402, that no audio adjustment is desired for a particular timeframe of the first audio signals 302. Additionally, although illustrated as providing the adjusted first audio signals 404a and the adjusted second audio signals 404b directly to the transducer array 406, it should be understood that additional audio processing may occur prior to the adjusted first audio signals 404a and / or the adjusted second audio signals 404b being output for consumption by listeners. Other modifications of the process 400 are possible.V. Conclusion

[0137] The above discussions relating to enhanced dialogue rendering provide only some examples of operating environments within which functions and methods described below may be implemented. Other operating environments and configurations of media playback systems, playback devices, and network devices not explicitly described herein may also be applicable and suitable for implementation of the functions and methods.

[0138] The description above discloses, among other things, various example systems, methods, apparatus, and articles of manufacture including, among other components, firmware and / or software executed on hardware. It is understood that such examples are merely illustrative and should not be considered as limiting. For example, it is contemplated that any or all of the firmware, hardware, and / or software aspects or components can be embodied exclusively in hardware, exclusively in software, exclusively in firmware, or in any combination of hardware, software, and / or firmware. Accordingly, the examples provided are not the only ways to implement such systems, methods, apparatus, and / or articles of manufacture.

[0139] Additionally, references herein to “embodiment” means that a particular element, structure, or characteristic described in connection with the embodiment can be included in at least one example embodiment disclosed herein. The appearances of this phrase in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments mutually exclusive of other embodiments. As such, the embodimentsAtorney Docket No. SON00095WOU 1 Client Docket No. 24-0404p described herein, explicitly and implicitly understood by one skilled in the art, can be combined with other embodiments.

[0140] The specification is presented largely in terms of illustrative environments, systems, procedures, steps, logic blocks, processing, and other symbolic representations that directly or indirectly resemble the operations of data processing devices coupled to networks. These process descriptions and representations are typically used by those skilled in the art to most effectively convey the substance of their work to others skilled in the art. Numerous specific details are set forth to provide a thorough understanding of the present disclosure. However, it is understood to those skilled in the art that certain embodiments of the present disclosure can be practiced without certain, specific details. In other instances, well known methods, procedures, components, and circuitry have not been described in detail to avoid unnecessarily obscuring aspects of the embodiments. Accordingly, the scope of the present disclosure is defined by the appended claims rather than the foregoing description of embodiments.

[0141] When any of the appended claims are read to cover a purely software and / or firmware implementation, at least one of the elements in at least one example is hereby expressly defined to include a tangible, non-transitory medium such as a memory, DVD, CD, Blu-ray, and so on, storing the software and / or firmware.

Claims

Atorney Docket No. SON00095WOU 1 Client Docket No. 24-0404pCLAIMS1. A method comprising: receiving, via a media playback system, media content; determining a presence of dialogue in the media content, wherein the media content comprises first audio and at least second audio; after determining the presence of dialogue in the media content, generating adjusted second audio, wherein generating the adjusted second audio comprises reducing sound levels in a first frequency range of the second audio, the first frequency range comprising a set of frequencies associated with human speech; and causing playback of the first audio via a first transducer and playback of the adjusted second audio via a second transducer.

2. The method of claim 1, wherein determining the presence of the dialogue comprises identifying, in the media content, metadata flagging a segment of the first audio containing the dialogue.

3. The method of claim 1 or claim 2, wherein the playback system comprises one of a sound bar, a sound base, or a surround-sound speaker system.

4. The method of any of claims 1 through 3, wherein the method is executed on one or more processors of a playback device.

5. The method of claim 4, wherein: the first transducer is directed toward a listening area in a vicinity of the playback device; and the second transducer is directed away from the listening area.

6. The method of any of claims 1 through 5, wherein: the first audio is a center channel audio portion of the media content; and determining the presence of the dialogue comprises detecting, in the first audio, signals within the first frequency range.

7. The method of any of claims 1 through 6, wherein the first frequency range comprises frequencies within a range of between 90 Hz and 255 Hz.

8. The method of any of claims 1 through 7, further comprising, after determining the presence of dialogue in the media content, generating adjusted first audio, wherein generating theAtorney Docket No. SON00095WOU 1Client Docket No. 24-0404p adjusted first audio comprises at least one of i) emphasizing sound levels in the first frequency range of the first audio or ii) boosting all signals within the first frequency range by a voltage gain within a range from 1.8 decibels to 3.7 decibels.

9. The method of claim 8, wherein the voltage gain is selected based at least in part on one or more of a) a user setting, b) a type of the first transducer, or c) a zone setting of a multi-zone audio environment.

10. The method of claim 9, wherein determining the presence of the dialogue in the media content comprises extracting, from the first audio, a dialogue portion, wherein to extract comprises at least one of a) applying neural network analysis to the first audio to recognize speech content, or b) dividing the first audio into speech audio data and non-speech audio data.

11. The method of claim 10, wherein emphasizing the signals within the first frequency range comprises adding the extracted dialogue portion to the first audio to create a gain within the first frequency range matching the extracted dialogue portion.

12. The method of claim 10, further comprising: after dividing the first audio, comparing, by the hardware-based operations and / or software code executed by the one or more processors, the speech audio data to the non-speech audio data to determine an energy level of the extracted dialogue portion; wherein the adjusted second audio is generated responsive to the energy level being less than a threshold percentage of an energy level of the first audio; optionally wherein the threshold percentage is within a range of about 50% to about 80%.

13. The method of any of claims 1 to 12, wherein reducing the sound levels in the first frequency range of the second audio comprises attenuating all signals within the first frequency range by an attenuation within a range from 1.8 decibels to 3.7 decibels.

14. The method of any of claims 1 to 13, wherein determining the presence of the dialogue in the media content comprises recognizing, within a metadata portion of the media content, an indication of one or more time segments of the first audio including speech content, wherein the metadata portion optionally comprises closed caption data.

15. The method of any of claims 1 to 14, wherein: the first audio is a dialogue audio track; the second audio is a non-dialogue audio track; andAtorney Docket No. SON00095WOU 1 Client Docket No. 24-0404p determining the presence of the dialogue in the media content comprises identifying sound content provided on the dialogue audio track.

16. The method of any of claims 1 to 15, further comprising, after causing the playback of the first audio via the first transducer and the playback of the adjusted second audio via the second transducer: receiving subsequent media content, wherein the subsequent media content comprises subsequent first audio and at least subsequent second audio; determining an absence of dialogue in the subsequent media content; and maintaining, in the subsequent second audio, sound levels in the first frequency range of the subsequent second audio.

17. The method of claim 16, wherein: generating the adjusted second audio comprises filtering the second audio according to a speech-enhancing filtering setting; and maintaining the sound levels in the first frequency range comprises reverting to a default filtering setting; wherein, optionally, the adjusted second audio is generated by applying a scoop filter to the second audio.

18. A playback device comprising a network interface; a hardware input interface; a transducer array including a first transducer and a second transducer; at least one processor; and at least one non-transitory computer-readable medium comprising program instructions that are executable by the at least one processor such that the playback device is configured to receive, via at least one of the network interface or the hardware input interface, media content, wherein the media content comprises first audio and second audio, determine a presence of dialogue in the media content, generate, after determining the presence of dialogue in the media content, adjusted second audio, whereinAtorney Docket No. SON00095WOU 1Client Docket No. 24-0404p generating the adjusted second audio comprises reducing sound levels in a first frequency range of the second audio, the first frequency range comprising a set of frequencies associated with human speech, and cause playback of the first audio via the first transducer and playback of the adjusted second audio via the second transducer.

19. The playback device of claim 18, wherein the first audio is a center channel audio portion of the media content.

20. The playback device of claim 19, wherein determining the presence of the dialogue comprises detecting, in the first audio, signals within the first frequency range.

21. The playback device of any of claims 18 through 20, wherein the first frequency range comprises frequencies within a range of between 90 Hz and 255 Hz.

22. The playback device of claim any of claims 18 through 21, wherein: the at least one non-transitory computer-readable medium further comprises program instructions that are executable by the at least one processor such that the playback device is configured to generate, after determining the presence of dialogue in the media content, adjusted first audio; and generating the adjusted first audio comprises emphasizing sound levels in the first frequency range of the first audio.

23. The playback device of claim 22, wherein emphasizing the sound levels in the first frequency range of the first audio comprises boosting all signals within the first frequency range by a voltage gain within a range from 1.8 decibels to 3.7 decibels.

24. The playback device of claim 23, wherein the voltage gain is selected based at least in part on a user setting.

25. The playback device of claim 23 or 24, wherein the voltage gain is selected based at least in part on a type of the first transducer.

26. The playback device of any of claims 23 through 25, wherein the voltage gain is selected based at least in part on a zone setting of a multi-zone audio environment.

27. The playback device of any of claims 23 through 25, wherein determining the presence of the dialogue in the media content comprises extracting, from the first audio, a dialogue portion.Atorney Docket No. SON00095WOU 1Client Docket No. 24-0404p28. The playback device of claim 27, wherein extracting the dialogue portion comprises applying neural network analysis to the first audio to recognize speech content.

29. The playback device of claim 27 or 28, wherein: the playback device further comprises a media signal buffer configured to temporarily store at least the first audio; and extracting the dialogue portion comprises dividing the first audio into speech audio data and non-speech audio data.

30. The playback device of claim 29, wherein emphasizing the signals within the first frequency range comprises adding the extracted dialogue portion to the first audio to create a gain within the first frequency range matching the extracted dialogue portion.

31. The playback device of claim 29 or 30, wherein: the at least one non-transitory computer-readable medium further comprises program instructions that are executable by the at least one processor such that the playback device is configured to, after dividing the first audio, compare the speech audio data to the nonspeech audio data to determine an energy level of the extracted dialogue portion; and the adjusted second audio is generated responsive to the energy level being less than a threshold percentage of an energy level of the first audio.

32. The playback device of claim 31, wherein the threshold percentage is within a range of about 50% to about 80%.

33. The playback device of any of claims 18 through 32, wherein the playback device comprises one of a sound bar, a sound base, or a surround-sound speaker system.

34. The playback device of any of claims 18 through 33, wherein reducing the sound levels in the first frequency range of the second audio comprises attenuating all signals within the first frequency range by an attenuation within a range from 1.8 decibels to 3.7 decibels.

35. The playback device of any of claims 18 through 34, wherein determining the presence of the dialogue in the media content comprises recognizing, within a metadata portion of the media content, an indication of one or more time segments of the first audio including speech content.

36. The playback device of claim 35, wherein the metadata portion comprises closed caption data.Attorney Docket No. SON00095WOU 1 Client Docket No. 24-0404p37. The playback device of any of claims 18 through 36, wherein: the first audio is a dialogue audio track and the second audio is a non-dialogue audio track; and determining the presence of the dialogue in the media content comprises identifying sound content provided on the dialogue audio track.

38. The playback device of any of claims 18 through 37, wherein the first transducer is directed toward a listening area in a vicinity of the playback device; and the second transducer is directed away from the listening area.

39. The playback device of any of claims 18 through 38, wherein the at least one non-transitory computer-readable medium further comprises program instructions that are executable by the at least one processor such that the playback device is configured to, after causing the playback of the first audio via the first transducer and the playback of the adjusted second audio via the second transducer: receive, via at least one of the network interface or the hardware input interface, subsequent media content, wherein the subsequent media content comprises subsequent first audio and subsequent second audio; determine an absence of dialogue in the subsequent media content; and maintain, in the subsequent second audio, sound levels in the first frequency range of the subsequent second audio.

40. The playback device of claim 39, wherein: generating the adjusted second audio comprises filtering the second audio according to a speech-enhancing filtering setting; and maintaining the sound levels in the first frequency range comprises reverting to a default filtering setting.

41. The playback device of claim 40, wherein generating the adjusted second audio comprises applying a scoop filter to the second audio.

42. A method for enhancing rendering of dialogue by a playback device, the method comprising: receiving, via at least one of a network interface of the playback device or a hardware input interface of the playback device, media content, wherein the media content comprises first audio and second audio;Attorney Docket No. SON00095WOU 1Client Docket No. 24-0404p in real-time, determining, by hardware-based operations coded as logic circuitry of at least a first portion of one or more processors and / or at least a second portion of the one or more processors executing software code stored to a non-transitory computing readable medium, a presence of dialogue in the media content; after determining the presence of dialogue in the media content, generating, by the hardwarebased operations and / or software code executed by the one or more processors, adjusted second audio, wherein generating the adjusted second audio comprises reducing sound levels in a first frequency range of the second audio, the first frequency range comprising a set of frequencies associated with human speech; and by the hardware-based operations and / or software code executed by the one or more processors, causing playback of the first audio via a first transducer of the playback device and playback of the adjusted second audio via a second transducer of the playback device.

43. The method of claim 42, wherein determining the presence of the dialogue comprises identifying, in the media content, metadata flagging a segment of the first audio containing the dialogue.

44. The method of claim 42 or claim 43, wherein the first portion of the one or more processors comprises at least one a programmable logic device.

45. The method of any of claims 42 through 44, wherein a playback device comprises the one or more processors.

46. The method of any of claims 42 through 45, wherein the first audio is a center channel audio portion of the media content.

47. The method of claim 46, wherein determining the presence of the dialogue comprises detecting, in the first audio, signals within the first frequency range.

48. The method of any of claims 42 through 47, wherein the first frequency range comprises frequencies within a range of between 90 Hz and 255 Hz.

49. The method of any of claims 42 through 48, further comprising: after determining the presence of dialogue in the media content, generating, by the hardwarebased operations and / or software code executed by the one or more processors, adjusted first audio, whereinAtorney Docket No. SON00095WOU 1Client Docket No. 24-0404p generating the adjusted first audio comprises emphasizing sound levels in the first frequency range of the first audio.

50. The method of claim 49, wherein emphasizing the sound levels in the first frequency range of the first audio comprises boosting all signals within the first frequency range by a voltage gain within a range from 1.8 decibels to 3.7 decibels.

51. The method of claim 50, wherein the voltage gain is selected based at least in part on a user setting.

52. The method of claim 50 or claim 51, wherein the voltage gain is selected based at least in part on a type of the first transducer.

53. The method of any of claims 50 through 52, wherein the voltage gain is selected based at least in part on a zone setting of a multi-zone audio environment.

54. The method of any of claims 50 through 53, wherein determining the presence of the dialogue in the media content comprises extracting, from the first audio, a dialogue portion.

55. The method of claim 54, wherein extracting the dialogue portion comprises applying neural network analysis to the first audio to recognize speech content.

56. The method of claim 54 or claim 55, wherein extracting the dialogue portion comprises dividing the first audio into speech audio data and non-speech audio data.

57. The method of claim 56, wherein emphasizing the signals within the first frequency range comprises adding the extracted dialogue portion to the first audio to create a gain within the first frequency range matching the extracted dialogue portion.

58. The method of claim 56 or claim 57, further comprising: after dividing the first audio, comparing, by the hardware-based operations and / or software code executed by the one or more processors, the speech audio data to the non-speech audio data to determine an energy level of the extracted dialogue portion; wherein the adjusted second audio is generated responsive to the energy level being less than a threshold percentage of an energy level of the first audio.

59. The method of claim 58, wherein the threshold percentage is within a range of about 50% to about 80%.

60. The method of any of claims 42 through 59, wherein the playback device comprises one of a sound bar, a sound base, or a surround-sound speaker system.Attorney Docket No. SON00095WOU 1Client Docket No. 24-0404p61. The method of any of claims 42 through 60, wherein reducing the sound levels in the first frequency range of the second audio comprises attenuating all signals within the first frequency range by an attenuation within a range from 1.8 decibels to 3.7 decibels.

62. The method of any of claims 42 through 61, wherein determining the presence of the dialogue in the media content comprises recognizing, within a metadata portion of the media content, an indication of one or more time segments of the first audio including speech content.

63. The method of claim 62, wherein the metadata portion comprises closed caption data.

64. The method of any of claims 42 through 63, wherein: the first audio is a dialogue audio track; the second audio is a non-dialogue audio track; and determining the presence of the dialogue in the media content comprises identifying sound content provided on the dialogue audio track.

65. The method of any of claims 42 through 64, wherein the first transducer is directed toward a listening area in a vicinity of the playback device; and the second transducer is directed away from the listening area.

66. The method of any of claims 42 through 65, further comprising, after causing the playback of the first audio via the first transducer and the playback of the adjusted second audio via the second transducer: by the hardware-based operations and / or software code executed by the one or more processors, receiving, via at least one of the network interface or the hardware input interface, subsequent media content, wherein the subsequent media content comprises subsequent first audio and subsequent second audio; determining, by the hardware-based operations and / or software code executed by the one or more processors, an absence of dialogue in the subsequent media content; and by the hardware-based operations and / or software code executed by the one or more processors, maintaining, in the subsequent second audio, sound levels in the first frequency range of the subsequent second audio.Attorney Docket No. SON00095WOU 1 Client Docket No. 24-0404p67. The method of claim 66, wherein: generating the adjusted second audio comprises fdtering the second audio according to a speech-enhancing filtering setting; and maintaining the sound levels in the first frequency range comprises reverting to a default filtering setting.

68. The method of claim 67, wherein generating the adjusted second audio comprises applying a scoop filter to the second audio.

Citation Information

Patent Citations

  • Voice control of a media playback system

    US10499146B2

  • Room association based on name

    US10712997B2

  • Transmission wire clamp

    US4579306A

  • System and method for synchronizing operations among a plurality of independently clocked digital data processing devices

    US8234395B2

  • Controlling and manipulating groupings in a multi-zone media system

    US8483853B1