Audio Device Coordination

CHASM addresses the fragmented use of audio devices by coordinating them through a dynamic management system, enhancing audio quality and user experience by leveraging multiple devices for optimal capture and playback.

JP7710002B2Active Publication Date: 2025-07-17DOLBY INTERNATIONAL AB +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023076015
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-20
Filing Date
2023-05-02
Publication Date
2025-07-17
Estimated Expiration
2040-07-28

AI Technical Summary

Technical Problem

Existing audio devices, including smart audio devices, lack efficient coordination and management systems that enable seamless integration and orchestration of multiple devices for optimal audio experiences, leading to fragmented and suboptimal use of audio capabilities.

Method used

The implementation of a Continuous Hierarchical Audio Session Manager (CHASM) that coordinates and manages audio devices through a discoverable opportunistically orchestrated distributed audio subsystem (DOODAD), allowing for dynamic routing, signal processing, and device utilization based on user location and preferences, enabling unified audio experiences across multiple devices.

Benefits of technology

CHASM facilitates seamless audio device coordination, improving audio quality, reducing echo issues, and enhancing user experience by opportunistically leveraging multiple devices for optimal audio capture and playback, regardless of their initial intended purposes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710002000191
    Figure 0007710002000191
  • Figure 0007710002000192
    Figure 0007710002000192
  • Figure 0007710002000193
    Figure 0007710002000193
Patent Text Reader

Abstract

To provide a system and method for controlling a plurality of smart audio devices by an audio session manager according to respective media engine capabilities, for the plurality of smart audio devices.SOLUTION: A continuous hierarchical audio session manager (CHASM) 401 controls a plurality of smart audio devices according to media engine capabilities of respective media engines, via audio session management control signals sent via respective smart audio device communication links to respective smart audio devices 420-422 using application control signals 430-432, respectively. The CHASM transmits the audio session management control signals to the respective smart audio devices. The CHASM functions as a gateway.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application claims priority based on U.S. Provisional Patent Application No. 62 / 949,998 filed on December 18, 2019, U.S. Provisional Patent Application No. 62 / 992,068 filed on March 19, 2020, European Patent Application No. 19217580.0 filed on December 18, 2019, Spanish Patent Application P201930702 filed on July 30, 2019, U.S. Provisional Patent Application No. 62 / 971,421 filed on February 7, 2020, U.S. Provisional Patent Application No. 62 / 705,410 filed on June 25, 2020, U.S. Provisional Patent Application No. 62 / 880,114 filed on July 30, 2020, U.S. Provisional Patent Application No. 62 / 705,351 filed on June 23, 2020, U.S. Provisional Patent Application No. 62 / 880,115 filed on July 30, 2019, U.S. Provisional Patent Application No. 62 / 705,143 filed on June 12, 2020, U.S. Provisional Patent Application No. 62 / 880,118 filed on July 30, 2019, U.S. Patent Application No. 16 / 929,215 filed on July 15, 2020, U.S. Provisional Patent Application No. 62 / 705,883 filed on July 20, 2020, U.S. Provisional Patent Application No. 62 / 880,121 filed on July 30, 2019, and U.S. Provisional Patent Application No. 62 / 705,884 filed on July 20, 2020, and incorporates by reference in their entirety the disclosures of each application into this application.

[0002] The present disclosure relates to a system for coordinating (orchestrating) and implementing an audio device that may include a smart audio device. And a method.

Background Art

[0003] Audio devices (including, but not limited to, smart audio devices) are widely used and are becoming a common element in many households. Existing systems and methods for controlling audio devices provide benefits, but improved systems and methods are desired.

[0004] [Notation and Naming] Throughout the present disclosure, including the claims, the terms "speaker" and "loudspeaker" are used interchangeably to represent any acoustic radiation transducer (or set of transducers) driven by a single speaker feed. A typical headset includes two speakers. A speaker may be implemented to include multiple transducers (such as a woofer and a tweeter) that are driven by a single common speaker feed or multiple speaker feeds. In some examples, the speaker feed(s) may, in some cases, undergo different processing in different circuit branches connected to different transducers.

[0005] Throughout the present disclosure, including the claims, the expression "performing" an operation (such as filtering, scaling, transforming, or applying gain to a signal or data) on a signal or data is used in a broad sense to mean performing the operation directly on the signal or data or performing the operation on a processed version of the signal or data (such as a pre-filtered or pre-processed version of the signal before undergoing the execution of the operation).

[0006] Throughout the present disclosure, including the claims, the term "system" is used in a broad sense to mean a device, system, or subsystem. For example, implementing a decoder The subsystem to be described may be referred to as a decoder system, and a system including such a subsystem (for example, a system that generates X output signals in response to a plurality of inputs, where the subsystem generates M of the inputs and the other X - M inputs are received from an external source) may also be referred to as a decoder system.

[0007] Throughout this disclosure, including the claims, the term "processor" is used in a broad sense to mean a system or device that is programmable or otherwise configurable (e.g., by software or firmware) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chip sets), digital signal processors configured (e.g., programmed and / or otherwise) to perform pipelined processing on audio or other sound data, programmable general purpose processors or computers, and programmable microprocessor chips or chip sets, among others.

[0008] Throughout this disclosure, including the claims, the terms "connecting" or "connected" are used such that they may mean either a direct connection or an indirect connection. Thus, if a first device is connected to a second device, that connection may be by a direct connection or by an indirect connection through other devices and connections.

[0009] As used herein, a "smart device" is an electronic device that is configured to communicate with one or more other devices (or networks) via various wireless protocols such as Bluetooth, Zigbee, Near Field Communication, WiFi, Light Fidelity (LiFi), 3G, 4G, 5G, etc., and is operable to some extent interactively and / or autonomously. Some notable types of smart devices are smartphones, smart cars, smart thermostats, smart doorbells, smart locks, smart refrigerators, tablets and tablets, smart watches, smart bands, smart keychains, and smart audio devices. The term "smart device" may also refer to a device that exhibits some properties of ubiquitous computing such as artificial intelligence.

[0010] As used herein, the expression "smart audio device" is used to represent a smart device that is a single-purpose audio device or a multi-purpose audio device (e.g., an audio device that implements at least some aspects of a virtual assistant function). A single-purpose audio device is a device (e.g., a smart speaker, a television (TV), or a mobile phone) that includes or is connected to at least one microphone (and optionally also includes or is connected to at least one speaker and / or at least one camera) and is generally or primarily designed to achieve a single purpose. For example, a TV can typically (is considered to be able to) play audio from program material, but in most cases, modern TVs run some operating system on which multiple applications, including applications for watching TV, are locally executed. Similarly, the audio input and output of a mobile phone can do many things, but these are provided by the applications running on the phone. In this sense, a single-purpose audio device having speakers (singular or plural) and microphones (singular or plural) is often configured to execute local applications and / or services for directly using the speakers and microphones. There are also single-purpose audio devices configured to be grouped to achieve audio playback across a zone, i.e., a user-defined area.

[0011] One common type of multi-purpose audio device implements at least some aspects of a virtual assistant function, but other aspects of the virtual assistant function can be implemented by one or more other devices, such as one or more servers to which the multi-purpose audio device is configured to communicate. Such a multi-purpose audio device may be referred to herein as a "virtual assistant." A virtual assistant is a device (e.g., a smart speaker or a virtual assistant integrated device) that includes or is connected to (and optionally includes or is connected to at least one speaker and / or at least one camera) at least one microphone. In some examples, the virtual assistant is cloud-enabled in some sense or, otherwise, provides the ability to utilize multiple devices (different from the virtual assistant) for applications that are not fully implemented within or on the virtual assistant itself. In other words, at least some aspects of the virtual assistant function, such as a speech recognition function, can be implemented (at least in part) by one or more servers or other devices to which the virtual assistant can communicate via a network such as the Internet. Multiple virtual assistants may cooperate, for example, in a very discrete and conditionally defined manner. For example, two or more virtual assistants may cooperate in the sense that one of them (e.g., the one most confident in hearing the wake word) responds to that wake word. In some embodiments, multiple connected virtual assistants may form a kind of aggregate managed by one main application. That one main application can be a virtual assistant (or can include or implement a virtual assistant).

[0012] As used herein, the term "wake word" is used in a broad sense to mean any sound (e.g., a word uttered by a human, or some other sound). A smart audio device is configured to wake up in response to the detection of a sound (a "hearing", using at least one microphone included in or connected to the smart audio device, or at least one other microphone). In this context, "awake" means that the device enters a state of waiting for a sound command (i.e., listening ). In some examples, what may be referred to herein as a "wake word" may include multiple words, such as a phrase. As used herein, the term "wake word" is used in a broad sense to represent any sound (e.g., a word spoken by a person, or some other sound). Here, a smart audio device is configured to be awakened (awake) in response to the detection of that sound (a "hearing"), using at least one microphone within or connected to the smart audio device, or at least one other microphone. In this situation, "awakened" means that the device enters a state of waiting for a sound command (in other words, listening). In some cases, what may be referred to herein as a "wake word" may include more than one word, e.g., a phrase.

[0013] As used herein, the expression "wake word detector" refers to a device (or software including instructions for configuring the device) configured to continuously search for a match between real-time sound (e.g., speech) features and a learned model. Typically, a wake word event is triggered each time a wake word detector determines that the probability of a wake word being detected exceeds a pre-defined threshold. For example, the threshold may be a predetermined threshold adjusted to provide a good compromise between the false acceptance rate and the false rejection rate. After a wake word event, the device enters a state of listening for commands (sometimes referred to as the "awakened" state or the "attentiveness" " state), in which the received commands can be passed to a larger and more computationally intensive recognizer. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION

[0014] [Summary] In one class of embodiments, an audio device (which may include a smart audio device) is coordinated using a Continuous Hierarchical Audio Session Manager (CHASM). Some disclosures In an implementation example, at least some aspects of CHASM can be implemented by what is referred to herein as a "smart home hub." According to some examples, CHASM can be implemented by certain devices in an audio environment. In some cases, CHASM can be implemented at least in part via software that can be executed by one or more devices in an audio environment. In some embodiments, a device (e.g., a smart audio device) may also be referred to herein as a Discoverable Opportunistically Orchestrated Distributed Audio Subsystem (DOODAD), a network-connectable element or subsystem (e.g., a network-connectable media engine and device property descriptor), and a plurality (e.g., a large number) of devices (e.g., smart audio devices including DOODADs or other devices) are managed collectively by CHASM or otherwise moved in a manner that achieves an orchestrated function (e.g., replacing what is known or intended for a device when first purchased). Herein, the architecture of the development of a CHASM-enabled audio system, as well as the control language appropriate for the representation and control of audio functions, are described. Also described herein are the language of orchestration and the functional elements and functional differences for dealing with a collective audio system without directly referring to audio devices (or roots). Also described are the audio persistent sessions, destinations, prioritizations, and routings, as well as the requests for acknowledgment, which are specific to the idea of orchestrating and routing audio to and from people and places. including a network-connectable element or subsystem (e.g., a network-connectable media engine and device property descriptor), and a plurality (e.g., a large number) of devices (e.g., smart audio devices including DOODADs or other devices) are managed collectively by CHASM or otherwise moved in a manner that achieves an orchestrated function (e.g., replacing what is known or intended for a device when first purchased). Herein, the architecture of the development of a CHASM-enabled audio system, as well as the representation and control of audio functions. Also described herein are the language of orchestration and the functional elements and functional differences for dealing with a collective audio system without directly referring to audio devices (or roots). Also described are the audio persistent sessions, destinations, prioritizations, and routings, as well as the requests for acknowledgment, which are specific to the idea of orchestrating and routing audio to and from people and places.

Means for Solving the Problems

[0015] Aspects of the present disclosure include a system configured (e.g., programmed) to perform an embodiment of the disclosed method or any of its steps, and a tangible non-transitory computer-readable medium (e.g., a disk or other tangible storage medium) implementing non-transitory storage of data storing code (e.g., code executable to perform it) for performing an embodiment of the disclosed method or any of its steps. For example, some embodiments may be programmed (and / or otherwise configured) using software or firmware to perform any of various operations on data according to one or more of the disclosed methods or its steps, and / or may be, or may include, a programmable general-purpose processor, a digital signal processor, or a microprocessor. Such a general-purpose processor can be, or can include, a computer system including an input device, a memory, and a processing subsystem programmed (and / or otherwise configured) to perform more of the disclosed method (or its steps) in response to data asserted thereto.

[0016] At least some aspects of the present disclosure can be implemented via a method. In some cases, the method can be implemented, at least in part, by a control system as disclosed herein. Some such methods can include audio session management for an audio system of an audio environment.

[0017] Some such methods include establishing a first smart audio device communication link between the audio session manager of an audio system and at least a first smart audio device. In some examples, the first smart audio device may be either a single-purpose audio device or a multi-purpose audio device, or may include either. In some such examples, the first smart audio device includes one or more loudspeakers. Some such methods include establishing a first application communication link between the audio session manager and a first application device that executes a first application.

[0018] Some such methods include determining, by the audio session manager, one or more first media engine capabilities of a first media engine of the first smart audio device. In some examples, the first media engine is configured to manage one or more audio media streams received by the first smart audio device and perform first smart audio device signal processing on the one or more audio media streams according to a first media engine sample clock.

[0019] In some such examples, the method includes receiving, by an audio session manager, a first application control signal from a first application via a first application communication link. Some such methods include controlling a first smart audio device according to first media engine capabilities. According to some implementations, controlling is performed by the audio session manager via a first audio session management control signal transmitted to the first smart audio device via a first smart audio device communication link. In some such examples, the audio session manager transmits the first audio session management control signal to the first smart audio device without referring to a first media engine sample clock.

[0020] In some implementations, the first application communication link may be established in response to a first route start request from a first application device. According to some examples, the first application control signal may be transmitted from the first application without referring to a first media engine sample clock. In some examples, the first audio session management control signal may cause the first smart audio device to delegate control of the first media engine to the audio session manager.

[0021] According to some examples, a device other than the audio session manager or the first smart audio device may be configured to execute the first application. However, in some cases, the first smart audio device may be configured to execute the first application.

[0022] In some examples, the first smart audio device may include a specific-purpose audio session manager. According to some such examples, the audio session manager may communicate with a specific-purpose audio session manager via a first smart audio device communication link. According to some such examples, the audio session manager may obtain one or more first media engine capabilities from the specific-purpose audio session manager. for the specific purpose of the audio session manager to obtain.

[0023] According to some implementation examples, the audio session manager may function as a gateway for all applications that control the first media engine, regardless of whether the application operates on the first smart audio device or on another device.

[0024] Some such methods may also include, at least, establishing a first audio stream corresponding to a first audio source. The first audio stream may include a first audio signal. In some such examples, establishing at least the first audio stream may include causing the first smart audio device to establish at least the first audio stream via a first audio session management control signal transmitted to the first smart audio device via a first smart audio device communication link.

[0025] In some examples, such a method may also include a rendering process that causes the first audio signal to be rendered as a first rendered audio signal. In some examples, the rendering process may be performed by the first smart audio device in response to a first audio session management control signal.

[0026] Some such methods may also include establishing a smart audio device - to - device communication link between a first smart audio device and each of one or more other smart audio devices in an audio environment via a first audio session management control signal. Some such methods may also include causing the first smart audio device to transmit raw microphone signals, processed microphone signals, rendered audio signals, and / or unrendered audio signals to one or more other smart audio devices via the smart audio device - to - device communication link.

[0027] In some examples, such methods may also include establishing a second smart audio device communication link between an audio session manager and at least a second smart audio device of a home audio system. The second smart audio device may be either a single - purpose audio device or a multi - purpose audio device, or may include either. The second smart audio device may include one or more microphones. Some such methods may also include determining, by the audio session manager, one or more second media engine capabilities of a second media engine of the second smart audio device. The second media engine may be configured to receive microphone data from one or more microphones and perform second smart audio device signal processing on the microphone data. Some such methods may also include controlling the second smart audio device via a second audio session manager control signal transmitted to the second smart audio device via the second smart audio device communication link according to the second media engine capabilities.

[0028] According to some such examples, controlling a second smart audio device may also include causing the second smart audio device to establish a smart audio device-to-device communication link between the second smart audio device and the first smart audio device. In some examples, controlling a second smart audio device includes causing the second smart audio device to transmit processed and / or unprocessed microphone data from a second media engine to a first media engine via the smart audio device-to-device communication link. This may include causing the second smart audio device to transmit processed and / or unprocessed microphone data from a second media engine to a first media engine via the smart audio device-to-device communication link.

[0029] In some examples, controlling a second smart audio device may include receiving, by an audio session manager, a first application control signal from a first application via a first application communication link, and determining a second audio session manager control signal according to the first application control signal.

[0030] Alternatively, or in addition, some audio session management methods include receiving, from a first device implementing a first application and by a device implementing an audio session manager, a first route start request for starting a first route for a first audio session. In some examples, the first route start request indicates a first audio source and a first audio environment destination, and the first audio environment destination corresponds to at least a first person within the audio environment, but the first audio environment destination does not indicate an audio device.

[0031] Some such methods include establishing a first route in response to a first route initiation request by a device implementing an audio session manager. According to some examples, establishing the first route includes determining a first position of at least a first person within an audio environment, determining at least one audio device for a first stage of a first audio session, and starting or scheduling the first audio session.

[0032] According to some examples, the first route initiation request may include a first audio session priority. In some cases, the first route initiation request may include a first connectivity mode. For example, the first connectivity mode may be a synchronous connection mode, a transactional connection mode, or a scheduled connection mode.

[0033] In some implementations, the first route initiation request may include an indication of whether approval will be required from at least the first person. In some cases, the first route initiation request may include a first audio session goal. For example, the first audio session goal may include intelligibility, audio quality, spatial fidelity, audibility, inaudibility, and / or privacy.

[0034] Some such methods may include determining a first persistent unique audio session identifier for the first route. Such methods may include sending the first persistent unique audio session identifier to a first device.

[0035] According to some examples, establishing a first route may include causing at least one device in the environment to establish at least a first media stream corresponding to the first route. The first media stream includes a first audio signal. Some such methods may include causing the first audio signal to be rendered to a first rendered audio signal.

[0036] Some such methods may include determining a first orientation of a first person with respect to a first stage of an audio session. Some such examples according to which the first audio signal is rendered to a first rendered audio signal may include determining a first reference spatial mode corresponding to the first position and the first orientation of the first person and determining a first relative activation of a loudspeaker in the audio environment corresponding to the first reference spatial mode. rendering the first audio signal to a first rendered audio signal may include determining a first reference spatial mode corresponding to the first position and the first orientation of the first person and determining a first relative activation of a loudspeaker in the audio environment corresponding to the first reference spatial mode.

[0037] Some such methods may include determining a second position and / or a second orientation of the first person with respect to a second stage of the first audio session. Some such methods may include determining a second reference spatial mode corresponding to the second position and / or the second orientation and determining a second relative activation of a loudspeaker in the audio environment corresponding to the second reference spatial mode.

[0038] According to some examples, the method may include receiving, from a second device implementing a second application and by a device implementing an audio session manager, a second route start request for starting a second route for a second audio session. The second route start request may indicate a second audio source and a second audio environment destination. The second audio environment destination may correspond to at least a second person within the audio environment. In some examples, the second audio environment destination does not indicate an audio device.

[0039] Some such methods may include establishing, by a device implementing an audio session manager, a second route corresponding to the second route start request. In some implementations, establishing the second route may include determining a first position of at least a second person within the audio environment, determining at least one audio device for a first stage of the second audio session, and starting the second audio session. In some examples, establishing the second route may include at least establishing a second media stream corresponding to the second route. The second media stream includes a second audio signal. Some such methods may include causing the second audio signal to be rendered to a second rendered audio signal.

[0040] Some such methods may include changing the rendering process for a first audio signal based at least in part on at least one of a second audio signal, a second rendered audio signal, or a characteristic thereof to generate a modified first rendered audio signal. According to some examples, changing the rendering process for the first audio signal may include warping the rendering of the first audio signal away from the rendering position of the second rendered audio signal. Alternatively, or in addition, changing the rendering process for the first audio signal may include changing the loudness of one or more signals of the first rendered audio signal in response to the loudness of one or more signals of the second audio signal or the second rendered audio signal.

[0041] In some examples, the first route start request may indicate at least a first area of the audio environment as a first route source or a first route destination. In some implementations, the first route start request may indicate at least a first service (e.g., an online content providing service such as a music providing service or a podcast providing service) as the first audio source.

[0042] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described herein. Memory devices may include, but are not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, and the like. Thus, some innovative aspects of the subject matter described in this disclosure can be implemented within one or more non-transitory media having stored software.

[0043] For example, software may include instructions for controlling one or more devices to perform one or more methods including audio session management for an audio system in an audio environment. Some such methods include establishing a first smart audio device communication link between an audio session manager of the audio system and at least a first smart audio device. In some examples, the first smart audio device may be either a single-purpose audio device or a multi-purpose audio device, or may include either. In some such examples, the first smart audio device includes one or more loudspeakers. Some such methods include establishing a first application communication link between the audio session manager and a first application device that executes a first application.

[0044] Some such methods include determining, by the audio session manager, one or more first media engine capabilities of a first media engine of the first smart audio device. In some examples, the first media engine is configured to manage one or more audio media streams received by the first smart audio device and perform first smart audio device signal processing on the one or more audio media streams according to a first media engine sample clock.

[0045] In some such examples, the method includes receiving, by an audio session manager, a first application control signal from a first application via a first application communication link. Some such methods include controlling a first smart audio device according to first media engine capabilities. According to some implementations, controlling is performed by the audio session manager via a first audio session management control signal transmitted to the first smart audio device via a first smart audio device communication link. In some such examples, the audio session manager transmits the first audio session management control signal to the first smart audio device without referring to a first media engine sample clock.

[0046] In some implementations, the first application communication link may be established in response to a first route start request from a first application device. According to some examples, the first application control signal may be transmitted from the first application without referring to a first media engine sample clock. In some examples, the first audio session management control signal may cause the first smart audio device to delegate control of the first media engine to the audio session manager.

[0047] According to some examples, a device other than the audio session manager or the first smart audio device may be configured to execute the first application. However, in some cases, the first smart audio device may be configured to execute the first application.

[0048] In some examples, the first smart audio device may include an audio session manager for a particular purpose. According to some such examples, the audio session manager may communicate with an audio session manager for a particular purpose via a first smart audio device communication link. According to some such examples, the audio session manager may obtain one or more first media engine capabilities from an audio session manager for a particular purpose.

[0049] According to some implementations, the audio session manager may function as a gateway for all applications that control the first media engine, regardless of whether the application operates on the first smart audio device or on another device.

[0050] Some such methods may also include establishing, at least, a first audio stream corresponding to a first audio source. The first audio stream may include a first audio signal. In some such examples, establishing at least the first audio stream may include causing the first smart audio device to establish at least the first audio stream via a first audio session management control signal transmitted to the first smart audio device via a first smart audio device communication link.

[0051] In some examples, such a method may also include a rendering process that causes the first audio signal to be rendered to a first rendered audio signal. In some examples, the rendering process may be performed by the first smart audio device in response to a first audio session management control signal.

[0052] Some such methods may also include establishing a smart audio device - to - device communication link between the first smart audio device and each of one or more other smart audio devices in the audio environment via a first audio session management control signal. Some such methods may also include causing the first smart audio device to transmit raw microphone signals, processed microphone signals, rendered audio signals, and / or unrendered audio signals to one or more other smart audio devices via the smart audio device - to - device communication link.

[0053] In some examples, such methods may also include establishing a second smart audio device communication link between the audio session manager and at least a second smart audio device of the home audio system. The second smart audio device may be either a single - purpose audio device or a multi - purpose audio device, or may include either. The second smart audio device may include one or more microphones. Some such methods may also include determining, by the audio session manager, one or more second media engine capabilities of a second media engine of the second smart audio device. The second media engine may be configured to receive microphone data from one or more microphones and perform second smart audio device signal processing on the microphone data. Some such methods may also include controlling the second smart audio device via a second audio session manager control signal transmitted to the second smart audio device via the second smart audio device communication link, according to the second media engine capabilities.

[0054] According to some such examples, controlling the second smart audio device It may also include causing a second smart audio device to establish a smart audio device - to - device communication link between the second smart audio device and the first smart audio device. In some examples, controlling the second smart audio device may include causing the second smart audio device to transmit processed and / or unprocessed microphone data from the second media engine to the first media engine via the smart audio device - to - device communication link to the first smart audio device.

[0055] In some examples, controlling the second smart audio device may include receiving, by an audio session manager, a first application control signal from a first application via a first application communication link, and determining a second audio session manager control signal according to the first application control signal.

[0056] Alternatively, or in addition, the software may include instructions for controlling one or more devices to perform one or more other methods including audio session management for an audio system of an audio environment. Some such audio session management methods include receiving, from a first device implementing a first application and by a device implementing an audio session manager, a first route start request for starting a first route for a first audio session. In some examples, the first route start request indicates a first audio source and a first audio environment destination, and the first audio environment destination corresponds to at least a first person within the audio environment, but the first audio environment destination does not indicate an audio device.

[0057] Some such methods include establishing a first route in response to a first route start request by a device implementing an audio session manager. According to some examples, establishing the first route includes determining a first position of at least a first person in an audio environment, determining at least one audio device for a first stage of a first audio session, and starting or scheduling the first audio session.

[0058] According to some examples, the first route start request may include a first audio session priority. In some cases, the first route start request may include a first connection mode. For example, the first connection mode may be a synchronous connection mode, a transaction connection mode, or a scheduled connection mode.

[0059] In some implementations, the first route start request may include an indication of whether approval will be required from at least the first person. In some cases, the first route start request may include a first audio session goal. For example, the first audio session goal may include intelligible, audio quality, spatial fidelity, audible, inaudible, and / or privacy.

[0060] Some such methods may include determining a first persistent unique audio session identifier for the first route. Such methods may include sending the first persistent unique audio session identifier to a first device.

[0061] According to some examples, establishing the first route may include causing at least one device in the environment to establish at least a first media stream corresponding to the first route, the first media stream including a first audio signal. Some such methods may include causing the first audio signal to be rendered as a first rendered audio signal to be rendered.

[0062] Some such methods may include determining a first orientation of a first person with respect to a first stage of an audio session. According to some such examples, causing a first audio signal to be rendered to a first rendered audio signal may include determining a first reference spatial mode corresponding to a first position and a first orientation of the first person, and determining a first relative activation of loudspeakers in an audio environment corresponding to the first reference spatial mode.

[0063] Some such methods may include determining a second position and / or a second orientation of a first person with respect to a second stage of a first audio session. Some such methods may include determining a second reference spatial mode corresponding to the second position and / or the second orientation, and determining a second relative activation of loudspeakers in an audio environment corresponding to the second reference spatial mode.

[0064] According to some examples, the method may include receiving, from a second device implementing a second application and by a device implementing an audio session manager, a second route start request for starting a second route for a second audio session. The second route start request may indicate a second audio source and a second audio environment destination. The second audio environment destination may correspond to at least a second person in the audio environment. In some examples, the second audio environment destination does not indicate an audio device.

[0065] Some such methods may include establishing a second route corresponding to a second route start request by a device implementing an audio session manager. In some implementation examples, establishing the second route may include determining a first position of at least a second person within the audio environment, determining at least one audio device for a first stage of a second audio session, and starting the second audio session. In some examples, establishing the second route may include at least establishing a second media stream corresponding to the second route. The second media stream includes a second audio signal. Some such methods may include causing the second audio signal to be rendered to a second rendered audio signal.

[0066] Some such methods may include changing a rendering process for a first audio signal based at least in part on at least one of the second audio signal, the second rendered audio signal, or their characteristics to generate a modified first rendered audio signal. According to some examples, changing the rendering process for the first audio signal may include warping the rendering of the first audio signal away from the rendering position of the second rendered audio signal. Alternatively, or in addition, changing the rendering process for the first audio signal may include changing the loudness of one or more signals of the first rendered audio signal in response to the loudness of one or more signals of the second audio signal or the second rendered audio signal.

[0067] In some examples, the first route start request may indicate at least a first area of the audio environment as the first route source or the first route destination. In some implementation examples, the first route start request may indicate at least a first service (e.g., an online content providing service such as a music providing service or a podcast service) as the first audio source.

[0068] In some implementation examples, the device (or system) may include an interface system and a control system. The control system may include one or more general-purpose single or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or combinations thereof.

[0069] In some implementation examples, the control system may be configured to implement one or more of the methods disclosed herein. Some such methods may include audio session management for an audio system of an audio environment. According to some such examples, the control system may be configured to implement what may be referred to herein as an audio session manager.

[0070] Some such methods include establishing a first smart audio device communication link between an audio session manager (e.g., a device implementing the audio session manager) and at least a first smart audio device of the audio system. In some examples, the first smart audio device may be either a single-purpose audio device or a multi-purpose audio device, or may include any of them. In some such examples, the first smart audio device includes one or more loudspeakers. Some such methods include establishing a first application communication link between the audio session manager and a first application device executing a first application.

[0071] Some such methods include determining, by the audio session manager, one or more first media engine capabilities of a first media engine of the first smart audio device. In some examples, the first media engine is configured to manage one or more audio media streams received by the first smart audio device and perform first smart audio device signal processing on the one or more audio media streams according to a first media engine sample clock.

[0072] In some such examples, the method includes receiving, by an audio session manager, a first application control signal from a first application via a first application communication link. Some such methods include controlling a first smart audio device according to first media engine capabilities. According to some implementations, controlling is performed by an audio session manager via a first audio session management control signal transmitted to the first smart audio device via a first smart audio device communication link. In some such examples, the audio session manager transmits the first audio session management control signal to the first smart audio device without referring to a first media engine sample clock.

[0073] In some implementations, the first application communication link can be established in response to a first route start request from a first application device. According to some examples, the first application control signal can be transmitted from the first application without referring to a first media engine sample clock. In some examples, the first audio session management control signal can cause the audio session manager to proxy control of the first media engine to the first smart audio device. In some examples, a device other than the audio session manager or the first smart audio device can be configured to execute the first application. However, in some cases, the first smart audio device can be configured to execute the first application.

[0074] According to some examples, a device other than the audio session manager or the first smart audio device can be configured to execute the first application. However, in some cases, the first smart audio device can be configured to execute the first application.

[0075] In some examples, the first smart audio device may include an audio session manager for a particular purpose. According to some such examples, the audio session manager may communicate with the audio session manager for a particular purpose via a first smart audio device communication link. According to some such examples, the audio session manager may obtain one or more first media engine capabilities from the audio session manager for a particular purpose.

[0076] According to some implementations, the audio session manager may function as a gateway for all applications that control the first media engine, regardless of whether the application operates on the first smart audio device or on another device.

[0077] Some such methods may also include establishing at least a first audio stream corresponding to a first audio source. The first audio stream may include a first audio signal. In some such examples, establishing at least the first audio stream may include causing the first smart audio device to establish at least the first audio stream via a first audio session management control signal transmitted to the first smart audio device via the first smart audio device communication link.

[0078] In some examples, such a method may also include a rendering process that causes the first audio signal to be rendered to a first rendered audio signal. In some examples, the rendering process may be performed by the first smart audio device in response to a first audio session management control signal.

[0079] Some such methods may also include establishing, via a first audio session management control signal, a smart audio device - to - device communication link between a first smart audio device and each of one or more other smart audio devices in an audio environment. Some such methods may also include causing the first smart audio device to transmit, via a smart audio device - to - device communication link, raw microphone signals, processed microphone signals, rendered audio signals, and / or unrendered audio signals to one or more other smart audio devices.

[0080] In some examples, such methods may also include establishing a second smart audio device communication link between an audio session manager and at least a second smart audio device of a home audio system. The second smart audio device may be either a single - purpose audio device or a multi - purpose audio device, or may include either. The second smart audio device may include one or more microphones. Some such methods may also include determining, by the audio session manager, one or more second media engine capabilities of a second media engine of the second smart audio device. The second media en gine may be configured to receive microphone data from one or more microphones and perform second smart audio device signal processing on the microphone data. Some such methods may also include controlling the second smart audio device, by the audio session manager, via a second audio session manager control signal transmitted to the second smart audio device via the second smart audio device communication link, according to the second media engine capabilities.

[0081] According to some such examples, controlling a second smart audio device may also include causing the second smart audio device to establish a smart audio device - to - device communication link between the second smart audio device and the first smart audio device. In some examples, controlling the second smart audio device may include causing the second smart audio device to transmit processed and / or unprocessed microphone data from a second media engine to a first media engine via the smart audio device - to - device communication link.

[0082] In some examples, controlling the second smart audio device may include receiving, by an audio session manager, a first application control signal from a first application via a first application communication link, and determining a second audio session manager control signal according to the first application control signal.

[0083] Alternatively, or in addition, the control system may be configured to implement one or more other audio session management methods. Some such audio session management methods include receiving, from a first device implementing a first application and by a device implementing an audio session manager, a first route start request for starting a first route for a first audio session. In some examples, the first route start request indicates a first audio source and a first audio environment destination, and the first audio environment destination corresponds to at least a first person within the audio environment, but the first audio environment destination does not indicate an audio device.

[0084] Some such methods include establishing a first route in response to a first route start request by a device implementing an audio session manager. According to some examples, establishing the first route includes determining a first position of at least a first person within the audio environment, determining at least one audio device for a first stage of a first audio session, and starting or scheduling a first audio session.

[0085] According to some examples, the first route start request may include a first audio session priority. In some cases, the first route start request may include a first connection mode. For example, the first connection mode may be a synchronous connection mode, a transaction connection mode, or a scheduled connection mode.

[0086] In some implementations, the first route start request may include an indication of whether approval will be required from at least the first person. In some cases, the first route start request may include a first audio session goal. For example, the first audio session goal may include intelligible, audio quality, spatial fidelity, audible, inaudible, and / or privacy.

[0087] Some such methods may include determining a first persistent unique audio session identifier for the first route. Such methods may include sending the first persistent unique audio session identifier to a first device. Some such methods may include causing at least one device in the environment to establish at least a first media stream corresponding to the first route, the first media stream including a first audio signal. Some such methods may include causing the first audio signal to be rendered to a first rendered audio signal.

[0088] According to some examples, establishing the first route may include causing at least one device in the environment to establish at least a first media stream corresponding to the first route, the first media stream including a first audio signal. Some such methods may include causing the first audio signal to be rendered to a first rendered audio signal.

[0089] Some such methods may include determining a first orientation of a first person with respect to a first stage of an audio session. According to some such examples, causing a first audio signal to be rendered to a first rendered audio signal may include determining a first reference spatial mode corresponding to a first position and a first orientation of the first person, and determining a first relative activation of loudspeakers in an audio environment corresponding to the first reference spatial mode.

[0090] Some such methods may include determining a second position and / or a second orientation of the first person with respect to a second stage of the first audio session. Some such methods may include determining a second reference spatial mode corresponding to the second position and / or the second orientation, and determining a second relative activation of loudspeakers in the audio environment corresponding to the second reference spatial mode.

[0091] According to some examples, the method may include receiving, from a second device implementing a second application and by a device implementing an audio session manager, a second route start request for starting a second route for a second audio session. The second route start request may indicate a second audio source and a second audio environment destination. The second audio environment destination may correspond to at least a second person in the audio environment. In some examples, the second audio environment destination does not indicate an audio device.

[0092] Some such methods may include establishing a second route corresponding to a second route start request by a device implementing an audio session manager. In some implementation examples, establishing the second route may include determining a first position of at least a second person within the audio environment, determining at least one audio device for a first stage of a second audio session, and starting the second audio session. In some examples, establishing the second route may include at least establishing a second media stream corresponding to the second route. The second media stream includes a second audio signal. Some such methods may include causing the second audio signal to be rendered to a second rendered audio signal.

[0093] Some such methods may include changing a rendering process for a first audio signal at least partially based on at least one of a second audio signal, a second rendered audio signal, or its characteristics to generate a modified first rendered audio signal. According to some examples, changing the rendering process for the first audio signal may include warping the rendering of the first audio signal away from the rendering position of the second rendered audio signal. Alternatively, or in addition, changing the rendering process for the first audio signal may include changing the loudness of one or more signals of the first rendered audio signal in response to the loudness of one or more signals of the second audio signal or the second rendered audio signal. Some such methods may include changing the loudness of one or more signals of the first rendered audio signal.

[0094] In some examples, the first route start request may indicate at least a first area of the audio environment as the first route source or the first route destination. In some implementations, the first route start request may indicate at least a first service (e.g., an online content providing service such as a music providing service or a podcast service) as the first audio source.

[0095] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will be apparent from the specification, drawings, and claims. Note that the relative dimensions of the figures below may not be drawn to exact scale.

Brief Description of the Drawings

[0096] Brief Description of the Drawings

Figure 1A

Figure 1B

Figure 1C

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 3A

Figure 3B

Figure 3C

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 10

Figure 11A

Figure 11B

Figure 11C

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18A

Figure 18B

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44

Figure 45

Figure 46

Figure 47

Figure 48A

Figure 48B

Figure 48C

Figure 48D

Figure 49

Figure 50A

Figure 50B

Figure 50C

Figure 51A

Figure 51B

Figure 52

Figure 53

Figure 54

Figure 55

Figure 56A

Figure 56B

Figure 57

Figure 58

Figure 59

Figure 60

[0097] Detailed description of the embodiment Numerous embodiments are disclosed. How to implement these embodiments will be apparent to those skilled in the art from the present disclosure.

[0098] Currently, designers view an audio device as a single interface point for audio that can be a mix of entertainment, communication, and information services. Using audio for notifications and voice control has the advantage of avoiding visual or physical intervention. The expanding device landscape is becoming more fragmented as more systems compete for our pair of ears. Wearable augmented audio is starting to become available, but does not appear to be converging towards an ideal general-purpose audio personal assistant, nor are many of the devices around us capable of being used for seamless capture, connection, and communication.

[0099] It is useful to develop services for bridging devices and better manage location, context, content, timing, and user preferences. Along with this, a set of standards, infrastructure, and APIs may enable better access to consolidated access to an audio space around the user. Consider an operating system that manages basic audio input / output and enables connection of an audio device to a particular application. This idea and design form the framework of an interactive audio transport, for example, enabling rapid and organic development of improvements and providing a service that provides an audio connection that is device-independent of others.

[0100] The scope of audio interaction includes real-time communication, asynchronous chat, alerts, speech transcription, history, archives, music, recommendations, reminders, facilitation, and context-aware assistance. Disclosed herein is a platform that facilitates an integrated approach and can implement intelligent media formats. The platform may include, or be implementable with, ubiquitous wearable audio and / or manage the positioning of the user, the selection of one or more (e.g., a set of) audio devices for optimal use, identity, privacy, timeliness, geographical location, and / or the infrastructure for transport, storage, retrieval, and algorithm execution. Some aspects of the present disclosure may include managing identity, priority (rank), and respect for the user's preferences, e.g., the desirability of listening and the value of being heard. The cost of unwanted audio is high. The inventors believe that the "Internet of Audio" can provide or implement an integral element of security and trust.

[0101] The categories of single-purpose and multi-purpose audio devices are not strictly orthogonal, but the speaker(s) and microphone(s) of an audio device (e.g., a smart audio device) can be assigned to the functions enabled, attached to (or implemented by) the smart audio device. However, typically, the speaker and / or microphone of an audio device are considered separately (distinct from the audio device) and there is no meaning in being added as a group.

[0102] As used herein, local audio devices (each of which may include a speaker and a microphone) are notified and made available to a collective audio platform that exists independently of any of the local audio devices in an abstract sense. Describe the category of audio device connectivity that can be enabled. Also, describe an embodiment that includes at least one discoverable opportunistically orchestrated distributed audio subsystem (DOODAD) that implements design techniques and a set of steps towards the realization of this idea for the orchestration and utilization of a collective set of audio devices.

[0103] To illustrate the application and results of some embodiments of the present disclosure, a simple example will be described with reference to FIG. 1A.

[0104] FIG. 1A shows an example of a person and smart audio devices in an audio environment. In this example, the audio environment 100 is a home audio environment.

[0105] In the scenario of FIG. 1A, a person (102) can capture the user's voice (103) using a microphone and has a smart audio device (101) capable of speaker playback (104). The user can speak from a considerable distance from device 101, which limits the duplex ability of device 101.

[0106] In FIG. 1A, the labeled elements are as follows. ● 100. An audio environment illustrating an example usage scenario of smart audio devices. A person 102 sitting on a chair is interacting with device 101. ● 101. A smart audio device capable of audio playback via speakers (singular or plural) and audio capture from microphones (singular or plural). ● 102. A person (also called a user or listener) participating in an audio experience using device 101. ● 103. Sound emitted by person 102 speaking to device 101. ● 104. Audio reproduced from the speakers (singular or plural) of device 101.

[0107] As such, in this example, FIG. 1A is a diagram of a person (102) at home having a smart audio device (101), which is a telephone device for communication in this case. The device 101 can output the audio heard by the person 102 and can also capture the sound (103) from the person 102. However, since the device 101 is at a certain distance from the person 102, the microphone on the device 101 has the problem of hearing the person 102 through the audio output from the device 101 (a problem known as the echo problem). Prior art for solving this problem typically uses echo cancellation. Echo cancellation is a form of processing that is greatly limited by full-duplex activity (double talk, i.e., two-way simultaneous audio).

[0108] FIG. 1B is a diagram of a modification of the scenario of FIG. 1A. In this example, a second smart audio device (105) is also present within the audio environment 100. The smart audio device 105 is placed near the person 102 but is designed to have the function of outputting audio. In some examples, the smart audio device 105 can implement a virtual assistant at least partially.

[0109] Currently, the person 102 (in FIG. 1A) has acquired the second smart audio device 105, but the device 105 (shown in FIG. 1B) can only perform a specific purpose (106) that is completely different from the purpose of the first device (101). In this example, the two smart audio devices (101 and 105) are in the same acoustic space as the user, yet they cannot share information and orchestrate experiences.

[0110] In FIG. 1B, the labeled elements are as follows. ● 100 - 104. Refer to FIG. 1A. ●105. Further smart audio devices capable of audio playback via a speaker(s) and audio capture from a microphone(s). ●106. Audio playback via the speaker(s) of device 105.

[0111] It may be possible to pair or shift an audio call from the phone (101) to this smart audio device (105), which has not been possible heretofore without user intervention and detailed configuration. Thus, the scenario illustrated in FIG. 1B is a situation where there are two independent audio devices, each performing a very specific application. In this example, the smart audio device 105 was purchased more recently than device 100. The smart audio device (105) is purchased for a specific purpose and, out of the box, is only useful for that specific purpose(s) and does not immediately add value to the device (101) already present in the audio environment 100 and being used as a communication device.

[0112] FIG. 1C is a diagram according to an embodiment of the present disclosure. In this example, the smart audio device 105 was purchased more recently than the smart audio device 100.

[0113] In the embodiment of FIG. 1C, the smart audio devices 101 and 105 are capable of orchestration. In this example of orchestration, the smart audio device (105) is positioned in a better location to pick up the voice (103B) of person 102 for a call involving the smart audio device (101) when the smart audio device 101 plays sound 104 from the speaker(s).

[0114] In FIG. 1C, the labeled elements are as follows. ●100 - 104. Refer to FIG. 1A. ●105. Refer to FIG. 1B. ● The microphone(s) of the smart audio device (105) of 103B is / are close to the user, so the sound emitted by the person 102 can be better captured by the smart audio device (105).

[0115] In Figure 1C, the new smart audio device 105 can be detected in some way (examples of which are described herein) such that the microphone within the device 105 can function to support the application that operated on the smart audio device (101). The new smart audio device 105 of Figure 1C is coordinated or orchestrated (in accordance with some aspects of the present disclosure) with the device 101 of Figure 1C, and as an excellent microphone for the situation illustrated in Figure 1C, the proximity of the smart audio device 105 to the person 102 is opportunistically detected or evaluated. In Figure 1C, the audio 104 is coming from the relatively more distant speakerphone 101. However, the audio 103B to be sent to the phone 101 is captured by the local smart audio device 105. In some embodiments, considering the routing complexity and the capabilities of the smart audio device 105, this opportunistic usage of components of different smart audio devices enables the phone 101 and / or the application to be used without being used. Rather, in some such examples, a hierarchical system can be implemented for discovery, routing, and utilization of such capabilities.

[0116] More details about the concept of the abstract Continuous Hierarchical Audio Session Manager (CHASM) will be described later, but some of its embodiments allow an application to be given audio capabilities without the application having to know the full details of the management device, device connectivity, synchronous device usage, and / or device leveling and tuning. In a sense, this approach allows the device (having at least one speaker and at least one microphone) that enables the application to function properly to relinquish control of the audio experience. However, when the number of speakers and, importantly, microphones in a room is much larger than the number of people, it can be seen that many solutions to audio problems may involve detecting the location of the device closest to the person involved (which may not be the device normally used for such an application).

[0117] One way of thinking about audio transducers (speakers and microphones) is that they can implement one step in the route that audio takes from a person's mouth to an application and a return step in the route from the application to a person's ear. In this sense, it can be understood that by opportunistically leveraging devices and interacting with the audio subsystem on any device to output audio or obtain input audio, any application that needs to send or capture audio from a user can be improved (or at least not made worse). Such decisions and routing, in some examples, are made when the device and user move, become available, or are removed from the system. , can be made continuously. In this regard, the Continuous Hierarchical Audio Session Manager (CHASM) disclosed in this specification is useful. In some implementation examples, can the discoverable opportunistically orchestrated Distributed Audio Subsystem (DOODAD) be collectively included within CHASM, or can the DOODAD be used collectively with CHASM?

[0118] Some of the disclosed embodiments implement the concept of an integrated audio system designed to route audio to and from people and locations. This departs from the conventional "device-centric" design generally regarding the input and output from audio devices and the batch management of devices.

[0119] Next, referring to FIGS. 2A - 2D, some exemplary embodiments of the present disclosure will be described. First, a device having a communication function will be described. In this case, for example, consider a doorbell audio intercom device. When activated, the doorbell audio intercom device starts a local application that creates a full-duplex audio link from a local device to a certain remote user. The basic function of the device in this mode is to manage the speaker and microphone and relay the speaker signal and the microphone signal via a two-way network stream. This may be referred to as a "link" in this specification.

[0120] FIG. 2A is a block diagram of a conventional system. In this example, the operation of the smart audio device 200A includes an application (205A) that transmits media (212) and control information (213) to and from the media engine (201). Both the media engine 201 and the application 205A are in the smart audio device It can be implemented by a control system of 200A. According to this example, the media engine 201 has the role of managing audio inputs (203) and outputs (204), and can be configured to perform signal processing and other real-time audio tasks. Also, the application 205A can have other network input and output connectivity (210).

[0121] In FIG. 2A, the labeled elements are as follows.

[0122] ● 200A. A specific-purpose smart audio device.

[0123] ● 201. A media engine that has the role of managing real-time audio media streams input from the application 205A and performing signal processing for microphone inputs and speaker outputs. Examples of signal processing can include acoustic echo cancellation, automatic gain control, wake word detection, limiters, noise suppression, dynamic beamforming, speech recognition, encoding / decoding to a volatile format, voice activity detection, and other classifiers. In this context, the term "real-time" can refer to the need for the processing of an audio block to be completed within the time it takes to sample the audio block from, for example, an analog-to-digital converter (ADC) implemented by the device 200A. For example, in a specific implementation example, an audio block can include 480 to 960 consecutive samples sampled at 48,000 samples per second and having a length of 10 to 20 ms.

[0124] ● 203. Microphone input. Input from one or more microphones that can detect acoustic information and interface with the media engine (201) by a plurality of ADCs.

[0125] ●204. Speaker output. Input from one or more speakers that can reproduce acoustic energy and are interfaced with the media engine (201) by a plurality of digital - to - analog converters (DACs) and / or amplifiers.

[0126] ●205A. Application (an "app") operating on device 200A. The app handles media coming from and going to the network, and has the role of sending and receiving media streams to and from the media engine 201. In this example, app 205A also manages the control information transmitted and received by the media engine. Examples of app 205A include:

[0127] 〇 Control logic within a webcam that connects to the Internet and streams packets from a microphone to a web service. 〇 A conference phone that interfaces with the user via a touch screen. With the touch screen, the user can dial a phone number, browse the contact list, change the volume, and start and end a call. 〇 A voice - driven application within a smart speaker that can play music from a music library.

[0128] ●210. Optional network connection (e.g., to the Internet via WiFi or Ethernet or 4G / 5G cellular radio waves) that connects device 200A to a network. The network can carry streaming media traffic.

[0129] ●212. Media streams streamed to and from the media engine (201). For example, app 205A within a dedicated teleconference device streams media to and from the network (210) can receive Real-time Transport Protocol (RTP) packets, extract the headers, and send the G.711 payload to the media engine 201 for processing and playback. In some such examples, the app 205A can receive a G.711 stream from the media engine 201 and have the role of packing RTP packets for upstream delivery via the network (210).

[0130] ● Control signals sent to and from the app 205A for controlling the media engine (201). For example, when the user presses the volume up button on the user interface, the app 205A sends control information to the media engine 201 to amplify the playback signal (204). In the specific purpose device 200A, the control of the media engine (201) is only performed by the local app (205A) that does not have the ability to control the media engine externally.

[0131] Figure 2B shows a variant of the device shown in Figure 2A. The audio device 200B can run a second application. For example, the second application can be an application that continuously streams audio from a doorbell device for a security system. In this case, the same audio subsystem (described with reference to Figure 2A) can be used with a second application that controls a network stream in only one direction.

[0132] In this example, FIG. 2B is a block diagram of a specific purpose smart audio device (200B) that can host two applications or "apps" (app 205B and app 206) that send information to media engine 202 (the media engine of device 200B) via a specific purpose audio session manager (SPASM). When using SPASM as an interface to apps 205B and 206, network media can, in this case, flow directly to media engine 202. Here, media 210 is media to and from the first app (206), and media 211 is media for the second app (205B).

[0133] As used herein, the term SPASM (or, specific purpose audio session manager) is used to represent an (on - device) element or subsystem configured to implement an audio chain for a single type of function. The device is manufactured to provide that single type of function. The SPASM may need to be reconfigured (including, for example, by disassembling the entire audio system) to implement a change in the device's mode of operation. For example, audio in most laptops is implemented as, or uses, an SPASM. Here, the SPASM is configured (and reconfigurable) to implement any desired single - purpose audio chain for a particular function.

[0134] In FIG. 2B, the labeled elements are as follows.

[0135] ● 200B. A smart audio device having a specific purpose audio session manager (SPASM) 207B that hosts two apps (app 205B and app 206)

[0136] ● 202 - 204. Refer to FIG. 2A.

[0137] ● 205B, 206. An app that operates on the local device 200B.

[0138] ● 207B. A specific-purpose audio session manager that manages audio processing and exposes the capabilities of the media engine (202) of the device 200B. The demarcation between each app (206 or 205B) and SPASM 207B indicates how different apps desire to use different audio capabilities (e.g., an app may require a different sampling rate or a different number of inputs and outputs). All audio capabilities are exposed and managed by the SPASM for different apps. A certain limitation of the SPASM is that the SPASM is designed for a specific purpose and can only perform operations that the SPASM knows.

[0139] ● 210. Media information that is streamed to and from the network for the first app (205B). The SPASM (207B) enables the media stream to be streamed directly to the media engine (202).

[0140] ● 211. Media information that is streamed to the network for the second app (206). In this example, app 206 does not have any media streams to receive.

[0141] ● 214. Control information that is sent to and from the SPASM (207B) and the media engine (202) to manage the functions of the media engine.

[0142] ● 215, 216. Control information that is sent to and from the apps (205B, 206) and the SPASM (207B).

[0143] Including SPASM207B as a separate subsystem of the device 200B in Figure 2B may seem to implement an artificial step in the design. In fact, it does not include what may be an unnecessary operation from the perspective of a single-purpose audio device design. Most of the value of CHASM (described later) is enabled as a network effect and scales as the square of the number of nodes in the network (in a sense). However, including SPASM (e.g., SPASM207B in Figure 2B) within a smart audio device has advantages and value, including the following:

[0144] - Abstraction of control by SPASM207B makes it easier for multiple applications to operate on the same device. - SPASM207B is closely connected to the audio device and reduces the latency between audio data over the network and physical input and output sounds by directly introducing network stream connectivity to SPASM207B. For example, SPASM207B can be in a lower layer (such as a lower OSI or TCP / IP layer) in the smart audio device 200B, closer to the device driver / data link layer or in the underlying physical hardware layer. If SPASM207B were implemented in a higher layer, for example, as an application operating within the device operating system, such an implementation example may suffer a penalty in latency. This is because audio data may need to be copied from the low-level layer back up through the operating system to the application layer. As a further possible bad feature of such an implementation example, the latency may be variable or unpredictable. - This design is prepared for greater interconnectivity at a lower audio level before the application level. - This design is prepared for greater interconnectivity at a lower audio level before the application level.

[0145] In a smart audio device where an operating system operates SPASM, in some examples, many apps may obtain shared access to the speaker(s) and microphone(s) of the smart audio device. By introducing SPASM that does not need to transmit and receive audio stream(s), according to some examples, the media engine can be optimized for very low latency. This is because the media engine is separated from the control logic. A device with SPASM enables an application to establish additional media streams (e.g., media streams 210 and 211 in FIG. 2B). This advantage is due to separating the media engine function from the control logic of SPASM. This configuration is in contrast to the situations shown in FIGS. 1A and 2A. In FIGS. 1A and 2A, the media engine is dedicated to a specific application and is stand-alone. In these examples, the device was not designed to have additional low-latency connectivity to / from additional devices made possible, for example, by including SPASM as illustrated in FIG. 2B. In some such examples, if the devices as illustrated in FIGS. 1A and 2A were designed to be stand-alone, it would not be possible to easily update, for example, application 205A to provide orchestration. However, in some examples, device 200B is designed to be orchestration-ready. In some examples, many apps may obtain shared access to the speaker(s) and microphone(s) of the smart audio device. By introducing SPASM that does not need to transmit and receive audio stream(s), according to some examples, the media engine can be optimized for very low latency. This is because the media engine is separated from the control logic. A device with SPASM enables an application to establish additional media streams (e.g., media streams 210 and 211 in FIG. 2B). This advantage is due to separating the media engine function from the control logic of SPASM. This configuration is in contrast to the situations shown in FIGS. 1A and 2A. In FIGS. 1A and 2A, the media engine is dedicated to a specific application and is stand-alone. In these examples, the device was not designed to have additional low-latency connectivity to / from additional devices made possible, for example, by including SPASM as illustrated in FIG. 2B. In some such examples, if the devices as illustrated in FIGS. 1A and 2A were designed to be stand-alone, it would not be possible to easily update, for example, application 205A to provide orchestration. However, in some examples, device 200B is designed to be orchestration-ready.

[0146] Next, referring to FIG. 2C, to enable advertising and control (smart au Describe aspects of some of the disclosed embodiments of the SPASM of the diode device. With SPASM, other devices and / or systems within the network can better understand the audio capabilities of the device using the protocol, and, where applicable from a security and usability perspective, enable an audio stream to be directly connected to that device and played back on a speaker(s) or obtained from a microphone(s). In this case, it can be seen that a second application for setting up the ability to continuously stream audio from the device does not need to operate the application locally, for example, to control a media engine (e.g., 202) to stream output the monitoring stream (e.g., 211) as described above.

[0147] Figure 2C is a block diagram of an implementation example of one disclosure. In this example, at least one of two or more apps (e.g., app 205B in Figure 2C) is implemented by a device other than the smart audio device 200C, for example, by one or more servers implementing a cloud-based service, by another device in the audio environment where the smart audio device 200C exists, etc. Therefore, another controller (in this example, CHASM208C in Figure 2C) is required to manage the audio experience. In this implementation example, CHASM208C is a controller that fills the gap between the remote app(s) and the smart audio device (200C) with audio capabilities. In various embodiments, CHASM (e.g., CHASM208C in Figure 2C) can be implemented as a device, or a subsystem of a device (e.g., implemented in software). Here, CHASM or the device including it is different from one or more (e.g., many) smart audio devices. However, in some implementation examples, CHASM can be implemented via software that may be optionally executed by one or more devices in the audio environment. In some such implementation examples, CHASM can be implemented via software that may be optionally executed by one or more smart audio devices in the audio environment. In Figure 2C, CHASM208C coordinates with the SPASM of device 200C (i.e., SPASM207C in Figure 2C) to obtain access to the media engine (202) that controls the audio input (203) and output (204) of device 200C.

[0148] As used herein, the term "CHASM" is used to represent a manager (e.g., an audio session manager, e.g., the current audio session manager or a device implementing it) that can be made available to a plurality (e.g., one set) of devices (which may include, but are not limited to, smart audio devices). According to some implementations, CHASM can adjust routing and signal processing for at least one software application continuously (at least during the time when what is herein called the "root" is implemented). The application may or may not be implemented on any of the devices in the audio environment, depending on the particular implementation. In other words, CHASM can implement or be configured as an audio session manager (also called the "session manager" herein) for one or more software applications executed by one or more devices within the audio environment and / or one or more software applications executed by one or more devices outside the audio environment. The software application may sometimes be called an "app" herein.

[0149] In some cases, as a result of using CHASM, an audio device may be used for purposes not envisioned by the maker and / or manufacturer of that audio device. For example, a smart audio device (including at least one speaker and microphone) may enter a mode in which the smart audio device provides speaker feed signals and / or microphone signals to one or more other audio devices within the audio environment. This is because an app (e.g., implemented on a different device than the smart audio device) may request that CHASM (connected to the smart audio device) find and use all available speakers and / or microphones (or a selected group of available speakers and / or microphones) that may include speakers and / or microphones from multiple audio devices in the audio environment. In many such implementations, the application does not need to select the device, speaker, and / or microphone because CHASM will provide this functionality. In some cases, the application may not need to know which particular audio device is involved in implementing the commands provided to CHASM by the application (e.g., CHASM may not need to show the audio device to the application).

[0150] In Figure 2C, the labeled elements are as follows.

[0151] ● 200C. A special-purpose smart audio device that is a session manager for operating local app (206) and remote app (205B) via CHASM (208C).

[0152] ● 202 - 204. Refer to Figure 2A.

[0153] ●205B. For example, an app that operates remotely from device 200C on a server configured such that CHASM 208C communicates via the Internet, or on another device in the audio environment (a device different from device 200C, e.g., another smart audio device such as a mobile phone). In some examples, app 205B may be implemented on a first device and CHASM 208C may be implemented on a second device. Here, both the first device and the second device are different from device 200C.

[0154] ●206. An app that operates locally on device 200C.

[0155] ●207C. A SPASM that can manage control inputs from CHASM 208C in addition to interfacing with the media engine (202).

[0156] ●208C. A Continuous Hierarchical Audio Session Manager (CHASM) that enables an app (205B) to utilize the audio capabilities, inputs (203), and outputs (204) of the media engine (202) of device 200C. In this example, CHASM 208C is configured to do so via SPASM (207C) by obtaining at least partial control of the media engine (202) from SPASM 207C.

[0157] ●210 - 211. Refer to FIG. 2C.

[0158] ●217. Control information transmitted to and from SPASM (207B) and the media engine (202) to manage the functions of the media engine.

[0159] ●218. Control information transmitted to and from local app 206 and SPASM (207C) to implement local app 206. In some implementation examples, such control information may conform to an orchestration language as disclosed herein.

[0160] ● 219. Control information to and from CHASM (208C) and SPASM (207C) for controlling the functions of the media engine 202. Such control information may, in some cases, be the same as, or similar to, the control information 217. However, in some implementation examples, the control information 219 may have a lower level of detail. This is because, in some examples, device-specific details may be delegated to the SPASM 207C.

[0161] ● 220. Control information between the application (205B) and CHASM (208C). In some examples, this control information may be represented by what is referred to herein as the orchestration language.

[0162] The control information 217 may include, for example, a control signal from the SPASM207C to the media engine 202. This control signal has the effect of adjusting the output level of the output loudspeaker feed(s), for example, gain adjustment in decibels, or a linear scalar value. Such as changing the equalization curve(s) applied to the output loudspeaker feed(s). In some examples, the control information 217 from the SPASM207C to the media engine 202 may include a control signal having the effect of changing the equalization curve(s) applied to the output loudspeaker feed(s) by, for example, parametrically described (as a series combination of basic filter stages) or as a table listing gain values at specific frequencies to give a new equalization curve. In some examples, the control information 217 from the SPASM207C to the media engine 202 may include a control signal having the effect of changing the upmixing or downmixing process that renders multiple audio source feeds to the output loudspeaker feed by, for example, giving a mixing matrix used to combine the source feed to the loudspeaker feed. In some examples, the control information 217 from the SPASM207C to the media engine 202 may include a control signal having the effect of changing the dynamics processing applied to the output loudspeaker feed(s), such as changing the dynamic range of the audio content.

[0163] In some examples, the control information 217 from the SPASM207C to the media engine 202 may indicate a change to a set of media streams provided to the media engine. In some examples, the control information 217 from the SPASM207C to the media engine 202 may indicate the need to establish or terminate a media stream using another media engine or another source of media content (e.g., a cloud-based streaming service).

[0164] In some cases, the control information 217 may include a control signal from the media engine 202 to the SPASM 207C, such as wake word detection information. Such wake word detection information may, in some cases, include a wake word confidence value, or a message indicating that a probable wake word has been detected. In some examples, the wake word confidence value may be transmitted once per period (e.g., once every 100 ms, once every 150 ms, once every 200 ms, etc.).

[0165] In some cases, the control information 217 from the media engine 202 may include a voice recognition phone probability that enables the SPASM, CHASM, or another device (e.g., a cloud-based service device) to perform decoding (e.g., Viterbi decoding) to determine which command is being issued. In some cases, the control information 217 from the media engine 202 may include SPL information from a sound pressure level (SPL) meter. According to some such examples, the SPL information may be transmitted at time intervals such as once per second, once every half second, once every N seconds, or once every N milliseconds. In some such examples, the CHASM may be configured to determine, for example, whether the devices are in the same room and / or whether the devices are detecting the same sound, i.e., whether there is a correlation between the measurements of the SPL meters across multiple devices.

[0166] According to some examples, the control information 217 from the media engine 202 may include information obtained from a microphone feed that exists as a media stream available to the media engine, such as, for example, an estimate of background noise, an estimate of direction of arrival (DOA) information, an indication of voice presence via voice activity detection, the current echo cancellation ability, etc. In some such examples, the DOA information may perform an acoustic mapping of the audio devices in the audio environment, and in some such examples, may be provided to an upstream CHASM (or another device) configured to create an acoustic map of the audio environment. In some such examples, the DOA information may be associated with a wake word detection event. In some such implementations, the DOA information may be provided to an upstream CHASM (or another device) configured to perform an acoustic mapping to identify the location of the user who uttered the wake word.

[0167] In some examples, the control information 217 from the media engine 202 may include status information such as, for example, information regarding which active media streams are available, the time position within a linear-time media stream (e.g., a TV program, a movie, a streaming video), information related to current network capabilities such as the latency associated with an active media stream, reliability information (e.g., packet loss statistics), etc.

[0168] The design of the device 200C in FIG. 2C can be extended in various ways according to aspects of the present disclosure. It can be seen that the function of the SPASM 207C in FIG. 2C is to implement a set of hooks or functions for controlling the local media engine 202. Therefore, device 200C can be considered to be closer to a media engine (e.g., closer to the functionality of device 200A or device 200B) that connects an audio device, performs network streaming, signal processing, and responds to configuration commands from both the local application(s) of device 200C and also the audio session manager. In this case, it is important for the device to have information about itself (e.g., stored in a memory device and available to the audio session manager) to assist the audio session manager. A simple example of this information includes the number of speakers, the capabilities of the speaker(s), dynamic processing information, the number of microphones, microphone placement, and sensitivity information, etc.

[0169] FIG. 2D illustrates an example of multiple applications interacting with CHASM. In the example illustrated in FIG. 2D, all applications, including an app local to the smart audio device (e.g., app 206 stored in the memory of smart audio device 200D), need to interface with CHASM 208D to provide functions involving the smart audio device 200D. In this example, since CHASM 208D inherits the interface, the smart audio device 200D only needs to notify its property 207D or make the property available to CHASM 208D, and SPASM becomes unnecessary. Thereby, CHASM 208D serves as the main controller for orchestrating experiences such as audio sessions for the applications.

[0170] In FIG. 2D, the labeled elements are as follows.

[0171] ● 200D. A smart audio device implementing local app 206. CHASM 208B operates the local app (206) and the remote app (205B).

[0172] ●Refer to FIGS. 2A (202 - 204).

[0173] ●205B. An application that operates remotely from the smart audio device 200D (in other words, on a device separate from the smart audio device 200D) (and, in this example, also operates remotely from the CHASM 208D). In some examples, the application 205B can be executed by a device configured such that the CHASM 208D communicates via, for example, the Internet or a local network. In some examples, the application 205B can be stored on another device in the audio environment, such as another smart audio device. According to some implementations, the application 205B can be stored on a mobile device that can be moved into or out of the audio environment, such as a mobile phone.

[0174] ●206. An application that operates locally on the smart audio device 200D. However, control information (223) is sent to and / or from the CHASM 208D.

[0175] ●207D. Property descriptor. In a state where the CHASM 208D is in charge of managing the media engine 202, the smart audio device 200D can substitute SPASM instead of a simple property descriptor. In this example, the descriptor 207D indicates to the CHASM the capabilities of the media engine 202, such as the number of inputs and outputs, possible sample rates, and signal processing components. In some examples, the descriptor 207D is, for example, the type, size of one or more loudspeakers, and Data indicating numbers, data corresponding to the capabilities of one or more loudspeakers, data regarding dynamic processing that the media engine 202 is to apply to the audio data before the audio data is reproduced by one or more loudspeakers, etc., can indicate data corresponding to one or more loudspeakers of the smart audio device 200D. In some examples, the descriptor 207D indicates whether the media engine 202 (or, more generally, the control system of the smart audio device 200D) is configured to provide functions related to the coordination of audio devices in the environment, such as rendering the audio data to be reproduced by the smart audio device 200D and / or other audio devices in the audio environment, whether the device that currently provides the functions of the CHASM 208D (e.g., in accordance with the CHASM software stored in the device's memory) has been turned off or, otherwise, if the functions have been stopped, whether the control system of the smart audio device 200D can implement the functions of the CHASM 208D, etc.

[0176] ●208D. In this example, CHASM208D functions as a gateway for all apps (regardless of whether they are local or remote) to interact with the media engine 202. Even local apps (e.g., app 206) obtain access to the local media engine 202 via CHASM208D. In some cases, CHASM208D can be implemented only within a single device of the audio environment, for example, via CHASM software stored in a wireless router, smart speaker, etc. However, in some implementation examples, multiple devices of the audio environment can be configured to implement at least some aspects of the CHASM function. In some examples, the control system of one or more other devices within the audio environment, such as one or more smart audio devices of the audio environment, can implement the function of CHASM208D when the device currently providing the function of CHASM208D is turned off or otherwise stops functioning.

[0177] ●210 - 211. Refer to FIG. 2C.

[0178] ●221. Control information transmitted (e.g., to and from) between CHASM208D and the media engine 202.

[0179] ●222. Data transmitted from the property descriptor 207D to CHASM to indicate the capabilities of the media engine 202.

[0180] ●223. Control information transmitted between the local app 206 and CHASM208.

[0181] ●224. Control information transmitted between the remote app 205B and CHASM208D.

[0182] Next, further embodiments will be described. To implement some such embodiments, first, a single device (such as a communication device) is designed and coded for a specific purpose. An example of such a device is the smart audio device 101 of FIG. 1C. The smart audio device 101 can be implemented as illustrated in FIG. 3C. As background, implementation examples of the device 101 of FIGS. 1A and 1B are also described.

[0183] FIG. 3A is a block diagram illustrating details of the device 101 of FIG. 1A according to an example. The user's voice 103 is captured by the microphone 303, and the local app 308A manages the network stream 317 received via the network interface of the device 101, manages the media stream 341, and has the role of providing the control signal 3 40 to the media engine 301A.

[0184] In FIG. 3A, the labeled elements are as follows. 101, 103 - 104. Refer to FIG. 1A. 301A. A media engine having the role of managing the real - time audio media stream input from the app 308A. 303. Microphone. 304. Loudspeaker. 308A. Local app. 317. Media stream to and from the network. 340. Control information transmitted to and from the app 308A and the media engine 301A. 341. Media stream transmitted to and from the app 308A.

[0185] Figure 3B illustrates the details of the embodiment of FIG. 1B according to an example. In this example, implementation examples of the device 105 of FIG. 1B and the device 101 of FIG. 1B are shown. In this case, both devices are designed with the aim of being "orchestration compliant" in the sense that there is a general-purpose or flexible media engine controlled via the abstraction of SPASM.

[0186] In FIG. 3B, the output 106 of the second device 105 is not related to the output 104 of the first device 101, and the input 103 to the microphone 303 of the device 101 may be able to capture the output 106 of the device 105. In this example, the devices 105 and 101 may not function in an orchestrated manner.

[0187] In FIG. 3B, the labeled elements are as follows. 101, 103 - 106. Refer to FIG. 1B. 301, 303 - 304. Refer to FIG. 3A. 302. The media engine of device 105. 305. The microphone of device 105. 306. The loudspeaker of device 105. 308B. The local app of device 101. 312B. SPASM for device 101. 314B. SPASM for device 105. 317. Media stream to and from the network. 320. The local app for device 105. 321. Control information transmitted between app 308B and SPASM 312B. 322. Control information transmitted between SPASM 312B and media engine 301. 323. Control information transmitted (to and from) between app 320 and SPASM 314B. 324. Control information transmitted (to and from) between SPASM 314B and media engine 302. 325. Media stream from the network into the media engine 302.

[0188] Figure 3C is a block diagram illustrating an example of CHASM that orchestrates two audio devices in an audio environment. Based on the above discussion, it should be understood that the situation in which the system of FIG. 3B is used while CHASM operates on a variant of device 101 and device 105 or on another device will be better managed by CHASM (for example, similar to the embodiment of FIG. 3C). When using CHASM 307 in some examples, the application 308 on the phone 101 stops directly controlling its audio device 101 and delegates all audio control to CHASM 307. According to some such examples, the signal from the microphone 305 may include less echo than the signal from the microphone 303. In some examples, CHASM 307 may infer, based on the presumption that the microphone 305 is closer to the person 102, that the signal from the microphone 305 includes less echo than the signal from the microphone 303. In some such examples, CHASM may utilize routing the raw or processed microphone signal from the microphone 305 on the device 105 to the phone device 101 as a network stream. These microphone signals may be used preferentially over the signal from the local microphone 303 to achieve a better voice communication experience for the person 102.

[0189] According to some implementation examples, CHASM307 can ensure that this continues to be the best configuration, for example, by monitoring the position of person 102, monitoring the position of device 101 and / or device 105, etc. In some examples, CHASM307 can ensure that this continues to be the best configuration. According to some such examples, CHASM307 can ensure that this continues to be the best configuration through the exchange of low-rate (e.g., low bitrate) data and / or metadata. In the state where only a small amount of information is shared between devices, for example, the position of person 102 can be tracked. When information is exchanged between devices at a low bitrate, considering the limited bandwidth may not be much of a problem. Examples of low bitrate information that can be exchanged between devices are, for example, microphones described with reference to the "follow me" implementation example Including, but not limited to, information obtained from a microphone signal. An example of low-bitrate information that may be useful in determining which device's microphone has a higher voice-to-echo ratio is, during a period, e.g., during the last 1 second, an estimated value of the SPL caused by the sound emitted by the local loudspeaker on each of a plurality of audio devices in the audio environment. An audio device that radiates more energy from the loudspeaker(s) may capture less of the other sounds in the audio environment that exceed the echo caused by the loudspeaker(s). Another example of low-bitrate information that may be useful in determining which device's microphone has a higher voice-to-echo ratio is the amount of energy in the echo prediction of the acoustic echo canceller of each device. A high amount of predicted echo energy indicates that the microphone(s) of the audio device may be overwhelmed by the echo. In some such examples, there may be some echoes that the acoustic echo canceller will not be able to cancel (assuming that the acoustic echo canceller has already converged at this time). In some examples, CHASM307 may continue to be prepared to control device 101 to resume using the microphone signal from the local microphone 303 if information is provided somehow that there is a problem with microphone 305 or that microphone 305 does not exist.

[0190] Figure 3C illustrates an example of the underlying system used to coordinate app 308 operating on device 101 with respect to device 105. In this example, CHASM307 causes the user's voice 103B to be captured within microphone 305 of device 105 and causes the captured audio 316 to be used in the media engine 301 of device 101, while the loudspeaker output 104 comes from the first device 101. Thereby, the experience is orchestrated for user 102 across devices 101 and 105.

[0191] In FIG. 3C, the labeled elements are as follows. 101, 103B, 104, 105. Refer to FIG. 1C. 301 to 306. Refer to FIG. 3B. 307. CHASM. 309. Control information between CHASM 307 and media engine 302. 310. Control information between CHASM 307 and media engine 301. 311. Control information between (to and from) CHASM 307 and application 308. 312C. Device property descriptor of device 101. 313. Control information and / or data from device property descriptor 312C to CHASM 307. 314C. Device property descriptor of device 105. 315. Control information and / or data from device property descriptor 314C to CHASM 307. 316. Media stream from device media engine 302 to device media engine 301. 317. Media stream to and from the network to media engine 301.

[0192] In some embodiments, when a DOODAD is included in a CHASM (e.g., CHASM 307) to interact with smart audio devices (e.g., to send and receive control information to and from each of the respective smart audio devices), and / or when a DOODAD is provided to operate with a CHASM (e.g., CHASM 401 of FIG. 4 described below, or CHASM 307) (e.g., as a subsystem of devices such as devices 101 and 105, where these devices are separate from the CHASM-implementing devices such as the device implementing CHASM 307), the need for a SPASM (in each of one or more smart audio devices) is replaced by the operation of the smart audio device that notifies (the CHASM) of the audio capabilities and an application that delegates to a single abstract control point (e.g., the CHASM) for the audio functions.

[0193] FIG. 4 is a block diagram illustrating another disclosed embodiment. The design of FIG. 4 introduces important abstractions. The important abstractions are that in some implementation examples, the application does not need to directly select or control, and in some cases, the application may not be given information regarding which particular audio device is involved in performing functions related to the application, the specific capabilities of such audio devices, etc.

[0194] FIG. 4 is a block diagram of a system including three separate physical audio devices 420, 421, and 422. In this example, each device implements a discoverable opportunistically orchestrated distributed audio subsystem (DOODAD) and is controlled by CHASM 401 that operates applications (410 - 412). According to this example, CHASM 401 is configured to manage media requests for each of applications 410 - 412.

[0195] In FIG. 4, the labeled elements are as follows: 400. An example of an audio system orchestrated across three different devices. Devices 420, 421, and 422 each implement a DOODAD (DUDAD). Each of devices 420, 421, and 422 that implement a DOODAD may itself be referred to as a DOODAD. In this example, each DOODAD is different from the smart audio device 101 of FIG. 3C (or implemented thereby) in that it is a DOODAD in FIG. 4 or the device implementing it does not implement the relevant application, but device 101 implements application 308. 401. CHASM 410 - 412. In this example, applications with different audio requests. In some examples, each of applications 410 - 412 may be stored on, or executable by, a device in the audio environment or a mobile device that may be placed in the audio environment. 420 - 422. Smart audio devices implementing a discoverable opportunistically orchestrated distributed audio subsystem (DOODAD), with each DOODAD operating within a separate physical smart audio device 430, 431, and 432. Control information sent between (to and from) applications 410, 411, and 412 and to CHASM 401 433, 434, and 435. Control information sent to and from DOODADs 420 - 422 and CHASM 401 440, 441, and 442. Media engines 450, 451, and 452. Device property descriptors. In this example, DOODADs 420 - 422 are configured to provide these device property descriptors to CHASM 401 460 - 462. Loudspeakers 463 - 465. Microphones 470. A media stream that exits the media engine 442 and heads towards the network, 471. A media stream from the media engine 441 to the media engine 442, 472. A media stream from the media engine 441 to the media engine (440), 473. A media stream from a cloud - based service for providing media over a network, such as a cloud - based service provided to the media engine 441 via the Internet by one or more servers in a data center, 474. A media stream from the media engine 441 to a cloud - based service, 477. One or more cloud - based services that may include one or more music streaming services, movie streaming services, television show streaming services, podcast providers, etc.

[0196] Figure 4 is a block diagram of a system in which multiple audio devices may perform routing to create an audio experience. According to this example, the media engine 441 of the smart audio device 421 that provides the media stream 471 to the media engine 442 and the media stream 472 to the media engine 440 has received audio data corresponding to the media stream 473. According to some such implementation examples, the media engine 441 may process the media stream 473 according to the sample clock of the media engine 441. Such a sample clock is an example of what may be referred to herein as the "media engine sample clock".

[0197] In some such examples, CHASM401 may provide instructions and information to the media engine 441 via the control information 434 regarding the processing related to the acquisition and processing of the media stream 473. Such instructions and information are as follows in this specification This is an example of what can be referred to as an "audio session management control signal".

[0198] However, in some implementation examples, CHASM401 can send an audio session management control signal without referring to the media engine sample clock of media engine 441. Such examples can be advantageous, for example, because CHASM401 does not need to synchronize the transmission to the audio device in the audio environment of the media. Instead, in some implementation examples, any such synchronization can be delegated to another device such as smart audio device 421 in the above example.

[0199] According to some such implementation examples, CHASM401 may provide to media engine 441 an audio session management control signal regarding obtaining and processing media stream 473 in response to control information 430, 431, or 432 from application 410, application 411, or application 412. Such control information is an example of what can be referred to as an "application control signal" in this specification. According to some implementation examples, the application control signal can be sent from the application to CHASM401 without referring to the media engine sample clock of media engine 441.

[0200] In some examples, CHASM401 may provide audio processing information to media engine 441, along with instructions for processing audio corresponding to the processed media stream 473 accordingly. The audio processing information includes, but is not limited to, rendering information. However, in some implementations, a device implementing CHASM401 (or a device implementing a similar function such as the functions of the smart home hub described elsewhere in this specification) may be configured to provide at least some audio processing functions. Some examples are given below. In some such implementations, CHASM401 may be configured to receive and process audio data and provide the processed (e.g., rendered) audio data to an audio device in an audio environment.

[0201] Figure 5 is a flow diagram including blocks of an audio session management method according to some implementations. The blocks of method 500 are not necessarily performed in the order described, similar to other methods described herein. In some implementations, one or more of the blocks of method 500 may be performed simultaneously. For example, in some cases, blocks 505 and 510 may be performed simultaneously. Further, some implementations of method 500 may include more or fewer blocks than those illustrated and / or described. The blocks of method 500 may be performed by one or more devices. Such a device may be (or may include) a control system such as the control system 610 described below and illustrated in Figure 6 or one of the other disclosed control system examples.

[0202] According to some implementation examples, the blocks of method 500 can be performed, at least in part, by a device that implements what is referred to herein as an audio session manager, such as CHASM. In some such examples, the blocks of method 500 can be performed, at least in part, by CHASM208C, CHASM208D, CHASM307, and / or CHASM401 as described above with reference to FIGS. 2C, 2D, 3C, and 4. More specifically, in some implementation examples, the functionality of the "audio session manager" called in the blocks of method 500 can be performed, at least in part, by CHASM208C, CHASM208D, CHASM307, and / or CHASM401.

[0203] According to this example, block 505 is a first app that executes a first application Including establishing a first application communication link between a cation device and an audio session manager of an audio environment. In some examples, the first application communication link can be generated via any suitable wireless communication protocol suitable for use within the audio environment. Such wireless communication protocols include Zigbee, Apple's Bonjour (Rendezvous), WiFi, Bluetooth, Bluetooth Low Energy (Bluetooth LE), 5G, 4G, 3G, General Packet Radio Service (GPRS), Amazon Sidewalk, custom protocols within Nordic's RF24L01 chip, and the like. In some examples, the first application communication link can be established in response to a "handshake" process. The "handshake" process can be initiated, in some examples, via a "handshake start" sent by the first application device to a device implementing the audio session manager. In some examples, the first application communication link can be established in response to what may be referred to herein as a "route start request" from the first application device. For convenience, the route start request from the first application device may be referred to herein as the "first route start request" to indicate that the route start request corresponds to the "first application device". In other words, the term "first" may or may not have a temporal meaning in this context depending on the particular implementation example.

[0204] In one such example, the first application communication link can be established between the device on which application 410 of FIG. 4 is running and CHASM 401. In some such examples, the first application communication link can be established in response to CHASM 401 receiving a first route start request from the device on which application 410 is running. The device on which application 410 is running can be, for example, an audio device in a smart audio environment. In some cases, the device on which application 410 is running can be a mobile phone. Application 410 can be used to access media such as music, TV shows, movies, etc. via CHASM 401. In some cases, the media is available for streaming via a cloud-based service.

[0205] Various examples of what is meant by "route" are described in detail below. Generally, a route indicates parameters of an audio session that will be managed by an audio session manager. A route start request can indicate, for example, an audio source and an audio environment destination. The audio environment destination can, in some cases, correspond to at least one person within the audio environment. In some cases, the audio environment destination can correspond to an area or zone of the audio environment.

[0206] However, in most cases, the audio environment destination will not indicate any specific audio device that will be involved in playing the media in the audio environment. Instead, an application (such as application 410) may provide, for example, a route start request stating that a particular type of media should be made available to a particular person within the audio environment. In various disclosed implementations, the audio session manager will have the role of determining, for example, which audio devices will be involved in obtaining, rendering, and playing the audio data related to the media, such as determining which audio devices will be involved in the route. In some implementations, the audio session manager will determine whether the audio devices involved in the route have changed (e.g., in response to a determination that the person who is the intended recipient of the media has changed location). and will have roles such as updating the corresponding data structure. A detailed example will be described below.

[0207] In this example, block 510 includes receiving, by the audio session manager, a first application control signal from a first application via a first application communication link. Referring back to FIG. 4. In some examples, the application control signal may correspond to control information 430 transmitted (to and from) between app 410 and CHASM 401. In some examples, the first application control signal may be transmitted after the audio session manager (e.g., CHASM 401) has started the route. However, in some cases, the first application control signal may correspond to a first route start request. In some such examples, blocks 505 and 510 may occur at least partially simultaneously.

[0208] According to this example, block 515 includes establishing a first smart audio device communication link between an audio session manager and at least a first audio device of a smart audio environment. In this example, the first smart audio device may be either a single-purpose audio device or a multi-purpose audio device, or may include any of them. According to this implementation example, the first smart audio device includes one or more loudspeakers.

[0209] In some examples, as described above, the first application control signal and / or the first route start request do not indicate any particular audio device that will be involved in the route. According to some such examples, method 500 may include processing before block 515 to determine which audio device of the audio environment will be involved in the route at least initially (e.g., by the audio session manager).

[0210] For example, CHASM401 of FIG. 4 may determine that audio devices 420, 421, and 422 will be involved in the route at least initially. In the example illustrated in FIG. 4, the first smart audio device communication link of block 515 may be established between an audio session manager (CHASM401 in this example) and smart audio device 421. The first smart audio device communication link may correspond to the dashed line shown in FIG. 4 between CHASM401 and smart audio device 421. Control information 434 is transmitted via the first smart audio device communication link. In some such examples, the first smart audio device communication link may be generated via any suitable wireless communication protocol appropriate for use within the audio environment. Such wireless communication protocols include Apple Airplay, Miracast, Blackfire, Bluetooth 5, Real-time Transport Protocol (RTP), and the like.

[0211] In the example shown in FIG. 5, block 520 includes determining, by an audio session manager, one or more first media engine capabilities of a first media engine of a first smart audio device. According to this example, the first media engine is configured to manage one or more audio media streams received by the first smart audio device and perform first smart audio device signal processing on the one or more audio media streams according to a first media engine sample clock. In the above example, block 520 may include, for example, providing device property descriptor 451 to CHASM401 so that CHASM401 receives information regarding one or more capabilities of media engine 441 from smart audio device 421. According to some implementations, block 520 may include CHASM401 receiving information regarding the capabilities of one or more loudspeakers of smart audio device 421. In some examples, CHASM401 may have determined some or all of this information in advance, for example, prior to blocks 505, 510, and / or 515 of method 500.

[0212] According to this example, block 525 includes controlling the first smart audio device according to the first media engine capabilities via a first audio session management control signal transmitted by the audio session manager to the first smart audio device via a first smart audio device communication link. According to some examples, the first audio session management control signal may cause the first smart audio device to delegate control of the first media engine to the audio session manager. In this example, the audio session manager transmits the first audio session management control signal to the first smart audio device without referring to the first media engine sample clock. In some such examples, the first application control signal may be transmitted from the first application to the audio session manager without referring to the first media engine sample clock.

[0213] In one example of block 525, CHASM401 may control media engine 441 to receive media stream 473. In some such examples, via the first audio session management control signal, CHASM401 may provide media engine 441 with a Universal Resource Locator (URL) corresponding to the website from which media stream 473 may be received, along with an instruction to start media stream 473. According to some such examples, CHASM401 may also, via the first audio session management control signal, provide media engine 441 with instructions to provide media stream 471 to media engine 442 and media stream 472 to media engine 440.

[0214] In some such examples, CHASM401 may provide to media engine 441, via a first audio session management control signal, audio processing information (including but not limited to rendering information), along with instructions to process the audio corresponding to media stream 473 accordingly. For example, CHASM401 may provide to media engine 441, for example, an indication that smart audio device 420 will receive a speaker feed signal corresponding to the left channel, that smart audio device 421 will play a speaker feed signal corresponding to the center channel, and that smart audio device 422 will receive a speaker feed signal corresponding to the right channel.

[0215] Various other examples of rendering are disclosed herein. Some of these may be CHASM401 or another audio session manager that conveys different types of audio processing information to the smart audio devices. For example, in some implementations, one or more devices in the audio environment may be configured to implement flexible rendering such as Center of Mass Amplitude Panning (CMAP) and / or Flexible Virtualization (FV). In some such implementations, a device configured to implement flexible rendering may be provided with the position of a set of audio devices, the estimated current listener position, and the estimated current listener orientation. A device configured to implement flexible rendering may be configured to render audio for a set of audio devices in the environment according to the position of the set of audio devices, the estimated current listener position, and the estimated current listener orientation. Some detailed examples are described below.

[0216] In the above example of method 500 described with reference to FIG. 4, a device other than the audio session manager or the first smart audio device is configured to execute the first application. However, in some examples, as described above with reference to FIGS. 2C and 2D, the first smart audio device may be configured to execute the first application.

[0217] According to some such examples, for example, as described above with reference to FIG. 2C, the first smart audio device may include a specific-purpose audio session manager. In some such implementations, the audio session manager may communicate with the specific-purpose audio session manager via the first smart audio device communication link. In some examples, the audio session manager may obtain one or more first media engine capabilities from the specific-purpose audio session manager. According to some such examples, the audio session manager may function as a gateway for all applications that control the first media engine, regardless of whether the application operates on the first smart audio device or on another device.

[0218] As described above, in some examples, method 500 may include establishing at least a first audio stream (e.g., media stream 473 of FIG. 4) corresponding to a first audio source. In some examples, the first audio source may be one or more servers configured to provide a cloud-based media streaming service such as a music streaming service, a television show and / or a movie streaming service. The first audio stream may include a first audio signal. In some such implementations, establishing at least the first audio stream may include causing the first smart audio device to establish at least the first audio stream via a first audio session management control signal transmitted to the first smart audio device via a first smart audio device communication link.

[0219] In some examples, method 500 may include a rendering process that causes the first audio signal to be rendered to a first rendered audio signal. In some such implementations, the rendering process may be performed by the first smart audio device in response to a first audio session management control signal. In the above example, media engine 441 may render the audio signal corresponding to media stream 473 to a speaker feed signal in response to a first audio session management control signal.

[0220] According to some examples, method 500 may include causing a first smart audio device to establish a smart audio device - to - device communication link between the first smart audio device and each of one or more other smart audio devices in an audio environment via a first audio session management control signal. In the example described above with reference to FIG. 4, media engine 441 may establish a wired or wireless smart audio device - to - device communication link with media engines 440 and 442. In the example described above with reference to FIG. 3C, media engine 302 may establish a wired or wireless smart audio device - to - device communication link to provide media stream 316 to media engine 301.

[0221] In some examples, method 500 may include causing a first smart audio device to transmit, via a smart audio device - to - device communication link, one or more raw microphone signals, processed microphone signals, rendered audio signals, or unrendered audio signals to one or more other smart audio devices. In the example described above with reference to FIG. 4, a smart audio device - to - device communication link may be used to provide a rendered audio signal or an unrendered audio signal via media streams 471 and 472. In some such examples, media stream 471 may include a speaker feed signal for media engine 442, and media stream 472 may include a speaker feed signal for media engine 440. In the example described above with reference to FIG. 3C, via media stream 316, media engine 302 may provide a raw microphone signal or a processed microphone signal to media engine 301.

[0222] ​According to some examples, method 500 may include establishing a second smart audio device communication link between an audio session manager and at least a second audio device of the smart audio environment. In some such examples, the second smart audio device may be a single-purpose audio device or a multi-purpose audio device. In some cases, the second smart audio device may include one or more microphones. Some such methods may include determining, by the audio session manager, one or more second media engine capabilities of a second media engine of the second smart audio device. The second media engine may be configured to receive microphone data, for example, from one or more microphones and perform second smart audio device signal processing on the microphone data.

[0223] For example, referring to FIG. 3C, the “first smart audio device” may be the smart audio device 101. According to some such examples, the “second smart audio device” may be the smart audio device 105. Control signal 310 may be provided using the “first smart audio device communication link” and control signal 309 may be provided using the “second smart audio device communication link”. CHASM 307 may determine one or more media engine capabilities of the media engine 302 based at least in part on the device property descriptor 314c.

[0224] Some such methods may include controlling a second smart audio device via a second audio session manager control signal transmitted to the second smart audio device via a second smart audio device communication link according to the second media engine capabilities. In some cases, controlling the second smart audio device may include causing the second smart audio device to establish a smart audio device-to-device communication link between the second smart audio device and the first smart audio device (e.g., the smart audio device-to-device communication link used to provide media stream 316). Some such examples may include causing the second smart audio device to transmit at least one of processed or unprocessed microphone data (e.g., processed or unprocessed microphone data from microphone 305) from the second media engine to the first media engine via the smart audio device-to-device communication link.

[0225] In some examples, controlling the second smart audio device may include receiving, by the audio session manager, a first application control signal from the first application via a first application communication link. FIG. In the 3C example, CHASM307 receives a control signal 311 from application 308, which is in this case a telephone application. Some such examples may include determining a second audio session manager control signal according to a first application control signal. For example, referring again to FIG. 3C, CHASM307 may be configured to optimize the speech-to-echo ratio (SER) for a conference call provided according to control signal 311 from application 308. CHASM307 may determine that it can improve the SER for a remote conference by using microphone 305 instead of microphone 303 to capture the speech of person 102 (see FIG. 1C). This determination may, in some examples, be based on an estimated value of the position of person 102. Some detailed examples of estimating the position and / or orientation of a person within an audio environment are disclosed herein.

[0226] FIG. 6 is a block diagram illustrating an example of components of an apparatus that can implement various aspects of the present disclosure. Similar to other figures provided herein, the types and numbers of elements illustrated in FIG. 6 are provided by way of example only. Other implementations may include more, fewer, and / or different types and numbers of elements. According to some examples, apparatus 600 may be, or may include, a smart audio device configured to perform at least some of the methods disclosed herein. In other implementations, apparatus 600 may be, or may include, another device configured to perform at least some of the methods disclosed herein, such as a laptop computer, a mobile phone, a tablet device, a smart home hub, etc. In some such implementations, apparatus 600 may be, or may include, a server. In some such implementations, apparatus 600 may be configured to implement what may be referred to herein as CHASM.

[0227] In this example, device 600 includes interface system 605 and control system 610. Interface system 605 can be configured to communicate with one or more devices that are running or configured to run a software application in some implementation examples. Such a software application may also be referred to herein as an "application" or simply an "app". Interface system 605 can be configured to exchange control information and related data related to the application in some implementation examples. Interface system 605 can be configured to communicate with one or more other devices in an audio environment in some implementation examples. The audio environment can be a home audio environment in some examples. Interface system 605 can be configured to exchange control information and related data with audio devices in the audio environment in some implementation examples. The control information and related data relate to one or more applications that device 600 is configured to communicate with in some examples.

[0228] Interface system 605 can be configured to receive audio data in some implementation examples. The audio data can include an audio signal that is intended to be reproduced by at least some speakers in the audio environment. The audio data can include one or more audio signals and associated spatial data. The spatial data can include, for example, channel data and / or spatial metadata. Interface system 605 can be configured to provide the rendered audio signal to at least some of a set of loudspeakers in the environment. Interface system 605 can be configured to receive input from one or more microphones within the environment in some implementation examples.

[0229] Interface system 605 includes one or more network interfaces and / or or may include one or more external device interfaces (such as one or more Universal Serial Bus (USB) interfaces). According to some implementation examples, the interface system 605 may include one or more wireless interfaces. The interface system 605 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 605 may include one or more interfaces between the control system 610 and a memory system, such as the optional memory system 615 illustrated in FIG. 6. However, the control system 610 may include a memory system in some cases.

[0230] The control system 610 may include, for example, a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gates or transistor logic, and / or discrete hardware components.

[0231] In some implementation examples, the control system 610 may be present within multiple devices. For example, a portion of the control system 610 may be present within a device within one of the environments shown herein, and another portion of the control system 610 may be present within a device outside of an environment such as a server, a mobile device (e.g., a smartphone or a tablet computer). In other examples, a portion of the control system 610 may be present within a device within one of the environments shown herein, and another portion of the control system 610 may be present within one or more other devices of the environment. For example, the control system functionality may be distributed across multiple smart audio devices of the environment, or may be shared by an orchestrating device (such as what may be referred to herein as a smart home hub) and one or more other devices of the environment. The interface system 605 may also, in some such examples, be present within multiple devices.

[0232] In some implementation examples, the control system 610 may be configured, at least in part, to perform the methods disclosed herein. According to some examples, the control system 610 may be configured to implement an audio session management method.

[0233] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described herein. Such memory devices may include, but are not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, and the like. One or more non-transitory media may be present, for example, within the optional memory system 615 and / or control system 610 illustrated in FIG. 6. Accordingly, various innovative aspects of the subject matter described in this disclosure can be implemented within one or more non-transitory media having stored software. The software may include instructions for controlling at least one device to implement, for example, an audio session management method. In some examples, the software may include instructions for controlling audio devices in one or more audio environments to acquire, process, and / or provide audio data. The software may be executable by one or more components of a control system, such as the control system 610 of FIG. 6, for example.

[0234] In some examples, the apparatus 600 may include the optional microphone system 620 illustrated in FIG. 6. The optional microphone system 620 may include one or more microphones. In some implementations, one or more of the microphones may be part of or associated with another device, such as a speaker of a speaker system, a smart audio device, etc. In some examples, the apparatus 600 may not include the microphone system 620. However, in some such implementations, the apparatus 600 may be configured to receive microphone data for one or more microphones within an audio environment via the interface system 610.

[0235] According to some implementation examples, the apparatus 600 may include the optional loudspeaker system 625 illustrated in FIG. 6. The optional speaker system 625 may include one or more loudspeakers. The loudspeakers may also be referred to as "speakers" herein. In some examples, at least some of the loudspeakers of the optional loudspeaker system 625 may be optionally positioned. For example, at least some of the speakers of the optional loudspeaker system 625 may be arranged at positions that do not conform to any standard predetermined speaker layout, such as Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4, Dolby 9.1, Hamasaki 22.2, etc. In some such examples, at least some of the speakers of the optional loudspeaker system 625 may be arranged at positions convenient for the space (for example, a position where there is a space for accommodating the loudspeaker), but may not be arranged in any standard predetermined speaker layout.

[0236] In some implementation examples, the apparatus 600 may include the optional sensor system 630 illustrated in FIG. 6. The optional sensor system 630 may include one or more cameras, touch sensors, gesture sensors, motion detectors, etc. According to some implementation examples, the optional sensor system 630 may include one or more cameras. In some implementation examples, the camera may be a self-standing camera. In some examples, one or more cameras of the optional sensor system 630 may be present within a smart audio device that may be a single-purpose audio device or a virtual assistant. In some such examples, one or more cameras of the optional sensor system 630 may be present within a TV, a mobile phone, or a smart speaker. In some examples, the apparatus 600 may not include the sensor system 630. However, in some such implementation examples, the apparatus 600 may be configured to receive sensor data for one or more sensors in the audio environment via the interface system 610.

[0237] In some implementation examples, device 600 may include the optional display system 635 illustrated in FIG. 6. The optional display system 635 may include one or more displays, such as one or more light emitting diode (LED) displays. In some cases, the optional display system 635 may include one or more organic light emitting diode (OLED) displays. In some examples in which device 600 includes the display system 635, the sensor system 630 may include a touch sensor system and / or a gesture sensor system proximate to one or more displays of the display system 635. According to some such implementation examples, the control system 610 may be configured to control the display system 635 to provide one or more graphical user interfaces (GUIs).

[0238] According to some examples, device 600 may be or include a smart audio device. In some such implementation examples, device 600 may be or implement (at least in part) a wake word detector. For example, device 600 may be or implement (at least in part) a virtual assistant.

[0239] Now referring again to FIG. 4. According to some examples, the system of FIG. 4 may be implemented such that CHASM 401 provides abstractions to applications 410, 411, and 412, and the applications 410, 411, and 412 can achieve a presentation (e.g., an audio session) by generating routes in an orchestration language. Using this language, the applications 410, 411, and 412 can command CHASM 401 via control links 430, 431, and 432.

[0240] Referring to FIGS. 1A, 1B, and 1C above, it is contemplated that the description in the language of orchestration of an application and its resulting situation can be the same in the case of FIG. 1A and the case of FIG. 1C.

[0241] To explain examples of syntax and examples of the orchestration language, first, some examples considering the situations of FIGS. 1A and 1C are given.

[0242] In some examples, what is referred to herein as the "root" of the orchestration language may include instructions for a media source (including, but not limited to, an audio source) and a media destination. The media source and the media destination can be specified, for example, in a root start request sent by an application to CHASM. According to some implementation examples, the media destination may be an audio environment destination or may include it. The audio environment destination can, in some cases, correspond to at least one person present within the audio environment for at least some time. In some cases, the audio environment destination can correspond to one or more areas or zones of the audio environment. Some examples of audio environment zones are disclosed herein. However, the audio environment destination generally does not include any particular audio device of the audio environment that will be involved in the root. By generalizing the orchestration language (including, but not limited to, the details required from the application to establish the root), the details for root implementation examples can be determined by CHASM and updated as needed.

[0243] In some examples, a route may include other information such as audio session priority, connection mode, one or more audio session goals or criteria. In some implementation examples, the route will have a corresponding code or identifier, which may be referred to herein as an audio session identifier. In some cases, the audio session identifier may be a persistent and unique audio session identifier.

[0244] In the first example above, the corresponding "route" may include a route from person X (e.g., user 102 in FIG. 1C) to the network (e.g., the Internet) in synchronous mode with high priority, and a route from the network to person X in synchronous mode with high priority. Such terms are similar to natural language when stating that one "wants to connect a phone call to person X", but are quite different from saying the following (see the situation in FIG. 1A including device 101):

[0245] Connect this device Mic (i.e., the microphone of device 101) to the processing (noise / echo cancellation), Connect the processing output (noise / echo) to the network, Connect the network to the processing input (dynamics), Connect the processing output (dynamics) to this device speaker, Connect the processing output (dynamics) to the processing input (reference).

[0246] At this point, if device 105 should be introduced (i.e., if the execution of the phone application should be performed in the situation of FIG. 1C including device 105 and device 101), it can be seen that executing a phone application according to such a list of commands requires details of the device and the necessary processing (echo and noise cancellation) may need to be completely changed. For example, as follows:

[0247] Connect the device Mic (i.e., the microphone of device 105) to the network, Connect the network to the input processing (noise / echo cancellation), Connect the processing output (noise / echo) to the network, Connect the network to the processing input (dynamics), Connect the processing output (dynamics) to this device speaker (i.e., the speaker of device 101), Connect the processing output (dynamics) to the processing input (reference).

[0248] The details of where it is best to perform signal processing, how to connect the signals, and generally what is the best output for the user (who may be located at a known or unknown location) can, in some cases, be pre-calculated for a limited number of use cases, but may involve optimizations that become unmanageable for a large number of devices and / or a large number of simultaneous audio sessions. The inventors have recognized that it is better to provide a framework that enables better connectivity, capabilities, knowledge, and control of smart audio devices (including by orchestrating or coordinating the devices), and to generate a portable and effective syntax for controlling the devices.

[0249] Some of the disclosed embodiments use approaches and languages that are design - effective and very general. There are certain aspects of the language that are best understood when considering an audio device as part of a route rather than as a specific end - point audio device (e.g., in embodiments where the system includes CHASM as described herein rather than SPASM as described herein). Aspects of some embodiments include one or more of the following: ROUTE SPECIFICATION SYNTAX, PERSISTENT UNIQUE SESSION IDENTIFIER, and CONTINUOUS NOTION OF DELIVERY, ACKNOWLEDGEMENT, and / or QUALITY.

[0250] The route - specification syntax (which addresses the need to specify elements, either explicitly or implicitly, for any generated route) may include: 〇 Source (person / device / automatic decision) and, thus, implicit permissions, 〇 Priority regarding how important this desired audio routing is with respect to other audio that may already be in progress or that may come later, 〇 Destination (ideally a person or a set of persons and, optionally, generalizable to a location), 〇 Mode of connectivity in terms of synchronization, transaction, or schedule scope, 〇 Degree to which a message must be acknowledged, or requirements for certainty regarding delivery confidence, and / or 〇A sense of what is the most important aspect of the content being listened to (comprehensible, quality, spatial fidelity, consistency, or perhaps inaudible). This last point may include the concept of negative roots, which is not only the interest in listening and being listened to, but also the interest in controlling what cannot be heard and / or what cannot be listened to. Some such examples include keeping one or more areas of the audio environment relatively quiet, such as the "do not wake the baby" implementation example described in detail below. Other such examples may include preventing others in the audio environment from overhearing a private conversation, for example, by playing "white noise" on one or more nearby loudspeakers, increasing the playback level of one or more other audio contents on one or more nearby loudspeakers, etc. For example, it may include preventing others in the audio environment from overhearing a private conversation, for example, by playing "white noise" on one or more nearby loudspeakers, increasing the playback level of one or more other audio contents on one or more nearby loudspeakers, etc.

[0251] Aspects of the persistent unique session identifier may include the following. As an important aspect of some embodiments, in some examples, the audio session corresponding to the root persists until it is completed or closed. For example, this enables the system to monitor the ongoing audio session (e.g., via CHASM) and determine which individual set of connectivity needs to be changed, rather than requiring the application to end or release the audio session to change the routing. The persistent unique session identifier may involve control and status aspects that enable the system to implement message or poll-driven management once generated. For example, the control of the audio session or route may be as follows, or may include them: - End - Move the destination, and - Raise or lower the priority.

[0252] The items that can be queried about an audio session or route can be or include the following: - Whether it is in a fixed position, - How well the stated goals are implemented among competing priorities, - How much of a sense or confidence the user has that they heard / approved the audio, - How good the quality is (e.g., for different goals of fidelity, space, intelligibility, information, attention, consistency, or inaudibility), - If desired, query the actual route layer down for which audio device is in use.

[0253] Aspects of the continuous concepts of delivery, approval, and / or quality can include the following. There may be a degree of a networking socket approach (and session layer) sense, but especially when considering the number of audio activities that can be routed simultaneously or queued, etc., audio routing can be very different. Also, since the destination can be at least one person and there can be uncertainty about the position of the person with respect to the audio device that can potentially be routed in some cases, it can be useful to have a very continuous confidence. Networking can include or be related to links that are either (possibly incoming or not) DATAGRAMS and (guaranteed to be incoming) STREAMS. In the case of audio, there can be a sense of whether things are audible or not, and / or a sense of whether we think we can or cannot hear someone.

[0254] These items are introduced in some embodiments of an orchestration language that can have some aspects of simple networking. On top of this (in some embodiments) there are presentation and application layers (e.g., for use when implementing an application example of a "phone call").

[0255] Embodiments of the orchestration language may relate to aspects related to the Session Initiation Protocol (SIP) and / or the Media Server Markup Language (MSML) (e.g., device-centric, continuous, and autonomous adaptive routing based on a current set of audio sessions . SIP is a transmission protocol used to initiate, maintain, and terminate sessions that may include voice, video, and / or messaging applications. In some cases, SIP may be used to transmit and control Internet telephony communication sessions, such as for voice calls, video calls, for private IP telephone systems, for instant messaging over Internet Protocol (IP) networks, for mobile calls, etc. SIP is a text-based protocol that defines the format of messages and the communication sequence of participants. SIP includes elements of the Hypertext Transfer Protocol (HTTP) and the Simple Mail Transfer Protocol (S MTP). Calls established using SIP may, in some cases, include multiple media streams, but for applications that exchange data as the payload of SIP messages (e.g., for text messaging applications), separate streams are not necessary at all.

[0256] MSML is described in Request for Comments (RFC) 5707. Using MSML, various types of services are controlled on an IP media server. According to MSML, a media server is a device specialized in controlling and / or operating media streams such as a Real-Time Transport Protocol media stream. According to MSML, an application server is separated from the media server and is configured to establish and disconnect call connections. According to MSML, an application server is configured to establish a control "tunnel" via SIP or IP. The application server uses the control "tunnel" to communicate requests and responses encoded in MSML with the media server.

[0257] Using MSML, how a multimedia session interacts with a media server can be defined and services can be applied to individual users or groups of users. Using MSML, media server features such as video layout and audio mixing can be controlled, a sidebar conference or personal mix can be generated, and properties of media streams can be set, etc.

[0258] Some embodiments do not require a user to be able to control a group of audio devices by issuing specific commands. However, it is contemplated that some embodiments can effectively achieve all desired presentations at the application layer without referring to the device itself.

[0259] FIG. 7 is a block diagram illustrating blocks of CHASM according to an example. FIG. 7 illustrates an example of CHASM 401 shown in FIG. 4. FIG. 7 illustrates CHASM 401 that receives routes from multiple applications using an orchestration language and stores information regarding the routes in a routing table 701. The elements of FIG. 7 include the following: 401: CHASM, 430: Commands from the first application using the orchestration language (application 410 in FIG. 4), and responses from CHASM 401, 431: Commands from the second application using the orchestration language (application 411 in FIG. 4), and responses from CHASM 401, 432: Commands from the third application using the orchestration language (application 412 in FIG. 4), and responses from CHASM 401, 703: Commands from additional applications using the orchestration language (not shown in FIG. 4), and responses from CHASM 401, 701: Routing table maintained by CHASM 401, 702: Optimizer, also referred to herein as an audio session manager, that continuously controls a plurality of audio devices based on current routing information, 435: Commands from CHASM 401 to the first audio device (audio device 420 in FIG. 4), and responses from the first audio device, 434: Commands from CHASM 401 to the second audio device (audio device 421 in FIG. 4), and responses from the second audio device, 435: Commands from CHASM 401 to the third audio device (audio device 422 in FIG. 4), and responses from the third audio device.

[0260] FIG. 8 shows details of the routing table shown in FIG. 7 according to an example. The elements of FIG. 8 include:

[0261] 701: Table of routes maintained by CHASM. According to this example, each route has the following fields, ● ID or "Persistent Unique Session Identifier", ● Record of which application requested the route, ● Source, ● In this example, a destination that may include one or more persons or one or more locations, but does not include an audio device, ● Priority, ● In this example, a connection mode selected from a list of modes including synchronous mode, scheduled mode, and transaction mode, ● Indication of whether approval is required, ● Which audio quality aspect(s) (referred to herein as audio session goal(s)) should be prioritized. In some examples, the audio session manager or prioritizer 702 will optimize the audio session according to the audio session goal(s), ● In this example, the route in the routing table 701 requested by application 410 and assigned ID 50. This route specifies that Alex (destination) wants to listen to Spotify with a priority of 4. In this example, the priority is an integer value, and the highest priority is 1. The connection mode is synchronous. This means ongoing in this example. In this case, Alex is not required to confirm or approve whether the corresponding music is provided to Alex. In this example, the only specified audio session goal is music quality,

[0262] 801: The route in the routing table 701 requested by application 410 and assigned ID 50. This route specifies that Alex (destination) wants to listen to Spotify with a priority of 4. In this example, the priority is an integer value, and the highest priority is 1. The connection mode is synchronous. This means ongoing in this example. In this case, Alex is not required to confirm or approve whether the corresponding music is provided to Alex. In this example, the only specified audio session goal is music quality, ● In this example, Alex is not required to confirm or approve whether the corresponding music is provided to Alex. In this example, the only specified audio session goal is music quality,

[0263] 802: The route in the routing table 701 requested by app 811 and the assigned ID 51. Angus is supposed to hear a timer alarm with priority 4. This audio session is scheduled for a future time. The future time is stored by CHASM401 but not shown in the routing table 701. In this example, Angus is required to approve having heard the alarm. In this example, the only specified audio session goal to increase the likelihood of Angus hearing the alarm is to be audible.

[0264] 803: The route in the routing table 701 requested by app 410 and the assigned ID 52. The destination is "infant", but the underlying audio session goal near the infant in the audio environment is to be inaudible. Therefore, this is an implementation example of "Don't wake the infant!", and a detailed example will be described later. This audio session has a priority of 2 (more important than almost anything). The connection mode is synchronous (in progress). Approval from the unwoken infant is not required. In this example, the only specified audio session goal at the location of the infant is to be inaudible. is. Thus, this is an implementation example of "Don't wake the infant!", and a detailed example will be described later. This audio session has a priority of 2 (more important than almost anything). The connection mode is synchronous (in progress). Approval from the unwoken infant is not required. In this example, the only specified audio session goal at the location of the infant is to be inaudible.

[0265] 804: The route in the routing table 701 requested by app 411 and the assigned ID 53. In this example, app 411 is a phone app. In this case, George is on the phone. Here, the priority of the audio session is 3. The connection mode is synchronous (in progress). Approval that George is still on the phone is not required. For example, George may intend to ask the virtual assistant to end the call if he is ready to end the call. In this example, the only specified audio session goal is to be understandable (understandability).

[0266] 805: The route in routing table 701 requested by app 412 and assigned ID 54. In this example, the underlying purpose of the audio session is to notify Richard that the plumber is at the front door and needs to talk to him. The connection mode is a transaction. Considering the priority of other audio sessions, play the message to Richard as soon as possible. In this example, Richard has just put the baby to bed and Richard is still in the baby's room. Considering route 803 with a higher priority, the CHASM audio session manager will wait until Richard leaves the baby's room until the message corresponding to route 805 is delivered. In this example, approval is required. In this case, Richard is required to verbally approve that he has heard the message and is on his way to meet the plumber. According to some examples, if Richard does not approve within a specified amount of time, the CHASM audio session manager may cause this message to be provided to all audio devices in the audio environment (excluding any audio devices in the baby's room if there are any in some examples) until Richard responds. In this example, the only specified audio session goal is that the voice be understandable (understandability), and Richard hears and understands the message.

[0267] The route in the routing table 701 requested by the fire alarm system application 806 and the assigned ID 55. The purpose underlying this route is to sound a fire alarm to evacuate from the house under certain circumstances (e.g., in accordance with the response from the smoke detection sensor). This route has the highest possible priority. Waking up infants is also allowed. The connection mode is synchronous. Approval is not required. In this example, the only specified audio session goal is to be audible. According to this example, CHASM will control all audio devices in the audio environment to play the alarm at a high volume so that all persons in the audio environment can hear the alarm and evacuate.

[0268] In some implementation examples, the audio session manager (e.g., CHASM) will maintain information corresponding to each route within one or more memory structures. According to some such implementation examples, the audio session manager can be configured to update the information corresponding to each route according to the changing state within the audio environment (e.g., a person changing position within the audio environment) and / or a control signal from the audio session manager 702. For example, referring to route 801, the audio session manager can include the following information or store and update a memory structure corresponding thereto. store and update.

[0269]

Table 1

[0270] The information shown in Table 1 is in a format that can be read by humans for the purpose of providing an illustration. The actual format used by the audio session manager to store such information (e.g., the position of the destination and the orientation of the destination) may or may not be understandable by humans depending on the specific implementation example.

[0271] In this example, the audio session manager is configured to monitor the position and orientation of Alex, the destination for route 801, and determine which audio device will be involved in providing audio content for route 801. According to some such examples, the audio session manager can be configured to determine the position of the audio device, the position of the person, and the orientation of the person according to a certain method (details will be described later). When the information in Table 1 changes, in some implementation examples, the audio session manager will send the corresponding command / control signal to the device that is rendering the audio from the media stream for route 801, and update the memory structure as illustrated via Table 1.

[0272] Figure 9A represents an example of a context-free grammar for a route start request in the language of orchestration. In some examples, Figure 9A can represent the grammar of a route request from an application to CHASM. For example, a route start request can be triggered according to, for example, the user selecting and interacting with an application, such as selecting the corresponding icon for the application via a voice command on a mobile phone.

[0273] In this example, element 901, in combination with elements 902A, 902B, 902C, and 902D, enables the definition of a root source. As illustrated by elements 902A, 902B, 902C, and 902D, in this example, the root source may be or may include one or more persons, one or more services, and audio environment locations. The service may be, for example, a cloud-based media streaming service, an external doorbell, or an in-home service that provides an audio feed from a doorbell-related audio device. In some implementations, the service may be specified according to a URL (e.g., a Spotify URL), a name of the service, an IP address of a home doorbell, etc. The audio environment location may, in some implementations, correspond to an audio environment zone described later. In some examples, the audio environment location source corresponds to one or more microphones within the zone. The comma in element 902D indicates that multiple sources may be specified. For example, a root request may indicate "root from Roger, Michael" or "root from Spotify" or "root from the kitchen" or "root from Roger and the kitchen", etc. obtain.

[0274] In this example, element 903, in combination with elements 904A, 904B, 904C, and 904D, enables the definition of a root destination. In this implementation, the root destination may be or may include one or more persons, one or more services, and audio environment locations. For example, a root request may indicate "root to David" or "root to the kitchen" or "root to the deck" or "root to Roger and the kitchen", etc.

[0275] In this example, only one connection mode can be selected per route. According to this implementation example, the connection mode options are synchronous, scheduled, or transactional. However, in some implementation examples, multiple connection modes can be selected per route. For example, in some such implementation examples, a route start request may indicate that the route can be both scheduled and transactional. For example, a route start request may indicate that a message should be delivered to David at a scheduled time and that David should reply to the message. Although not illustrated in FIG. 9A, in some implementation examples, a specific message (e.g., a pre-recorded message) may be included in the route start request.

[0276] In this example, the audio session goal is referred to as a "trait". This example, according to which one or more audio session goals can be indicated in a route start request via a combination of quality 907 and one or more traits 908A. A comma 908B indicates, according to this example, that one or more traits can be specified. However, in an alternative implementation example, only a single audio session goal can be indicated in a route start request.

[0277] FIG. 9B gives an example of an audio session goal. According to this example, the "trait" list 908A allows the specification of one or more important qualities. In some implementation examples, a route start request may specify multiple traits, for example, in descending order. For example, a route start request may specify (quality = intelligible, spatially faithful). This means that intelligible is the most important trait, followed by spatially faithful. A route start request may specify (quality = audible). This means that the only audio session goal is that a person can, for example, hear an alarm.

[0278] A root start request may specify (quality = inaudible). This means that the only audio session goal is that the person specified as the root destination (e.g., an infant) does not hear the audio being played in the audio environment. This is an example of a root start request for an "do not wake the infant" implementation example.

[0279] In another example, a root start request may specify (quality = audible, privacy). This may mean that, for example, the primary audio session goal is for the person specified as the root destination to hear the distributed audio, but the secondary audio session goal is to limit the extent to which other people can hear the audio distributed and / or exchanged along the route, for example, during a confidential phone call. As described elsewhere in this specification, the latter audio session goal can be achieved by playing white noise or other masking noise between the root destination and one or more other people in the audio environment, increasing the volume of other audio being played near one or more other people, etc.

[0280] Now refer to FIG. 9A. In this example, a root start request may specify a priority via elements 909 and 910. In some examples, the priority may be indicated via one of a finite number of integers (e.g., 3, 4, 5, 6, 7, 8, 9, 10, etc.). In some examples, 1 may indicate the highest priority.

[0281] According to this example, a root start request may, if necessary, specify an approval via element 911. For example, a root start request may mean "telling Michael that Richard is ready for dinner and getting approval". In response, in some examples, the audio session manager may attempt to determine Michael's location. For example, CHASM may infer that Michael is in the garage because the last place where Michael's voice was detected was the garage. Accordingly, the audio session manager may cause a notification such as "Dinner is ready. Please confirm that you have heard this message" to be played via one or more loudspeakers in the garage. If Michael responds, the audio session manager may report or play that response to Richard. If there is no response from Michael to the garage notification (e.g., after 10 seconds), the audio session manager may cause the notification to be made at the second most likely location for Michael, e.g., the place where Michael spends a lot of time, the last place where Michael was heard before the previous utterance in the garage. Suppose that place is Michael's bedroom. If there is no response from Michael to the notification in Michael's bedroom (e.g., after 10 seconds), the audio session manager may cause the notification to be played on many loudspeakers in the environment, but will follow other constraints such as "do not wake the baby".

[0282] Figure 10 illustrates a flow for a request to change a route according to an example. The route change request may be sent by an application and received by an audio session manager, for example. The route change request may be caused, for example, according to a user selecting and interacting with an application.

[0283] In this example, ID1002 refers to a persistent unique audio session number or code that the audio session manager might have pre - given to the application in response to a root start request. According to this example, the change of the connection mode can be made via element 1003 and elements 1004A, 1004B, or 1004C. Alternatively, elements 1004A, 1004B, and 1004C can be bypassed if a change in the connection mode is not desired.

[0284] According to this example, one or more audio session targets can be changed via elements 1005, 1006A, and 1006B. Alternatively, elements 1005, 1006A, and 1006B can be bypassed if a change in the audio session target is not desired.

[0285] In this example, the root priority can be changed via elements 1007 and 1008. Alternatively, elements 1007 and 1008 can be bypassed if the root priority is not desired.

[0286] According to this example, element 1009 or element 1011 can be used to make a change to the approval request. For example, element 1009 indicates that approval can be added if approval was not previously required for the root. Conversely, element 1011 indicates that approval can be removed if approval was previously required for the root. The semicolon of element 1010 indicates the end of the request to change the root.

[0287] Figures 11A and 11B show a further flow for a request to change the root An example is illustrated. FIG. 11C illustrates an example of a flow for deleting a route. A route change or deletion request may be sent, for example, by an application and received by an audio session manager. A route change request may be caused, for example, in accordance with a user selecting and interacting with an application. In FIGS. 11A and 11B, "sink" refers to a route destination. Similar to other flow diagrams disclosed herein, the operations illustrated in FIGS. 11A-11C are not necessarily performed in the order shown. For example, in some implementations, a route ID may be specified earlier in the flow, for example, at the start of the flow.

[0288] FIG. 11A illustrates flow 1100A for adding a source or destination. In some cases, one or more sources or destinations may be added. In this example, a route source may be added via elements 1101 and 1102A. According to this example, a route destination may be added via elements 1101 and 1102B. In this example, the route to which a route source or destination is added is shown via elements 1103 and 1104. According to this example, a person may be added as a source or destination via element 1105A. In this example, a service may be added as a source or destination via element 1105B. According to this example, a location may be added as a source or destination via element 1105C. Element 1106 indicates the end of the flow for adding one or more sources or destinations.

[0289] Figure 11B illustrates flow 1100B for removing a source or destination. In some cases, one or more sources or destinations can be removed. In this example, the root source can be removed via elements 1107 and 1108A. According to this example, the root destination can be removed via elements 1107 and 1108B. In this example, the route through which the root source or destination is removed is shown via elements 1109 and 1110. According to this example, a person can be removed as a source or destination via element 1111A. In this example, a service can be removed as a source or destination via element 1111B. According to this example, a location can be removed as a source or destination via element 1111C. Element 1112 indicates the end of the flow for removing one or more sources or destinations.

[0290] Figure 11C illustrates a flow for removing a route. Here, element 1113 indicates removal. The route ID specified via element 1114 indicates the route to be removed. Element 1115 indicates the end of the flow for removing one or more sources or destinations.

[0291] FIG. 12 is a flowchart including blocks of an audio session management method according to some implementation examples. According to this example, method 1200 is an audio session management method for an audio environment having a plurality of audio devices. The blocks of method 1200 are not necessarily performed in the order shown, similar to other methods described herein. In some implementation examples, one or more of the blocks of method 1200 may be performed simultaneously. Further, some implementation examples of method 1200 may include more or fewer blocks than those illustrated and / or described. The blocks of method 1200 may be performed by one or more devices. Those devices may be a control system such as (or may include) the control system 610 described below illustrated in FIG. 6, or one of the other disclosed control system examples.

[0292] In this example, block 1205 includes receiving, from a first device implementing a first application and by a device implementing an audio session manager (e.g., CHASM), a first route start request for starting a first route for a first audio session. According to this example, the first route start request indicates a first audio source and a first audio environment destination. Here, the first audio environment destination corresponds to at least a first person within the audio environment. However, in this example, the first audio environment destination does not indicate an audio device. According to some examples, the first route start request may indicate at least a first area of the audio environment as the first route source or the first route destination. In some cases, the first route start request may indicate at least a first service as the first audio source.

[0293]

[0294] In this implementation example, block 1210 includes establishing a first route corresponding to a first root start request by a device implementing an audio session manager. In this example, establishing the first route includes determining a first position of at least a first person within the audio environment, determining at least one audio device for a first stage of a first audio session, and starting or scheduling the first audio session.

[0295] According to some examples, the first root start request may include a first audio session priority. In some cases, the first root start request may include a first connection mode. The first connection mode may be, for example, a synchronous connection mode, a transaction connection mode, or a scheduled connection mode. In some examples, the first root start request may indicate a plurality of connection modes.

[0296] In some implementation examples, the first root start request may include an indication of whether approval will be requested from at least the first person. In some examples, the first root start request may include a first audio session goal. The first audio session goal may include, for example, intelligible, audio quality, spatial fidelity, and / or inaudible.

[0297] As described elsewhere in this specification, in some implementation examples, a route may have an associated audio session identifier. The associated audio session identifier may be, in some implementation examples, a persistent unique audio session identifier. Thus, some implementation examples of method 1200 may include determining a first persistent unique audio session identifier for the first route (e.g., by an audio session manager) and sending the first persistent unique audio session identifier to a first device (the device running the first application).

[0298] In some implementation examples, establishing a first route may include causing at least one device in the environment to establish at least a first media stream corresponding to the first route, the first media stream including a first audio signal. Some implementation examples of method 1200 may include causing the first audio signal to be rendered to a first rendered audio signal. In some examples, method 1200 may include an audio session manager that causes the first audio signal to be rendered to a first rendered audio signal at another device in the audio environment. However, in some implementation examples, the audio session manager may be configured to receive the first audio signal and render the first audio signal to a first rendered au dio signal.

[0299] As described elsewhere herein, in some implementation examples, an audio session manager (e.g., CHASM) may monitor the state of the audio environment, such as the position and / or orientation of one or more persons in the audio environment, the position of audio devices in the audio environment. For example, for the use case of "don't wake the baby", the audio session manager (e.g., optimizer 702 in FIG. 7) may determine or at least estimate where the baby is. The audio session manager may know where the baby is based on the user's expression text (e.g., "don't wake the baby. The baby is in bedroom 1") sent in the "orchestration language" from the relevant application. Alternatively, or in addition, the audio session manager may determine where the baby is based on previous expression inputs or previous detection of the baby's crying (e.g., as described below). In some examples, the audio session manager receives this constraint (e.g., via an "inaudible" audio session goal) and may implement the constraint by, for example, causing the sound pressure level at the baby's position to be less than a threshold decibel level (e.g., 50 dB).

[0300] Some examples of method 1200 may include determining a first orientation of a first person with respect to a first stage of an audio session. According to some such examples, causing a first audio signal to be rendered to a first rendered audio signal may include determining a first reference spatial mode corresponding to the first position and the first orientation of the first person, and determining a first relative activation of loudspeakers in an audio environment corresponding to the first reference spatial mode. Some detailed examples are described below.

[0301] In some cases, the audio session manager may determine that the first person has changed position and / or orientation. Some examples of method 1200 may include determining at least one of a second position or a second orientation of the first person, determining a second reference spatial mode corresponding to at least one of the second position or the second orientation, and determining a second relative activation of loudspeakers in an audio environment corresponding to the second reference spatial mode.

[0302] As described elsewhere in this disclosure, the audio manager may in some cases be tasked with establishing and implementing multiple routes at once. Some examples of method 1200 may include receiving, from a second device implementing a second application and by a device implementing the audio session manager, a second route start request to start a second route for a second audio session. The first route start request may indicate a second audio source and a second audio environment destination. In some examples, the second audio environment destination may correspond to at least a second person in the audio environment. However, in some cases, the second audio environment destination may not indicate any particular audio device associated with the second route.

[0303] Some such examples of method 1200 may include establishing a second route corresponding to a second route start request by a device implementing an audio session manager. In some cases, establishing the second route may include determining a first position of at least a second person within the audio environment, determining at least one audio device for a first stage of a second audio session, and starting the second audio session.

[0304] According to some examples, establishing the second route may include at least establishing a second media stream corresponding to the second route. The second media stream may include a second audio signal. Some such examples of method 1200 may include causing the second audio signal to be rendered to a second rendered audio signal.

[0305] Some examples of method 1200 may include changing a rendering process for a first audio signal based at least in part on at least one of a second audio signal, a second rendered audio signal, or a characteristic thereof to generate a modified first rendered audio signal. Changing the rendering process for the first audio signal may include, for example, warping the rendering of the first audio signal away from the rendering position of the second rendered audio signal. Alternatively, or in addition, changing the rendering process for the first audio signal may include changing the loudness of one or more signals of the first rendered audio signal in response to the loudness of one or more signals of the second audio signal or the second rendered audio signal.

[0306] FIG. 13 is a flow diagram including blocks of an audio session management method according to some implementation examples. According to this example, method 1300 is an audio session management method for an audio environment having a plurality of audio devices. The blocks of method 1300 are not necessarily performed in the order shown, similar to other methods described herein. In some implementation examples, one or more of the blocks of method 1300 may be performed simultaneously. Further, some implementation examples of method 1300 may include more or fewer blocks than those illustrated and / or described. The blocks of method 1300 may be performed by one or more devices that may be (or include) a control system such as the control system 610 described below illustrated in FIG. 6 or one of the other disclosed control system examples.

[0307] In this example, block 1305 includes receiving, from a first device implementing a first application and by a device implementing an audio session manager (e.g., CHASM), a first route start request for starting a first route for a first audio session. According to this example, the first route start request indicates a first audio source and a first audio environment destination. Here, the first audio environment destination corresponds to at least a first area of the audio environment. However, in this example, the first audio environment destination does not indicate an audio device.

[0308] According to some examples, the first route start request may indicate at least a first person in the audio environment as a first route source or a first route destination. In some cases, the first route start request may indicate at least a first service as a first audio source.

[0309] In this implementation example, block 1310 includes establishing a first route corresponding to a first root start request by a device implementing an audio session manager. In this example, establishing the first route includes determining at least one audio device within a first area of the audio environment for a first stage of a first audio session, and starting or scheduling the first audio session.

[0310] According to some examples, the first root start request may include a first audio session priority. In some cases, the first root start request may include a first connection mode. The first connection mode can be, for example, a synchronous connection mode, a transaction connection mode, or a scheduled connection mode. In some examples, the first root start request may indicate a plurality of connection modes.

[0311] In some implementation examples, the first root start request may include an indication of whether approval will be required from at least a first person. In some examples, the first root start request may include a first audio session target. The first audio session target can include, for example, intelligible, audio quality, spatial fidelity, and / or inaudible.

[0312] Some implementation examples of method 1300 may include determining a first persistent unique audio session identifier for the first route (e.g., by an audio session manager), and sending the first persistent unique audio session identifier to a first device (a device running a first application).

[0313] In some implementation examples, establishing a first route may include causing at least one device in the environment to establish at least a first media stream corresponding to the first route, the first media stream including a first audio signal. Some implementation examples of method 1300 may include causing the first audio signal to be rendered to a first rendered audio signal. In some examples, method 1300 may include causing an audio session manager to render the first audio signal to a first rendered audio signal to another device in the audio environment. However, in some implementation examples, the audio session manager may be configured to receive the first audio signal and render the first audio signal to a first rendered audio signal.

[0314] As described elsewhere herein, in some implementation examples, an audio session manager (e.g., CHASM) may monitor the state of the audio environment, such as the location of one or more audio devices in the audio environment.

[0315] Some examples of method 1300 may include performing a first loudspeaker auto-location process that automatically determines a first position of each audio device of a plurality of audio devices within a first area of the audio environment at a first time. Some In such examples, the rendering process may be at least partially based on the first position of each audio device. Some such examples may include storing the first position of each audio device within a data structure associated with the first route.

[0316] In some cases, the audio session manager may determine that at least one audio device within a first area has a changed position. Some such examples may include performing a second loudspeaker automatic positioning process that automatically determines the changed position, and updating a rendering process based at least in part on the changed position. Some such implementations may include storing the changed position in a data structure associated with a first route.

[0317] In some cases, the audio session manager may determine that at least one additional audio device has moved into the first area. Some such examples may include performing a second loudspeaker automatic positioning process that automatically determines an additional audio device position of the additional audio device, and updating a rendering process based at least in part on the position of the additional audio device. Some such implementations may include storing the position of the additional audio device in a data structure associated with a first route.

[0318] As described elsewhere herein, in some examples, a first route start request may indicate at least a first person as a first route source or a first route destination. Some examples of method 1300 may include determining a first orientation of the first person with respect to a first stage of an audio session. According to some such examples, causing a first audio signal to be rendered to a first rendered audio signal may include determining a first reference spatial mode corresponding to the first position and the first orientation of the first person, and determining a first relative activation of loudspeakers within an audio environment corresponding to the first reference spatial mode. Some detailed examples are described below.

[0319] In some cases, the audio session manager may determine that the first person has changed position and / or orientation. Some examples of method 1300 may include determining at least one of a second position or a second orientation of the first person, determining a second reference spatial mode corresponding to at least one of the second position or the second orientation, and determining a second relative activation of loudspeakers within the audio environment corresponding to the second reference spatial mode.

[0320] As described elsewhere in this disclosure, the audio manager may in some cases be tasked with establishing and implementing multiple routes at once. Some examples of method 1300 may include receiving, from a second device implementing a second application and by a device implementing the audio session manager, a second route start request to start a second route for a second audio session. The first route start request may indicate a second audio source and a second audio environment destination. In some examples, the second audio environment destination may correspond to at least a second person within the audio environment. However, in some cases, the second audio environment destination may not indicate any particular audio device associated with the second route.

[0321] Some such examples of method 1300 may include establishing, by a device implementing the audio session manager, a second route corresponding to the second route start request. In some cases, establishing the second route may include determining a first position of at least a second person within the audio environment, determining at least one audio device for a first stage of the second audio session, and starting the second audio session.

[0322] According to some examples, establishing a second route may include, at least, establishing a second media stream corresponding to the second route. The second media stream may include a second audio signal. Some such examples of method 1300 may include causing the second audio signal to be rendered to a second rendered audio signal.

[0323] Some examples of method 1300 may include changing a rendering process for a first audio signal, at least in part, based on at least one of the second audio signal, the second rendered audio signal, or its characteristics, to generate a modified first rendered audio signal. Changing a rendering process for a first audio signal may include, for example, warping the rendering of the first audio signal away from the rendering position of the second rendered audio signal. Alternatively, or in addition, changing a rendering process for a first audio signal may include changing the loudness of one or more signals of the first rendered audio signal in response to the loudness of one or more signals of the second audio signal or the second rendered audio signal.

[0324] FIG. 14 is a flow diagram including blocks of an audio session management method according to some implementation examples. According to this example, method 1400 is an audio session management method for an audio environment having a plurality of audio devices. The blocks of method 1400 are not necessarily performed in the order shown, similar to other methods described herein. In some implementation examples, one or more of the blocks of method 1400 may be performed simultaneously. Further, some implementation examples of method 1400 may include more or fewer blocks than those illustrated and / or described. The blocks of method 1400 may be performed by one or more devices that can be (or include) a control system such as the control system 610 described below illustrated in FIG. 6 or one of the other disclosed control system examples.

[0325] In this example, at block 1405, application 410 of FIG. 4 commands CHASM 401 using an orchestration language. Block 1405 may include, for example, application 410 that sends a root start request or a root change request to CHASM 401.

[0326] According to this example, CHASM 401 determines optimal media engine control information that may respond to the commands received from application 410. In this example, the optimal media engine control information is at least partially based on the position of the listener in the audio environment, the availability of audio devices in the audio environment, and the audio session priority indicated in the commands from application 410. In some cases, the optimal media engine control information may be at least partially based on the media engine capabilities determined by CHASM 401 via, for example, device property descriptors shared by the relevant audio device(s). According to some examples, the optimal media engine control information may be at least partially based on the orientation of the listener.

[0327] In this case, block 415 includes transmitting control information to one or more audio device media engines. The control information may correspond to the audio session management control signals described above with reference to FIG. 5.

[0328] According to this example, block 1420 represents CHASM 401 that monitors the state within the audio environment and possible further communications from application 410 regarding this particular route to determine if there have been any significant changes such as a change in root priority, a change in the location(s) of the audio device(s), or a change in the listener's location. If so, the process returns to block 1410, and the processing at block 1410 is performed according to the new parameter(s). If not, CHASM 401 continues the monitoring process at block 1420.

[0329] FIG. 15 is a flowchart including blocks of an automatic setup process for one or more audio devices newly introduced into an audio environment according to some implementation examples. In this example, some or all of the audio devices are new audio devices. The blocks of method 1500 are not necessarily performed in the order shown, similar to other methods described herein. In one implementation example, one or more of the blocks of method 1500 may be performed simultaneously. Further, some implementation examples of method 1500 may include more or fewer blocks than those illustrated and / or described. The method 1500 blocks may be performed by one or more devices that may be (or may include) a control system such as control system 610 described below and illustrated in FIG. 6 or another disclosed control system example. The method 1500 blocks may be performed by one or more devices that may be (or may include) a control system such as control system 610 described below and illustrated in FIG. 6 or another disclosed control system example.

[0330] In this example, in block 1505, a new audio device is unpacked and powered on. In the example of block 1510, each of the new audio devices enters discovery mode to search for other audio devices and, in particular, to search for the CHASM of the audio environment. If an existing CHASM is discovered, the new audio device can be configured to communicate with the CHASM, such as to share information about the capabilities of each new audio device with the CHASM.

[0331] However, according to this example, no existing CHASM is discovered. Thus, in this example of block 1510, one of the new audio devices configures itself as the CHASM. In this example, the new audio device with the most available computing power and / or the greatest connectivity will configure itself as the new CHASM 401.

[0332] In this example, in block 1515, all of the new non-CHASM audio devices communicate with another new audio device that is the newly designated CHASM 401. According to this example, the new CHASM 401 starts the "Setup" application, which is the application 412 in FIG. 4 in this example. In this case, the setup application 412 can be configured to interact with the user, for example, via audio and / or visual prompts, to guide the user through the setup process.

[0333] According to this example, in block 1520, the setup application 412 indicates "Set up all new devices" and sends a command with the highest level of priority to the CHASM 401 in the language of orchestration.

[0334] In this example, at block 1525, CHASM 401 interprets the instructions from the setup application 412 and determines that a new acoustic mapping calibration is required. According to this example, the acoustic mapping process is initiated at block 1525 and completed at block 1530 via communication between CHASM 401 and the media engines of the new non-CHASM audio devices (in this case, media engines 440, 441, and 442 of FIG. 4). As used herein, the term "acoustic mapping" includes the estimation of the positions of all discoverable loudspeakers in an audio environment. The acoustic mapping process may include, for example, a loudspeaker automatic positioning process as described in detail below. In some cases, the acoustic mapping process may include the process of discovering loudspeaker capability information and / or individual loudspeaker dynamics processing information.

[0335] According to this example, at block 1535, CHASM 401 sends a confirmation to application 412 that the setup process has been completed. In this example, at block 1540, application 412 indicates to the user that the setup process has been completed.

[0336] FIG. 16 is a flowchart including blocks of a process for installing a virtual assistant application according to some implementation examples. In some cases, method 1700 may be performed after the setup process of method 1500. In this example, the Method 1600 includes installing a virtual assistant application in the context of the audio environment illustrated in FIG. 4. The blocks of method 1600 are not necessarily performed in the order shown, similar to the other methods described herein. In one implementation, one or more of the blocks of method 1600 may be performed simultaneously. Further, some implementations of method 1600 may include more or fewer blocks than those illustrated and / or described. The blocks of method 1600 may be performed by one or more devices that may be (or may include) a control system such as control system 610 described below and illustrated in FIG. 6 or another disclosed example of a control system.

[0337] In this example, at block 1605, a new application 411 called "Virtual Assisting Liaison" or VAL is installed by the user. According to some examples, block 1605 may include downloading application 411 to an audio device such as a mobile phone via the Internet from one or more servers.

[0338] According to this implementation, at block 1610, application 411 instructs CHASM401 to continuously listen for a new wake word "Hey Val" as a persistent audio session with the highest priority in the language of orchestration. In this example, at block 1615, CHASM401 interprets the instruction from application 411 and instructs the wake word detector to configure to listen for the wake word "Hey Val" and give a callback to CHASM401 whenever the wake word "Hey Val" is detected, for media engines 440, 441, and 442. In this implementation, at block 1620, media engines 440, 441, and 442 continue to attempt to listen for the wake word.

[0339] ​In this example, at block 1625, CHASM 401 receives from media engines 440 and 441 a callback indicating that the wake word "Hey Val" has been detected. In response, CHASM 401 attempts to listen for commands from media engines 440, 441, and 442 for a threshold period (in this example, 5 seconds) after the wake word was first detected, and if a command is detected, instructs the media engines to "duck" or reduce the volume of the audio within the area where the command was detected.

[0340] According to this example, at block 1630, media engines 440, 441, and 442 all detect a command and send to CHASM 401 the voice audio data and probability corresponding to the detected command. In this example, at block 1630, CHASM 401 sends to application 411 the voice audio data and probability corresponding to the detected command.

[0341] In this implementation example, at block 1635, application 411 receives the voice audio data and probability corresponding to the detected command and sends these data to a cloud-based voice recognition application for processing. In this example, at block 1635, the cloud-based voice recognition application sends the results of the voice recognition processing to application 411. The results of the voice recognition processing include, in this example, one or more words corresponding to the command. Here, at block 1635, application 411 instructs CHASM 401 to end the voice recognition session in the orchestration language. According to this example, CHASM 401 instructs media engines 440, 441, and 442 to stop listening for commands.

[0342] FIG. 17 is a flow diagram including blocks of an audio session management method according to some implementation examples. According to this example, method 1700 is an audio session management method for implementing a music application within the audio environment of FIG. 4. In some cases, method 1700 may be performed after the setup process of method 1500. In some examples, method 1700 may be performed before or after the process for installing the virtual assistant application described above with reference to FIG. 16. The blocks of method 1700 are not necessarily performed in the order shown, similar to other methods described herein. In one implementation example, one or more of the blocks of method 1700 may be performed simultaneously. Further, some implementation examples of method 1700 may include more or fewer blocks than those illustrated and / or described. The blocks of method 1700 may be performed by one or more devices that can be (or include) a control system such as the control system 610 described below and illustrated in FIG. 6 or one of the other disclosed control system examples.

[0343] In this example, at block 1705, the user provides an input to a music application operating on a device within the audio environment. In this case, the music application is application 410 of FIG. 4. According to this example, application 410 operates on a smartphone, and the input is provided via a user interface of the smartphone such as a touch and / or gesture sensor system.

[0344] According to this example, in block 1710, application 410 instructs CHASM401 to initiate a route to the user who is interacting with application 410 via a smartphone from a cloud-based music service via a root start request of the orchestration language in this example. In this example, the root start request indicates a synchronous mode where the audio session target is the best music playback quality, no approval is requested, and the priority is 4, using the current favorite playlist of the user of the cloud-based music service.

[0345] In this example, in block 1715, CHASM401 determines which audio devices in the audio environment will be involved in the route according to the instruction received in block 1710. The determination can be based at least in part on a pre-determined acoustic map of the audio environment where the audio device is currently available, the capabilities of the available audio devices, and the estimated current location of the user. In some examples, the determination in block 1715 can be based at least in part on the estimated current orientation of the user. In some implementations, also, in block 1715, a nominal or initial listening level can be selected. This level can be based at least in part on the estimated proximity of the user to one or more audio devices, the ambient noise level within the user's area, etc.

[0346] According to this example, in block 1720, CHASM401 sends control information to a selected audio device media engine (media engine 441 in this example) to obtain a media bitstream corresponding to the route requested by application 410. In this example, CHASM401 provides the media engine 441 with the HTTP address of a cloud-based music provider, for example, the HTTP address of a specific server hosted by the cloud-based music provider. According to this implementation example, in block 1725, media engine 441 obtains the media bitstream from a cloud-based music provider, from the location of one or more assigned servers in this example.

[0347] In this example, block 1730 includes playing the music corresponding to the media stream obtained in block 1725. According to this example, CHASM401 has determined that at least loudspeaker 461, and in some examples also loudspeaker 460 and / or loudspeaker 462 are involved in the playback of the music. In some such examples, CHASM401 instructs media engine 441 to render the audio data from the media stream and provide the rendered speaker feed signal to media engine 440 and / or media engine 442.

[0348] Figure 18A is a block diagram of a minimal version of an embodiment. N program streams ( JPEG0007710002000002.jpg59) are illustrated. The first program stream is explicitly labeled with space. A corresponding group of N audio signals of the program streams are fed through corresponding renderers. Each renderer feeds its corresponding program stream to a common set of M arbitrarily spaced loudspeakers ( They are individually configured to play through JPEG0007710002000003.jpg510). Also, the renderer may be referred to herein as the "rendering module". The rendering module and mixer 1830a may be implemented via software, hardware, firmware, or some combination thereof. In this example, the rendering module and mixer 1830a are implemented by a control system 610a, which is an example of the control system 610 described above with reference to FIG. 6. According to some implementations, the functions of the rendering module and mixer 1830a may be implemented, at least in part, according to instructions from a device (e.g., CHASM) that implements what is referred to herein as the audio session manager. In some such examples, the functions of the rendering module and mixer 1830a may be implemented, at least in part, according to instructions from CHASM208C, CHASM208D, CHASM307, and / or CHASM401 described above with reference to FIGS. 2C, 2D, 3C, and 4. In an alternative example, a device that implements the audio session manager may also implement the functions of the rendering module and mixer 1830a.

[0349] In the example shown in FIG. 18A, each of the N renderers outputs a total of M loudspeaker feeds across all N renderers for synchronous playback on a set of M loudspeakers. According to this implementation example, information about the layout of the M loudspeakers in the listening environment is provided to all the renderers. This information is indicated by the dashed line fed back from the loudspeaker block. Thereby, the renderer can be appropriately configured to perform playback via the speakers. This layout information may or may not be transmitted from one or more of the speakers themselves, depending on the particular implementation example. According to some examples, the layout information may be provided by one or more smart speakers configured to determine the relative positions of each of the M loudspeakers in the listening environment. Some such automatic positioning methods may be based on the direction-of-arrival method or the time of arrival (TOA) method. In other examples, this layout information may be determined by another device and / or input by the user. According to some examples, loudspeaker specification information about the capabilities of at least some of the M loudspeakers in the listening environment may be provided to all the renderers. Such loudspeaker specification information may include impedance, frequency response, sensitivity, power rating, number and position of individual drivers, etc. According to this example, information from the rendering of one or more of the additional program streams is fed to the renderer of the main spatial stream such that the rendering can be dynamically changed as a function of this information. This information is represented by the dashed line returning from renderer blocks 2 to N to renderer block 1.

[0350] Figure 18B illustrates another (more capable) embodiment having additional features. In this example, the rendering module and mixer 1830b are implemented via a control system 610, which is an example of the control system 610 described above with reference to FIG. 6. According to some implementation examples, the functions of the rendering module and mixer 1830b may be implemented at least in part according to instructions from a device (e.g., CHASM) that implements what is referred to herein as an audio session manager. In some such examples, the functions of the rendering module and mixer 1830b may be implemented at least in part according to instructions from CHASM208C, CHASM208D, CHASM307, and / or CHASM401 described above with reference to FIGS. 2C, 2D, 3C, and 4. In alternative examples, also, a device implementing the audio session manager may implement the functions of the rendering module and mixer 1830b.

[0351] In FIG. 18B, the dashed line extending vertically between all of the N renderers represents the idea that any one of the N renderers can contribute to a dynamic change in any of the remaining N - 1 renderers. In other words, rendering any one of the N program streams can be dynamically changed as a function of one or more combinations of rendering of any of the remaining N - 1 program streams. Additionally, any one or more of the program streams can be a spatial mix, and the rendering of any program stream can be dynamically changed as a function of any of the other program streams, whether or not that program stream is spatial. For example, as described above, loudspeaker layout information can be provided to the N renderers. In some examples, loudspeaker specification information can be provided to the N renderers. In some implementation examples, the microphone system 620a is a set of K microphones in the listening environment ( It may include JPEG0007710002000004.jpg59). In some examples, the microphone(s) can be attached to or cooperate with one or more of the loudspeakers. These microphones can feedback both the audio signals captured by them, represented by solid lines, and additional configuration information (e.g., their positions), represented by dashed lines, to a set of N renderers. Then, any of the N renderers can be dynamically changed as a function of this additional microphone input. Various examples are provided herein.

[0352] Examples of information obtained from the microphone input and then used to dynamically change any of the N renderers include, but are not limited to, the following. ● Detection of the utterance of a specific word or term by a user of the system. ● Estimated values of the positions of one or more users of the system. ● Estimated values of the loudness of any combination of N program streams at specific positions within the listening space. ● Estimated values of the loudness of other ambient sounds such as background noise within the listening environment.

[0353] FIG. 19 is a flowchart showing an overview of an example of a method that can be provided by an apparatus or system as illustrated in FIGS. 6, 18A, or 18B. The blocks of method 1900 are not necessarily performed in the order shown, similar to other methods described herein. Further, such a method may have more or fewer blocks than those illustrated and / or described. It may be included. The blocks of method 1900 may be (or may include) performed by one or more devices such as the control system 610, control system 610a or control system 610b illustrated in FIGS. 6, 18A and 18B, or one of the other disclosed control system examples. According to some implementations, the blocks of method 1900 may be performed at least in part in accordance with instructions from a device (eg, CHASM) that implements what is referred to herein as an audio session manager. In some such examples, the blocks of method 1900 may be performed at least in part in accordance with instructions from CHASM208C, CHASM208D, CHASM307 and / or CHASM401 described above with reference to FIGS. 2C, 2D, 3C and 4. In alternative examples, also, a device implementing an audio session manager may implement the blocks of method 1900.

[0354] In this implementation example, block 1905 includes receiving a first audio program stream via an interface system. In this example, the first audio program stream includes a first audio signal scheduled to be reproduced by at least some speakers of the environment. Here, the first audio program stream includes first spatial data. According to this example, the first spatial data includes channel data and / or spatial metadata. In some examples, block 1905 includes a first rendering module of a control system that receives a first audio program stream via an interface system.

[0355] According to this example, block 1910 includes rendering a first audio signal for playback via speakers in the environment to generate a first rendered audio signal. Some examples of method 1900 include receiving loudspeaker layout information, as described above for example. Some examples of method 1900 include receiving loudspeaker specification information, as described above for example. In some examples, the first rendering module may generate the first rendered audio signal based at least in part on the received loudspeaker layout information and / or the received loudspeaker specification information.

[0356] In this example, block 1915 includes receiving a second audio program stream via an interface system. In this implementation example, the second audio program stream includes a second audio signal scheduled to be played back by at least some of the speakers in the environment. According to this example, the second audio program stream includes second spatial data. The second spatial data includes channel data and / or spatial metadata. In some examples, block 1915 includes a second rendering module of a control system that receives the second audio program stream via the interface system.

[0357] According to this implementation example, block 1920 includes rendering a second audio signal for playback via speakers in the environment to generate a second rendered audio signal. In some examples, the second rendering module may generate the second rendered audio signal based at least in part on the received loudspeaker layout information and / or the received loudspeaker specification information.

[0358] In some cases, some or all of the speakers in an environment may be placed in positions that do not conform to any of the predetermined speaker layouts of any standard, such as Dolby 5.1, Dolby 7.1, Hamasaki 22.2, etc. In some such examples, at least some of the speakers in the environment may be placed in positions that are convenient for the furniture, walls, etc. of the environment (e.g., positions where there is space to accommodate a loudspeaker), but may not be placed in any of the predetermined speaker layouts of any standard. That is, they may not be arranged in any of the predetermined speaker layouts of any standard.

[0359] Thus, in some implementation examples, block 1910 or block 1920 may include flexibly rendering to speakers placed at arbitrary positions. Some such implementation examples may include center of mass amplitude panning (CMAP), flexible virtualization (FV), or a combination of both. At a high level, both of these techniques render a set of one or more audio signals, each having a desired perceived spatial position, for playback on a set of two or more speakers. Here, the relative activation of the set of speakers is a function of a model of the perceived spatial position of the audio signal being played on the speakers and the proximity of the desired perceived spatial position of the audio signal to the position of the speakers. This model ensures that the audio signal is heard by the listener near its intended spatial position, and the proximity term controls which speakers are used to achieve this spatial impression. In particular, the proximity term acts favorably on the activation of the speakers near the desired perceived spatial position of the audio signal. For both CMAP and FV, this functional relationship is conveniently obtained from a cost function described as the sum of two terms: a term for the spatial aspect and a term for the proximity.

Number

[0360] Here, the set JPEG0007710002000006.jpg56 represents the positions of a set of M loudspeakers, JPEG0007710002000007.jpg52 represents the desired perceived spatial position of the audio signal, and g represents the M - dimensional vector of speaker activations. For CMAP, each activation in the vector is a gain per speaker, while for FV, each activation represents a filter (in this second case, g can be considered equivalent to a vector of complex numbers at a specific frequency, and different gs are calculated over multiple frequencies to form the filter). The optimal vector of activations is found by minimizing a cost function over the activations.

Number

[0361] Depending on the definition of the cost function, even if the relative levels between the components of JPEG0007710002000009.jpg57 are appropriate, it is difficult to control the absolute level of the optimal activations obtained from the above minimization. To address this problem, normalization of JPEG0007710002000010.jpg57 can be performed later so that the absolute level of the activations is controlled. For example, it may be desirable to normalize the vector to obtain unit length, which is in line with the commonly used constant power panning rules. power panning rules)

Number

[0362] The exact behavior of the flexible rendering algorithm depends on two terms of the cost function, JPEG0007710002000012.jpg511 and Determined by the specific construction of JPEG0007710002000013.jpg515. For CMAP, JPEG0007710002000014.jpg511 places the perceived spatial position of an audio signal reproduced from a set of loudspeakers at the centroid of the positions of the loudspeakers weighted by the associated activation gain JPEG0007710002000015.jpg54 (components of vector g). It is obtained from a model.

Number

[0363] Then, manipulate Equation 3 to obtain a spatial cost that represents the squared error between the desired audio position and the audio position generated by the activated loudspeakers.

Number

[0364] In the case of FV, the spatial term of the cost function is defined differently. Here, the goal is to generate a binaural response b corresponding to the audio object positions JPEG0007710002000018.jpg52 at the listener's left and right ears. Conceptually, b is a 2×1 vector of filters (one filter per ear), but more conveniently, it is treated as a 2×1 vector of complex values at a specific frequency. Proceeding with this representation at a specific frequency, the desired binaural response can be retrieved from a set of HRTFs indexed by the object position.

Number

[0365] At the same time, the 2×1 binaural response e generated by the loudspeakers at the listener's ears is It is modeled as the product of the acoustic propagation matrix H of 2×M and the vector g of M×1 of the composite speaker activation values.

Number

[0366] The acoustic propagation matrix H is modeled based on the positions of a set of loudspeakers relative to the position of the listener. It is modeled based on JPEG0007710002000021.jpg56. Finally, the spatial component of the cost function is defined as the mean squared error between the desired binaural response (Equation 5) and the binaural response generated by the loudspeakers (Equation 6).

Number

[0367] For convenience, both spatial terms of the cost function for CMAP and FV defined in Equations 4 and 7 can be rearranged into a matrix quadratic form as a function of the speaker activation g.

Number

[0368] For this purpose, a second term of the cost function, JPEG0007710002000028.jpg515, can be defined as the weighted sum of the squared absolute values of the speaker activations by distance. This can be concisely expressed in matrix form as follows.

Equation

Equation

[0369] The distance penalty function can take many forms, but the following is a useful parameterization.

Equation

[0370] Combining the two terms of the cost function defined in Equations 8 and 9a yields the total cost function.

Number

[0371] Setting the derivative of this cost function with respect to g equal to zero and solving for g gives the optimal speaker activation solution.

Number

[0372] In general, the optimal solution in Equation 11 can produce speaker activations with negative values. For CMAP construction of a flexible renderer, such negative activations may not be desired, and thus Equation (11) can be minimized such that all activations remain positive.

[0373] Pairing a flexible rendering method (implemented according to some embodiments) with a set of wireless smart speakers (or other smart audio devices) can generate a highly capable and user-friendly spatial audio rendering system. Considering the interaction with such a system, it becomes apparent that it may be desirable to dynamically change the spatial rendering in order to optimize for other purposes that may arise during the use of the system. To achieve this goal, one class of embodiments extends an existing flexible rendering algorithm (where speaker activation is a function of the above spatial and proximity terms) using one or more properties of the audio signal being rendered, a set of speakers, and / or one or more additional dynamically configurable functions that depend on other external inputs. According to some embodiments, the existing cost function for flexible rendering given in Equation 1 is extended using these one or more additional dependencies described below.

Equation

[0374] In Equation 12, the term JPEG0007710002000043.jpg831 represents an additional cost term. Here, JPEG0007710002000044.jpg55 represents one or more properties of a set of the audio signal being rendered (e.g., the audio signal of an object-based audio program), JPEG0007710002000045.jpg56 represents one or more properties of a set of the speakers on which the audio is being rendered, JPEG0007710002000046.jpg55 represents one or more additional external inputs. Each term JPEG0007710002000047.jpg831 is the set Returns the cost as a function of activation g for one or more combinations of properties of the audio signal, speaker, and / or external input, comprehensively represented by JPEG0007710002000048.jpg821. Set JPEG0007710002000049.jpg821 includes at least JPEG0007710002000050.jpg55, JPEG0007710002000051.jpg56, or It should be understood to include only one element from any of JPEG0007710002000052.jpg55.

[0375] Examples of JPEG0007710002000053.jpg55 include, but are not limited to: ● The desired perceived spatial position of the audio signal, ● The level of the audio signal (possibly time-varying), and / or ● The spectrum of the audio signal (possibly time-varying).

[0376] Examples of JPEG0007710002000054.jpg56 include, but are not limited to: ● The position of the loudspeaker in the listening space, ● The frequency response of the loudspeaker, ● The playback level limit of the loudspeaker, ● Parameters of the dynamics processing algorithm in the speaker, such as limiter gain, ● Measured or estimated values of acoustic transmission from each speaker to other speakers, ● A measure of the echo canceller ability on the speaker, and / or ● The relative synchronization of the speakers with respect to each other.

[0377] Examples of JPEG0007710002000055.jpg55 include, but are not limited to: ● The position of one or more listeners or speakers within the playback space, ● Measured or estimated values of acoustic transmission from each loudspeaker to the listening position, ● Measured or estimated values of acoustic transmission from the speaker to a set of loudspeakers, ● The position of some other landmark within the playback space, and / or ● Measured or estimated values of acoustic transmission from each speaker to some other landmark within the playback space.

[0378] Using the new cost function defined in Equation 12, an optimal set of activations can be found via minimization with respect to g and the possible post-normalizations described above in Equations 2a and 2b.

[0379] Similar to the proximity cost defined in Equations 9a and 9b, each of the new cost function terms It is convenient to express JPEG0007710002000056.jpg831 as the weighted sum of the squared absolute values of the speaker activations.

Number

Number

[0380] By combining Equations 13a and 13b with the CMAP and FV cost functions given in Equation 10 represented as matrix quadratics, a potentially advantageous implementation example of the general extended cost function (for some embodiments) given in Equation 12 is produced.

Number

[0381] Using this definition of the new cost function term, the total cost function remains a matrix quadratic form, and the optimal set of activations JPEG0007710002000062.jpg57 can be found as follows through the differentiation of Equation 14.

Number

[0382] Weight term Each one of JPEG0007710002000064.jpg56 is considered as a function of a given continuous penalty value for each one of the loudspeakers. JPEG0007710002000065.jpg839 It is useful to consider it as a function. In one exemplary embodiment, this penalty value is the distance from the object (rendering target) to the loudspeaker of interest. In another exemplary embodiment, this penalty value represents the inability of a given loudspeaker to reproduce a certain frequency. Based on this penalty value, the weight term JPEG0007710002000066.jpg89 can be parameterized as follows.

Number

[0383] When all speakers receive a penalty, it is often convenient to subtract a minimum penalty from all weight terms in post - processing so that at least one of those speakers does not receive a penalty. [Number]

[0384] As described above, there are many possible use cases that can be realized using the new cost function terms described in this specification (and similar new cost function terms used according to other embodiments). Next, three examples are used to explain more specific details. That is, moving the audio towards the listener or speaker, moving the audio away from the listener or speaker, and moving the audio away from the landmark.

[0385] In the first example, what is referred to as "gravity" in this specification is used to pull the audio towards a certain position. This position can be, in some examples, the position of the listener or speaker, the position of the landmark, the position of the furniture, etc. This position can be referred to as the "gravity position" or "gravity source position" in this specification. As used in this specification, "gravity" is a factor that favors relatively higher loudspeaker activation in the vicinity of the gravity position. According to this example, the weight JPEG0007710002000081.jpg55 takes the form of Equation 17. In Equation 17, JPEG0007710002000082.jpg55 is a continuous penalty value given by the distance of the i-th speaker from the fixed gravity source position JPEG0007710002000083.jpg83, and JPEG0007710002000084.jpg53 is a threshold value given by the maximum value of these distances over all speakers.

Number

Number

[0386] To illustrate the use case of "pulling" the audio towards the listener or speaker, specifically, JPEG0007710002000087.jpg54 = 20, JPEG0007710002000088.jpg54 = 3, and Set JPEG0007710002000089.jpg83 to a vector corresponding to a listener / speaker position of 180 degrees. JPEG0007710002000090.jpg54, JPEG0007710002000091.jpg54, and These values of JPEG0007710002000092.jpg83 are only examples. In other implementation examples, JPEG0007710002000093.jpg54 is within the range of 1 to 100, JPEG0007710002000094.jpg54 can be within the range of 1 to 25.

[0387] In a second example, "repulsive force" is used to "push" the audio away from a certain position. This position can be the listener's position, the speaker's position, or another position such as the position of a landmark or furniture. This position can be referred to as the "repulsive force position" or "repulsive position" in this specification. As used in this specification, "repulsive force" is a factor advantageous for relatively lower loudspeaker activation closer to the repulsive force position. According to this example, similar to the attractive force in Equation 19, a fixed repulsive position For JPEG0007710002000095.jpg83 JPEG0007710002000096.jpg55 and JPEG0007710002000097.jpg53 are defined.

Number

Number

[0388] To illustrate the use case of pushing the audio away from the listener or speaker, specifically, JPEG0007710002000100.jpg54 = 5, JPEG0007710002000101.jpg54 = 2, and JPEG0007710002000102.jpgSet 83 to a vector corresponding to the listener / speaker position of 180 degrees. JPEG0007710002000103.jpg54, JPEG0007710002000104.jpg54, and JPEG0007710002000105.jpgThese values of 83 are just examples.

[0389] Returning now to FIG. 19. In this example, block 1925 includes changing the rendering process for the first audio signal based at least in part on at least one of the second audio signal, the second rendered audio signal, or its characteristics to generate a modified first rendered audio signal. Various examples of changing the rendering process are disclosed herein. The "characteristics" of the rendered signal may be, for example, either silent or in the presence of one or more additional rendered signals, but may include loudness or audibility estimated or measured at the intended listening position. Other examples of characteristics include the intended spatial position of the constituent signals of the associated program stream, the position of the loudspeaker at which the signal is rendered, the relative activation of the loudspeaker as a function of the intended spatial position of the constituent signals, and any other parameters or states related to the rendering of the signal, such as any other parameter related to the rendering algorithm utilized to generate the rendered signal. In some examples, block 1925 may be performed by the first rendering module.

[0390] According to this example, block 1930 includes changing the rendering process for the second audio signal based at least in part on at least one of the first audio signal, the first rendered audio signal, or its characteristics to generate a modified second rendered audio signal. In some examples, block 1930 may be performed by a second rendering module.

[0391] In some implementations, changing the rendering process for the first audio signal may include warping the rendering of the first audio signal away from the rendering position of the second rendered audio signal and / or changing the loudness of one or more signals of the first rendered audio signal in response to the loudness of one or more signals of the second audio signal or the second rendered audio signal. Alternatively, or in addition, changing the rendering process for the second audio signal may include warping the rendering of the second audio signal away from the rendering position of the first rendered audio signal and / or changing the loudness of one or more signals of the second rendered audio signal in response to the loudness of one or more signals of the first audio signal or the first rendered audio signal. Some examples are given below with reference to the figures after FIG. 3.

[0392] However, other types of rendering processing changes are also within the scope of the present disclosure. For example, in some cases, changing the rendering processing for the first audio signal or the second audio signal may include making spectral changes, audible-based changes, or dynamic range changes. These changes may or may not be related to loudness-based rendering depending on the specific example. For example, in the above case where the main spatial stream is rendered in an open-plan living area and a secondary stream consisting of cooking tips is rendered in an adjacent kitchen, it may be desirable to ensure that the cooking tips remain audible in the kitchen. This can be achieved by estimating how loud the stream of cooking tips rendered in the kitchen would be in the absence of interfering first signals, then estimating the loudness when the first signal is present in the kitchen, and finally dynamically changing the loudness and dynamic range of both streams over multiple frequencies to ensure the audibility of the second signal in the kitchen.

[0393] In the example shown in FIG. 19, block 1935 includes at least mixing the changed first rendered audio signal and the changed second rendered audio signal to generate a mixed audio signal. Block 1935 can be performed, for example, by mixer 1830b illustrated in FIG. 18B.

[0394] According to this example, block 1940 includes providing the mixed audio signal to at least some speakers of the environment. Some examples of method 1900 may include playing the mixed audio signal by speakers.

[0395] As shown in FIG. 19, some implementation examples may provide more than two rendering modules. Some such implementation examples may provide N rendering modules. Here, N is an integer greater than 2. Thus, some such implementation examples may include one or more additional rendering modules. In some such examples, each of the one or more additional rendering modules may be configured to receive an additional audio program stream via an interface system. The additional audio program stream may include an additional audio signal scheduled to be played by at least one speaker of the environment. Some such implementation examples may include rendering an additional audio signal for playback via at least one speaker of the environment to generate an additional rendered audio signal, and changing the rendering process for the additional audio signal based at least in part on at least one of the first audio signal, the first rendered audio signal, the second audio signal, the second rendered audio signal, or their characteristics to generate a modified additional rendered audio signal. According to some such examples, the mixing module may be configured to mix the modified additional rendered audio signal using at least the modified first rendered audio signal and the modified second rendered audio signal to generate a mixed audio signal.

[0396] As described above with reference to FIGS. 6 and 18B, some implementation examples may include a microphone system that includes one or more microphones within a listening environment. In some such examples, a first rendering module may be configured to change a rendering process for a first audio signal based at least in part on a first microphone signal from the microphone system. The "first microphone signal" may be received from a single microphone or two or more microphones, depending on the particular implementation example. In some such implementation examples, a second rendering module may be configured to change a rendering process for a second audio signal based at least in part on the first microphone signal.

[0397] As described above with reference to FIG. 18B, in some cases, the position of one or more microphones may be known or provided to a control system. According to some such implementation examples, the control system estimates a first sound source position based on the first microphone signal and, based at least in part on the first sound source position, may be configured to change a rendering process for at least one of the first audio signal or the second audio signal. The first sound source position may be estimated, for example, according to triangulation processing based on each of three or more microphones having known positions, or based on each of a group of microphones. Alternatively, or in addition, the first sound source position may be estimated according to the amplitude of signals received from two or more microphones. The microphone that generates the highest amplitude signal may be presumed to be closest to the first sound source position. In some such examples, the first sound source position may be set to the position of the closest microphone. In some such examples, the first sound source position may be associated with the position of a zone. Here, the zone is selected from two or more microphones by a processed signal via a pre-trained classifier such as a Gaussian mixer model.

[0398] ​In some such implementation examples, the control system may be configured to determine whether the first microphone signal corresponds to ambient noise. Some such implementation examples may include changing a rendering process for at least one of the first audio signal or the second audio signal, at least partially based on whether the first microphone signal corresponds to ambient noise. For example, if the control system determines that the first microphone signal corresponds to ambient noise, changing the rendering process for the first audio signal or the second audio signal may include increasing the level of the rendered audio signal such that the loudness of the signal perceived at the intended listening position in the presence of noise is substantially equal to the loudness of the signal perceived in the absence of noise.

[0399] In some examples, the control system may be configured to determine whether the first microphone signal corresponds to a human voice. Some such implementations may include changing the rendering process for at least one of the first audio signal or the second audio signal, at least in part based on whether the first microphone signal corresponds to a human voice, such as a wake word. For example, if the control system determines that the first microphone signal corresponds to a human voice, changing the rendering process for the first audio signal or the second audio signal may include reducing the loudness of the rendered audio signal played by a speaker closer to the first sound source, compared to the loudness of the rendered audio signal played by a speaker farther from the first sound source. Changing the rendering process for the first audio signal or the second audio signal may alternatively or additionally include warping the intended position of the constituent signals of the associated program stream away from the first sound source and / or applying a penalty to the use of a speaker closer to the first sound source compared to a speaker farther from the first sound source, by changing the rendering process.

[0400] In some implementation examples, if the control system determines that the first microphone signal corresponds to a human voice, the control system may be configured to reproduce the first microphone signal at one or more speakers close to a position in an environment different from the first sound source position. In some such examples, the control system may be configured to determine whether the first microphone signal corresponds to the cry of a child. According to some such implementation examples, the control system may be configured to reproduce the first microphone signal at one or more speakers close to a position in the environment corresponding to the estimated position of a caregiver such as a parent, relative, guardian, childcare service provider, teacher, nurse, etc. In some examples, the process of estimating the estimated position of the caregiver may be triggered by a voice command such as "<wake word, don't wake the baby>". The control system determines the position of the nearest smart audio device implementing the virtual assistant by triangulation based on, for example, DOA information provided by three or more local microphones Accordingly, the position of the speaker (caregiver) could be estimated. According to some implementation examples, the control system may know in advance the position of the baby's room (and / or the listening device therein) and then be able to perform appropriate processing.

[0401] According to some such examples, the control system may be configured to determine whether the first microphone signal corresponds to a command. If the control system determines that the first microphone signal corresponds to a command, in some cases, the control system may be configured to determine a response to the command and control at least one speaker close to the first sound source position to reproduce the response. In some such examples, after the control system controls at least one speaker close to the first sound source position to reproduce the response, the control system may be configured to return to the unchanged rendering process for the first audio signal or the second audio signal.

[0402] In some implementation examples, the control system may be configured to execute commands. For example, the control system may be or include a virtual assistant configured to control audio devices, televisions, household appliances, etc. according to commands.

[0403] In the case of this definition of the minimal and more capable multi-stream rendering system illustrated in FIGS. 6, 18A, and 18B, for many useful scenarios, dynamic management of the synchronized playback of multiple program streams can be achieved. Several examples will be described below with reference to FIGS. 20 and 21.

[0404] First, consider the above example that includes synchronously playing a spatial movie soundtrack in the living room and the tips for cooking in the kitchen following the living room. The spatial movie soundtrack is an example of the "first audio program stream" referred to above, and the audio of the cooking tips is an example of the "second audio program stream" referred to above. FIGS. 20 and 21 illustrate an example of a floor plan of a continuous living space. In this example, the living space 2000 includes a living room in the upper left, a kitchen in the middle lower, and a bedroom in the lower right. The boxes and circles 2005a to 2005h distributed throughout the living space represent a set of eight loudspeakers installed at convenient positions in the space, but do not adhere to any standard predetermined layout (arranged arbitrarily). In FIG. 20, only the spatial movie soundtrack is played, and all the loudspeakers in the living room 2010 and the kitchen 2015 are used to generate optimized spatial playback around the listener sitting on the sofa 2025 facing the television 2030, taking into account the capabilities and layout of the loudspeakers. This optimal playback of the movie soundtrack is visually represented by a cloud 2035a located within the range of the active loudspeakers.

[0405] In FIG. 21, for the second listener 2020b, cooking tips are simultaneously rendered and played back via a single loudspeaker 2005g in the kitchen 2015. The playback of this second program stream is visually represented by a cloud 2140 emanating from the loudspeaker 2005g. As shown in FIG. 20, if these cooking tips are played back simultaneously without changing the rendering of the movie soundtrack, the audio from the movie soundtrack coming from the speaker within or near the kitchen 2015 would interfere with the second listener's ability to understand the cooking tips. Instead, in this example, the rendering of the spatial movie soundtrack is dynamically changed as a function of the rendering of the cooking tips. Specifically, the rendering of the movie soundtrack is shifted away from the speakers near the location of the rendering of the cooking tips (kitchen 2015). This shift is visually represented in FIG. 21 by a smaller cloud 2035b being pushed away from the speakers near the kitchen. If the playback of the cooking tips stops while the movie soundtrack is still playing, in some implementations, the rendering of the movie soundtrack can be dynamically shifted back to the original optimal configuration seen in FIG. 20. Such dynamic shifts in the rendering of the spatial movie soundtrack can be achieved via many of the disclosed methods. Such dynamic shifts in the rendering of the spatial movie soundtrack can be achieved via many of the disclosed methods.

[0406] Many spatial audio mixes include a plurality of component audio signals designed to be reproduced at specific locations within the listening space. For example, Dolby 5.1 and 7.1 surround sound mixes consist of, respectively, six and eight signals intended to be reproduced on speakers at predetermined canonical positions around the listener. Object-based audio formats, such as Dolby Atmos, consist of component audio signals having associated metadata that describes a probably time-varying 3D position in the listening space where the audio is intended to be rendered. Assuming that a renderer of a spatial movie soundtrack can render individual audio signals at arbitrary positions with respect to an arbitrary set of loudspeakers, the dynamic shifts to rendering illustrated in FIGS. 20 and 21 can be achieved by warping the intended positions of the audio signals within the spatial mix. For example, the 2D or 3D coordinates associated with an audio signal can be pushed away from the position of a speaker in the kitchen or alternatively pulled towards the upper left corner of the living room. The result of such a warp is, here, that the use of speakers closer to the kitchen is reduced because the warped positions of the audio signals of the spatial mix are further away from this position. This method achieves the goal of making the second audio stream more understandable to a second listener, but it is done at the expense of significantly altering the intended spatial balance of the movie soundtrack for the first listener.

[0407] A second way to achieve a dynamic shift for spatial rendering can be realized using a flexible rendering system. In some such implementations, the flexible rendering system can be a hybrid of CMAP, FV, or both, as described above. Some such flexible rendering systems attempt to reproduce a spatial mix using all constituent signals that are perceived to come from the intended positions. While doing so for each signal of the mix, in some examples, the activation of loudspeakers in the vicinity of the desired position of that signal is prioritized. In some implementations, additional terms can be dynamically added to the optimization of the rendering. The additional terms penalize the use of a given loudspeaker based on other criteria. For the example here, something that could be called a "repulsive force" is dynamically placed at the position of the kitchen, which can effectively push away the rendering of a spatial movie soundtrack by imposing a high penalty on the use of loudspeakers near this position. As used herein, the term "repulsive force" can refer to a factor corresponding to a relatively lower speaker activation within a particular position or area of the listening environment. In other words, the term "repulsive force" can refer to a factor that is advantageous for the activation of speakers that are relatively farther away from a particular position or area corresponding to the "repulsive force". However, according to some such implementations, the renderer can still attempt to reproduce the intended spatial balance of the mix using the remaining speakers with a lower given penalty. Therefore, this approach can be considered an excellent way to achieve a dynamic shift in rendering compared to simply warping the intended positions of the constituent signals of the mix.

[0408] The above scenario of shifting the rendering of a spatial movie soundtrack away from the cooking tips within the kitchen is the optimum of the multi-stream renderer illustrated in FIG. 18A. It can be achieved using a limited version. However, this scenario can be improved using a more capable system as illustrated in FIG. 18B. Shifting the rendering of the spatial movie soundtrack improves the intelligibility of the cooking tips in the kitchen, but the movie soundtrack can still be significantly audible in the kitchen. Depending on the instantaneous state of both streams, the cooking tips can be masked by the movie soundtrack. For example, a loud moment in the movie soundtrack masks a soft moment in the cooking tips. To address this issue, a dynamic change to the rendering of the cooking tips as a function of the rendering of the spatial movie soundtrack can be added. For example, a method can be performed to dynamically modify the audio signal over frequency and time to preserve the loudness perceived in the presence of interfering signals. In this scenario, an estimated value of the perceived loudness of the shifted movie soundtrack at the location of the kitchen is generated and fed into such a process as an interfering signal. Then, the level of the cooking tips that varies over time and frequency can be dynamically changed to maintain the loudness perceived above this interference. This results in better maintained intelligibility for a second listener. The estimated value required for the loudness of the movie soundtrack in the kitchen can be generated from the speaker feed of the soundtrack rendering, signals from microphones within or near the kitchen, or a combination thereof. This process of maintaining the perceived loudness of the cooking tips will generally result in raising the level of the cooking tips, and in some cases, the overall loudness can become undesirably high. To address this issue, yet another rendering change can be employed. The interfering spatial movie soundtrack can be dynamically reduced as a function of the cooking tips with modified loudness that becomes overly large in the kitchen. Finally, there is a possibility that an external noise source can interfere simultaneously with the audibility of both program streams. For example, a blender can be used while cooking in the kitchen.Estimated values of the loudness of this environmental noise source in both the living room and the kitchen can be generated from a microphone connected to the rendering system. This estimated value can be added to the estimated value of the loudness of the sound track in the kitchen, for example, to affect the loudness change of the cooking tips. At the same time, the rendering of the sound track in the living room can be further changed as a function of the estimated value of the environmental noise in the living room to maintain the perceived loudness of the sound track in the living room in the presence of this environmental noise. This way, better audibility is maintained for the listener in the living room.

[0409] It can be seen that this use case of the disclosed multi - stream renderer uses many interrelated changes for two program streams to optimize synchronous playback. In summary, these changes to the stream can include the following.

[0410] ● Spatial movie sound track 〇 Spatial rendering shifted away from the kitchen as a function of the cooking tips being rendered in the kitchen 〇 Dynamic reduction of loudness as a function of the loudness of the cooking tips being rendered in the kitchen 〇 Dynamic increase of loudness as a function of the estimated value of the loudness of interfering blender noise from the kitchen in the living room

[0411] ● Cooking tips 〇 Dynamic increase of loudness as a function of the combined estimated value of the loudness of both the movie sound track and the blender noise in the kitchen

[0412] A second use case of the disclosed multi-stream renderer involves synchronously playing a spatial program stream, such as music, while a smart voice assistant responds to a query by a user. In existing smart speakers where playback has generally been restricted to monaural or stereo playback via a single device, interaction with the voice assistant typically consists of the following steps. 1) Playback of music 2) The user utters the voice assistant wake word 3) The smart speaker recognizes the wake word and reduces (lowers) the music by a significant amount 4) The user utters a command to the smart assistant (i.e., "Play the next song") 5) The smart speaker affirms this by recognizing the command and playing, via the speaker, some voice response (i.e., "OK, playing the next song") mixed over the music at the lowered volume, and then executes the command 6) The smart speaker returns the music to its original volume

[0413] Figures 22 and 23 illustrate an example of a multi-stream renderer that provides simultaneous playback of a spatial music mix and a voice assistant response. When playing spatial audio via a number of orchestrated smart speakers, some embodiments improve upon the above series of events. Specifically, the spatial mix can be shifted away from one or more of the speakers appropriately selected to relay the response from the voice assistant. Generating this space for the voice assistant response means that the reduction of the spatial mix may be less, or perhaps not reduced at all, compared to the existing state of the events listed above. Figures 22 and 23 illustrate this scenario. In this example, the modified series of events can occur as follows. 1) For the user cloud 2035c in Figure 22, a spatial music program stream is being played via a number of orchestrated smart speakers. 2) User 2020c vocalizes the voice assistant wake word. 3) One or more smart speakers (e.g., speaker 2005d and / or speaker 2005f) recognize the wake word and determine the position of user 2020c or which speaker(s) (singular or plural) user 2020c is closest to using relevant recordings from the microphone(s) that cooperate with the one or more smart speakers. 4) The rendering of the spatial music mix is shifted away from the position as if it will be rendered near the position determined in the previous step by the voice assistant response program stream (cloud 2035d in FIG. 23). 5) The user vocalizes a command to the smart assistant (e.g., the smart speaker that operates the smart assistant / voice assistant software). 6) The smart speaker recognizes the command, synthesizes the corresponding response program stream, and renders the response near the user's position (cloud 2340 in FIG. 23). 7) Once the voice assistant response is complete, the rendering of the spatial music program stream is shifted back to its original state (cloud 2035c in FIG. 22).

[0414] In addition to optimizing the synchronous playback of the spatial music mix and the voice assistant response, shifting the spatial music mix can also improve the ability of a set of speakers to understand the listener in step 5. This is because as the music is shifted away from the speakers near the listener, the voice-to-other ratio of the cooperating microphones improves.

[0415] The same as described for the above scenario using a spatial movie mix and cooking tips Likewise, this scenario can be further optimized than by shifting the rendering of the spatial mix as a function of the voice assistant response. Shifting the spatial mix alone may not be sufficient to fully enable the user to understand the voice assistant response. A simple solution is again to reduce the spatial mix by a certain amount, but less than what is required using the current state of the event. Alternatively, the loudness of the voice assistant response program stream can be dynamically increased as a function of the loudness of the spatial music mix program stream to maintain audibility of the response. As an extension, the loudness of the spatial music mix can also be dynamically reduced if this increase process for the response stream becomes overly large.

[0416] Figures 24, 25, and 26 illustrate a third use case for the disclosed multi - stream renderer. This example involves attempting to hear if an infant is crying while leaving the infant asleep in an adjacent room, and at the same time managing the synchronized playback of a spatial music mix program stream and a comfort noise program stream. Figure 24 illustrates the starting point where a spatial music mix (represented by cloud 2035e) is optimally playing on all speakers within the living room 2010 and kitchen 2015 for a number of people at a party. In Figure 25, the infant 2510 is currently attempting to sleep in the adjacent bedroom 2505 shown in the lower right. To assist in ensuring this, the spatial music mix is dynamically shifted away from the bedroom to minimize leakage into the bedroom, as illustrated by cloud 2035f, while still maintaining a reasonable experience for the people at the party. At the same time, a second program stream (represented by cloud 2540) including soothing white noise is played from speaker 2005h within the infant's room to mask any remaining leakage from the music in the adjacent room. To ensure complete masking, the loudness of this white noise stream can be dynamically changed in some examples as a function of an estimated value of the loudness of the spatial music leaking into the infant's room. This estimated value can be generated from the speaker feeds of the rendering of the spatial music, signals from microphones within the infant's room, or a combination thereof. Also, the loudness of the spatial music mix can be attenuated as a function of the noise with its loudness changed if it becomes overly loud. This is similar to the loudness processing between the spatial movie mix and the cooking tip in the first scenario. Finally, a microphone within the infant's room (e.g., a microphone that works in conjunction with speaker 2005h and which in some implementations could be a smart speaker) can be configured to record audio from the infant (spatial music and white noise and any canceling sounds that can be picked up from the white noise).Next, the combination of these processed microphone signals can function as a third program stream that can be simultaneously reproduced near the listener 2020d (who can be a parent or other caregiver) in the living room 2010 when a crying sound is detected (e.g., by machine learning via a pattern matching algorithm, etc.). FIG. 26 illustrates the reproduction of this additional stream using the cloud 2650. In this case, the spatial music mix can further be shifted away from the speaker near the parent who is reproducing the baby's crying sound. This is illustrated by the shape of the cloud 2035g, which is changed with respect to the shape of the cloud 2035f in FIG. 25. The program stream of the baby's crying sound can have its loudness changed as a function of the spatial music stream so that the baby's crying sound remains audible to the listener 2020d. The related changes that optimize the synchronous reproduction of the three program streams considered in this example can be summarized as follows.

[0417] ●Spatial music mix in the living room 〇Spatial rendering shifted away from the room to reduce transmission to the baby's room 〇Dynamic reduction of loudness as a function of the white noise rendered in the baby's room Dynamic reduction of loudness 〇Spatial rendering shifted away from the parent as a function of the baby's crying sound rendered on the speaker near the parent

[0418] ●White noise 〇Dynamic increase of loudness as a function of the estimated loudness of the music stream leaking into the baby's room

[0419] ●Recording of the baby's crying sound 〇Dynamic increase of loudness as a function of the estimated loudness of the music mix at the position of the parent or other caregiver

[0420] Next, examples of how some of the above embodiments can be implemented will be described.

[0421] In FIG. 18A, each of the render blocks 1...N can be implemented as an instance of any single-stream renderer such as the above CMAP, FV, or hybrid renderer. Constructing a multi-stream renderer in this way has several convenient and useful properties.

[0422] First, when rendering is performed in this hierarchical arrangement and each instance of the single-stream renderer is configured to operate in the frequency / transform domain (e.g., QMF), mixing of the streams can also occur in the frequency / transform domain, and for M channels, the inverse transform only needs to be performed once. This is significantly more efficient than performing N×M inverse transforms and mixing in the time domain.

[0423] FIG. 27 illustrates a frequency / transform domain example of the multi-stream renderer shown in FIG. 18A. In this example, after an orthogonal mirror analysis filter bank (QMF) is applied to each of the program streams 1 to N, each program stream is received by a corresponding one of the rendering modules 1 to N. According to this example, the rendering modules 1 to N operate in the frequency domain. After mixer 1830a mixes the outputs of the rendering modules 1 to N, inverse synthesis filter bank 2735a converts the mix to the time domain and provides the mixed speaker feed signal in the time domain to loudspeakers 1 to M. In this example, the orthogonal mirror filter bank, the rendering modules 1 to N, the mixer 1830a, and the inverse filter bank 2735a are components of the control system 610c.

[0424] FIG. 28 illustrates a frequency / transform domain example of the multi-stream renderer shown in FIG. 18B. Similar to FIG. 27, after an orthogonal mirror filter bank (QMF) is applied to each of program streams 1 to N, each program stream is received by a corresponding one of rendering modules 1 to N. According to this example, rendering modules 1 to N operate in the frequency domain. In this implementation example, the time-domain microphone signal from the microphone system 620b is also provided to the orthogonal mirror filter bank, and rendering modules 1 to N receive the microphone signal in the frequency domain. After mixer 1830b mixes the outputs of rendering modules 1 to N, inverse filter bank 2735b converts the mix into the time domain and provides the mixed speaker feed signal in the time domain to loudspeakers 1 to M. In this example, the orthogonal mirror filter bank, rendering modules 1 to N, mixer 1830b, and inverse filter bank 2735b are components of the control system 610d.

[0425] Referring to FIG. 29, another exemplary embodiment will be described. Similar to the other figures given herein, the types and numbers of elements illustrated in FIG. 29 are given only by way of example. Other implementation examples may include more, fewer, and / or different types and numbers of elements. FIG. 29 illustrates a floor plan of a listening environment that is a living space in this example. According to this example, the environment 2000 includes a living room 2010 in the upper left, a kitchen 2015 in the middle lower, and a bedroom 2505 in the lower right. The boxes and circles distributed throughout the living space represent a set of loudspeakers 2005a-2005h in some implementation examples that are installed in convenient positions in the space, at least some of which may be smart speakers in some implementation examples and do not adhere to any standard predetermined layout (are arbitrarily arranged). In some examples, the loudspeakers 2005a-2005h may be coordinated to implement one or more of the disclosed embodiments. For example, in some embodiments, the loudspeakers 2005a-2005h may be coordinated according to commands from a device implementing an audio session manager that may be CHASM in some examples. In some such examples, the audio processing disclosed above includes, but is not limited to, the flexible rendering function disclosed above, and may be implemented at least in part according to instructions from CHASM 208C, CHASM 208D, CHASM 307, and / or CHASM 401 described above with reference to FIGS. 2C, 2D, 3C, and 4. In this example, the environment 2000 includes cameras 2911a-2911e distributed throughout the environment. In some implementation examples, also, one or more smart audio devices within the environment 2000 may include one or more cameras. The one or more smart audio devices may be single-purpose audio devices or voice assistants.In some such examples, one or more cameras (which may be the cameras of the optional sensor system 630 of FIG. 6) may be present within or on the television 2030, within a mobile phone, or within a smart speaker such as one or more of the loudspeakers 2005b, 2005d, 2005e, or 2005h. Although cameras 2911a - 2911e are not shown in all of the figures of the listening environments presented in this disclosure, each listening environment including but not limited to environment 2000 may include one or more cameras in some implementation examples.

[0426] FIGS. 30, 31, 32, and 33 illustrate examples of flexibly rendering spatial audio in a reference space mode for a plurality of different listening positions and orientations within the living space illustrated in FIG. 29. FIGS. 30 - 33 illustrate this ability for four example listening positions. In each example, an arrow 3005 pointing in the direction of person 3020a represents the position of the front sound stage (towards which person 3020a is facing). In each example, arrow 3010a represents the left surround field and arrow 3010b represents the right surround field.

[0427] In FIG. 30, for person 3020a sitting on the living room sofa 2025, the reference space mode is determined (e.g., by a device implementing an audio session manager), and the spatial audio is flexibly rendered. In the example shown in FIG. 30, considering the loudspeaker capabilities and layout, all of the loudspeakers within the living room 2010 and kitchen 2015 are used to generate an optimized spatial reproduction of the audio data around the listener 3020a. This optimal reproduction is visually represented by a cloud 3035 that exists within the range of the active loudspeakers.

[0428] According to some implementation examples, a control system configured to implement an audio session manager (such as the control system 610 in FIG. 6) may be configured to determine the estimated listening position and / or the estimated orientation of the reference spatial mode according to the reference spatial mode data received via an interface system such as the interface system 605 in FIG. 6. Some examples will be described later. In some such examples, the reference spatial mode data may include microphone data from a microphone system (such as the microphone system 120 in FIG. 6). According to some such examples, the reference spatial mode data may include microphone data corresponding to wake words and voice commands such as "[Wake word], set the TV to the front sound stage". Alternatively, or in addition, microphone data may be used to triangulate the user's position according to the sound of the user's voice, for example, via direction of arrival (DOA) data. For example, three or more of the loudspeakers 2005a to 2005e may use microphone data and, via DOA data, triangulate the position of the person 3020a sitting on the living room sofa 2025 according to the sound of the voice of the person 3020a. According to the position of the person 3020a, the orientation of the person 3020a may be estimated. When the person 3020a is in the position shown in FIG. 30, it may be estimated that the person 3020a is facing the TV 2030.

[0429]

[0430] Alternatively, or in addition, the position and orientation of the person 3020a may be determined according to image data from a camera system (such as the sensor system 130 in FIG. 6).

[0431] ​In some examples, the position and orientation of person 3020a can be determined according to user input obtained via a graphical user interface (GUI). According to some such examples, the control system can be configured to control a display device (e.g., a display device of a mobile phone) to present a GUI that enables person 3020a to input their position and orientation.

[0432] In FIG. 31, a reference space mode is determined for person 3020a sitting in the reading chair 3115 in the living room, and spatial audio is rendered flexibly. In FIG. 32, a reference space mode is determined for person 3020a standing next to the kitchen counter 330, and spatial audio is rendered flexibly. In FIG. 33, a reference space mode is determined for person 3020 sitting at the breakfast table 340, and spatial audio is rendered flexibly. As indicated by arrow 3005, it can be seen that the orientation of the front sound stage does not necessarily correspond to any particular loudspeaker within environment 2000. As the position and orientation of the listener change, the roles of the speakers that render the various components of the spatial mix also change.

[0433] In each of FIGS. 30 - 33, the person 3020a listens to the intended spatial mix for each of the illustrated positions and orientations. However, the experience may be sub - optimal for additional listeners within the space. FIG. 34 illustrates an example of a reference spatial mode rendering when two listeners are in different positions in the listening environment. FIG. 34 illustrates a reference spatial mode rendering for a person 3020a on a sofa and a person 3020b standing in the kitchen. The rendering is optimal for the person 3020a, but for the person 3020b, considering that person's position, mainly signals from the surround field will be audible, while signals from the front sound stage will be barely audible. In this case, and in other cases where multiple people move around unpredictably within the space (e.g., a party), a more appropriate rendering mode for such a dispersed audience is needed. Examples of such a dispersed - space rendering mode are described with reference to FIGS. 4B - 9 on pages 27 - 43 of U.S. Provisional Patent Application No. 62 / 705,351, filed on June 23, 2020 (inventive name: "ADAPTABLE SPATIAL AUDIO PLAYBACK"), which application is incorporated herein by reference.

[0434] FIG. 35 illustrates an example of a GUI for receiving user input regarding the position and orientation of a listener. According to this example, the user has previously specified several possible listening positions and corresponding orientations. The positions of the loudspeakers corresponding to each position and the corresponding orientation have already been input and stored during the setup process. Several examples are disclosed herein. A detailed example of the automatic position processing of the audio device will be described below. For example, a listening environment layout GUI may be provided, and the user may be prompted to touch the position corresponding to the possible listening position and the position of the speaker, or may be prompted to say the name of the possible listening position. In this example, at the time illustrated in FIG. 35, the user may already have provided user input to the GUI 3500 regarding the user's position by touching the virtual button "sofa in the living room". Considering the L-shaped sofa 2025, there are two possible forward-facing positions, so the user is prompted to indicate which direction they are facing.

[0435] FIG. 36 illustrates an example of the geometric relationship between three audio devices in an environment. In this example, the environment 3600 is a room including a television 3601, a sofa 3603, and five audio devices 3605. According to this example, the audio devices 3605 are within positions 1-5 of the environment 3600. In this implementation example, each of the audio devices 3605 includes a microphone system 3620 having at least three microphones and a speaker system 3625 including at least one speaker. In some implementation examples, each microphone system 3620 includes an array of microphones. According to some implementation examples, each of the audio devices 3605 may include an antenna system including at least three antennas.

[0436] As with other examples disclosed herein, the type, number, and arrangement of the elements illustrated in FIG. 36 are merely examples. Other implementation examples may have different types, numbers, and arrangements of elements, such as more or fewer audio devices 3605, audio devices 3605 at different positions, and the like.

[0437] In this example, triangle 3610a has vertices at positions 1, 2, and 3. Here, triangle 3610a has sides 12, 23a, and 13a. According to this example, the angle between sides 12 and 23 is JPEG0007710002000106.jpg54, and the angle between sides 12 and 13a is JPEG0007710002000107.jpg54, and the angle between sides 23a and 13a is JPEG0007710002000108.jpg54. These angles can be determined according to the DOA data, although further details will be described later.

[0438] In some implementation examples, only the relative lengths of the sides of the triangle can be determined. In alternative implementation examples, the actual lengths of the sides of the triangle can be estimated. According to some such implementation examples, the actual lengths of the sides of the triangle are generated by an audio device located at a vertex of a triangle and estimated according to the arrival time of sound detected by an audio device located at a vertex of another triangle, for example, according to the TOA data. Alternatively, or in addition, the lengths of the sides of the triangle can be generated by an audio device located at a vertex of a triangle and estimated according to the electromagnetic wave detected by an audio device located at a vertex of another triangle. For example, the lengths of the sides of the triangle can be estimated according to the signal strength of the electromagnetic wave generated by an audio device located at a vertex of a triangle and detected by an audio device located at a vertex of another triangle and can be. In some implementation examples, the lengths of the sides of the triangle can be estimated according to the phase shift of the detected electromagnetic wave.

[0439] Figure 37 illustrates another example of the geometric relationship between three audio devices within the environment illustrated in Figure 36. In this example, triangle 3610b has vertices at positions 1, 3, and 4. Here, triangle 3610b has sides 13b, 14, and 34a. According to this example, the angle between sides 13b and 14 is JPEG0007710002000109.jpg54, and the angle between sides 13b and 34a is JPEG0007710002000110.jpg54, and the angle between sides 34a and 14 is JPEG0007710002000111.jpg54.

[0440] By comparing Figures 36 and 37, it can be seen that the length of side 13a of triangle 3610a should be equal to the length of side 13b of triangle 3610b. In some implementation examples, even if the length of the side of a certain triangle (for example, triangle 3610a) is assumed to be correct, the length of the side shared by the adjacent triangle will be restricted to this length.

[0441] Figure 38 illustrates both of the triangles illustrated in Figures 36 and 37, but omits the corresponding audio devices and other equipment of the environment. Figure 38 illustrates the estimated values of the side lengths and angular orientations of triangles 3610a and 3610b. In the example shown in Figure 38, the length of side 13b of triangle 3610b is restricted to the same length as side 13a of triangle 3610a. The lengths of the other sides of triangle 3610b are scaled in proportion to the resulting change in the length of side 13b. The resulting triangle 3610b’ is illustrated in Figure 38. Triangle 3610b’ is adjacent to triangle 3610a.

[0442] According to some implementation examples, the lengths of the sides of other triangles adjacent to triangles 3610a and 3610b can all be determined similarly until the positions of all the audio devices within environment 3600 are determined.

[0443] Some examples of the location of the audio device may proceed as follows. Each audio device may report the DOA of all other audio devices in the environment (e.g., a room) based on the sound generated by all other audio devices in the environment (e.g., according to an instruction from a device implementing an audio session manager such as CHASM). The Cartesian coordinates of the i-th audio device may be expressed as JPEG0007710002000112.jpg519. Here, the superscript T indicates a vector transpose. Assuming there are M audio devices in the environment, JPEG0007710002000113.jpg517.

[0444] Figure 39 illustrates an example of estimating the interior angles of a triangle formed by three audio devices. In this example, the audio devices are i, j, and k. The DOA of the sound source from device j observed from device i may be expressed as JPEG0007710002000114.jpg54. The DOA of the sound source from device k observed from device i may be expressed as JPEG0007710002000115.jpg55. In the example shown in Figure 39, JPEG0007710002000116.jpg54 and JPEG0007710002000117.jpg55 are measured from axis 3905a. The orientation of axis 3905a is arbitrary and may correspond to, for example, the orientation of audio device i. The interior angle a of triangle 3910 may be expressed as JPEG0007710002000118.jpg519. It can be seen that the calculation of the interior angle a does not depend on the orientation of axis 3905a.

[0445] In the example shown in Figure 39, JPEG0007710002000119.jpg55 and JPEG0007710002000120.jpg55 is measured from axis 3905b. The orientation of axis 3905b is arbitrary and may correspond, for example, to the orientation of audio device j. The interior angle b of triangle 3910 can be expressed as JPEG0007710002000121.jpg520. Similarly, in this example, JPEG0007710002000122.jpg55 and JPEG0007710002000123.jpg55 are measured from axis 3905c. The interior angle c of triangle 3910 can be expressed as JPEG0007710002000124.jpg519.

[0446] If there is a measurement error, JPEG0007710002000125.jpg525. For example, robustness can be improved by predicting each angle from the other two angles and then averaging them as follows.

Number

[0447] In some implementation examples, the length of the edge JPEG0007710002000127.jpg512 can be calculated (with a maximum scaling error) by applying the sine law. In some examples, an arbitrary value such as 1 can be assigned to the length of a certain edge. For example, Set it as JPEG0007710002000128.jpg59, and by placing vector JPEG0007710002000129.jpg517 at the origin, the positions of the remaining two vertices can be calculated as follows.

Number

[0448] However, any rotation may be allowed.

[0449] According to some implementation examples, the process of triangular parameterization is of size a superset of JPEG0007710002000131.jpg814 JPEG0007710002000132.jpg52 can be repeated for all possible subsets of the three audio devices in the environment listed. In some examples, JPEG0007710002000133.jpg53 may represent the l-th triangle. According to the implementation example, the triangles may not be enumerated in a specific order. The triangles may overlap and may not be perfectly aligned due to possible errors in the estimated values of the DOA and / or side lengths.

[0450] Figure 40 is a flowchart showing an overview of an example of a method that can be performed by an apparatus as illustrated in Figure 6. The blocks of method 4000 are not necessarily performed in the order shown, similar to other methods described herein. Furthermore, such a method may include more or fewer blocks than those illustrated and / or described. In this implementation example, method 4000 includes estimating the position of speakers in the environment. The blocks of method 4000 can be performed by one or more devices that can be (or include) the device 600 illustrated in Figure 6. According to some implementation examples, the blocks of method 4000 can be performed at least in part by a device implementing an audio session manager (e.g., CHASM) and / or in accordance with instructions from a device implementing an audio session manager. In some such examples, the blocks of method 4000 can be performed at least in part by CHASM208C, CHASM208D, CHASM307, and / or CHASM401 described above with reference to Figures 2C, 2D, 3C, and 4. According to some implementation examples, the blocks of method 4000 can be performed as part of the setup process of method 1500 described above with reference to Figure 15.

[0451] In this example, block 4005 includes obtaining direction of arrival (DOA) data for each of a plurality of audio devices. In some examples, the plurality of audio devices may include all of the audio devices in an environment, such as all of the audio devices 3605 illustrated in FIG. 36.

[0452] However, in some cases, the plurality of audio devices may include only a subset of all of the audio devices in the environment. For example, the plurality of audio devices may include all of the smart speakers in the environment, but may not include one or more of the other audio devices in the environment.

[0453] The DOA data may be obtained in various ways depending on the particular implementation. In some cases, determining the DOA data may include determining the DOA data for at least one of the plurality of audio devices. For example, determining the DOA data may include receiving microphone data from each of a plurality of audio device microphones corresponding to a single audio device of the plurality of audio devices, and determining the DOA data for the single audio device based at least in part on the microphone data. Alternatively, or in addition, determining the DOA data may include receiving antenna data from one or more antennas of a single audio device of the plurality of audio devices, and determining the DOA data for the single audio device based at least in part on the antenna data.

[0454] In some such examples, a single audio device itself may determine the DOA data. According to some such implementations, each of the plurality of audio devices may determine its own DOA data. However, in other implementations, another device, which may be a local or remote device, may determine the DOA data for one or more of the audio devices in the environment. DOA data can be determined for a diode device. According to some implementation examples, a server can determine DOA data for one or more audio devices in an environment.

[0455] According to this example, block 4010 includes determining an interior angle for each of a plurality of triangles based on the DOA data. In this example, each of the plurality of triangles has vertices corresponding to the positions of three of the audio devices among the audio devices. Some such examples have been described above.

[0456] FIG. 41 illustrates an example where each audio device in an environment is a vertex of a plurality of triangles. The sides of each triangle correspond to the distances between two of the audio devices 3605.

[0457] In this implementation example, block 4015 includes determining the length of each side of each triangle (the sides of the triangle can also be referred to as "edges" herein). According to this example, the length of the side is at least partially based on the interior angle. In some cases, the length of the side can be calculated by determining the first length of the first side of the triangle and then determining the lengths of the second and third sides of the triangle based on the interior angle of the triangle. Some such examples have been described above.

[0458] According to some such implementation examples, determining the first length can include setting the first length to a predetermined value. However, determining the first length can, in some examples, be based on arrival time data and / or received signal strength data. The arrival time data and / or received signal strength data can, in some implementation examples, correspond to sound waves from a first audio device in the environment detected by a second audio device in the environmen...

Claims

1. An audio session management method for an audio environment having a plurality of audio devices, comprising: Receiving, by a device implementing an audio session manager, a first route start request from a first device implementing a first application to start a first route for a first audio session, wherein the first route start request indicates a first audio source and a first audio environment destination, and the first audio environment destination corresponds to at least a first area of the audio environment and does not indicate an audio device; Establishing, by the device implementing the audio session manager, the first route corresponding to the first route start request, comprising: Determining, for a first stage of the first audio session, at least one audio device within the first area of the audio environment; Starting or scheduling the first audio session; Establishing the first route; Determining a first persistent unique audio session identifier for the first route and transmitting the first persistent unique audio session identifier to the first device. An audio session management method as described above.

2. The audio session management method according to claim 1, wherein the first route start request includes a first audio session priority.

3. The audio session management method according to claim 1 or 2, wherein the first route start request includes a first connection mode.

4. The audio session management method according to claim 3, wherein the first connection mode is a synchronous connection mode, a transaction connection mode, or a scheduled connection mode.

5. The audio session management method according to any one of claims 1 to 4, wherein the first route start request indicates a first person and includes an indication of whether approval is required from at least the first person.

6. The first route start request is an audio session management method according to any one of claims 1 to 5, including a first audio session target.

7. The first audio session target is an audio session management method according to claim 6, including one or more of understandable, audio quality, spatial fidelity, or inaudible.

8. The step of establishing the first route includes the step of causing at least one device in the environment to establish at least a first media stream corresponding to the first route, the first media stream including a first audio signal, which is an audio session management method according to any one of claims 1 to 7.

9. The audio session management method according to claim 8 further includes a rendering process for rendering the first audio signal into a first rendered audio signal.

10. Performing a first loudspeaker automatic positioning process for automatically determining a first position of each audio device among a plurality of audio devices in the first area of the audio environment at a first time point, wherein the rendering process is at least partially based on the first position of each audio device; Storing the first position of each audio device in a data structure associated with the first route; The audio session management method according to claim 9 further includes these steps.

11. Determining that at least one audio device in the first area has a changed position; Performing a second loudspeaker automatic positioning process for automatically determining the changed position; Updating the rendering process at least partially based on the changed position; Storing the changed position in the data structure associated with the first route; The audio session management method according to claim 10 further includes these steps.

12. Determining that at least one additional audio device has moved to the first area; Performing a second loudspeaker automatic positioning process for automatically determining an additional audio device position of the additional audio device; Updating a rendering process based at least in part on the additional audio device location; Storing the additional audio device location in the data structure associated with the first route; The audio session management method according to claim 10, further comprising. **Claim 13** The audio session management method according to any one of claims 1 to 12, wherein the first route start request indicates at least a first person as a first route source or a first route destination. **Claim 14** The audio session management method according to any one of claims 1 to 13, wherein the first route start request indicates at least a first service as the first audio source. **Claim 15** An apparatus configured to perform the method according to any one of claims 1 to 14. **Claim 16** A system configured to perform the method according to any one of claims 1 to 14. **Claim 17** One or more non-transitory media having encoded software, the software including instructions for controlling one or more devices to perform the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Content playback device

    JP2016100741A

  • satellite volume control

    JP2016528757A

  • Audio Response Playback

    JP2019509679A