Adjust audio and non-audio features based on noise metrics and speech intelligibility metrics
By determining noise indicators and voice intelligibility indicators in the content stream processing system and performing compensation processes, the problem of poor noise compensation effect in the prior art is solved, and the intelligibility and user experience of audio signals are improved in the noise environment.
Patent Information
- Application Number
- CN202080085359.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-30
- Filing Date
- 2020-12-09
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-12-09
AI Technical Summary
The prior art has the problem of poor noise compensation when adjusting audio and non-audio features to improve the comprehensibility of content streams, especially in television contexts where traditional noise compensation algorithms are difficult to overcome ambient noise.
Compensation processes are performed in response to noise metrics and speech intelligibility metrics, including processing to change audio data and applying non-audio-based compensation methods such as controlling closed subtitles systems and haptic display systems.
It realizes improving the intelligibility of the audio signal in a noisy environment and enhancing the user experience, especially in television devices, where the quality of audio reproduction is improved through effective noise compensation and voice enhancement.
Smart Images

Figure CN114830233B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 945,299, filed on December 9, 2019, U.S. Provisional Patent Application No. 63 / 198,158, filed on September 30, 2020, and U.S. Provisional Patent Application No. 63 / 198,160, filed on September 30, 2020, all of which are incorporated herein by reference in their entireties. Technical field
[0003] The present disclosure relates to systems and methods for adjusting audio and / or non - audio characteristics of a content stream. Background art
[0004] Audio and video devices (including but not limited to televisions and associated audio devices) are widely deployed. Although existing systems and methods for controlling audio and video devices provide benefits, improved systems and methods would still be desirable.
[0005] Symbols and terms
[0006] Throughout this disclosure, including in the claims, the terms "speaker", "loudspeaker", and "audio reproduction transducer" are used synonymously to denote any sound - producing transducer (or group of transducers) driven by a single speaker feed. A typical set of headphones includes two speakers. A speaker may be implemented to include multiple transducers (e.g., a woofer and a tweeter), which may be driven by a single common speaker feed or multiple speaker feeds. In some examples, one or more speaker feeds may undergo different processing in different circuit branches coupled to different transducers.
[0007] Throughout this disclosure, including in the claims, the expression of performing an operation on a "pair" of signals or data (e.g., filtering, scaling, transforming, or applying gain to signals or data) is used in a broad sense to denote performing an operation directly on the signals or data or on a processed version of the signals or data (e.g., a version of the signal that has undergone preliminary filtering or pre - processing before the operation is performed on it).
[0008] Throughout this disclosure, including in the claims, the expression "system" is used in a broad sense to denote a device, system, or subsystem. For example, a subsystem implementing a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M of the inputs and the other X - M inputs are received from an external source) may also be referred to as a decoder system.
[0009] Throughout this disclosure, including in the claims, the term "processor" is used in a broad sense to refer to a system or device that is programmable or otherwise configurable (e.g., using software or firmware) to perform operations on data (e.g., audio or video or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chip sets), digital signal processors that are programmed and / or otherwise configured to perform pipelined processing on audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chip sets.
[0010] Throughout this disclosure, including in the claims, the term "coupled" or "is coupled" is used to mean a direct or indirect connection. Thus, if a first device is coupled to a second device, the connection can be implemented by a direct connection or by an indirect connection via other devices and connections.
[0011] As used herein, an "intelligent device" is an electronic device that can operate interactively and / or autonomously to some extent and is typically configured to communicate with one or more other devices (or networks) via various wireless protocols such as Bluetooth, Zigbee, near field communication, Wi-Fi, Li-Fi (Light Fidelity), 3G, 4G, 5G, etc. A variety of notable types of intelligent devices are smartphones, smart cars, smart thermostats, smart doorbells, smart locks, smart refrigerators, phablets and tablet computers, smart watches, smart bracelets, smart keychains, and smart audio devices. The term "intelligent device" can also refer to a device that exhibits some properties of pervasive computing such as artificial intelligence.
[0012] As used herein, the expression "smart audio device" is used to refer to an intelligent device that is a single-purpose audio device or a multi-purpose audio device (e.g., an audio device that implements at least some aspects of a virtual assistant function). A single-purpose audio device is a device that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker and / or at least one camera) and is largely or primarily designed to achieve a single purpose (e.g., a television (TV)). For example, while a TV can typically play (and is considered capable of playing) audio from program material, in most cases, modern TVs run some operating system on which applications (including applications for watching TV) run locally. In this sense, a single-purpose audio device having one or more speakers and one or more microphones is typically configured to run local applications and / or services to directly use the one or more speakers and one or more microphones. Some single-purpose audio devices can be configured to be combined together to play audio over a defined area or a user-configured area.
[0013] A common type of multi-purpose audio device is an audio device that implements at least some aspects of a virtual assistant function, although other aspects of the virtual assistant function may be implemented by one or more other devices (e.g., one or more servers that the multi-purpose audio device is configured to communicate with). Such a multi-purpose audio device may be referred to herein as a "virtual assistant". A virtual assistant is a device (e.g., a smart speaker or a voice assistant integrated device) that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker and / or at least one camera). In some examples, the virtual assistant may provide the ability to use multiple devices (different from the virtual assistant) for applications that are in some sense cloud-enabled or otherwise not fully implemented in or on the virtual assistant itself. In other words, at least some aspects of the virtual assistant function (e.g., the speech recognition function) may be (at least partially) implemented by one or more servers or other devices, and the virtual assistant may communicate with the one or more servers or other devices via a network such as the Internet. Virtual assistants may sometimes work together, e.g., in a discrete and conditionally defined manner. For example, two or more virtual assistants may work together in the sense that one of them (e.g., the virtual assistant that is most certain to have heard the wake word) responds to the wake word. In some embodiments, connected virtual assistants may form a kind of cluster that can be managed by a main application, which may be (or implement) the virtual assistant.
[0014] As used herein, the term "wake word" is used in a broad sense to denote any sound (e.g., a word spoken by a human or other sound) in response to the detection of which (the "hearing" of the sound using at least one microphone included in or coupled to the smart audio device, or at least one other microphone) the smart audio device is configured to wake up. In this context, "waking up" means that the device enters a state of waiting (in other words, listening) for a voice command. In some instances, what may be referred to herein as a "wake word" may include more than one word, e.g., a phrase.
[0015] As used herein, the expression "wake word detector" denotes a device (or software including instructions for configuring the device) configured to continuously search for an alignment between real-time sound (e.g., speech) features and a training model. Generally, whenever the wake word detector determines that the probability of having detected a wake word exceeds a predefined threshold, a wake word event is triggered. For example, the threshold may be a predefined threshold tuned to give a reasonable trade-off between the false acceptance rate and the false rejection rate. After a wake word event, the device may enter a state of listening for commands and passing the received commands to a larger, more computationally intensive recognizer (which may be referred to as the "wake" state or the "focus" state).
[0016] As used herein, the terms "program stream" and "content stream" refer to a collection of one or more audio signals and, in some instances, a collection of one or more video signals, at least a portion of which audio and video signals are intended to be heard together as a whole. Examples include music compilations, movie soundtracks, movies, television shows, the audio portion of a television show, podcasts, live voice calls, synthetic voice responses from a smart assistant, etc. In some instances, a content stream can include multiple versions of at least a portion of an audio signal, e.g., the same dialogue in more than one language. In such instances, only one version (e.g., the version corresponding to a single language) is intended to be reproduced at a time of the audio data or a portion thereof. SUMMARY OF THE INVENTION
[0017] At least some aspects of the present disclosure can be implemented via one or more audio processing methods including, but not limited to, content stream processing methods. In some instances, the one or more methods can be implemented at least in part by a control system and / or via instructions (e.g., software) stored on one or more non-transitory media. Some such methods involve receiving, by a control system and via an interface system, a content stream including video data and audio data corresponding to the video data. Some such methods involve determining, by a control system, a noise metric and / or a speech intelligibility metric. Some such methods involve performing, by a control system in response to the noise metric and / or the speech intelligibility metric, a compensation process. In some examples, performing the compensation process involves one or more of the following: changing the processing of the audio data, where changing the processing of the audio data does not involve applying a broadband gain increase to the audio signal; or applying a non-audio-based compensation method. In some examples, the non-audio-based compensation method can involve controlling a haptic display system and / or controlling a vibrating surface.
[0018] Some such methods involve processing, by a control system, the video data and providing, by the control system, the processed video data to at least one display device of the environment. Some such methods involve rendering, by a control system, audio data for reproduction via an audio reproduction transducer set of the environment to produce a rendered audio signal. Some such methods involve providing, via the interface system, the rendered audio signal to at least some of the audio reproduction transducers of the audio reproduction transducer set of the environment.
[0019] In some examples, the speech intelligibility metric can be at least partially based on one or more of the following: speech transmission index (STI), common intelligibility scale (CIS), C50 (the ratio of the sound energy received between 0 ms and 50 ms after the initial sound to the sound energy arriving later than 50 ms), the reverberation of the environment, the frequency response of the environment, the playback characteristics of one or more audio reproduction transducers of the environment, or the ambient noise level.
[0020] According to some embodiments, a speech intelligibility metric may be at least partially based on one or more user characteristics of a user. The one or more user characteristics may include, for example, at least one of the user's native language, the user's accent, the user's location in the environment, the user's age, and / or the user's capabilities. The user's capabilities may include, for example, the user's hearing ability, the user's language proficiency, the user's accent comprehension level, the user's vision, and / or the user's reading comprehension.
[0021] According to some examples, a non-audio based compensation method may involve controlling a closed captioning system, a lyric captioning system, or a dialogue captioning system. In some such examples, controlling the closed captioning system, the lyric captioning system, or the dialogue captioning system may be at least partially based on the user's hearing ability, the user's language proficiency, the user's vision, and / or the user's reading comprehension. According to some examples, controlling the closed captioning system, the lyric captioning system, or the dialogue captioning system may involve at least partially controlling at least one of a font or a font size based on the speech intelligibility metric.
[0022] In some instances, controlling the closed captioning system, the lyric captioning system, or the dialogue captioning system may involve at least partially determining whether to filter out some speech-based text based on the speech intelligibility metric. In some embodiments, controlling the closed captioning system, the lyric captioning system, or the dialogue captioning system may involve at least partially determining whether to simplify or paraphrase at least some speech-based text based on the speech intelligibility metric.
[0023] In some examples, controlling the closed captioning system, the lyric captioning system, or the dialogue captioning system may involve at least partially determining whether to display text based on a noise metric. In some instances, determining whether to display text may involve applying a first noise threshold to determine that text will be displayed and applying a second noise threshold to determine that text display will stop.
[0024] According to some embodiments, audio data may include audio objects. In some such embodiments, changing the processing of the audio data may involve at least partially determining which audio objects will be rendered based on at least one of a noise metric or a speech intelligibility metric. In some examples, changing the processing of the audio data may involve changing the rendering location of one or more audio objects to improve intelligibility in the presence of noise. According to some embodiments, a content stream may include audio object priority metadata. In some examples, changing the processing of the audio data may involve selecting high-priority audio objects based on the priority metadata and rendering the high-priority audio objects without rendering at least some other audio objects.
[0025] In some examples, changing the processing of audio data can involve applying one or more speech enhancement methods at least in part based on a noise metric and / or a speech intelligibility metric. The one or more speech enhancement methods can include, for example, reducing the gain of non-speech audio and / or increasing the gain of speech frequencies.
[0026] According to some embodiments, changing the processing of audio data can involve changing one or more of an upmixing process, a downmixing process, a virtual bass process, a bass distribution process, an equalization process, a crossover filter, a delay filter, a multi-band limiter, or a virtualization process at least in part based on a noise metric and / or a speech intelligibility metric.
[0027] Some embodiments can involve transmitting audio data from a first device to a second device. Some such embodiments can involve transmitting at least one of a noise metric, a speech intelligibility metric, or echo reference data from the first device to the second device or from the second device to the first device. In some instances, the second device can be a hearing aid, a personal sound amplification product, a cochlear implant, or a headset.
[0028] Some embodiments can involve: receiving, by a second device control system, a second device microphone signal; receiving, by the second device control system, audio data and at least one of the following: a noise metric, a speech intelligibility metric, or echo reference data; determining, by the second device control system, one or more audio data gain settings and one or more second device microphone signal gain settings; applying, by the second device control system, the audio data gain settings to the audio data to produce gain-adjusted audio data; applying, by the second device control system, the second device microphone signal gain settings to the second device microphone signal to produce gain-adjusted second device microphone signals; mixing, by the second device control system, the gain-adjusted audio data and the gain-adjusted second device microphone signals to produce mixed second device audio data; providing, by the second device control system, the mixed second device audio data to one or more second device transducers; and reproducing, by the one or more second device transducers, the mixed second device audio data. Some such examples can involve controlling, by the second device control system, the relative levels of the gain-adjusted audio data and the gain-adjusted second device microphone signals in the mixed second device audio data at least in part based on a noise metric.
[0029] Some examples can involve receiving, by a control system and via an interface system, a microphone signal. Some such examples can involve determining, by the control system, a noise metric at least in part based on the microphone signal. In some instances, the microphone signal can be received from a device including at least one microphone of an environment and at least one audio reproduction transducer of an audio reproduction transducer group.
[0030] Some of the disclosed methods involve receiving, by a first control system and via a first interface system, a content stream that includes video data and audio data corresponding to the video data. Some such methods involve determining, by the first control system, a noise metric and / or a speech intelligibility metric. Some such methods involve determining, by the first control system, a compensation process to be performed in response to at least one of the noise metric or the speech intelligibility metric. In some examples, performing the compensation process involves one or more of the following: changing the processing of the audio data, where changing the processing of the audio data does not involve applying a broadband gain increase to the audio signal; or applying a non-audio-based compensation method.
[0031] Some such methods involve determining, by the first control system, compensation metadata corresponding to the compensation process. Some such methods involve generating, by the first control system, encoded compensation metadata by encoding the compensation metadata. Some such methods involve generating, by the first control system, encoded video data by encoding the video data. Some such methods involve generating, by the first control system, encoded audio data by encoding the audio data. Some such methods involve transmitting an encoded content stream that includes the encoded compensation metadata, the encoded video data, and the encoded audio data from a first device to at least a second device.
[0032] In some instances, the audio data can include speech data as well as music and effects (M&E) data. Some such methods can involve separating, by the first control system, the speech data from the M&E data, determining, by the first control system, speech metadata that permits extraction of the speech data from the audio data, and generating, by the first control system, encoded speech metadata by encoding the speech metadata. In some such examples, transmitting the encoded content stream can involve transmitting the encoded speech metadata to at least the second device.
[0033] According to some embodiments, the second device can include a second control system configured to decode the encoded content stream. In some such embodiments, the second device can be one of a plurality of devices to which the encoded audio data has been transmitted. In some instances, the plurality of devices can have been selected at least in part based on the speech intelligibility for a user category. In some examples, the user category can be defined by known or estimated hearing ability, known or estimated language level, known or estimated accent comprehension level, known or estimated visual acuity, and / or known or estimated reading comprehension.
[0034] In some embodiments, the compensation metadata may include a plurality of options that can be selected by the second device and / or by the user of the second device. In some such examples, two or more of the plurality of options may correspond to noise levels that may occur in the environment in which the second device is located. In some examples, two or more of the plurality of options may correspond to speech intelligibility metrics. In some such examples, the encoded content stream may include speech intelligibility metadata. Some such examples may involve selecting, by the second control system and at least partially based on the speech intelligibility metadata, one of the two or more options. According to some embodiments, each of the plurality of options may correspond to a known or estimated hearing ability, known or estimated language level, known or estimated accent comprehension level, known or estimated visual acuity, and / or known or estimated reading comprehension ability of the user of the second device. In some examples, each of the plurality of options may correspond to a speech enhancement level.
[0035] According to some examples, the second device may correspond to a specific playback device. In some such examples, the specific playback device may be a specific television or a specific device associated with a television.
[0036] Some embodiments may involve receiving, by the first control system and via the first interface system, a noise metric and / or a speech intelligibility metric from the second device. In some such examples, the compensation metadata may correspond to the noise metric and / or the speech intelligibility metric.
[0037] Some examples may involve determining, by the first control system and at least partially based on the noise metric or the speech intelligibility metric, whether the encoded audio data will correspond to all of the received audio data or only to a portion of the received audio data. In some examples, the audio data may include audio objects and corresponding priority metadata indicating the priority of the audio objects. Some such examples in which it is determined that the encoded audio data will correspond to only a portion of the received audio data may also involve selecting, at least partially based on the priority metadata, that portion of the received audio data.
[0038] In some embodiments, non-audio-based compensation methods may involve controlling a closed captioning system, a lyrics captioning system, or a dialogue captioning system. In some such examples, controlling a closed captioning system, a lyrics captioning system, or a dialogue captioning system may involve controlling the font and / or font size at least in part based on a speech intelligibility metric. In some embodiments, controlling a closed captioning system, a lyrics captioning system, or a dialogue captioning system may involve determining whether to filter out some speech-based text, determining whether to simplify at least some speech-based text, and / or determining whether to paraphrase at least some speech-based text at least in part based on a speech intelligibility metric. According to some embodiments, a closed captioning system, a lyrics captioning system, or a dialogue captioning system may involve determining whether to display text at least in part based on a noise metric.
[0039] In some examples, changing the processing of audio data may involve applying one or more speech enhancement methods at least in part based on a noise metric and / or a speech intelligibility metric. The one or more speech enhancement methods may include, for example, reducing the gain of non-speech audio and / or increasing the gain of speech frequencies.
[0040] According to some embodiments, changing the processing of audio data may involve changing one or more of an upmixing process, a downmixing process, a virtual bass process, a bass distribution process, an equalization process, a crossover filter, a delay filter, a multi-band limiter, or a virtualization process at least in part based on a noise metric and / or a speech intelligibility metric.
[0041] Some of the disclosed methods involve receiving, by a first control system and via a first interface system of a first device, a content stream including received video data and received audio data corresponding to the video data. Some such methods involve receiving, by the first control system and via the first interface system, a noise metric and / or a speech intelligibility metric from a second device. Some such methods involve determining, by the first control system and at least in part based on the noise metric and / or the speech intelligibility metric, whether to reduce the complexity level of transmitted encoded audio data corresponding to the received audio data and / or text corresponding to the received audio data. Some such methods involve selecting the transmitted encoded audio data and / or text based on the determination process. Some such methods involve transmitting an encoded content stream including encoded video data and the transmitted encoded audio data from the first device to the second device.
[0042] According to some embodiments, determining whether to reduce the complexity level can involve determining whether the transmitted encoded audio data will correspond to all of the received audio data or only to a portion of the received audio data. In some embodiments, the audio data can include audio objects and corresponding priority metadata indicating the priority of the audio objects. According to some such embodiments, it can be determined that the encoded audio data will correspond to only a portion of the received audio data. Some such embodiments can involve selecting the portion of the received audio data at least in part based on the priority metadata. In some examples, for a closed captioning system, a karaoke captioning system, or a dialogue captioning system, determining whether to reduce the complexity level can involve determining whether to filter out some speech-based text, determining whether to simplify at least some speech-based text, and / or determining whether to paraphrase at least some speech-based text.
[0043] Some or all of the operations, functions, and / or methods described herein can be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media can include memory devices as described herein, including but not limited to random access memory (RAM) devices, read only memory (ROM) devices, and the like. Accordingly, some innovative aspects of the subject matter described in this disclosure can be implemented via one or more non-transitory media having software stored thereon. For example, the software can include instructions for controlling one or more devices to perform one or more of the disclosed methods.
[0044] At least some aspects of the present disclosure can be implemented via a device and / or via a system including multiple devices. For example, one or more devices can be capable of performing at least in part the methods disclosed herein. In some embodiments, the device is or includes an audio processing system having an interface system and a control system. The control system can include one or more general single-chip or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or combinations thereof. In some examples, the control system can be configured to perform one or more of the disclosed methods.
[0045] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions of the following drawings may not be drawn to scale. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 An example of a noise compensation system is shown.
[0047] Figure 2 is a block diagram showing an example of components of an apparatus capable of implementing various aspects of the present disclosure.
[0048] Figure 3A is a flowchart outlining an example of the disclosed method.
[0049] Figure 3B shows examples of the speech transmission index (STI) and common intelligibility scale (CIS) metrics for measuring speech intelligibility.
[0050] Figure 4 shows an example of a system in which a closed captioning system is controlled based on a noise estimate.
[0051] Figure 5 shows an example of a graph related to the control of a closed captioning system.
[0052] Figure 6 shows an example of an intelligibility metric evaluation module.
[0053] Figure 7A shows an example of a closed captioning system controlled by an intelligibility metric.
[0054] Figure 7B shows an example of an audio description renderer controlled by an intelligibility metric.
[0055] Figure 8 shows an example of an echo predictor module.
[0056] Figure 9 shows an example of a system configured to determine an intelligibility metric based at least in part on a playback process.
[0057] Figure 10 shows an example of a system configured to determine an intelligibility metric based at least in part on an ambient noise level.
[0058] Figure 11 shows an example of a system configured to modify an intelligibility metric based at least in part on one or more user capabilities.
[0059] Figure 12 shows an example of a caption generator.
[0060] Figure 13 shows an example of a caption modifier module configured to change closed captions based on an intelligibility metric.
[0061] Figure 14 shows a further example of a non-audio compensation process and system that can be controlled based on a noise estimator.
[0062] Figure 15 Shows an example of a noise compensation system.
[0063] Figure 16 Shows an example of a system configured to perform speech enhancement in response to detected ambient noise.
[0064] Figure 17 Shows an example of a graph corresponding to elements of a system limited by microphone characteristics.
[0065] Figure 18 Shows an example of a system in which a hearing aid is configured to communicate with a television.
[0066] Figure 19 Shows an example of the hybrid and speech enhancement components of a hearing aid.
[0067] Figure 20 Is a graph showing an example of ambient noise levels.
[0068] Figure 21 Shows an example of an encoder block and a decoder block according to one embodiment.
[0069] Figure 22 Shows an example of an encoder block and a decoder block according to another embodiment.
[0070] Figure 23 Shows some examples of decoder-side operations that can be performed in response to receiving Figure 21 the encoded audio bitstream shown in
[0071] Figure 24 Shows some examples of decoder-side operations that can be performed in response to receiving Figure 22 the encoded audio bitstream shown in
[0072] Figure 25 Shows an example of an encoder block and a decoder block according to another embodiment.
[0073] Figure 26 Shows an example of an encoder block and a decoder block according to another embodiment.
[0074] Figure 27 Shows some alternative examples of decoder-side operations that can be performed in response to receiving Figure 21 the encoded audio bitstream shown in
[0075] Figure 28 Shows Figure 24 and Figure 27 an enhanced version of the system shown in
[0076] Figure 29Shows an example of an encoder block and a decoder block according to another embodiment.
[0077] Figure 30 Shows an example of an encoder block and a decoder block according to another embodiment.
[0078] Figure 31 Shows the relationships between various disclosed use cases.
[0079] Figure 32 Is a flowchart outlining an example of the disclosed method.
[0080] Figure 33 Is a flowchart outlining an example of the disclosed method.
[0081] Figure 34 Shows an example of a floor plan of an audio environment, in which the audio environment is a living space.
[0082] Like reference numerals and names in the various figures indicate like elements. Detailed Description
[0083] Voice assistants are becoming increasingly common. To enable voice assistants, television (TV) and soundbar manufacturers have started adding microphones to their devices. The added microphones may provide an input regarding background noise, which may be input into a noise compensation algorithm. However, applying traditional noise compensation algorithms in the TV context involves some technical challenges. For example, the drivers typically used in TVs have only a limited amount of capabilities. Applying traditional noise compensation algorithms via the drivers typically used in TVs may not be entirely satisfactory, in part because these drivers may not be able to overcome the noise within the listening environment, such as the noise within a room.
[0084] The present disclosure describes alternative methods for improving the experience. Some disclosed embodiments relate to determining a noise metric and / or a speech intelligibility metric and determining a compensation process in response to at least one of the noise metric or the speech intelligibility metric. According to some embodiments, the compensation process can be determined (at least in part) by one or more local devices of the audio environment. Alternatively or additionally, the compensation process can be determined (at least in part) by one or more remote devices (such as one or more devices implementing cloud-based services). In some examples, the compensation process can involve changing the processing of received audio data. According to some such examples, changing the processing of the audio data does not involve applying a broadband gain increase to the audio signal. In some examples, the compensation process can involve applying a non-audio-based compensation method, such as controlling a closed captioning system, a karaoke captioning system, or a dialogue captioning system. Some disclosed embodiments provide satisfactory noise compensation regardless of whether the corresponding audio data is reproduced via a relatively capable audio reproduction transducer or via a relatively less capable audio reproduction transducer, although in some examples, the type of noise compensation may be different for each case.
[0085] Figure 1 An example of a noise compensation system is shown. System 100 is configured to adjust the volume of the overall system based on a noise estimate to ensure that a listener can understand the audio in the presence of noise. In this example, system 100 includes a loudspeaker 108, a microphone 105, a noise estimator 104, and a gain adjuster 102.
[0086] In this example, the gain adjuster 102 is receiving an audio signal 101 from a file, a streaming service, etc. The gain adjuster 102 can be configured, for example, to apply a gain adjustment algorithm such as a broadband gain adjustment algorithm.
[0087] In this example, the signal 103 is sent to the loudspeaker 108. According to this example, the signal 103 is also provided to the noise estimator 104 and is the reference signal for the noise estimator. In this example, the signal 106 is also sent from the microphone 105 to the noise estimator 104.
[0088] According to this example, the noise estimator 104 is a component configured to estimate the noise level in the environment including the system 100. In some examples, the noise estimator 104 may include an echo canceller. However, in some embodiments, the noise estimator 104 may simply measure the noise when a signal corresponding to silence is sent to the loudspeaker 108. In this example, the noise estimator 104 is providing a noise estimate 107 to the gain adjuster 102. Depending on the particular embodiment, the noise estimate 107 may be a broadband estimate or a spectral estimate of the noise. In this example, the gain adjuster 102 is configured to adjust the output level of the loudspeaker 108 based on the noise estimate 107.
[0089] As noted above, loudspeakers in televisions typically have rather limited capabilities. Thus, the type of volume adjustment provided by the system 100 will generally be limited by the speaker protection components (e.g., limiters and / or compressors) of such loudspeakers.
[0090] The present disclosure provides various methods that can overcome at least some of the possible disadvantages of the system 100, as well as devices and systems for implementing the presently disclosed methods. Some such methods may be based on one or more noise metrics. Alternatively or additionally, some such methods may be based on one or more speech intelligibility metrics. The various disclosed methods provide one or more compensation processes in response to one or more noise metrics and / or one or more speech intelligibility metrics. Some such compensation processes involve changing the processing of audio data. In many of the disclosed methods, changing the processing of audio data does not involve applying a broadband gain increase to the audio signal. Alternatively or additionally, some such compensation processes may involve one or more non-audio-based compensation methods.
[0091] Figure 2 is a block diagram showing an example of the components of a device capable of implementing various aspects of the present disclosure. Like other Figure 1 as Figure 2The types and quantities of the elements shown are provided only as examples. Other embodiments may include more, fewer, and / or different types and quantities of elements. According to some examples, device 200 may be or may include a television configured to perform at least some of the methods disclosed herein. In some embodiments, device 200 may be or may include a television control module. Depending on the particular embodiment, the television control module may or may not be integrated into the television. In some embodiments, the television control module may be a device separate from the television, and in some instances, the television control module may be sold separately from the television or sold as an additional or optional device that a purchased television may include. In some embodiments, the television control module may be obtainable from a content provider (such as a provider of television programs, movies, etc.). In other embodiments, device 200 may be or may include another device configured to perform at least some of the methods disclosed herein at least in part, such as a laptop computer, a cellular phone, a tablet device, a smart speaker, etc.
[0092] According to some alternative embodiments, device 200 may be or may include a server. In some such examples, device 200 may be or may include an encoder. Accordingly, in some instances, device 200 may be a device configured to be used within an audio environment (such as a home audio environment), whereas in other instances, device 200 may be a device configured to be used in the "cloud", such as a server.
[0093] In this example, device 200 includes interface system 205 and control system 210. In some embodiments, interface system 205 may be configured to communicate with one or more other devices in the audio environment. In some examples, the audio environment may be a home audio environment. In some embodiments, interface system 205 may be configured to exchange control information and associated data with the audio devices in the audio environment. In some examples, the control information and associated data may relate to one or more software applications that device 200 is executing.
[0094] In some embodiments, interface system 205 may be configured to receive a content stream or to provide a content stream. The content stream may include an audio signal. In some examples, the content stream may include video data and audio data corresponding to the video data. In some instances, the audio data may include spatial data, such as channel data and / or spatial metadata. In some embodiments, interface system 205 may be configured to receive input from one or more microphones in the environment.
[0095] The interface system 205 may include one or more network interfaces and / or one or more external device interfaces (such as one or more Universal Serial Bus (USB) interfaces). According to some embodiments, the interface system 205 may include one or more wireless interfaces. The interface system 205 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 205 may include one or more interfaces between the control system 210 and a memory system (such as the optional memory system 215 shown in Figure 2 ). However, in some instances, the control system 210 may include the memory system.
[0096] For example, the control system 210 may include a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.
[0097] In some embodiments, the control system 210 may reside in more than one device. For example, in some embodiments, a part of the control system 210 may reside in a device within one of the environments depicted herein, and another part of the control system 210 may reside in a device outside the environment (such as a server, a mobile device (e.g., a smart phone or a tablet computer), etc.). In other examples, a part of the control system 210 may reside in a device within one of the environments depicted herein, and another part of the control system 210 may reside in one or more other devices of the environment. For example, the control system functions may be distributed across multiple smart audio devices of the environment, or may be shared by an orchestration device (such as a device that may be referred to herein as a smart home hub) and one or more other devices of the environment. In other examples, a part of the control system 210 may reside in a device (such as a server) implementing a cloud-based service, and another part of the control system 210 may reside in another device (such as another server, a memory device, etc.) implementing a cloud-based service. In some examples, the interface system 205 may also reside in more than one device.
[0098] In some embodiments, the control system 210 may be configured to at least partially execute the methods disclosed herein. According to some examples, the control system 210 may be configured to implement methods for content stream processing.
[0099] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. One or more non-transitory media may reside, for example, in Figure 2 the optional memory system 215 shown in and / or in the control system 210. Accordingly, various innovative aspects of the subject matter described in this disclosure may be implemented in one or more non-transitory media on which software is stored. The software may include, for example, instructions for controlling at least one device to process a content stream, encode a content stream, decode a content stream, etc. For example, the software may be executable by one or more components of a control system (such as Figure 2 the control system 210 of).
[0100] In some examples, the apparatus 200 may include Figure 2 the optional microphone system 220 shown in. The optional microphone system 220 may include one or more microphones. In some embodiments, one or more of the microphones may be part of or associated with another device (such as a speaker in a speaker system, a smart audio device, etc.). In some examples, the apparatus 200 may not include the microphone system 220. However, in some such embodiments, the apparatus 200 may still be configured to receive microphone data from one or more microphones in the audio environment via the interface system 205. In some such embodiments, a cloud-based implementation of the apparatus 200 may be configured to receive microphone data or a noise metric at least partially corresponding to the microphone data from one or more microphones in the audio environment via the interface system 205.
[0101] According to some embodiments, the apparatus 200 may include Figure 2 the optional loudspeaker system 225 shown in. The optional loudspeaker system 225 may include one or more loudspeakers, which may also be referred to herein as "speakers" or more generally as "audio reproduction transducers". In some examples (e.g., cloud-based implementations), the apparatus 200 may not include the loudspeaker system 225.
[0102] In some embodiments, the apparatus 200 may include Figure 2The optional sensor system 230 shown in FIG. The optional sensor system 230 may include one or more touch sensors, gesture sensors, motion detectors, etc. According to some embodiments, the optional sensor system 230 may include one or more cameras. In some embodiments, the camera may be a standalone camera. In some examples, one or more cameras in the optional sensor system 230 may reside in a smart audio device, which may be a single-purpose audio device or a virtual assistant. In some such examples, one or more cameras in the optional sensor system 230 may reside in a television, a mobile phone, or a smart speaker. In some examples, the device 200 may not include the sensor system 230. However, in some such embodiments, the device 200 may still be configured to receive sensor data from one or more sensors in the audio environment via the interface system 205.
[0103] In some embodiments, the device 200 may include Figure 2 The optional display system 235 shown in FIG. The optional display system 235 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some instances, the optional display system 235 may include one or more organic light-emitting diode (OLED) displays. In some examples, the optional display system 235 may include one or more displays of a television. In other examples, the optional display system 235 may include a laptop display, a mobile device display, or another type of display. In some examples in which the device 200 includes the display system 235, the sensor system 230 may include a touch sensor system and / or a gesture sensor system proximate to one or more displays of the display system 235. According to some such embodiments, the control system 210 may be configured to control the display system 235 to present one or more graphical user interfaces (GUIs).
[0104] According to some such examples, the device 200 may be or may include a smart audio device. In some such embodiments, the device 200 may be or may include a wake word detector. For example, the device 200 may be or may include a virtual assistant.
[0105] Figure 3A is a flowchart outlining an example of the disclosed method. Like other methods described herein, the blocks of method 300 need not be performed in the order indicated. Additionally, such a method may include more or fewer blocks than those shown and / or described.
[0106] Method 300 may be performed by, such as Figure 2performed by the apparatus or system of apparatus 200 shown and described above. In some examples, the blocks of method 300 may be performed by one or more devices within an audio environment (e.g., by a television or a television control module). In some embodiments, the audio environment may include one or more rooms of a home environment. In other examples, the audio environment may be another type of environment, such as an office environment, a car environment, a train environment, a street or sidewalk environment, a park environment, etc. However, in alternative embodiments, at least some of the blocks of method 300 may be performed by a device (such as a server) implementing a cloud-based service.
[0107] In this embodiment, block 305 involves receiving, by a control system and via an interface system, a content stream including video data and audio data corresponding to the video data. In some such embodiments, the control system and the interface system may be Figure 2 the control system 210 and the interface system 205 shown and described above. According to some embodiments, block 305 may involve receiving an encoded content stream. In such an embodiment, block 305 may involve decoding the encoded content stream. The content stream may correspond, for example, to a movie, a television program, a musical performance, a music video, etc. In some instances, the complexity of the video data may be relatively lower than that of a typical movie or television program. For example, in some instances, the video data may correspond to lyrics, song titles, pictures of one or more performers, etc. In some alternative embodiments, block 305 may involve receiving a content stream including audio data but not including corresponding video data.
[0108] In this example, block 310 involves determining, by the control system, at least one of a noise metric or a speech intelligibility metric (SIM). According to some examples, determining the noise metric may involve receiving, by the control system, microphone data from one or more microphones of the audio environment in which the audio data will be rendered and determining the noise metric by the control system at least in part based on the microphone signals.
[0109] Some such examples may involve receiving microphone data from one or more microphones of the audio environment in which the control system resides. In some such embodiments, the microphone signals may be received from a device including at least one microphone of the environment and at least one audio reproduction transducer of a group of audio reproduction transducers. For example, a device including at least one microphone and at least one audio reproduction transducer may be or may include a smart speaker. However, some alternative examples may involve receiving microphone data, a noise metric, or a speech intelligibility metric from one or more devices in the audio environment that are not in the same location as the control system.
[0110] According to some examples, determining a noise metric and / or a SIM may involve identifying ambient noise in a received microphone signal and estimating a noise level corresponding to the ambient noise. In some such examples, determining a noise metric may involve determining whether the noise level is above or below one or more thresholds.
[0111] In some examples, determining a noise metric and / or a SIM may involve determining one or more metrics corresponding to reverberation of the environment, a frequency response of the environment, playback characteristics of one or more audio reproduction transducers of the environment, etc. According to some examples, determining a SIM may involve determining one or more metrics corresponding to a speech transmission index (STI), a common intelligibility scale (CIS), or C50, where C50 is measured as a ratio of early-arriving sound energy (arriving between 0 ms and 50 ms) to late-arriving sound energy (arriving later than 50 ms).
[0112] In some examples, speech intelligibility may be measured by reproducing a known signal and measuring the quality of the signal when it arrives at each of a plurality of measurement locations in an audio environment. The IEC 60268-16 standard for STI defines how to measure any degradation of a signal.
[0113] Figure 3B Examples of STI and CIS scales for measuring speech intelligibility are shown. As shown by bar graph 350, STI and CIS scales for speech intelligibility may be displayed as a single number from 0 (unintelligible) to 1 (excellent intelligibility).
[0114] According to some embodiments, a SIM may be at least partially based on one or more user characteristics, e.g., characteristics of a person who is a user of a television or other device that will be used to reproduce a received content stream. In some examples, one or more user characteristics may include at least one of a user's native language, a user's accent, a user's age, and / or a user's capabilities. A user's capabilities may also include, for example, a user's hearing ability, a user's language level, a user's accent comprehension level, a user's vision, and / or a user's reading comprehension.
[0115] In some examples, one or more user characteristics may include a user's location in the environment, which may affect speech intelligibility. For example, if a user is not located on the central axis relative to a speaker, the speech intelligibility metric may be reduced because the mix will have more left / right channels than the central channel. If they are in an ideal listener position, in some embodiments, the intelligibility metric may remain unchanged.
[0116] In some such embodiments, the user may have previously provided input regarding one or more such user characteristics. According to some such examples, the user may have previously provided input via a graphical user interface (GUI) provided on a display device in accordance with a command from the control system.
[0117] Alternatively or additionally, the control system may have inferred one or more of the user characteristics based on the user's past behavior, such as one or more languages selected by the user for the reproduced content, the user's demonstrated ability to understand language and / or regional accent (e.g., as evidenced by instances where the user has selected a closed captioning system, a karaoke captioning system, or a dialogue captioning system), the relative complexity of the language used in the content selected by the user (e.g., whether the language corresponds to the speech of a television program for a preschool audience, the speech of a movie for a teenage audience, the speech of a documentary for a college-educated audience, etc.), the playback volume selected by the user (e.g., for the portion of the reproduced content corresponding to speech), whether the user has previously used a visual impairment device (such as a tactile display system), and the like.
[0118] According to this example, block 315 involves the control system performing a compensation process in response to at least one of a noise metric or a speech intelligibility metric. In this example, performing the compensation process involves changing the processing of the audio data and / or applying a non-audio-based compensation method. According to this embodiment, changing the processing of the audio data does not involve applying a broadband gain increase to the audio signal.
[0119] In some embodiments, the non-audio-based compensation method involves at least one of the following: controlling a tactile display system or controlling a vibrating surface. Some examples are described below.
[0120] According to Figure 3A the example shown in, block 320 involves the control system processing the received video data. In this example, block 325 involves the control system providing the processed video data to at least one display device of the environment. In some embodiments, block 320 may involve decoding the encoded video data. In some examples, block 320 may involve formatting the video data according to the aspect ratio, settings, etc. of the display device of the environment (such as a television, a laptop computer, etc.) on which the video data will be displayed.
[0121] In some examples, the non-audio-based compensation method may involve controlling a closed captioning system, a karaoke captioning system, or a dialogue captioning system. According to some such examples, block 320 may involve controlling a closed captioning system, a karaoke captioning system, or a dialogue captioning system to include text in the displayed video data.
[0122] According to some embodiments, controlling a closed captioning system, a karaoke captioning system, or a dialogue captioning system can involve determining whether to display text at least in part based on a noise metric. According to some such examples, determining whether to display text can involve applying a first noise threshold to determine that text will be displayed and applying a second noise threshold to determine that text will cease to be displayed. Some examples are described below.
[0123] According to some examples, controlling a closed captioning system, a karaoke captioning system, or a dialogue captioning system can be at least in part based on a user's hearing ability, a user's language proficiency, a user's visual acuity, and / or a user's reading comprehension. In some such examples, controlling a closed captioning system, a karaoke captioning system, or a dialogue captioning system can involve controlling at least one of a font or a font size at least in part based on a speech intelligibility metric and / or a user's visual acuity.
[0124] According to some embodiments, method 300 can involve determining whether to reduce a complexity level of audio data or corresponding text. In some such examples in which a non-audio-based compensation method involves controlling a closed captioning system, a karaoke captioning system, or a dialogue captioning system, method 300 can involve determining whether to filter out, simplify, and / or paraphrase at least some speech-based text at least in part based on at least one of a speech intelligibility metric and / or a user's ability such as a user's reading comprehension.
[0125] In this example, block 330 involves rendering, by a control system, audio data for reproduction via an audio reproduction transducer set of an environment to produce a rendered audio signal. According to this embodiment, block 335 involves providing the rendered audio signal to at least some of the audio reproduction transducers of the audio reproduction transducer set of the environment via an interface system.
[0126] In some examples in which the audio data includes audio objects and in which performing a compensation process involves changing a processing of the audio data, block 330 can involve determining which audio objects will be rendered at least in part based on at least one of a noise metric or a speech intelligibility metric. In some such examples in which the content stream includes audio object priority metadata, changing the processing of the audio data can involve selecting high-priority audio objects based on the priority metadata and rendering the high-priority audio objects without rendering other audio objects.
[0127] In some examples in which the audio data includes audio objects and in which performing a compensation process involves changing a processing of the audio data, block 330 can involve changing a rendering location of one or more audio objects to improve intelligibility in the presence of noise.
[0128] According to some examples in which performing the compensation process involves changing the processing of audio data, method 300 may involve applying one or more speech enhancement methods at least in part based on at least one of a noise metric or a speech intelligibility metric. In some such examples, the one or more speech enhancement methods may include reducing the gain of non-verbal audio and / or increasing the gain of speech frequencies (e.g., audio frequencies in the range of 50 Hz to 2 kHz). Other embodiments may be configured to increase the gain of other audio frequency ranges corresponding to speech frequencies (e.g., audio frequencies in the range of 300 Hz to 3400 Hz, audio frequencies in the range of 50 Hz to 3400 Hz, audio frequencies in the range of 50 Hz to 500 Hz, etc.).
[0129] Alternatively or additionally, changing the processing of audio data may involve changing one or more of an upmixing process, a downmixing process, a virtual bass process, a bass distribution process, an equalization process, a crossover filter, a delay filter, a multiband limiter, or a virtualization process at least in part based on at least one of a noise metric or a speech intelligibility metric. Some examples are described below.
[0130] Some embodiments of method 300 may involve transmitting audio data from a first device to a second device. Some such embodiments may involve transmitting at least one of a noise metric, a speech intelligibility metric, or echo reference data from the first device to the second device or from the second device to the first device. In some such examples, the second device may be a hearing aid, a personal sound amplification product, a cochlear implant, or a headset.
[0131] Some examples of method 300 may involve receiving, by a second device control system, a second device microphone signal, and receiving, by the second device control system, audio data and at least one of a noise metric, a speech intelligibility metric, or an echo reference data. Some such embodiments may involve determining, by the second device control system, one or more audio data gain settings and one or more second device microphone signal gain settings. Some such embodiments may involve applying, by the second device control system, the audio data gain settings to the audio data to produce gain-adjusted audio data. In some such examples, method 300 may involve applying, by the second device control system, the second device microphone signal gain settings to the second device microphone signal to produce gain-adjusted second device microphone signal. Some such embodiments may involve mixing, by the second device control system, the gain-adjusted audio data and the gain-adjusted second device microphone signal to produce mixed second device audio data. Some such examples may involve providing, by the second device control system, the mixed second device audio data to one or more second device transducers and reproducing, by the one or more second device transducers, the mixed second device audio data. Some such embodiments may involve controlling, by the second device control system, a relative level of the gain-adjusted audio data and the gain-adjusted second device microphone signal in the mixed second device audio data based at least in part on the noise metric. Some examples are described below.
[0132] Figure 4 Examples of systems are shown in which a closed captioning system is controlled based on at least one of a noise estimate or an intelligibility estimate. Similar to other Figure 1 aspects provided herein, Figure 4 the types and numbers of elements shown are provided only as examples. Other embodiments may include more, fewer, and / or different types and numbers of elements. In this example, system 400 is configured to turn on and off a closed captioning system based on a noise estimate (which, depending on the particular embodiment, may be configured to provide closed captions, lyrics captions, or dialogue captions). In this example, if the estimated noise level is too high, the closed captions are displayed. If the estimated noise level is too low, then in this example the closed captions are not displayed. This allows system 400 to respond in noisy environments where, in some examples, the loudspeakers used to reproduce the speech may be too limited to overcome the noise. In other instances (e.g., when content is being provided to a hearing-impaired user), such embodiments may also be advantageous.
[0133] In this example, control system 210 (referred to above with reference to Figure 2An example of the described control system 210) includes a noise estimator 104, a closed caption system controller 401, and a video display control 403. According to this embodiment, the closed caption system controller 401 is configured to receive a multi-band signal and determine whether to turn on closed captions based on the noise in the band corresponding to speech. To prevent the closed captions from turning on and off too frequently, in some embodiments, the closed caption system controller 401 may implement a certain amount of hysteresis, whereby the threshold for turning on the closed captions and the threshold for turning off the closed captions may be different. This can be advantageous in various contexts, such as in the case of a periodic noise source such as a fan that hovers around the noise threshold (which would cause the text to flicker in a single-threshold system). In some embodiments, the threshold for turning on the closed captions is lower than the threshold for turning off the closed captions. According to some embodiments, the closed caption system controller 401 may be configured to allow the turning on and off of the text only when new text should be displayed on the screen.
[0134] In Figure 4 the example shown, the closed caption system controller 401 is sending an enable control signal 402 to the video display control 403 to enable the display of closed captions. If the closed caption system controller 401 determines that the display of the closed captions should be stopped, then in this example, the closed caption system controller 401 will send a disable control signal to the video display control 403.
[0135] According to this example, the video display control 403 is configured to superimpose the closed captions on the video frame of the content being displayed on the television 405 in response to having received the enable control signal 402. In this embodiment, the television 405 is an example of the optional display system 235 described above with reference to Figure 2 In this example, the video display control 403 is sending a video frame 404 with superimposed closed captions to the television 405. The closed captions 406 are shown as being displayed on the television 405. In this example, the video display control 403 is configured to stop superimposing the closed captions on the video frame of the content in response to having received the disable control signal.
[0136] Figure 5An example of a graph related to the control of a closed captioning system is shown. In this embodiment, graph 500 shows an example of the behavior of a closed captioning system that estimates a set of noises to show how the closed captioning is turned on and off. According to this example, when the average noise level is higher than a first threshold (in this example, threshold 506), the closed captioning is turned on, and if the average noise level is lower than a second threshold (in this example, threshold 505), the closed captioning is turned off. According to some examples, the average value can be measured during a time interval in the range of 1 second or 2 seconds. However, in other embodiments, the average value can be measured across longer or shorter time intervals. In some alternative embodiments, the threshold can be based on the maximum noise level, the minimum noise level, or the median noise level.
[0137] In this example, the vertical axis 501 indicates the sound pressure level (SPL) and the horizontal axis 502 indicates the frequency. According to this example, threshold 506 is the average level above which the noise must be for the control system or a part thereof (in this example, Figure 4 the closed captioning system controller 401) to turn on the closed captioning. In this example, threshold 505 is the average level below which the noise must be for the closed captioning system controller 401 to turn off the closed captioning.
[0138] In some examples, threshold 506 and / or threshold 505 can be adjustable according to user input. Alternatively or additionally, threshold 506 and / or threshold 505 can be automatically adjusted by a control system (such as Figure 2 or Figure 4 control system 210) based on information previously obtained about one or more capabilities of the user (e.g., according to the user's hearing acuity). For example, if the user has a hearing impairment, threshold 506 can be made relatively low.
[0139] According to some examples, threshold 506 and / or threshold 505 can correspond to the capabilities of one or more audio reproduction transducers of the environment and / or the limitations imposed on the playback volume. In some such examples, since an upper limit on the playback volume is imposed, the closed captioning system controller 401 can turn on the closed captioning, where the compensation for the playback volume may not exceed this upper limit. In some such instances, the noise level can reach a point where the playback level may not be further increased to compensate for the ambient noise, resulting in the ambient noise (at least partially) masking the playback content. In such cases, closed captions, dialogue captions, or lyric captions can be desirable or even necessary in order for the user to understand the dialogue or other speech.
[0140] Curve 503 shows an example of a noise level that is on average below threshold 505, where threshold 505 is the closed captioning off threshold. In this embodiment, the closed captioning system controller 401 will turn off the closed captioning regardless of the previous state. Curve 504 shows an example of a noise level that is on average above threshold 506, where threshold 506 is the closed captioning on threshold. In this embodiment, the closed captioning system controller 401 will turn on the closed captioning regardless of the previous state. According to this example, curve 507 shows an example of a noise level that is on average between the on and off thresholds of the closed captioning. In this embodiment, if the closed captioning was previously on before the noise estimate entered this region, the closed captioning system controller 401 will keep the closed captioning on. Otherwise, the closed captioning system controller 401 will turn off the closed captioning. In cases where the noise estimate hovers around a single threshold that is used to turn the closed captioning on or off, such a hysteresis-based embodiment has the potential advantage of preventing closed captioning from flickering.
[0141] Figure 6 Shows an example of an intelligibility metric evaluation module. Figure 6 Shows an example of a subsystem that receives an incoming audio stream and then obtains a measure of the intelligibility of the speech present in that stream. According to this embodiment, the intelligibility metric evaluation module 602 is configured to estimate speech intelligibility. The intelligibility metric evaluation module 602 may be implemented via a control system (such as the control system 210 described in reference Figure 2 As described). In this example, the intelligibility metric evaluation module 602 is receiving content stream 601. Here, content stream 601 includes audio data, which in some instances may be or may include audio data corresponding to speech. According to this embodiment, the intelligibility metric evaluation module 602 is configured to output a speech intelligibility metric (SIM) 603. In this example, the intelligibility metric evaluation module 602 is configured to output a time-varying information stream indicative of speech intelligibility.
[0142] Depending on the particular embodiment, the intelligibility metric evaluation module 602 may estimate speech intelligibility in various ways. There are currently multiple methods for estimating speech intelligibility, and more such methods are expected in the future.
[0143] According to some embodiments, the intelligibility metric assessment module 602 may estimate speech intelligibility by analyzing audio data according to one or more methods for determining the speech intelligibility of young children and / or persons with speech or hearing impairments. In some such examples, the intelligibility of each word of a speech sample may be evaluated and an overall score for the speech sample may be determined. According to some such examples, the overall score for the speech sample may be a ratio I / T, which is determined by dividing the number of intelligible words I by the total number of words T. In some such examples, whether a word is intelligible or not may be determined based on an automatic speech recognition (ASR) confidence score. For example, a word having an ASR confidence score at or above a threshold may be considered intelligible, while a word having an ASR confidence score below the threshold may be considered unintelligible. According to some examples, the text corresponding to the speech may be provided to the control system as the "ground truth" of the actual words of the speech. In some such examples, the overall speech intelligibility score for the speech sample may be a ratio C / T, which is determined by dividing the number of words C correctly identified by the ASR process by the total number of words T.
[0144] In some embodiments, the intelligibility metric assessment module 602 may estimate speech intelligibility by analyzing audio data according to published metrics such as the speech transmission index (STI), the common intelligibility scale (CIS), or C50, where C50 is measured as the ratio of the sound energy arriving early (arriving between 0 ms and 50 ms) to the sound energy arriving late (arriving later than 50 ms).
[0145] According to some embodiments, the SIM may be at least partially based on one or more user characteristics, e.g., characteristics of a person who is a user of a television or other device that will be used to reproduce a received content stream. In some examples, the one or more user characteristics may include at least one of the user's native language, the user's accent, the user's location in the environment, the user's age, and / or the user's capabilities. The user's capabilities may include, for example, the user's hearing ability, the user's language level, the user's accent comprehension level, the user's vision, and / or the user's reading comprehension.
[0146] In some such embodiments, the user may have previously provided input regarding one or more such user characteristics. According to some such examples, the user may have previously provided input according to a command from the control system via a graphical user interface (GUI) provided on a display device.
[0147] Alternatively or additionally, the control system may have inferred one or more user characteristics based on the user's past behavior, such as one or more languages the user has selected for the reproduced content, the user's demonstrated ability to understand languages and / or regional accents (e.g., as evidenced by instances where the user has selected a closed captioning system, a karaoke captioning system, or a dialogue captioning system), the relative complexity of the language used in the content selected by the user, the playback volume selected by the user (e.g., for the portion of the reproduced content corresponding to speech), whether the user has previously used a visual impairment device (such as a tactile display system), etc.
[0148] According to some alternative embodiments, the intelligibility metric assessment module 602 may estimate speech intelligibility via a machine learning-based method. In some such examples, the intelligibility metric assessment module 602 may estimate the speech intelligibility of the input audio data via a neural network that has been trained on a set of content for which the intelligibility is known.
[0149] Figure 7A An example of a closed captioning system controlled by an intelligibility metric is shown. Like other Figure 1 here, Figure 7A The types and quantities of elements shown are provided only as examples. Other embodiments may include more, fewer, and / or different types and quantities of elements. For example, in this example and other disclosed examples, the functions of a closed captioning system, a karaoke captioning system, or a dialogue captioning system may be described. Unless otherwise specified in this disclosure, the examples provided for any one such system are intended to apply to all such systems.
[0150] According to this example, Figure 7A A system 700 is shown in which an intelligibility metric is used to change the displayed captions. In this example, the system 700 includes an intelligibility metric assessment module 602 and a caption display renderer 703, both of which are implemented via the control system 210 described above with reference to Figure 2 In some embodiments, the system 700 may use knowledge of a display sensor or user capabilities (such as vision) to change the rendering of the font used in the closed captioning system.
[0151] According to this embodiment, the intelligibility metric assessment module 602 is configured to output a speech intelligibility metric (SIM) 603. In this example, the intelligibility metric assessment module 602 is configured to output a time-varying information stream including the speech intelligibility metric 603, as described above with reference to Figure 6 above.
[0152] In Figure 7AIn the example shown, subtitle display renderer 703 is receiving video stream 704 that includes video frames. In some embodiments, video stream 704 may include metadata that describes closed captions and / or descriptive text (such as "[Music playing]"). In this example, subtitle display renderer 703 is configured to obtain video frames from video stream 704, overlay closed captions on the frames, and output modified video frames 705 with the overlaid closed captions and / or descriptive text. According to this embodiment, input subtitle text is embedded within video stream 704 as metadata.
[0153] In some embodiments, subtitle display renderer 703 may be configured to change the content being displayed based on a plurality of factors. These factors may include, for example, user capabilities and / or display capabilities. User capabilities may include, for example, visual acuity, language level, accent comprehension level, reading ability, and / or mental state. To make the text easier to read for a particular user, subtitle display renderer 703 may change the font and / or font size based on one or more factors (e.g., increase the font size).
[0154] According to some examples, the factor may include an external input. In some instances, subtitle display renderer 703 may change the font and / or font size to ensure that the text is readable in an illumination environment having a particular light intensity or color spectrum. In some such examples, the environment may be measured via a light sensor, color sensor, or camera, and the corresponding sensor data may be provided to a control system 210, e.g., to subtitle display renderer 703. In Figure 7A One such example is shown in which subtitle display renderer 703 is shown receiving image-based information 706 that may be used to change the closed captions. Image-based information 706 may include, for example, user vision information, lighting conditions within a room, capabilities of a display on which the closed captions will be shown, etc.
[0155] According to some examples, subtitle display renderer 703 may respond to a changing illumination condition of the environment in a hysteresis-based manner, similar to the hysteresis-based response to a noise condition in the environment described above with reference to Figure 5 For example, some embodiments may involve a first light intensity threshold that triggers a change from the normal state of the closed captions (such as enlarging the text and / or bolding letters) and a second light intensity threshold that causes a return to the normal state of the closed captions. In some examples, the first light intensity threshold may correspond to a lower light intensity than the second light intensity threshold.
[0156] If the user has a low language ability and / or a low reading ability, and corresponding image-based information 706 has been provided to the subtitle display renderer 703, then in some embodiments, the subtitle display renderer 703 may modify the text to simplify the meaning and make the text easier to understand than a verbatim transcription of the audio data corresponding to the speech.
[0157] Figure 7B An example of an audio description renderer controlled by an intelligibility metric is shown. In this embodiment, the control system 210 is configured to provide the functionality of the audio description renderer 713. According to this example, the speech intelligibility metric 603 is used by the audio description renderer 713 to optionally mix an audio-based description of the content (referred to herein as an audio description) into the input audio stream 714 to produce an output audio stream 715 that includes the audio description.
[0158] According to some embodiments, the modified video frame 705 output from the subtitle display renderer 703 is optionally input to the audio description renderer 713. In some such instances, the modified video frame 705 may include subtitles and / or descriptive text. In some such examples, the audio description in the output audio stream 715 may be synthesized by the audio description renderer 713 from the subtitles and / or by analyzing the video and / or audio content input to the audio description renderer 713. In some embodiments, the audio description may be included in the input to the audio description renderer 713. According to some embodiments, the input audio description may be mixed with the audio description synthesized by the audio description renderer 713. In some examples, the mixing ratio may be at least partially based on the speech intelligibility metric 603.
[0159] In some embodiments, the audio description renderer 713 may be configured to vary the mixed content based on a plurality of factors. These factors may include, for example, user ability and / or display ability. User ability may include, for example, visual acuity, language level, accent comprehension level, reading ability, and / or mental state. In some embodiments, the outputs from the subtitle renderer 703 and the audio description renderer 713 may be used together to improve the user's understanding.
[0160] Figure 8An example of an echo predictor module is shown. The reverberation of an audio environment can have a significant impact on speech intelligibility. In some embodiments, the control system determining whether to turn on or off closed captions can be at least partially based on one or more metrics corresponding to the reverberation of the audio environment. According to some examples, such as in the case of extreme reverberation, the control system determining whether to turn on or off closed captions can be entirely based on one or more metrics corresponding to the reverberation of the audio environment. In some alternative embodiments, the control system determining whether to turn on or off closed captions can be partially based on one or more metrics corresponding to the reverberation of the audio environment, but can also be based on content-based speech intelligibility metrics, ambient noise, and / or other speech intelligibility metrics.
[0161] In Figure 8 the example shown, data 801 is provided to the echo predictor module 803. In this example, data 801 includes information about the characteristics of the audio environment. For example, data 801 can be obtained from sensors (such as microphones) in the audio environment. In some instances, at least a portion of data 801 can include direct user input.
[0162] In this example, a content stream 802 is also provided to the echo predictor module 803. The content stream 802 includes content to be presented in the audio environment, for example, via one or more display devices and audio reproduction transducers of the environment. The content stream 802 can correspond, for example, to the content stream 601 described above with reference to Figure 6 According to some embodiments, the content stream 802 can include a rendered audio signal.
[0163] According to this example, the echo predictor module 803 is configured to calculate the impact that room echo and reverberation will have on the intelligibility of a given piece of content. In this example, the echo predictor module 803 is configured to output a metric 804 of speech intelligibility, which can be a time-varying metric in some instances. For example, if the audio environment is highly reverberant, it will be difficult to understand speech and the metric 804 of intelligibility will generally be low. If the room is highly anechoic, the metric 804 of intelligibility will generally be high.
[0164] Figure 9 An example of a system configured to determine an intelligibility metric based at least in part on playback processing is shown. Similar to other Figure 1 as provided herein, Figure 9 the types and quantities of the elements shown are provided only as examples. Other embodiments can include more, fewer, and / or different types and quantities of elements.
[0165] In this example, Figure 9 includes the following elements:
[0166] 901: Input audio data;
[0167] 902: An equalization (EQ) filter (optional) for flattening the response of an audio reproduction transducer;
[0168] 903: A crossover and / or delay filter (optional);
[0169] 904: A multi-band limiter (optional) for avoiding non-linear behavior in a loudspeaker;
[0170] 905: A wideband limiter;
[0171] 906: An audio reproduction transducer (e.g., a loudspeaker);
[0172] 907: Characteristics of the EQ filter (e.g., frequency response, delay, ringing, etc.);
[0173] 908: Characteristics of crossover and delay;
[0174] 909: The current state of the multi-band limiter (e.g., the amount of limiting being applied by the multi-band limiter);
[0175] 910: The current state of the limiter (e.g., the amount of limiting being applied by the limiter);
[0176] 911: An environmental intelligibility metric module that takes into account the playback characteristics of a device with the current content;
[0177] 912: An intelligibility metric; and
[0178] 913: A reference of the audio reproduction transducer feed signal that can optionally be used to determine the intelligibility metric.
[0179] In some devices, it may not be possible to reproduce audio at the volume requested by the user. This may be due to the capabilities of the audio reproduction transducer, amplifier, or some other implementation details that result in limited dynamic headroom within the system.
[0180] For example, the frequency response of the system may be non-flat (e.g., due to the audio reproduction transducer itself, or in the case of a multi-driver due to the crossover and / or the placement of the audio reproduction transducer). A non-flat frequency response can reduce the intelligibility of speech. This situation can be alleviated by using the equalization filter 902 to flatten the frequency response. However, in some instances, even after applying the equalization filter, the device may still have a non-flat response. By measuring the speech intelligibility after applying the equalization filter, true speech intelligibility can be achieved, and then this speech intelligibility can be used to turn on and off closed captions.
[0181] Some embodiments may incorporate one or more multi-band limiters 904 and / or wide-band limiters 905 to ensure that components of the system are protected from exceeding their linear range.
[0182] In some multi-band systems, captions may be turned on only when limiting occurs for audio data within the voice frequencies (e.g., audio data between 50 Hz and 2 kHz or audio data within other publicly disclosed voice frequency ranges). In some alternative examples, system 900 may determine one or more intelligibility metrics based on audio data 913 after applying a limiter and just before sending the audio to speaker 906.
[0183] In some instances, a speaker may be driven into its non-linear region to obtain increased volume. Driving a speaker into its non-linear region results in distortion, such as intermodulation or harmonic distortion. According to some examples, in these cases, captions may be turned on whenever any non-linear behavior (estimated via a model or determined via direct microphone-based measurements (e.g., a linear echo canceller)) is present. Alternatively or additionally, the measured or modeled non-linear behavior may be analyzed to determine the intelligibility of speech in the presence of distortion.
[0184] Figure 10 An example of a system configured to determine an intelligibility metric based at least in part on an ambient noise level is shown. Like other Figure 1 presented herein, Figure 10 the types and numbers of elements shown are provided only as examples. Other embodiments may include more, fewer, and / or different types and numbers of elements.
[0185] In this example, Figure 10 includes the following elements:
[0186] 1001: A microphone configured to measure ambient noise;
[0187] 1002: A microphone signal;
[0188] 1003: A background noise estimator;
[0189] 1004: A background noise estimate;
[0190] 1005: Input content, including audio data and / or input metadata;
[0191] 1006: A playback processing module that may be configured for decoding, rendering, and / or post-processing;
[0192] 1007: An audio reproduction transducer feed and echo reference;
[0193] 1008: An audio reproduction transducer for the forward playback content;
[0194] 1009: An environmental intelligibility metric module configured to determine an intelligibility metric based at least in part on background environmental noise; and
[0195] 1010: An intelligibility metric.
[0196] In some examples, the ambient noise level can be used as an alternative to a pure speech intelligibility metric relative to the level of an input audio signal corresponding to speech. In some such embodiments, the environmental intelligibility metric module 1009 can be configured to combine the ambient noise level with a speech intelligibility metric or level to produce a combined intelligibility metric.
[0197] According to some embodiments, if the intelligibility level of speech is low and the ambient noise level is high, the environmental intelligibility metric module 1009 can be configured to output an intelligibility metric 1010 indicating that a compensation process will be enabled. According to some embodiments, the compensation process can involve altering the processing of the audio data in one or more ways disclosed herein. Alternatively or additionally, in some embodiments, the compensation process can involve applying a non-audio-based compensation method, such as enabling closed captions. In some such examples, if the ambient noise level is low and the speech intelligibility is high, the environmental intelligibility metric module 1009 can be configured to output an intelligibility metric 1010 indicating that closed captions will remain off. In some embodiments, if the speech intelligibility is high and the ambient noise level is high, depending on the combined intelligibility of the speech, the environmental intelligibility metric module 1009 can be configured to output an intelligibility metric 1010 indicating whether closed captions will be turned on or off.
[0198] Figure 11 An example of a system configured to modify an intelligibility metric based at least in part on one or more user capabilities is shown. Like other Figure 1 as Figure 11 The types and quantities of elements shown herein are provided only as examples. Other embodiments can include more, fewer, and / or different types and quantities of elements.
[0199] In this example, Figure 11 includes the following elements:
[0200] 1101: An input intelligibility metric or stream of intelligibility metrics, such as calculated by another process (e.g., one of the other disclosed methods for determining an intelligibility metric). In some examples, the input intelligibility metric can correspond to an output speech intelligibility metric (SIM) 603, which is determined by the above reference Figure 6Output by the intelligibility metric evaluation module 602 described;
[0201] 1102: User profile. Examples of user profiles can include profiles such as "Hearing Impaired", "No Adjustment Needed" (e.g., having average hearing ability), or "Superhuman" (e.g., having abnormally good hearing ability);
[0202] 1103: An intelligibility metric modifier configured to modify the input intelligibility metric 1101 at least in part based on one or more user capabilities; and
[0203] 1104: An adjusted intelligibility metric that takes into account the user profile.
[0204] The intelligibility metric modifier 1103 can be implemented, for example, via a control system (such as Figure 2 control system 210). There are various ways in which the intelligibility metric modifier 1103 can be used to modify the input intelligibility metric 1101. In some examples, if the user profile indicates that the user has a hearing impairment, the intelligibility metric modifier 1103 can be configured to decrease the input intelligibility metric 1101. The amount of decrease can correspond to the degree of the hearing impairment. For example, if the input intelligibility metric 1101 is 0.7 on a scale from 0 to 1.0 and the user profile indicates that the degree of the hearing impairment is mild, then in one example, the intelligibility metric modifier 1103 can be configured to decrease the input intelligibility metric 1101 to 0.6. In another example, if the input intelligibility metric 1101 is 0.8 on a scale from 0 to 1.0 and the user profile indicates that the degree of the hearing impairment is moderate, then in one example, the intelligibility metric modifier 1103 can be configured to decrease the input intelligibility metric 1101 to 0.6 or 0.5.
[0205] In some embodiments, if the user profile indicates that the user is "Superhuman" (having abnormally good hearing ability, abnormally good language level, abnormally good accent understanding level, etc.), the intelligibility metric modifier 1103 can be configured to increase the input intelligibility metric 1101. For example, if the input intelligibility metric 1101 is 0.5 on a scale from 0 to 1.0 and the user profile indicates that the user has abnormally good hearing ability, then in one example, the intelligibility metric modifier 1103 can be configured to increase the input intelligibility metric 1101 to 0.6.
[0206] In some examples, a user profile may include a frequency-based hearing profile of the user. According to some such examples, the intelligibility metric modifier 1103 may be configured to determine whether to change the input intelligibility metric 1101 at least in part based on the frequency-based hearing profile. For example, if the frequency-based hearing profile indicates that the user has normal hearing ability in the frequency range corresponding to speech, the intelligibility metric modifier 1103 may determine that the input intelligibility metric 1101 will not be changed. In another example, if the input intelligibility metric 1101 is 0.8 on a scale from 0 to 1.0 and the frequency-based hearing profile indicates that the user has a moderate level of hearing impairment in the frequency range corresponding to speech, the intelligibility metric modifier 1103 may be configured to reduce the input intelligibility metric 1101 to 0.6 or 0.5.
[0207] In some alternative embodiments, the user frequency response profile may be directly applied to the input audio. For example, referring to Figure 6 , in some such embodiments, the control system may be configured to multiply the user's hearing response to the frequency domain representation of the audio portion of the content 601 in the frequency domain and then input the result into the intelligibility metric assessment 602.
[0208] Referring again to Figure 11 , examples of some other ways in which the intelligibility metric modifier 1103 can be used to modify the input intelligibility metric 1101 are described below.
[0209] User characteristics
[0210] Language
[0211] Speech intelligibility and the need for closed captions may be related to whether the user is a native speaker of the language of the speech the user is listening to. Thus, in some embodiments, user input regarding language comprehension ability and / or regional accent comprehension ability (and / or data obtained during the user's previous viewing / listening events, such as instances where the user manually turned on dialogue subtitles) may be used as a modifier of the speech intelligibility metric. For example, if the user appears to be proficient in a certain language and the content will be presented in that language, the intelligibility metric modifier 1103 may determine that no change will be made to the input intelligibility metric 1101.
[0212] However, if the user appears to have limited capabilities in that language, the intelligibility metric modifier 1103 can determine that the input intelligibility metric 1101 will be reduced proportionally to the limitations of the user's language level. For example, if the user appears to have little or no understanding of the language, the intelligibility metric modifier 1103 can determine that the input intelligibility metric 1101 will be reduced to zero. In some instances, the intelligibility metric modifier 1103 can include metadata with an adjusted intelligibility metric 1104 that indicates that dialogue subtitles should be presented and indicates the language in which the dialogue subtitles will be presented. For example, if the content voice is in English, the user profile indicates that the user has little or no understanding of English, and the user is proficient in French, the metadata can indicate that French dialogue subtitles should be presented.
[0213] In another example, if the user appears to have a moderate level of language listening comprehension, the intelligibility metric modifier 1103 can determine that the input intelligibility metric 1101 will be reduced by one half. In some such examples, the user profile can indicate that the user's reading comprehension of the same language is sufficient to understand the text corresponding to the content's voice. In some such examples, the intelligibility metric modifier 1103 can be configured to include metadata with an adjusted intelligibility metric 1104 that indicates that dialogue subtitles should be presented and indicates the language in which the dialogue subtitles will be presented. For example, if the content voice is in English and the user profile indicates that the user's reading comprehension of English is sufficient to understand the text corresponding to the English voice of the content, the metadata can indicate that English dialogue subtitles should be presented.
[0214] In some instances, a user may have a spoken level of a language without having a reading level, and vice versa. Some listeners may prefer closed captions and / or dialogue subtitles, while other listeners may not. For example, some users may prefer dubbed audio over dialogue subtitles. Thus, in some embodiments, user input and / or observed user behavior regarding such preferences can be used to automatically accommodate such differences. In some embodiments, user preference data can indicate the user's primary preferred language, secondary preferred language, etc. and the user's preference for dialogue subtitles or dubbed audio.
[0215] In one example, a user may have a French proficiency of 100%, an English proficiency of 50%, and a German proficiency of 25%. If the received broadcast video content has an English native audio, with options for language tracks with French and German dubbing and dialogue subtitles in all three languages, then according to some embodiments, an intelligibility metric can be used to select (1) an audio playback track, (2) a dialogue subtitle track, or (3) a combination of both. In this example, a 50% level may be sufficient to default to the English audio and use French dialogue subtitles as an aid. In some instances, the user may have explicitly indicated a preference for listening to the English audio when the received broadcast has an English audio. In other instances, the user's selection of the English audio can be recorded and / or used to update user preference data. The user may prefer to listen to the English audio and watch French dialogue subtitles rather than experiencing the French audio with dubbing, so that the content can be experienced as created with better lip-sync, the voices of the original actors, etc. Each country has regional preferences for dialogue subtitles and dubbing (most Americans prefer dialogue subtitles, while most Germans prefer dubbing). In some embodiments, if no specific user preference data is available, country-based or region-based defaults can be used to determine whether to present dialogue subtitles or dubbing.
[0216] In some embodiments, the control system can be configured to select a combination of an audio playback track and a dialogue subtitle track that achieves the highest estimated level of user intelligibility. Alternatively or additionally, some embodiments can involve at least selecting an acceptable minimum level of intelligibility, e.g., determined by an intelligibility metric threshold.
[0217] Accent
[0218] The accent of the content, along with whether the user is accustomed to the accent, can affect the user's speech intelligibility. There are various methods for determining whether the user will be accustomed to the accent. For example, in some embodiments, a quick user input regarding the user's preference can be provided to the intelligibility metric modifier 1103. In other cases, the location of one or more devices for playback (e.g., the location of the TV) can be compared with a data structure that includes one or more sets of known regional accents and corresponding locations. If the user may not be accustomed to the accent (e.g., a listener in Canada watching an Australian TV program), then the intelligibility metric modifier 1103 can be configured to lower the input intelligibility metric 1101 corresponding to a Canadian listener watching a Canadian program.
[0219] User ability
[0220] In some cases, a user may suffer from conditions that make text difficult to read or reduce attention (such as dyslexia or ADHD). In some such examples, the intelligibility metric modifier 1103 may be configured to include metadata with an adjusted intelligibility metric 1104 that indicates that closed captions should be turned off because the user will not benefit from closed captions. In other embodiments, the intelligibility metric modifier 1103 may be configured to include metadata with an adjusted intelligibility metric 1104 that indicates that the text of the closed captions should be simplified and / or the font size of the text should be increased in response to the user's condition. In some embodiments, the dialogue captions may include less text (a simplified version of the speech), and the closed captions may include more text (a complete or substantially complete version of the speech). According to some such embodiments, simplifying the text may involve presenting dialogue captions instead of closed captions.
[0221] Age and reading comprehension
[0222] The age and / or reading comprehension of the listener can affect the determination of whether the speech intelligibility metric should be modified and / or whether closed captions should be used. For example, the control system can determine not to turn on closed captions for a television program watched only by people who cannot read (such as young children).
[0223] If a viewer has difficulty understanding rapid speech (a common characteristic of the elderly), then according to some examples, the intelligibility metric modifier 1103 may be configured to reduce the input intelligibility metric 1101 at least in part based on the speech rate of the content (e.g., the rhythm of the speech in the audio). For example, the speech rate of the content can be determined based on the number of words per unit of time. In some such embodiments, if the speech rate of the content is at or above a threshold level, the control system can turn on the closed captions.
[0224] Hearing profile
[0225] In some cases, a listener may have lost some ability to hear certain frequencies. According to some embodiments, such a condition can be used as a basis for changing the speech intelligibility (e.g., if the listener has lost some ability to hear speech frequencies). For example, the control system can be configured to apply a mathematical representation of a person's hearing profile (e.g., a representation of the frequency response of the person's ear, including hearing loss) to the input audio before calculating the speech intelligibility metric. Such embodiments can increase the probability of turning on and off closed captions at the appropriate times. The hearing profile can be provided to the control system by the user (e.g., via an interactive test process or via user input) or from one or more other devices (such as the user's hearing aid or cochlear implant).
[0226] Figure 12Shows an example of a caption generator. The caption generator 1202 can be implemented via a control system (e.g., Figure 2 's control system 210). In this example, the caption generator 1202 is configured to automatically synthesize captions 1203 corresponding to the speech in the input audio stream 1201 at least in part based on an ASR process. According to this example, the caption generator 1202 is also configured to modify the content of the captions 1203 based on an input intelligibility metric 1204. Depending on the specific implementation, the type of the intelligibility metric 1204 can vary. In some examples, the intelligibility metric 1204 can be an adjusted intelligibility metric that takes into account a user profile, such as the adjusted intelligibility metric 1104 described above with reference to Figure 11 . Alternatively or additionally, in some examples, the intelligibility metric 1204 can be an adjusted intelligibility metric that takes into account the characteristics of the audio environment (such as ambient noise and / or reverberation). One such example is the intelligibility metric 1010 described above with reference to Figure 10 .
[0227] In some embodiments, if the intelligibility metric 1204 indicates that the intelligibility is medium, captions indicating descriptive text such as "[Playing music]" can be omitted and only speech captions can be included. In some such examples, as the intelligibility decreases, more descriptive text can be included.
[0228] Figure 13 Shows an example of a caption modifier module configured to change captions based on an intelligibility metric. According to this example, the caption modifier module 1302 is configured to receive the captions 1203 output from the Figure 12 's caption generator 1202. In this example, the captions 1203 are included within a video stream. According to this example, the caption modifier module 1302 is also configured to receive the intelligibility metric 1204 and determine whether and how to modify the captions based on the intelligibility metric 1204.
[0229] In some embodiments, the caption modifier module 1302 can be configured to increase the font size to improve text intelligibility when the intelligibility metric 1204 is low. In some such examples, the caption modifier module 1302 can also be configured to change the font type to improve text intelligibility when the intelligibility metric 1204 is low.
[0230] According to some examples, the caption modifier module 1302 can be configured to apply a caption "filter" that can potentially reduce the number of captions in the modified caption stream 1303 depending on the intelligibility metric 1204. For example, if the intelligibility metric 1204 is low, the caption modifier module 1302 may not filter out many (and in some instances may not filter out any) captions. If the intelligibility metric 1204 is high, the caption modifier module 1302 can determine that the required number of captions has been reduced. For example, the caption modifier module 1302 can determine that descriptive captions such as "[Music playing]" are not needed, but speech captions are needed. Thus, the descriptive captions will be filtered out, but the speech captions will remain in the modified caption stream 1303.
[0231] In some embodiments, the caption modifier module 1302 can be configured to receive user data 1305. According to some such embodiments, the user data 1305 can indicate one or more of the user's native language, the user's accent, the user's location in the environment, the user's age, and / or the user's capabilities. Data related to one or more of the user's capabilities can include data related to the user's hearing ability, the user's language level, the user's accent comprehension level, the user's vision, and / or the user's reading comprehension.
[0232] According to some examples, if the user data 1305 indicates that the user's capabilities are low and / or if the intelligibility metric 1204 is low, the caption modifier module 1302 can be configured to simplify or paraphrase (e.g., using a language engine) the text of the closed captions to increase the likelihood that the user can understand the captions shown on the display. In some examples, the caption modifier module 1302 can be configured to simplify speech-based text for non-native speakers and / or to present the text in a relatively large font size for people with vision problems. According to some examples, if the user data 1305 indicates that the user's capabilities are high and / or if the intelligibility metric 1204 is high, the caption modifier module 1302 can be configured to leave the text of the closed captions unchanged.
[0233] In some examples, the caption modifier module 1302 can be configured to filter the text of the closed captions to remove specific phrases due to the user's preferences, age, etc. For example, the caption modifier module 1302 can be configured to filter out text corresponding to swear words and / or slang.
[0234] Figure 14 Further examples of non-audio compensation processes and systems that can be controlled based on a noise estimator are shown. These systems can be turned on and off in a manner similar to that described above for closed captioning system control.
[0235] In this example, the control system of the television 1401 incorporates a noise estimation system and / or a speech intelligibility system. The control system is configured to make an estimate of the ambient noise and / or speech intelligibility based on microphone signals received from one or more microphones of the environment 1400. In an alternative example, the control system for one or more elements of an audio system (e.g., a smart speaker including one or more microphones) can incorporate a noise estimation system. The control system can be an Figure 2 instance of the control system 210.
[0236] According to this example, the tactile display system 1403 is an electrically controlled braille display system configured to generate braille text. In this example, a signal 1402 has been transmitted from the control system indicating that the tactile display system 1403 should be turned on or that the braille text should be simplified. The signal 1402 can have been transmitted from the control system after the noise estimate has reached or exceeded a threshold (e.g., as described elsewhere herein). Such an implementation can allow a blind or visually impaired user to understand speech when the ambient noise becomes too high for the user to easily understand the audio version of the speech.
[0237] In this example, the seat vibrator 1405 is configured to vibrate the seat. For example, the seat vibrator 1405 can be used to at least partially compensate for the lack of low-frequency performance of the speakers within the television 1401. In this example, a signal 1404 has been sent from the control system indicating that the seat vibrator should be started. The control system can, for example, send the signal 1404 in response to determining that the noise estimate has reached or exceeded a threshold in a certain frequency band (e.g., the low-frequency band). According to some examples, if the noise estimation system determines that the ambient noise continues to increase, the control system will progressively route the low-frequency audio to the seat vibrator 1405.
[0238] As described elsewhere in this disclosure, some of the disclosed compensation processes that can be invoked by the control system in response to a noise metric or a speech intelligibility metric involve audio-based processing methods. In some such examples, the audio-based processing methods can also at least partially compensate for the limitations of one or more audio reproduction transducers, e.g., compensate for the limitations of a television audio reproduction transducer system. Some such audio processing can involve audio simplification (such as audio scene simplification) and / or audio enhancement. In some examples, audio simplification can involve removing one or more components of the audio data (e.g., leaving only the more important parts). Some audio enhancement methods can involve adding audio to the overall audio reproduction system, e.g., adding audio to a relatively more capable audio reproduction transducer of the environment.
[0239] Figure 15An example of a noise compensation system is shown. In this example, system 1500 includes a noise compensation module 1504 and a processing module 1502. In this instance, both the noise compensation module and the processing module are implemented via a control system 210. The processing module 1502 is configured to process input audio data 1501. In some examples, the input audio data can be an audio signal from a file or an audio signal from a streaming media service. The processing module 1502 can include, for example, an audio decoder, an equalizer, a multi-band limiter, a wide-band limiter, a renderer, an upmixer, a voice enhancer, and / or a bass distribution module.
[0240] In this example, the noise compensation module 1504 is configured to determine the ambient noise level and send a control signal 1507 to the processing module 1502 when the noise compensation module 1504 determines that the ambient noise level is at or above a threshold. For example, if the processing module 1502 includes a decoder, in some examples, the control signal 1507 can instruct the decoder to decode an audio stream with relatively low quality to save power. In some examples, if the processing module 1502 includes a renderer, in some examples, the control signal 1507 can instruct the renderer to only render high-priority audio objects when the noise level is high. In some examples, if the processing module 1502 includes an upmixer, in some examples, the control signal 1507 can instruct the upmixer to discard diffuse audio content so that only direct content will be reproduced.
[0241] In some embodiments, the processing module 1502 and the noise compensation module 1504 can reside in more than one device. For example, a certain audio can be upmixed to other devices based on noise estimation. Alternatively or additionally, in some embodiments (e.g., embodiments for blind or visually impaired persons), the control system 210 can request a high-noise-level audio description or a low-noise-level audio description from the source of the input audio data 1501. The high-noise-level audio description can correspond to relatively less content within the input audio data 1501, while the low-noise-level audio description can correspond to relatively more content within the input audio data 1501. In some embodiments, these audio streams can be included within a multi-stream audio codec (such as Dolby TrueHD). In an alternative embodiment, these audio streams can be decrypted by the control system 210 and then re-synthesized (e.g., for closed captions).
[0242] Figure 16An example of a system configured to perform speech enhancement in response to detected ambient noise is shown. In this example, system 1600 includes a speech enhancement module 1602 and a noise compensation module 1604, both of which are implemented via a control system 210. According to this embodiment, the noise compensation module 1604 is configured to determine the ambient noise level and send a signal 1607 to the speech enhancement module 1602 if the noise compensation module 1604 determines that the ambient noise level is at or above a threshold. In some embodiments, the noise compensation module 1604 may be configured to determine the ambient noise level and send a signal 1607 corresponding to the ambient noise level to the speech enhancement module 1602 regardless of whether the ambient noise level is at or above the threshold.
[0243] According to this example, the speech enhancement module 1602 is configured to process input audio data 1601, which in some examples may be an audio signal from a file or an audio signal from a streaming media service. For example, the audio data 1601 may correspond to video data such as a movie or a television program. In some embodiments, the speech enhancement module 1602 may be configured to receive an ambient noise estimate from the noise compensation module 1604 and adjust the amount of speech enhancement applied based on the ambient noise estimate. For example, if the signal 1607 indicates high ambient noise, then in some embodiments, the speech enhancement module 1602 will increase the amount of speech enhancement because speech intelligibility becomes more challenging in the presence of high ambient noise levels.
[0244] The type and degree of speech enhancement caused by configuring the speech enhancement module 1602 depends on the particular embodiment. In some examples, the speech enhancement module 1602 may be configured to reduce the gain of non-speech audio data. Alternatively or additionally, the speech enhancement module 1602 may be configured to increase the gain of speech frequencies.
[0245] In some embodiments, the speech enhancement module 1602 and the noise compensation module 1604 may be implemented in more than one device. For example, in some embodiments, the speech enhancement features of a hearing aid may be controlled (at least in part) based on a noise estimate from another device (e.g., a television).
[0246] According to some examples, the speech enhancement module 1602 may be used when one or more audio reproduction transducers of the environment reach their limits in order to remove audio from the system and thus enhance clarity. When the audio reproduction transducers do not reach the limits of their linear range, a speech enhancement of the speech frequency emphasis type may be used.
[0247] Figure 17Shows an example of a graph corresponding to elements of a system limited by audio reproduction transducer characteristics. The elements shown on graph 1700 are as follows:
[0248] Limiting line 1701 represents the upper limit of the audio reproduction transducer before limiting occurs. For example, limiting line 1701 can represent the limit determined by a microphone model (e.g., multi-band limiter tuning). In this simple example, limiting line 1701 is the same for all indicated frequencies, but in other examples, limiting line 1701 can have different levels corresponding to different frequencies.
[0249] According to this example, curve 1702 represents the output sound pressure level (SPL) that the audio reproduction transducer (the audio reproduction transducer corresponding to limiting line 1701) is generating at the microphone. In this example, curve 1703 represents the noise estimate at the microphone.
[0250] In this example, difference 1704 represents the difference between the noise estimate and the output sound pressure level, which in some instances can be a frequency-dependent difference. As difference 1704 gets smaller, in some embodiments, one or more features will be gradually enabled to increase the likelihood that the user can continue to understand and appreciate the content despite the ambient noise of the audio environment. According to some embodiments, these features can be gradually turned on on a per-band basis.
[0251] The following paragraphs describe examples of how the control system can control various components of the system at least in part based on difference 1704.
[0252] Voice enhancement
[0253] As the magnitude of difference 1704 decreases, the control system can increase the amount of voice enhancement. There are at least two forms of voice enhancement that can be controlled in this way:
[0254] · A de-voicing voice enhancer, where as the magnitude of difference 1704 decreases, the gain of non-voice channels / audio objects is attenuated (volume or level is decreased). In some such examples, the control system can make the de-voicing gain negatively correlated with the difference of difference 1704 (e.g., an inverse linear relationship).
[0255] · An enhanced voice enhancer that emphasizes voice frequencies (e.g., increases the level of voice frequencies). In some such examples, as the magnitude of difference 1704 decreases, the control system can apply more gain to voice frequencies (e.g., an inverse linear relationship). Some embodiments can be configured to enable both the de-voicing voice enhancer and the enhanced voice enhancer simultaneously.
[0256] Audio object rendering
[0257] As the magnitude of the difference 1704 decreases, in some embodiments, the control system may decrease the number of rendered audio objects. In some such examples, the audio objects to be rendered or discarded may be selected based on an audio object priority field within the object audio metadata. According to some embodiments, audio objects of interest (e.g., audio objects corresponding to speech) may be rendered relatively closer to the listener and / or relatively farther from noise sources within the environment.
[0258] Upmixing
[0259] As the magnitude of the difference 1704 decreases, in some embodiments, the control system may change the total energy within the mix. For example, as the difference 1704 decreases, the control system may cause the upmixing matrix to copy the same audio to all audio channels so that all of the audio channels within the audio channels act as a common mode. As the difference 1704 decreases, the control system may cause the spatial fidelity to be preserved (e.g., no upmixing occurs). Some embodiments may involve discarding the diffuse audio. Some such examples may involve rendering non-diffuse content to all audio channels.
[0260] Downmixing
[0261] As the magnitude of the difference 1704 decreases, in some embodiments, the control system may discard the less important channels within the mix. According to some alternative embodiments, the control system may cause the less important channels to be de-emphasized, e.g., by applying a negative gain to the less important channels.
[0262] Virtual bass
[0263] As the magnitude of the difference 1704 decreases within the bass band, in some embodiments, the control system may turn on and off a virtual bass algorithm that relies on the missing harmonic effect to attempt to overcome the loudness of the noise source.
[0264] Bass distribution
[0265] In some bass propagation / distribution methods, the bass in all channels may be extracted by a low-pass filter, summed into one channel, and then remixed as a common mode into all channels. According to some such methods, a high-pass filter may be used to allow non-bass frequencies to pass. In some embodiments, the control system may increase the cut-off frequency of the low-pass / high-pass combination as the magnitude of the difference 1704 decreases. As the magnitude of the difference 1704 increases, the control system may cause the cut-off frequency to approach zero and thus no bass will be propagated.
[0266] Virtualizer
[0267] As the magnitude of the difference 1704 decreases, in some examples, the control system can cause the virtualization to decrease until the virtualization is completely turned off. In some embodiments, the control system can be configured to calculate a virtualized version and a non-virtualized version of an audio stream and a cross-fade therebetween, where the weighting of each component corresponds to the magnitude of the difference 1704.
[0268] Figure 18 An example of a system in which a hearing aid is configured to communicate with a television is shown. In this example, system 1800 includes the following elements:
[0269] 1801: A television incorporating a noise compensation system;
[0270] 1802: A microphone that measures ambient noise;
[0271] 1803: A wireless transmitter configured to convert digital or analog audio into a wireless stream that can be received by hearing aid 1807;
[0272] 1804: A digital or analog audio stream. In some examples, the digital or analog audio stream can incorporate metadata to assist the hearing aid in changing the mixing of real-world audio and the wireless stream, changing the amount of speech enhancement, or applying some other noise compensation method;
[0273] 1805: Audio transmitted wirelessly;
[0274] 1806: A user with a hearing impairment;
[0275] 1807: A hearing aid. In this example, hearing aid 1807 is configured to communicate with television 1801 via a wireless protocol (e.g., Bluetooth);
[0276] 1808: A noise source.
[0277] Like other Figure 1 as Figure 18 The types and numbers of elements shown in are provided only as examples. Other embodiments can include more, fewer, and / or different types and numbers of elements. For example, in some alternative embodiments, a personal sound amplification product, a cochlear implant, or other hearing assistance device, or a headset can be configured to communicate with television 1801 and can also be configured to perform some or all of the operations described herein with reference to hearing aid 1807.
[0278] According to this example, the audio 1805 from the television 1801 is sent via a wireless protocol. The hearing aid 1807 is configured to mix the wireless audio with the ambient noise in the room to ensure that the user can hear the television audio while still being able to interact with the real world. The television 1801 incorporates a noise compensation system, which can be one of the types of noise compensation systems disclosed herein. In this instance, a noise source 1808 is present within the audio environment. The noise compensation system incorporated into the television 1801 is configured to measure the ambient noise and transmit information and / or control signals indicating the required mixing amount to the hearing aid 1807. For example, in some embodiments, the higher the ambient noise level, the higher the degree of mixing of the television audio signal with the signal from the hearing aid microphone. In some alternative embodiments, the hearing aid 1807 may incorporate some or all of the noise estimation system. In some such embodiments, the hearing aid 1807 may be configured to implement volume control within the hearing aid. In some such embodiments, the hearing aid 1807 may be configured to adjust the volume of the television.
[0279] Figure 19 An example of the mixing and speech enhancement components of a hearing aid is shown. In this example, the hearing aid 1807 includes the following elements:
[0280] 1901: Noise estimation of the audio environment. In this example, the noise estimation is provided by the noise compensation system of the television. In an alternative embodiment, the noise estimation is provided by the noise compensation system of another device (e.g., the hearing aid 1807);
[0281] 1902: Gain setting of the television audio stream;
[0282] 1903: Gain setting of the hearing aid microphone stream;
[0283] 1904: Audio in the television stream;
[0284] 1905: Audio in the hearing aid stream, which is provided by one or more hearing aid microphones in this example;
[0285] 1906: Gain-adjusted audio in the hearing aid audio stream before summation (which can be frequency-dependent gain in the case of a speech enhancer);
[0286] 1907: Gain-adjusted audio in the television audio stream before summation (which can be frequency-dependent gain in the case of a speech enhancer);
[0287] 1908: Summation block that produces the mixed audio stream;
[0288] 1909: The mixed audio stream to be played to the speaker of the hearing aid (or via cochlear implant electrodes in other examples);
[0289] 1910: An enhancement control module configured to adjust the gain (which can be frequency - dependent gain in some instances) based on the noise estimate 1901 sent by the television. In some embodiments, the enhancement control module can be configured to keep the volume level constant or within a certain volume level range and change the ratio of the hearing aid microphone and the television audio stream based on the noise estimate 1901, for example, by making the sum of the gains equal to a predetermined value such as one;
[0290] 1911: A gain application block for the television stream. In some embodiments (e.g., in the case of a simple mixer), the gain can be, for example, broadband gain. In some alternative embodiments (e.g., in the case of a speech enhancement module), the gain can be a set of frequency - dependent gains for each speech enhancement level. In some such embodiments, the control system 210 can be configured to access the set of frequency - dependent gains from a stored data structure (e.g., a lookup table);
[0291] 1912: A gain application block for the television stream. In some embodiments (e.g., in the case of a simple mixer), the gain can be, for example, broadband gain. In some alternative embodiments (e.g., in the case of a speech enhancement module), the gain can be a set of frequency - dependent gains for each speech enhancement level. In some such embodiments, the control system 210 can be configured to access the set of frequency - dependent gains from a stored data structure (e.g., a lookup table).
[0292] As with other Figure 1 like Figure 19 The types and numbers of elements shown are provided only as examples. Other embodiments can include more, fewer, and / or different types and numbers of elements. For example, in some alternative embodiments, the noise estimate can be calculated in the hearing aid 1807, for example, based on an echo reference sent by the television to the hearing aid 1807. Some alternative examples can relate to personal sound amplification products, cochlear implants, headsets, or buddy microphone devices (e.g., buddy microphone devices having a directional microphone and configured to allow the user to focus on a conversation with their buddy). Some such buddy microphone devices can also be configured to transmit the sound of a multimedia device (e.g., the audio corresponding to video data to be reproduced via the television).
[0293] Figure 20 is a graph showing an example of the ambient noise level. In this example, the graph 2000 shows what can be used as Figure 18 and Figure 19An example of the input environmental noise level of the hearing aid 1807.
[0294] In this example, the curve 2001 represents a low environmental noise level estimate. According to some embodiments, at this level, the control system 210 may cause the amount of mixing that occurs to be dominated by the audio from the outside world (e.g., the audio signal from one or more hearing aid microphones), where the television audio level is relatively low. Such embodiments may be configured to increase the likelihood that a hearing-impaired user will be able to carry on a conversation with others in a quiet environment.
[0295] According to this example, the curve 2002 represents a medium environmental noise level estimate. According to some embodiments, at this level, the control system 210 may cause the amount of television audio in the mix to be relatively more than in the scenario discussed with reference to the curve 2001. In some examples, at this level, the control system 210 may cause the amount of television audio in the mix to be the same as or approximately the same as the level of the audio signal from one or more hearing aid microphones.
[0296] In this example, the curve 2003 represents a high environmental noise level estimate. According to some embodiments, at this level, the control system 210 may maximize the amount of television audio in the mix. In some embodiments, there may be an upper limit (e.g., a user-set upper limit) on the proportion of television audio in the mix. Such embodiments may increase the likelihood that the hearing aid still allows a high-level acoustic signal (such as a shout) fed by the hearing aid microphone to be detected by the user in order to enhance user safety.
[0297] Some disclosed embodiments may relate to the operation of a device that will be referred to herein as an "encoder". Although the encoder may be illustrated by a single block, the encoder may be implemented via one or more devices. In some embodiments, the encoder may be implemented by one or more devices of a cloud-based service in a data center (such as one or more servers, data storage devices, etc.). In some examples, the encoder may be configured to determine a compensation process to be performed in response to a noise metric and / or a speech intelligibility metric. In some embodiments, the encoder may be configured to determine a speech intelligibility metric. Some such embodiments may involve the interaction between the encoder and a downstream "decoder", e.g., where the decoder provides an environmental noise metric to the encoder. Embodiments in which the encoder performs at least some of the disclosed methods (e.g., determining a compensation process or determining multiple alternative compensation processes) may potentially be advantageous because the encoder will generally have much more processing power than the decoder.
[0298] Figure 21Shows an example of an encoder block and a decoder block according to one embodiment. In this example, encoder 2101 is shown transmitting an encoded audio bitstream 2102 to decoder 2103. In some such examples, encoder 2101 may be configured to transmit the encoded audio bitstream to multiple decoders.
[0299] According to some embodiments, encoder 2101 and decoder 2103 may be implemented by separate instances of control system 210, while in other examples, encoder 2101 and decoder 2103 may be considered part of a single instance of control system 210, e.g., as components of a single system. Although encoder 2101 and decoder 2103 are shown as single blocks in Figure 21 , in some embodiments, encoder 2101 and / or decoder 2103 may include more than one component, such as modules and / or sub-modules configured to perform various tasks.
[0300] In some embodiments, decoder 2103 may be implemented via one or more devices of an audio environment such as a home audio environment. Some of the tasks that decoder 2103 may perform were described in the above paragraph with reference to Figures 2 to 20 . In some such examples, decoder 2103 may be implemented via a television of the audio environment, via a television control module of the audio environment, etc. However, in some examples, at least some of the functions of decoder 2103 may be implemented via one or more other devices of the audio environment (e.g., hearing aids, personal sound amplification products, cochlear implants, headphones, laptop computers, mobile devices, smart speakers, a smart home hub configured to communicate with decoder 2103 (e.g., via the Internet), and a television of the audio environment, etc.).
[0301] Some of the tasks that encoder 2101 may perform are described in the following paragraph. In some embodiments, encoder 2101 may be implemented via one or more devices of a cloud-based service in a data center (such as one or more servers, data storage devices, etc.). In Figure 21In the example shown, encoder 2101 has received or obtained an audio bitstream, has encoded the received audio bitstream, and is in the process of transmitting the encoded audio bitstream 2102 to decoder 2103. In some such examples, the encoded audio bitstream 2102 can be part of an encoded content stream that includes encoded video data (e.g., corresponding to a television program, movie, music performance, etc.). The encoded audio bitstream 2102 can correspond to the encoded video data. For example, the encoded audio bitstream 2102 can include speech (e.g., dialogue) corresponding to the encoded video data. In some embodiments, the encoded audio bitstream 2102 can include music and audio effects (M&E) corresponding to the encoded video data.
[0302] In some disclosed embodiments, encoder 2101 can be configured to determine a noise metric and / or a speech intelligibility metric. In some examples, encoder 2101 can be configured to determine a compensation process to perform in response to the noise metric and / or the speech intelligibility metric, e.g., as disclosed elsewhere herein. In some embodiments, encoder 2101 can be configured to determine a compensation process for one or more types of ambient noise curves. In some examples, each of the ambient noise curves can correspond to a category of ambient noise (e.g., traffic noise, train noise, rain, etc.). In some such examples, encoder 2101 can be configured to determine a plurality of compensation processes for each category of ambient noise. Each compensation process of the plurality of compensation processes can correspond to a different level of ambient noise, for example. For example, one compensation process can correspond to a low level of ambient noise, another compensation process can correspond to a medium level of ambient noise, and another compensation process can correspond to a high level of ambient noise. According to some such examples, encoder 2101 can be configured to determine compensation metadata corresponding to the compensation process and to provide the compensation metadata to decoder 2103. In some such embodiments, encoder 2101 can be configured to determine compensation metadata corresponding to each compensation process of the plurality of compensation processes. In some such examples, decoder 2103 (or another downstream device) can be configured to determine the category and / or level of ambient noise in the audio environment and to select a corresponding compensation process based on the compensation metadata received from encoder 2101. Alternatively or additionally, decoder 2103 can be configured to determine the audio environment location and to select a corresponding compensation process based on the compensation metadata received from encoder 2101. In some examples, encoder 2101 can be configured to determine speech metadata that allows extraction of speech data from the audio data and to provide the speech metadata to decoder 2103.
[0303] However, in Figure 21In the example shown, the encoder 2101 does not provide speech metadata or compensation metadata to the decoder 2103 or other downstream devices. Figure 21 The example shown herein may sometimes be referred to as "single - ended post - processing" or "use case 1" in this document.
[0304] Figure 22 An example of an encoder block and a decoder block according to another embodiment is shown. In Figure 22 In the example shown, the encoder 2101 provides compensation metadata 2204 to the decoder 2103. Figure 22 The example shown herein may sometimes be referred to as "dual - ended post - processing" or "use case 2" in this document.
[0305] In some examples, the compensation metadata 2204 may correspond to a process for changing the processing of audio data, for example, as described above. According to some such examples, changing the processing of audio data may involve applying one or more speech enhancement methods, for example, reducing the gain of non - speech audio or increasing the gain of speech frequencies. In some such examples, changing the processing of audio data does not involve applying a broadband gain increase to the audio signal. Alternatively or additionally, the compensation metadata 2204 may correspond to a process for applying non - audio - based compensation methods (e.g., controlling a closed - captioning system, a karaoke captioning system, or a dialogue captioning system).
[0306] Figure 23 Some examples of decoder - side operations that can be performed in response to receiving the Figure 21 encoded audio bitstream shown are shown. In this example, single - ended post - processing noise compensation (use case 1) is enhanced through local noise determination and compensation. Figure 23 The example shown herein may sometimes be referred to as "single - ended post - processing - noise compensation" or "use case 3" in this document.
[0307] In this example, the audio environment in which the decoder 2103 is located includes one or more microphones 2301 configured to detect ambient noise. According to this example, the decoder 2103 or one or more microphones 2301 are configured to calculate a noise metric 2302 based on ambient noise measurements made by the one or more microphones 2301. In this embodiment, the decoder 2103 is configured to use the noise metric 2302 to determine and apply appropriate noise compensation 2303 for local playback. If the noise compensation is insufficient (in this example, as determined according to the noise metric 2302), the decoder 2103 is configured to enable non - audio - based compensation methods. In Figure 23 In the example shown, the non - audio - based compensation method involves enabling a closed - captioning system, a karaoke captioning system, or a dialogue captioning system represented by the "dialogue caption" box 2304.
[0308] Figure 24 illustrates some examples of decoder-side operations that can be performed in response to receiving the Figure 22 encoded audio bitstream shown in. In this example, dual-ended post-processing noise compensation (use case 2) is enhanced through local noise determination and compensation. Figure 24 The example shown in may sometimes be referred to herein as "use case 4".
[0309] In this example, the audio environment in which decoder 2103 is located includes one or more microphones 2301 configured to detect ambient noise. According to this example, decoder 2103 or one or more microphones 2301 are configured to calculate a noise metric 2302 based on ambient noise measurements made by one or more microphones 2301.
[0310] According to some examples, compensation metadata 2204 may include multiple selectable options. In some examples, at least some of the selectable options may correspond to a noise metric or to a range of noise metrics. In some embodiments, decoder 2103 may be configured to use noise metric 2302 to automatically select appropriate compensation metadata 2204 received from encoder 2101. Based on this automatic selection, in some examples, decoder 2103 may be configured to determine and apply appropriate audio-based noise compensation 2303 for local playback.
[0311] If the noise compensation 2303 is insufficient (determined according to noise metric 2302 in this example), then decoder 2103 is configured to enable a non-audio-based compensation method. In the Figure 24 example shown in, the non-audio-based compensation method involves enabling a closed captioning system, a karaoke captioning system, or a dialogue captioning system represented by "dialogue caption" box 2304.
[0312] Figure 25 illustrates an example of an encoder block and a decoder block according to another embodiment. In the Figure 25 example shown in, encoder 2101 is configured to determine an intelligibility metric 2501 based on an analysis of speech in the audio bitstream. In this example, encoder 2101 is configured to provide intelligibility metric 2501 to decoder 2103. Figure 25 The example shown in may sometimes be referred to herein as "dual-ended post-processing - intelligibility metric" or "use case 5".
[0313] In Figure 25In the example shown, the decoder 2103 is configured to determine whether a user in a local audio environment is likely to understand speech in an audio bitstream based at least in part on one or more intelligibility metrics 2501 received from the encoder 2101. If the decoder 2103 determines that the user is unlikely to understand the speech (e.g., if the decoder 2103 determines that the intelligibility metric 2501 is below a threshold), the decoder 2103 is configured to enable a non-audio-based compensation method represented by the "caption" box 2304.
[0314] Figure 26 An example of an encoder block and a decoder block according to another embodiment is shown. In Figure 26 the example shown, similar to Figure 25 the example of, the encoder 2101 is configured to determine an intelligibility metric 2501 based on an analysis of speech in the audio bitstream. In this example, the encoder 2101 is configured to provide the intelligibility metric 2501 to the decoder 2103. However, in this example, the decoder 2103 or one or more microphones 2301 are configured to calculate a noise metric 2302 based on ambient noise measurements made by one or more microphones 2301. Figure 26 The example shown in may sometimes be referred to herein as "two-sided post-processing - noise compensation and intelligibility metric" or "use case 6".
[0315] In this embodiment, the decoder 2103 is configured to use the noise metric 2302 to automatically select appropriate compensation metadata 2204 received from the encoder 2101. Based on this automatic selection, in this example, the decoder 2103 is configured to determine and apply appropriate noise compensation 2303 for local playback.
[0316] In Figure 26 the example shown, the decoder 2103 is also configured to determine whether a user in a local audio environment is likely to understand speech in an audio bitstream based at least in part on the intelligibility metric 2501 received from the encoder 2101 and the noise metric 2302 after appropriate noise compensation 2303 has been applied. If the decoder 2103 determines that the user is unlikely to understand the speech (e.g., if the decoder 2103 determines that the intelligibility metric 2501 is below a threshold corresponding to a particular noise metric 2302, such as by querying a data structure of intelligibility metrics and corresponding noise metrics and thresholds), the decoder 2103 is configured to enable a non-audio-based compensation method represented by the "caption" box 2304.
[0317] Figure 27 An example is shown that may respond to receiving Figure 21Some alternative examples of decoder-side operations performed on the encoded audio bitstream shown in Figure 23 the noise compensation "Use Case 3" described above is further enhanced by a feedback loop. Figure 27 The example shown in
[0318] In this example, the audio environment in which the decoder 2103 is located includes one or more microphones 2301 configured to detect ambient noise. According to this example, the decoder 2103 or one or more microphones 2301 are configured to calculate a noise metric 2302 based on ambient noise measurements made by the one or more microphones 2301. In this example, the noise metric 2302 is provided to the encoder 2101.
[0319] According to this embodiment, the encoder 2101 is configured to determine whether to reduce the complexity level of the encoded audio data 2102 transmitted to the decoder 2103 based at least in part on the noise metric 2302. In some examples, if the noise metric 2302 indicates a high level of noise in the audio environment of the decoder 2103, the encoder 2101 may be configured to determine to transmit a less complex version of the encoded audio data 2102 to the decoder 2103. In some such examples, if the noise metric 2302 indicates a high level of noise in the audio environment of the decoder 2103, the encoder 2101 may be configured to transmit a lower quality, lower data rate audio bitstream that is more suitable for playback in a noisy environment.
[0320] According to some embodiments, the encoder 2101 may have access to multiple audio versions, for example, ranging from the lowest quality audio version to the highest quality audio version. In some such examples, the encoder 2101 may have previously encoded the multiple audio versions. According to some such examples, the encoder 2101 may be configured to receive a content stream including received video data and received audio data corresponding to the video data. In some such examples, the encoder 2101 may be configured to prepare multiple encoded audio versions corresponding to the received audio data (ranging from the lowest quality encoded audio version to the highest quality encoded audio version).
[0321] In some examples, determining whether to reduce the complexity level of the encoded audio data 2102 based at least in part on the noise metric 2302 may involve determining which encoded audio version to transmit to the decoder 2103.
[0322] In some examples, the received audio data may include audio objects. According to some such examples, the highest quality encoded audio version may include all of the audio objects in the audio objects of the received audio data. In some such examples, a lower quality encoded audio version may include less than all of the audio objects in the audio objects of the received audio data. According to some embodiments, the lower quality encoded audio version may include lossy compressed audio that includes fewer bits than the received audio data and may be transmitted at a lower bitrate than the bitrate of the received audio data. In some instances, the received audio data may include audio object priority metadata indicating the priority of the audio objects. In some such examples, the encoder 2101 may be configured to determine which audio objects will be in each of the encoded audio versions based at least in part on the audio object priority metadata.
[0323] In this embodiment, the decoder 2103 is configured to use the noise metric 2302 to determine and apply appropriate noise compensation 2303 for local playback. If the noise compensation is insufficient (in this example, as determined according to the noise metric 2302), the decoder 2103 is configured to enable a non-audio based compensation method. In Figure 27 the example shown, the non-audio based compensation method involves enabling a closed captioning system, a karaoke captioning system, or a dialogue captioning system represented by the "Dialogue Caption" box 2304. If the decoder 2103 has received a relatively low quality audio bitstream, audio based noise compensation and / or non-audio based noise compensation may be necessary. In some examples, a low quality audio bitstream may have been previously sent based on feedback from the decoder side regarding noise in the user's audio environment, information about the audio capabilities of the user's system, etc.
[0324] Figure 28 is shown Figure 24 and Figure 27 an enhanced version of the system shown in Figure 28 The example shown in Figure 27 may sometimes be referred to herein as "feedback-dual-sided post-processing" or "use case 8". According to this embodiment, the encoder 2101 is configured to provide the encoded audio data 2102 in response to the noise metric 2302 received from the decoder 2103 to the decoder 2103. In some examples, the encoder 2101 may be configured to select and provide the encoded audio data 2102 as described above with reference to
[0325] In this example, the encoder 2101 is also configured to provide compensation metadata 2204 in response to the noise metric 2302 received from the decoder 2103 to the decoder 2103. In some such examples, the decoder 2103 may be configured to simply apply an audio or non-audio compensation method corresponding to the compensation metadata 2204 received from the encoder 2101.
[0326] According to some alternative examples, the encoder 2101 may be configured to provide compensation metadata 2204 corresponding to various selectable compensation options, e.g., as described above with reference to Figure 24 However, in some such embodiments, the encoder 2101 may be configured to select all of the compensation options and the corresponding compensation metadata 2204 in the compensation options at least in part based on the noise metric 2302 received from the decoder 2103. In some embodiments, the previously transmitted compensation metadata 2204 may be adjusted or recalculated at least in part based on the noise metric 2302 received from the decoder 2103. In some such embodiments, the decoder 2103 may be configured to automatically select the appropriate compensation metadata 2204 received from the encoder 2101 using the noise metric 2302. Based on this automatic selection, in this example, the decoder 2103 may be configured to determine and apply an appropriate noise compensation 2303 for local playback. If the noise compensation 2303 is insufficient (in this example, as determined according to the noise metric 2302), then the decoder 2103 is configured to enable a non-audio-based compensation method. In the example shown in Figure 24 a non-audio-based compensation method involves enabling a closed captioning system, a karaoke caption system, or a dialogue caption system represented by the "Dialogue Caption" box 2304. According to some examples, the encoder 2101 may be configured to modify the closed captioning, karaoke captioning, or dialogue captioning at least in part based on the noise metric 2302 received from the decoder 2103. For example, if the noise metric 2302 indicates an ambient noise level at or above a threshold level, the encoder 2101 may be configured to simplify the text.
[0327] Figure 29 An example of an encoder block and a decoder block according to another embodiment is shown. Figure 29 An enhanced version of the example described above with reference to Figure 25 is shown. As in the example of Figure 25 the encoder 2101 is configured to determine a speech intelligibility metric 2501 based on an analysis of speech in an audio bitstream. In this example, the encoder 2101 is configured to provide one or more speech intelligibility metrics 2501 and compensation metadata 2204 to the decoder 2103.
[0328] However, inFigure 29 In the example shown in Figure 29 , the decoder 2103 is also configured to determine one or more speech intelligibility metrics 2901 and provide the one or more speech intelligibility metrics 2901 to the encoder 2101. Figure 29 The example shown in Figure 29 may sometimes be referred to herein as "feedback - intelligibility metric" or "use case 9".
[0329] According to some examples, the one or more speech intelligibility metrics 2901 may be at least partially based on one or more user characteristics corresponding to viewers and / or listeners in the audio environment in which the decoder 2103 resides. The one or more user characteristics may include, for example, at least one of the user's native language, the user's accent, the user's location in the environment, the user's age, and / or the user's capabilities. The user's capabilities may include, for example, the user's hearing ability, the user's language level, the user's accent comprehension level, the user's vision, and / or the user's reading comprehension.
[0330] In some embodiments, the encoder 2101 may be configured to select a compensation metadata and / or a quality level of the encoded audio data 2102 at least partially based on the one or more speech intelligibility metrics 2901. In some such examples, if the one or more speech intelligibility metrics 2901 are to indicate that the user has a high language level but a very slightly diminished hearing ability, the encoder 2101 may be configured to select and send a high - quality speech channel / object to increase intelligibility for local playback. According to some examples, if the one or more speech intelligibility metrics 2901 are to indicate that the user has a low language level and / or accent comprehension, the encoder 2101 may be configured to send compensation metadata 2204 corresponding to a non - audio - based compensation method (e.g., a method involving controlling a closed - captioning system, a karaoke - captioning system, or a dialogue - captioning system) to the decoder 2103. According to some examples, the encoder 2101 may be configured to modify the closed - captioning, karaoke - captioning, or dialogue - captioning at least partially based on the noise metric 2302 and / or the intelligibility metric 2901 received from the decoder 2103. For example, the encoder 2101 may be configured to simplify the text when the noise metric 2302 indicates an ambient noise level at or above a threshold level and / or when the intelligibility metric 2901 (or the updated intelligibility metric 2501) is below a threshold level.
[0331] In Figure 29In the example shown, the decoder 2103 is configured to determine whether a user in a local audio environment is likely to understand speech in an audio bitstream based at least in part on one or more speech intelligibility metrics 2901 and / or the intelligibility metric 2501 received from the encoder 2101. If the decoder 2103 determines that the user is unlikely to understand the speech (e.g., if the decoder 2103 determines that the intelligibility metric 2501 is below a threshold), the decoder 2103 is configured to enable a non-audio-based compensation method represented by the "dialogue subtitles" box 2304. In some embodiments, if one or more speech intelligibility metrics 2901 are to indicate that the user has a low level of language proficiency and / or accent comprehension, the encoder 2101 may be configured to send an encoded video stream that already includes closed captions, lyrics subtitles, or dialogue subtitles in the video stream.
[0332] Figure 30 An example of an encoder block and a decoder block according to another embodiment is shown. Figure 30 An enhanced version of the example described above with reference to Figure 28 and Figure 29 is shown. As in the example of Figure 29 , the encoder 2101 is configured to determine a speech intelligibility metric 2501 based on an analysis of speech in the audio bitstream. In this example, the encoder 2101 is configured to provide one or more speech intelligibility metrics 2501 and compensation metadata 2204 to the decoder 2103. As in the example shown in Figure 29 , the decoder 2103 is also configured to determine one or more speech intelligibility metrics 2901 and provide the one or more speech intelligibility metrics 2901 to the encoder 2101. Additionally, as in the example shown in Figure 28 , the decoder 2103 is also configured to determine a noise metric 2302 and transmit the noise metric 2302 to the encoder 2101. Figure 30 The example shown in
[0333] may sometimes be referred to herein as "two-sided post-processing - compensation and intelligibility metrics" or "use case 10". Figure 29 One or more speech intelligibility metrics 2901 may be determined as described above with reference to Figure 29 . In some embodiments, the encoder 2101 may be configured to select compensation metadata and / or the quality level of the encoded audio data 2102 based at least in part on the one or more speech intelligibility metrics 2901, for example, as described above with reference to
[0334] In some such examples, the decoder 2103 may be configured to simply apply an audio or non-audio compensation method corresponding to the compensation metadata 2204 received from the encoder 2101. According to some alternative examples, the encoder 2101 may be configured to provide compensation metadata 2204 corresponding to various selectable compensation options, e.g., as described above with reference to Figure 24 as described.
[0335] In some embodiments, if the decoder 2103 determines that the user is unlikely to understand the speech in the encoded audio signal 2102, the decoder 2103 may be configured to provide feedback to the encoder 2101. In some such examples, the decoder 2103 may be configured to request high-quality audio from the encoder 2101. In an alternative example, the decoder 2103 may be configured to send a request to the encoder 2101 to transmit an encoded video stream having closed captions, karaoke lyrics, or dialogue subtitles already included in the video stream. In some such examples, if the closed captions, karaoke lyrics, or dialogue subtitles are included in the corresponding video stream, the encoder 2101 may transmit a lower-quality version of the encoded audio data 2102.
[0336] Figure 31 Illustrates the relationships between the various disclosed use cases. Figure 31 Is summarized in one figure and on one page Figures 21 to 30 along with many of the foregoing explanatory paragraphs, including a comparison of the various use cases. For example, traversing from single-ended post-processing (“Use Case 1”) to double-ended post-processing (“Use Case 2”), Figure 31 indicates that compensation metadata is added in “Use Case 2”. This can also be seen by comparing Figure 21 and Figure 22 as seen. Traversing from “Use Case 2” to double-ended post-processing enhanced by local noise determination and compensation (“Use Case 4”), Figure 31 indicates that a noise metric and compensation metadata are added in “Use Case 4”. This can also be seen by comparing Figure 22 and Figure 24 as seen.
[0337] Figure 32 Is a flowchart outlining an example of the disclosed method. Like other methods described herein, the blocks of method 3200 need not be performed in the order indicated. Additionally, such a method may include more or fewer blocks than those shown and / or described.
[0338] Method 3200 may be performed by, such as Figure 2performed by the apparatus or system of apparatus 200 shown and described above. In some examples, the blocks of method 3200 may be performed by a device (such as a server) implementing a cloud-based service. According to some examples, method 3200 may be performed at least in part by the encoder 2101 described above with reference to Figures 21 to 31 However, in an alternative implementation, at least some of the blocks of method 3200 may be performed by one or more devices within the audio environment (e.g., by the decoder 2103 described above with reference to Figures 21 to 31 , by a television, or by a television control module).
[0339] In this implementation, block 3205 involves receiving, by a first control system and via a first interface system, a content stream including video data and audio data corresponding to the video data. For example, block 3205 may involve receiving the content stream from a content provider (such as a provider of television programs, movies, etc.) by the control system of the encoder 2101 described above with reference to Figures 21 to 31 or a control system of a similar encoding system.
[0340] In this example, block 3210 involves determining, by the first control system, a noise metric and / or a speech intelligibility metric. Block 3210 may involve any of the disclosed methods for determining a noise metric and / or a speech intelligibility metric. In some examples, block 3210 may involve receiving the noise metric and / or the speech intelligibility metric from another device (such as the decoder 2103 described above with reference to Figures 21 to 31 ). In some examples, block 3210 may involve determining the speech intelligibility metric by analyzing the audio data of the content stream corresponding to the speech.
[0341] According to this example, block 3215 involves determining, by the first control system, a compensation process to be performed in response to at least one of the noise metric or the speech intelligibility metric. In this example, the compensation process involves changing the processing of the audio data and / or applying a non-audio-based compensation method. According to this implementation, changing the processing of the audio data does not involve applying a broadband gain increase to the audio signal. Block 3215 may involve determining any of the disclosed compensation processes, including methods for changing the processing of the audio data and / or methods for applying a non-audio-based compensation method.
[0342] In some examples, a non-audio-based compensation method can involve controlling a closed captioning system, a karaoke captioning system, or a dialogue captioning system. In some such examples, controlling a closed captioning system, a karaoke captioning system, or a dialogue captioning system can involve controlling at least one of a font or a font size based at least in part on a speech intelligibility metric. In some such examples, controlling a closed captioning system, a karaoke captioning system, or a dialogue captioning system can involve one or more of the following: determining whether to filter out some speech-based text based at least in part on a speech intelligibility metric, determining whether to simplify at least some speech-based text, or determining whether to paraphrase at least some speech-based text. In some instances, controlling a closed captioning system, a karaoke captioning system, or a dialogue captioning system can involve determining whether to display text based at least in part on a noise metric.
[0343] According to some examples, changing the processing of audio data can involve applying one or more speech enhancement methods based at least in part on at least one of a noise metric or a speech intelligibility metric. In some such examples, the one or more speech enhancement methods can include reducing the gain of non-speech audio and / or increasing the gain of speech frequencies. In some instances, changing the processing of audio data can involve changing one or more of an upmixing process, a downmixing process, a virtual bass process, a bass distribution process, an equalization process, a crossover filter, a delay filter, a multiband limiter, or a virtualization process based at least in part on at least one of a noise metric or a speech intelligibility metric.
[0344] In this example, block 3220 involves the first control system determining compensation metadata corresponding to a compensation process. Here, block 3225 involves generating encoded compensation metadata by the first control system encoding the compensation metadata. In this example, block 3230 involves generating encoded video data by the first control system encoding the video data. According to this example, block 3235 involves generating encoded audio data by the first control system encoding the audio data.
[0345] In this embodiment, block 3240 involves transmitting an encoded content stream including the encoded compensation metadata, the encoded video data, and the encoded audio data from the first device to at least a second device. The first device can be, for example, the encoder 2101 described above with reference to Figures 21 to 31 what was described.
[0346] In some examples, the second device includes a second control system configured to decode the encoded content stream. The second device can be, for example, the decoder 2103 described above with reference to Figures 21 to 31 what was described.
[0347] According to some examples, the compensation metadata can include multiple options that can be selected by the second device or by a user of the second device. In some such examples, at least some (e.g., two or more) of the multiple options can correspond to noise levels that can occur in the environment in which the second device is located. Some such methods can involve automatically selecting one of the two or more options by the second control system and at least partially based on the noise level.
[0348] In some examples, at least some (e.g., two or more) of the multiple options can correspond to one or more speech intelligibility metrics. In some such examples, the encoded content stream can include speech intelligibility metadata. Some such methods can involve selecting one of the two or more options by the second control system and at least partially based on the speech intelligibility metadata. In some such examples, each of the multiple options can correspond to one or more of a known or estimated hearing ability, known or estimated language level, known or estimated accent comprehension level, known or estimated visual acuity, or known or estimated reading comprehension ability of a user of the second device. According to some examples, each of the multiple options can correspond to a speech enhancement level.
[0349] In some embodiments, the second device corresponds to a specific playback device (e.g., a specific television). Some such embodiments can involve receiving, by the first control system and via the first interface system, at least one of a noise metric or a speech intelligibility metric from the second device. In some examples, the compensation metadata can correspond to the noise metric and / or the speech intelligibility metric.
[0350] Some examples can involve determining, by the first control system and at least partially based on the noise metric or the speech intelligibility metric, whether the encoded audio data will correspond to all of the received audio data or only to a portion of the received audio data. In some such examples, the audio data includes audio objects and corresponding priority metadata indicating the priorities of the audio objects. According to some such examples in which it is determined that the encoded audio data will correspond only to a portion of the received audio data, the method can involve selecting that portion of the received audio data at least partially based on the priority metadata.
[0351] In some embodiments, the second device may be one of a plurality of devices to which the encoded audio data has been transmitted. According to some such embodiments, the plurality of devices may have been selected at least in part based on known or estimated speech intelligibility for a user category. In some instances, the user category may have been defined by one or more of known or estimated hearing ability, known or estimated language level, known or estimated accent comprehension level, known or estimated visual acuity, or known or estimated reading comprehension. According to some such examples, the user category may have been defined at least in part based on one or more assumptions about language level and / or accent comprehension level for a particular geographic region (e.g., a particular country or a particular region of a country).
[0352] In some embodiments, the audio data may include speech data as well as music and effects (M&E) data. Some such embodiments may involve separating the speech data from the M&E data by a first control system. Some such methods may involve determining, by the first control system, speech metadata that permits extraction of the speech data from the audio data and generating encoded speech metadata by encoding the speech metadata by the first control system. In some such embodiments, transmitting the encoded content stream may involve transmitting the encoded speech metadata to at least a second device.
[0353] Figure 33 is a flowchart outlining an example of the disclosed method. Like other methods described herein, the blocks of method 3300 need not be performed in the order indicated. Additionally, such a method may include more or fewer blocks than those shown and / or described.
[0354] Method 3300 may be performed by a device or system such as Figure 2 the device 200 shown and described above. In some examples, the blocks of method 3300 may be performed by a device (such as a server) implementing a cloud-based service. According to some examples, method 3300 may be performed at least in part by the encoder 2101 described above with reference to Figures 21 to 31 However, in alternative embodiments, at least some of the blocks of method 3300 may be performed by one or more devices within the audio environment (e.g., by the decoder 2103 described above with reference to Figures 21 to 31 or by a television or a television control module).
[0355] In this embodiment, block 3305 involves receiving, by the first control system and via a first interface system of the first device, a content stream including video data and audio data corresponding to the video data. For example, block 3305 may involve the above reference to Figures 21 to 31The control system of the described encoder 2101 or a control system of a similar encoding system receives a content stream from a content provider.
[0356] In this example, block 3310 involves the first control system receiving a noise metric and / or a speech intelligibility metric from a second device. In some examples, block 3310 can involve receiving a noise metric and / or a speech intelligibility metric from the decoder 2103 described above with reference to Figures 21 to 31 the decoder 2103 described above.
[0357] According to this example, block 3315 involves the first control system determining whether to reduce the complexity level of the transmitted encoded audio data and / or text corresponding to the received audio data, at least in part based on the noise metric or the speech intelligibility metric. In some examples, block 3315 can involve determining whether the transmitted encoded audio data will correspond to all of the received audio data or only a portion of the received audio data. In some examples, the audio data can include audio objects and corresponding priority metadata indicating the priority of the audio objects. In some such examples where it is determined that the encoded audio data will correspond to only a portion of the received audio data, block 3315 can involve selecting that portion of the received audio data, at least in part based on the priority metadata. In some examples, for a closed captioning system, a karaoke captioning system, or a dialogue captioning system, determining whether to reduce the complexity level can involve determining whether to filter out some speech-based text, determining whether to simplify at least some speech-based text, and / or determining whether to paraphrase at least some speech-based text.
[0358] In this example, block 3320 involves selecting a version of the encoded audio data and / or a version of the text to be transmitted based on the determination process of block 3315. Here, block 3325 involves transmitting an encoded content stream including the encoded video data and the transmitted encoded audio data from the first device to the second device. For instances where block 3320 involves selecting a version of the text to be transmitted, some implementations involve transmitting the version of the text to the second device.
[0359] Figure 34 An example of a floor plan of an audio environment is shown, where the audio environment is a living space. Similar to other Figure 1 as Figure 34 presented herein, the types and quantities of the elements shown are provided only as examples. Other implementations can include more, fewer, and / or different types and quantities of elements.
[0360] According to this example, the environment 3400 includes a living room 3410 at the upper left, a kitchen 3415 at the lower center, and a bedroom 3422 at the lower right. The boxes and circles distributed across the living space represent a set of loudspeakers 3405a to 3405h, and in some embodiments, at least some of the loudspeakers in the set can be smart speakers placed in convenient locations for the space but not following any standard prescribed layout (placed arbitrarily). In some examples, the television 3430 can be configured to implement at least in part one or more of the disclosed embodiments. In this example, the environment 3400 includes cameras 3411a to 3411e distributed throughout the environment. In some embodiments, one or more of the smart audio devices in the environment 3400 can also include one or more cameras. The one or more smart audio devices can be single-purpose audio devices or virtual assistants. In some such examples, one or more cameras of the optional sensor system 130 can reside in or on the television 3430, in a mobile phone, or in a smart speaker (e.g., one or more of the loudspeakers 3405b, 3405d, 3405e, or 3405h). Although the cameras 3411a to 3411e are not shown in every depiction of the environment 3400 presented in this disclosure, in some embodiments, each of the environments 3400 can still include one or more cameras.
[0361] Some aspects of the present disclosure include systems or devices configured (e.g., programmed) to perform one or more examples of the disclosed methods, and tangible computer-readable media (e.g., disks) storing code for implementing one or more examples of the disclosed methods or their steps. For example, some disclosed systems can be or include programmable general-purpose processors, digital signal processors, or microprocessors programmed with software or firmware to and / or otherwise configured to perform any of a variety of operations on data, including embodiments of the disclosed methods or their steps. Such general-purpose processors can be or include a computer system comprising an input device, a memory, and a processing subsystem programmed (and / or otherwise configured) to perform one or more examples of the disclosed method (or its steps) in response to data asserted thereto.
[0362] Some embodiments may be implemented as configurable (e.g., programmable) digital signal processors (DSPs) that are configured (e.g., programmed and otherwise configured) to perform the required processing on one or more audio signals, including the execution of one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed systems (or their elements) may be implemented as general-purpose processors (e.g., a personal computer (PC) or other computer system or microprocessor that may include an input device and memory), which are programmed with software or firmware and / or otherwise configured to perform any of the various operations (including one or more examples of the disclosed methods). Alternatively, elements of some embodiments of the system of the present invention are implemented as general-purpose processors or DSPs that are configured (e.g., programmed) to execute one or more examples of the disclosed methods, and the system further includes other elements (e.g., one or more loudspeakers and / or one or more microphones). A general-purpose processor configured to execute one or more examples of the disclosed methods may be coupled to an input device (e.g., a mouse and / or a keyboard), a memory, and a display device.
[0363] Another aspect of the present disclosure is a computer-readable medium (e.g., a disk or other tangible storage medium) that stores code (e.g., an encoder executable to execute one or more examples of the disclosed method or its steps) for performing one or more examples of the disclosed method or its steps.
[0364] Although specific embodiments of the present disclosure and applications of the present disclosure have been described herein, it will be apparent to those of ordinary skill in the art that many variations to the embodiments and applications described herein are possible without departing from the scope of the present disclosure as described and claimed herein. It should be understood that although certain forms of the present disclosure have been shown and described, the present disclosure is not limited to the specific embodiments or the specific methods described and shown.
Claims
1. A content stream processing method, comprising: Receiving, by a first control system and via a first interface system, a content stream including video data and audio data corresponding to the video data; Determining, by the first control system, a noise metric and a speech intelligibility metric, the speech intelligibility metric being based on the speech intelligibility for a user category, wherein the user category is defined by one or more of: a known or estimated language level, a known or estimated accent comprehension level, a known or estimated visual acuity, or a known or estimated reading comprehension; Determining, by the first control system, a compensation process to be performed in response to the noise metric and the speech intelligibility metric, wherein performing the compensation process involves: Changing the processing of the audio data, wherein changing the processing of the audio data does not involve applying a broadband gain increase to the audio signal; and Applying a non-audio-based compensation method, wherein the non-audio-based compensation method involves controlling at least in part a closed captioning system, a karaoke captioning system, or a dialogue captioning system based on the speech intelligibility metric, wherein controlling the closed captioning system, the karaoke captioning system, or the dialogue captioning system involves at least one of: Controlling the font; Controlling the font size; Determining whether to filter out some speech-based text; Determining whether to simplify at least some speech-based text; or Determining whether to paraphrase at least some speech-based text; Determining, by the first control system, compensation metadata corresponding to the compensation process; Generating encoded compensation metadata by encoding the compensation metadata by the first control system; Generating encoded video data by encoding the video data by the first control system; Generating encoded audio data by encoding the audio data by the first control system; and Transmitting an encoded content stream including the encoded compensation metadata, the encoded video data, and the encoded audio data from a first device to at least a second device.
2. The method according to claim 1, wherein, The audio data includes speech data and music and effects (M&E) data, and the content stream processing method further includes: Separating, by the first control system, the speech data from the M&E data; Determining, by the first control system, speech metadata that allows extraction of the speech data from the audio data; and Generating encoded speech metadata by encoding the speech metadata by the first control system, wherein transmitting the encoded content stream includes transmitting the encoded speech metadata to at least the second device.
3. The method according to claim 1 or claim 2, wherein The second device includes a second control system configured to decode the encoded content stream.
4. The method according to claim 3, wherein, The second device is one of a plurality of devices to which the encoded audio data has been transmitted.
5. The method according to claim 4, wherein, The plurality of devices have been selected at least in part based on the speech intelligibility for the user category.
6. The method according to claim 5, wherein The user category is further defined by a known or estimated hearing ability.
7. The method according to claim 3, wherein The compensation metadata includes a plurality of options that can be selected by the second device or by a user of the second device.
8. The method according to claim 7, wherein, Two or more of the plurality of options correspond to a noise level that can occur in the environment in which the second device is located.
9. The method according to claim 7, wherein, Two or more of the plurality of options correspond to a speech intelligibility metric.
10. The method according to claim 9, wherein, The encoded content stream includes speech intelligibility metadata, and the content stream processing method further includes selecting, by the second control system and at least in part based on the speech intelligibility metadata, one of the two or more options.
11. The method according to claim 7, wherein, Each of the plurality of options corresponds to one or more of the following of the user of the second device: known or estimated hearing ability, known or estimated language level, known or estimated accent comprehension level, known or estimated visual acuity, or known or estimated reading comprehension.
12. The method according to claim 7, wherein Each of the plurality of options corresponds to a speech enhancement level.
13. The method according to claim 1 or claim 2, wherein The second device corresponds to a specific playback device.
14. The method according to claim 13, wherein, The specific playback device is a specific television.
15. The method according to claim 13, further comprising: Receiving, by the first control system and via the first interface system from the second device, at least one of the noise metric or the speech intelligibility metric.
16. The method according to claim 15, wherein, The compensation metadata corresponds to at least one of the noise metric or the speech intelligibility metric.
17. The method according to claim 15, further comprising: Determining, by the first control system and at least in part based on the noise metric or the speech intelligibility metric, whether the encoded audio data will correspond to all of the received audio data or only a portion of the received audio data.
18. The method according to claim 17, wherein, The audio data includes audio objects and corresponding priority metadata indicating the priority of the audio objects, and wherein it is determined that the encoded audio data will correspond to only the portion of the received audio data, and the content stream processing method further includes selecting, at least in part based on the priority metadata, the portion of the received audio data.
19. The method according to claim 1 or 2, wherein Controlling the closed captioning system, the lyrics captioning system, or the dialogue captioning system involves determining whether to display text at least in part based on the noise metric.
20. The method according to claim 1 or claim 2, wherein Changing the processing of the audio data involves applying one or more speech enhancement methods at least in part based on at least one of the noise metric or the speech intelligibility metric.
21. The method according to claim 20, wherein, The one or more speech enhancement methods include at least one of the following: reducing the gain of non-speech audio or increasing the gain of speech frequencies.
22. The method according to claim 1 or claim 2, wherein Changing the processing of the audio data involves changing one or more of the following at least in part based on at least one of the noise metric or the speech intelligibility metric: upmixing process, downmixing process, virtual bass process, bass distribution process, equalization process, crossover filter, delay filter, multiband limiter, or virtualization process.
23. A content stream processing method, comprising: Receiving, by a first control system and via a first interface system of a first device, a content stream, the content stream including received video data and received audio data corresponding to the video data; Receive, by the first control system and via the first interface system, a noise metric and a speech intelligibility metric from a second device, wherein the speech intelligibility metric is based on speech intelligibility for a user category, and wherein the user category is defined by one or more of: a known or estimated language level, a known or estimated accent comprehension level, a known or estimated visual acuity, or a known or estimated reading comprehension ability; Determine, by the first control system, at least in part based on the noise metric and the speech intelligibility metric, whether to reduce a complexity level of transmitted encoded audio data or text corresponding to the received audio data, wherein determining whether to reduce the complexity level involves at least one of the following: Determine whether the transmitted encoded audio data will correspond to all of the received audio data or only a portion of the received audio data; For a closed captioning system, a karaoke captioning system, or a dialogue captioning system, determine whether to filter out some speech-based text; For a closed captioning system, a karaoke captioning system, or a dialogue captioning system, determine whether to simplify at least some speech-based text; Or For a closed captioning system, a karaoke captioning system, or a dialogue captioning system, determine whether to paraphrase at least some speech-based text; Select, based on the determination, at least one of the transmitted encoded audio data or text to be transmitted; And Transmit an encoded content stream including encoded video data and the selected encoded audio data from the first device to the second device.
24. The method according to claim 23, wherein, The audio data includes audio objects and corresponding priority metadata indicating the priority of the audio objects, and wherein, determining that the encoded audio data will correspond to only the portion of the received audio data, the content stream processing method further includes selecting the portion of the received audio data at least in part based on the priority metadata.
25. An apparatus configured to perform the method of any one of claims 1 to 9 or 11 to 24.
26. A system configured to perform the method of any one of claims 1 to 24.
27. One or more non-transitory media having software stored thereon, the software including instructions for controlling one or more devices to perform the method of any one of claims 1 to 24.
28. A computer program product including a computer program, the computer program including instructions for controlling one or more devices to perform the method of any one of claims 1 to 24.
Citation Information
Patent Citations
System and method for enhancing speech intelligibility for the hearing impaired
US20050086058A1
Automatic activation of closed captioning for low volume periods
WO2018112789A1