Intelligent speech or dialogue enhancement
Machine learning and AI are used to dynamically adjust audio playback systems to enhance speech intelligibility by detecting speech components and transitioning to a speech equalization mode, improving user experience in multimedia environments.
Patent Information
- Application Number
- JP2025502938
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-19
- Filing Date
- 2023-07-19
- Publication Date
- 2025-08-20
- Estimated Expiration
- 2043-07-19
AI Technical Summary
Existing audio playback systems struggle to dynamically adjust equalization settings to enhance speech intelligibility in multimedia content, leading to suboptimal user experience due to manual adjustments being slow and inconvenient, especially in environments with varying sound elements like dialogue, music, and sound effects.
Implementing machine learning and artificial intelligence to detect speech components in audio signals and seamlessly transition to a speech equalization mode, enhancing frequency bands and suppressing non-speech elements to improve linguistic understanding.
Enhances speech intelligibility by automatically adjusting equalization settings based on detected speech content, minimizing adverse effects on non-speech audio and reducing the need for manual user input.
Smart Images

Figure 2025527151000001_ABST
Abstract
Description
[Technical Field]
[0001] Aspects of the present disclosure generally relate to audio signal processing. [Background technology]
[0002] In recording and playback, an equalizer may perform equalization to adjust the magnitude of different frequency bands in an audio signal. For example, an equalizer may use filters to adjust bass and treble to enhance the listening experience. Equalization may be dynamically adjusted in real time by a user or may include one or more preset profiles for different genres of audio input (e.g., jazz, classical, pop, etc.).
[0003] In a home theater, different surround sound configurations and / or different recording or compression profiles may make speech or dialogue less understandable than intended, and a user may prefer to use an equalization profile to enhance the speech or dialogue sound component (e.g., by increasing the loudness of the relevant frequency bands). Summary of the Invention
[0004] All examples and features mentioned in this specification can be combined in any technically possible manner.
[0005] Aspects of the present disclosure provide a method for processing and generating audio signals. The method includes detecting speech components in an original set of audio signals using a trained machine learning network based on a mixture of audio content categories, the speech components consisting of sound elements that convey linguistic meaning. The method includes enhancing the detected speech components by transitioning from an original equalization mode to a speech equalization mode that enables a user to better understand the linguistic meaning therein. The method further includes, in the absence of detection of the speech components, outputting the original set of audio signals in the original equalization mode.
[0006] In an aspect, detecting the speech component using the trained machine learning network is based on calculating the root mean square of the respective energy level of each category of a mixture category of audio content. In some cases, the trained machine learning network includes a deep learning model that estimates the respective energy level of each category of a mixture category of audio content, the mixture category of audio content including a speech component, a music component, and a singing component. In some cases, detecting the speech component includes determining a ratio of the energy level of the speech component to an overall energy level above a threshold.
[0007] In some cases, the speech component includes sounds whose meaning is conveyed based on linguistic features, the musical component includes sounds that lack linguistic features, and the singing component includes a mixture of sounds that simultaneously include linguistic and musical components. In some cases, a trained machine learning network may be trained to identify categories and thresholds for each of a mixture of categories of audio content based on a known database of movie content.
[0008] In some cases, detecting speech components involves processing the ongoing audio signal at an advanced time prior to the transition or output action.
[0009] In aspects, enhancing the speech component includes gradually fading from the original equalization mode to a speech equalization mode, for example, the speech equalization mode includes at least one of increasing the magnitude or contrast of speech-related frequency bands, decreasing the magnitude of non-speech-related frequency bands or signal channels, or changing equalization or dynamic range compression settings on non-speech-related frequency bands or signal channels to improve intelligibility of the speech component.
[0010] In an aspect, enhancing the detected speech components is performed on a first device and outputting the original set of audio signals is performed on a second device. For example, the first device and the second device are paired in a short-range wireless communication network. In some cases, the method further includes extracting the detected speech components and playing the extracted speech components on a third device. In some cases, the first device, the second device, and the third device are configured to generate a mixed surround sound.
[0011] In an aspect, outputting the original set of audio signals includes determining loss of speech components.
[0012] An aspect of the present disclosure provides an apparatus for processing and generating audio signals. The apparatus includes a memory and a processor coupled to the memory. The processor and memory are configured to detect speech components in an original set of audio signals using a trained machine learning network based on a mixture of audio content categories, the speech components consisting of sound elements that convey linguistic meaning. The processor and memory are configured to enhance the detected speech components by transitioning from an original equalization mode to a speech equalization mode that enables a user to better understand the linguistic meaning therein. The processor and memory are further configured to output the original set of audio signals in the original equalization mode in the absence of detection of the speech components.
[0013] In aspects, the processor and memory are configured to enhance the detected speech components and output the original set of audio signals at the second device. In some cases, the device includes a sound bar configured to output surround sound, and the second device includes noise-canceling headphones. In some cases, the device includes noise-canceling headphones, and the second device includes a sound bar configured to output surround sound.
[0014] In an aspect, the processor and memory are configured to enhance the speech component by gradually fading from the original equalization mode to the speech equalization mode.
[0015] In an aspect, the processor and memory are configured to extract the detected speech components and play the extracted speech components at the third device. In some cases, the first device, the second device, and the third device are configured to generate mixed surround sound.
[0016] An aspect provides a method for audio signal processing, the method including: analyzing content of the audio signal prior to playback of the content to determine whether, during playback of the audio signal, one or more predefined conditions are met to indicate that the content includes speech; in response to determining that the one or more predefined conditions are met, automatically applying a first playback equalization to the audio signal, the first playback equalization being configured to enhance the speech in the content; and in response to determining that the one or more predefined conditions are not met, applying either i) no playback equalization or ii) a second playback equalization different from the first playback equalization to the audio signal.
[0017] In an aspect, the analyzing is performed using a trained machine learning model. In an aspect, the trained machine learning model comprises a deep learning model that estimates an energy level of the audio signal. In an aspect, the energy level of the audio signal includes energy levels of any combination of speech, a musical component of the audio signal, and a singing component of the audio signal.
[0018] In an aspect, the analyzing includes analyzing metadata associated with the content, the metadata indicating that the content includes speech. In an aspect, the analyzing includes analyzing a voice track of the audio signal, and the one or more predefined conditions include the voice track exceeding a threshold.
[0019] In an aspect, the audio signal comprises different channels, and analyzing includes analyzing the different channels. In an aspect, analyzing the different channels includes comparing correlation content between two of the different channels.
[0020] In an aspect, the analyzing includes analyzing a center channel of the audio signal.
[0021] In an aspect, automatically applying a first playback equalization to an audio signal configured to enhance speech in the content includes transitioning to the first playback equalization from either i) no playback equalization or ii) a second playback equalization.
[0022] In an aspect, automatically applying a first playback equalization configured to enhance speech in the content to the audio signal includes increasing the volume of speech in the content relative to other content in the audio signal.
[0023] In an aspect, automatically applying a first playback equalization to the audio signal configured to enhance speech in the content includes reducing a volume of non-speech content in the audio signal. In an aspect, automatically applying the first playback equalization further includes increasing a volume of speech in the content.
[0024] In aspects, the second playback equalization comprises at least one of low frequency enhancement or music playback enhancement.
[0025] In aspects, at least one of the first playback equalization or the one or more predefined conditions is configurable by a user.
[0026] In an aspect, the method further includes analyzing sounds in an environment in which the audio signal is to be reproduced to help determine whether a first reproduction equalization should be applied to the audio signal.
[0027] An aspect provides an apparatus for audio signal processing comprising a memory and a processor coupled to the memory, wherein the processor and the memory are configured to: during playback of the audio signal, analyze content of the audio signal prior to playback of the content to determine whether one or more predefined conditions are met to indicate that the content includes speech; in response to determining that the one or more predefined conditions are met, automatically apply a first playback equalization to the audio signal, the first playback equalization being configured to enhance the speech in the content; and in response to determining that the one or more predefined conditions are not met, i) apply no playback equalization or ii) apply a second playback equalization to the audio signal, the second playback equalization being different from the first playback equalization.
[0028] In an aspect, the processor and memory are configured to detect using a trained machine learning model. In an aspect, the audio signal comprises different channels, and the memory and processor are configured to detect by analyzing the different channels, and analyzing the different channels includes comparing correlation content between two of the different channels.
[0029] In an aspect, the processor and memory are configured to detect by analyzing a center channel of the audio signal.
[0030] In an aspect, the processor and memory are configured to automatically apply a first playback equalization to the audio signal, the first playback equalization being configured to enhance speech within the content, by transitioning from either i) no playback equalization or ii) a second playback equalization to the first playback equalization.
[0031] In aspects, the second playback equalization comprises at least one of low frequency enhancement or music playback enhancement. In aspects, at least one of the first playback equalization or the one or more predefined conditions is user configurable.
[0032] In an aspect, the processor and memory are further configured to analyze sounds within an environment in which the audio signal is to be reproduced to aid in determining whether to apply a first reproduction equalization to the audio signal.
[0033] An aspect provides a non-transitory computer-readable medium storing instructions that, when executed by a device for processing and generating an audio signal, cause the device to: analyze content of the audio signal during playback of the audio signal and prior to playback of the content; determine whether one or more predefined conditions are met to indicate that the content includes speech; in response to determining that the one or more predefined conditions are met, automatically apply a first playback equalization to the audio signal, the first playback equalization being configured to enhance the speech in the content; and in response to determining that the one or more predefined conditions are not met, i) apply no playback equalization or ii) apply a second playback equalization to the audio signal, the second playback equalization being different from the first playback equalization.
[0034] Two or more features described in this disclosure, including features described in the Summary of the Invention section, may be combined to form implementations not specifically described herein.
[0035] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims.
[0036] Like numbers refer to like elements, and "speech" and "dialogue" may be used interchangeably. [Brief explanation of the drawings]
[0037] [Figure 1] 1 illustrates an example of a system in which aspects of the present disclosure may be implemented. [Figure 2] 1 illustrates a block diagram of a machine learning network and associated components in accordance with some aspects of the present disclosure. [Figure 3A]1 illustrates an exemplary determination of different sound categories in an audio signal according to some aspects of the present disclosure. [Figure 3B] 1 illustrates an exemplary determination of different sound categories in an audio signal according to some aspects of the present disclosure. [Figure 4] 1 illustrates an example equalization process for speech in accordance with some aspects of the present disclosure. [Figure 5] 1 illustrates an exemplary process for generating a training dataset according to some aspects of the present disclosure. [Figure 6] FIG. 1 is a flow diagram illustrating example operations for automatic speech enhancement according to some aspects of the present disclosure. [Figure 7] FIG. 1 is a flow diagram illustrating example operations for automatic speech enhancement according to some aspects of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0038] Because users often react with significant delay to audio content, such as movies, that contain various types of audio signals (e.g., dialogue, music, singing, etc.), a user's manual equalization adjustments may be too slow. Additionally, manual selection of a speech-based equalization profile can be tedious and discouraging. Therefore, methods for automatically (without additional user input) updating equalization profiles to take speech content into account, as well as devices and systems configured to implement these methods, are desired.
[0039] The present disclosure provides processes, methods, systems, and devices for intelligently detecting speech or dialogue content in audio signals (e.g., by implementing machine learning and artificial intelligence) and smoothly transitioning to a speech equalization mode that allows a user to better understand the linguistic meaning of the speech content in the audio signals. For example, aspects of the present disclosure provide methods for processing and generating audio signals. In one example, the method detects speech components in an original set of audio signals using a machine learning network trained based on a mixture of categories of audio content. Regardless of how the speech components are detected, the detected speech components may be enhanced by transitioning from the original equalization mode to a speech equalization mode that allows a user to better understand the linguistic meaning therein.
[0040] Multimedia played in a home theater or entertainment center often includes different types of movies or television programs. These different types of multimedia may include sound reproduction that benefits from different equalization settings. As an example, a documentary may include substantial spoken narration. The spoken narration may benefit more from equalization settings for speech than, for example, a concert that includes substantial singing or music. Traditionally, equalization settings may be adjusted on high-fidelity (hi-fi) equipment by applying filters or amplification across different frequency bands from low to high. Hi-fi equipment may allow manual adjustment or recall of pre-configured equalization settings (e.g., pre-configured for jazz, classical, pop, concerts, etc.). However, such manual adjustment or recall may be inconvenient, especially when a user cycles through television channels or when a movie includes different scenes with different key sound elements (e.g., dialogue, music, sound effects, etc.).
[0041] For example, a user may manually select a pre-configured equalization profile that enhances speech and then recognize that the media includes a musical background. The sound quality of the music may be adversely affected by the selected equalization profile. Similarly, if a user selects a pre-configured equalization profile that enhances music or other sound effects, or in situations with complex, loud surround sound, the sound quality for dialogue may be adversely affected or may be less intelligible.
[0042] Aspects of the present disclosure overcome such challenges by detecting different categories of sounds and identifying whether or when dialogue or speech is the primary content being played. In some aspects, machine learning models are used to detect or identify different sound categories. To facilitate comprehension, aspects of the present disclosure implement a smooth transition to a speech-enhancing equalization profile so that a user can better understand the meaning therein. Once the speech content ends, aspects of the present disclosure provide for a smooth transition of playback to the original equalization profile or another preferred equalization profile.
[0043] Dialogue often suffers from a lack of intelligibility in A4V (audio-for-video) content. To improve dialogue intelligibility, sound products (e.g., speakers, sound bars, headphones, etc.) may include pre-configured equalization modes that enhance dialogue or speech intelligibility. When dialogue equalization modes are turned on, non-dialogue audio quality may be adversely affected. However, due to interlacing with other non-dialogue content, users may not timely or actually turn on dialogue modes at the appropriate times (e.g., on multiple different occasions in the same playback session).
[0044] The present disclosure provides the benefit of improved (e.g., better sounding) dialogue equalization modes. Generally, automatic equalization of audio content to enhance speech in the content, as variously described herein, allows equalization to be applied when speech is detected, thereby dynamically applying speech enhancement equalization based on the content itself, as opposed to statically turning equalization on or off. This allows for various benefits, such as more aggressive / desired speech enhancement and / or more aggressive / desired equalization of non-speech content (e.g., enhancing bass in action scenes in movies or applying music enhancement equalization when music is being played). In other words, dialogue enhancement, as variously described herein, may be applied only when one or more predefined conditions are met to indicate that there is known or high confidence that there is actual speech in the audio content. One exemplary method for determining whether audio being played contains speech (or whether it is likely to contain speech can be determined by preprocessing the audio data immediately before it is played) is to use an algorithm, such as a trained machine learning model. By implementing machine learning to detect actual speech / dialogue, dialogue tuning can be implemented as needed (e.g., taking into account background sounds and ambient noise). By using machine learning and artificial intelligence in dialogue recognition models, aspects of the present disclosure significantly reduce the misapplication of dialogue modes to non-dialogue modes, resulting in improved audio performance for music or action scenes. Automatic aspects of implementing the present disclosure require minimal input from the user and improve overall dialogue intelligibility.
[0045] In an aspect, the deep learning model estimates multiple energy levels in media audio content in real time. The energy levels may include speech energy, music energy, and singing energy. The system may also calculate the ratio of different sound categories compared to the total energy in the content signal to identify the dominant sound category.
[0046] A deep learning model (e.g., artificial intelligence) can be trained such that sound category decisions are not pre-programmed but are determined based on machine learning. In one example, to train the deep learning model, data is synthesized to emulate A4V content. A4V content can be emulated by mixing separate tracks of music, speech, sound effects, and singing. Using these energy levels, control logic is applied to seamlessly fade or transition in and out of interactive modes.
[0047] In addition to classifying content types, deep learning models can also be trained for content understanding (e.g., what the content is as well as what category such content belongs to). For example, in addition to estimating the musical or vocal energy level, the sound of the instruments used can also be estimated. In some cases, music or video genre classifiers can also be included in deep learning models to determine whether a sound may belong to a hip hop or action movie.
[0048] Based on the detection of content, aspects of the present disclosure can improve control logic for the playback device. For example, in addition to the output of the deep learning model, system states such as volume level and subwoofer availability can also be output. Environmental noise monitoring can also be implemented to improve control logic (e.g., for active noise cancellation).
[0049] The detection and / or measurement information can be used to calculate an equalization profile for the interactive mode, and thus a comprehensive machine learning model can be applied to estimate the salience or intelligibility of the speech.
[0050] In an aspect, a deep learning model is trained to calculate a confidence level for the determined content category. For example, the deep learning model can output a confidence measure of whether the content is, for example, speech or non-speech. The confidence parameter can help control aspects of equalization changes for the user's benefit. As an example, when the model indicates a sudden high probability or confidence level that the detected sound is speech, the ballistic time constant for switching equalization can occur quickly to switch to speech mode as quickly as possible. Additionally or alternatively, in an aspect, the confidence level of the determined content category affects how much equalization change is implemented. As an example, if the maximum change in channel center is 4 dB and the model indicates a high confidence level that the content category is speech, the control logic may add the full 4 dB for output by the playback device. If the maximum change in channel center is 4 dB and the model indicates a low confidence level that the content category is speech, the control logic may apply less than the full 4 dB for output by the playback device. In this manner, the amount of adjustment is based, at least in part, on the model's confidence level.
[0051] In some cases, to improve sound quality, speech intelligibility, and specialization, deep learning models detect or recognize sources in an audio signal and isolate or extract some source or type of content for a particular output channel. For example, speech content may be extracted from the rest of the soundtrack and played over a pair of synchronized open audio headphones. In this way, surround sound is maintained, allowing the user to clearly understand the speech content.
[0052] In some cases, aspects of the present disclosure are applied to a variety of speakers in different environments, including automotive audio, portable speakers, headphones, earphones, etc., as described below.
[0053] 1 illustrates an example of a system 100 in which aspects of the present disclosure may be practiced. As shown, system 100 includes one or more sound processing and playback devices 110 (e.g., wireless audio devices such as sound bars or smart speakers) communicatively coupled to a source device 120 (e.g., a computing device or user device such as a smartphone or tablet computer). One or more partner devices 112 (e.g., portable speakers, headsets, etc.) may be available to accept pairing requests from sound processing and playback device 110 or source device 120. Sound processing and playback device 110 may be paired with source device 120 and may receive content data (including audio signals) from source device 120. Sound processing and playback device 110 may also receive content data directly from network 130. Partner device 112 may be a battery-powered portable device suitable for mobile or privacy applications.
[0054] According to aspects of the present disclosure, sound processing and playback device 110 may receive an original set of audio signals from at least one of source device 120, network 130, or cloud 140 (via network 130). The content of the audio signals is analyzed prior to playback to determine whether one or more predefined conditions are met that indicate the content includes speech. The analysis may use a variety of different techniques, such as using a trained machine learning model as described herein, analyzing metadata associated with the content (e.g., if the metadata indicates that the content includes speech), analyzing the voice track of the audio signal (if present), analyzing different channels of the audio signal, and / or other techniques as may be understood based on this disclosure. Techniques that utilize metadata associated with the audio signal and / or its content may analyze the metadata immediately prior to playback to determine whether the metadata indicates that the audio content includes speech or otherwise meets predefined conditions, and if so, speech / dialogue-enhancing equalization (as variously described herein) may be automatically applied to assist intelligibility of the speech. The metadata may be in the form of text associated with the audio content (such as closed caption data or other subtitles), genre data (such as indicating that the content is a podcast or talk show), and / or other data that assists in determining whether speech is included in the content. If the audio signal or associated content includes a voice track, the energy level of the voice track may be analyzed to determine whether the audio content includes speech. If the audio signal includes different channels, they may be analyzed and / or compared to determine whether speech is likely occurring.For example, such analysis may include comparing the correlation content between two channels (such as the correlation content between the left and right channels of a stereo audio signal) and / or analyzing the center channel (e.g., in a 5.0, 5.1, 7.0, or 7.1 audio signal), since this is where the majority of dialogue from movies and television typically occurs. For example, center channel analysis may include determining when the center channel reproduction exceeds a threshold (e.g., a nominal threshold or a threshold relative to other channels) to determine that the content is likely to include speech, and therefore, speech-enhancement equalization should be automatically applied. Many different techniques will become apparent in light of this disclosure.
[0055] In response to determining that one or more predefined conditions are met, a first playback equalization configured to enhance speech in the content is applied automatically. As described herein, automatically may mean without user input. Thus, in response to determining that one or more predefined conditions are met, a first playback equalization configured to enhance speech in the content is applied without user input.
[0056] In an aspect, applying the first playback equalization to the audio signal comprises a transition from either no playback equalization or the second playback equalization to the first playback equalization. The transition comprises a gradual change from either no playback equalization or the second playback equalization to the first playback equalization. In an aspect, the second playback equalization comprises low frequency enhancement and / or music playback enhancement.
[0057] In an aspect, applying the first playback equalization to the audio signal includes increasing the volume of speech in the content relative to other content in the audio signal. In an aspect, the speech content is extracted using a machine learning algorithm. Additionally or alternatively, in an aspect, the speech content is taken from at least one of correlation content between two channels or a center channel (as described above). In an aspect, the speech content is taken from a speech component of the audio signal.
[0058] In an aspect, applying the first playback equalization to the audio signal includes reducing a volume of non-speech content within the audio signal. In an aspect, in addition to reducing the volume of the non-speech content, a volume of speech within the content is increased.
[0059] In aspects, the second playback equalization includes enhancing low frequency components of the content and / or enhancing musical playback of the content.
[0060] In some aspects, the sound processing and playback device 110 may analyze an original set of audio signals and detect speech components using a trained machine learning network (e.g., deep learning model 260 of FIG. 2). The machine learning network is trained to identify speech components based on mixed categories of audio content, such as dialogue, music, singing, and other sound categories (e.g., instruments or digital sound effects). Speech components may include any sound elements that convey linguistic meaning.
[0061] Upon detecting speech components, sound processing and reproduction device 110 can enhance the detected speech components by transitioning from the original equalization mode to a speech (or dialogue) equalization mode. The speech equalization mode can enable a user to better understand the linguistic meaning in the speech components. For example, the speech equalization mode can enhance frequency spectrum associated with dialogue and / or suppress non-speech frequency spectrum. When sound processing and reproduction device 110 does not detect speech components, or when the detected speech components are missing or interrupted in the arriving audio signal, sound processing and reproduction device 110 can output the original set of audio signals in the original equalization mode. Thus, sound processing and reproduction device 110 intelligently applies the speech equalization mode only when necessary to achieve speech enhancement and minimize adverse effects on non-speech audio content.
[0062] In some scenarios, playback device 110 provides different volume settings for different frequency bands. Dynamic equalization can adjust the overall system frequency response as a function of the detected mode and can provide a loudness compensation function that can de-emphasize lower bass content and emphasize higher frequency content to improve volume setting-related intelligibility.
[0063] In aspects, a predicted system-wide sound pressure level (SPL) is monitored and maintained as playback device 110 switches between equalization modes. By monitoring the SPL, a similar SPL is maintained between the two modes of equalization, which reduces perceived volume changes during non-dialogue content.
[0064] The sound processing and playback device 110 may further include hardware and circuitry, including a processor / processing system and memory, configured to implement one or more sound management or other capabilities, including, but not limited to, noise cancellation circuitry (not shown) and / or noise masking circuitry (not shown), body movement detection devices / sensors and circuitry (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetometers, etc.), geolocation circuitry, and other sound processing circuitry.
[0065] In one aspect, sound processing and playback device 110 is wirelessly connected to source device 120 or partner device 112 using one or more wireless communication methods, including but not limited to Bluetooth, Wi-Fi, Bluetooth Low Energy (BLE), other RF-based techniques, etc. In one aspect, sound processing and playback device 110 includes a transceiver that transmits and receives data via one or more antennas to exchange audio data and other information with source device 120.
[0066] In one aspect, sound processing and playback device 110 includes communications circuitry capable of transmitting and receiving audio data and other information from source device 120. Sound processing and playback device 110 also includes an incoming audio buffer, such as a render buffer, that buffers at least a portion of the incoming audio signal (e.g., audio packets) to allow time for retransmission of any missing or missing data packets from source device 120. For example, when sound processing and playback device 110 receives a Bluetooth transmission from source device 120, the communications circuitry typically buffers at least a portion of the incoming audio data in the render buffer before the audio is actually rendered and output as audio to at least one of the transducers (e.g., audio speakers) of sound processing and playback device 110. This is done to ensure that even if there is an RF collision that causes an audio packet to be lost during transmission, the lost audio packet has time to be retransmitted by source device 120 before it needs to be rendered by sound processing and playback device 110 for output by one or more acoustic transducers of sound processing and playback device 110.
[0067] Although one example of partner device 112 is shown as noise-canceling headphones, the techniques described herein apply to other wireless audio devices, such as wearable audio devices, including any audio output device that fits around, on, in, or near the ear (including open-ear audio devices worn on the user's head or shoulders), or other body part of the user, such as the head or neck. Partner device 112 may take any form, wearable or otherwise, including stand-alone devices (including automobile speaker systems), stationary devices (including portable devices such as battery-powered portable speakers), headphones, earphones, earpieces, headsets, goggles, headbands, earbuds, armbands, sports headphones, neckbands, or glasses with integrated speakers.
[0068] In one aspect, sound processing and playback device 110 is connected to source device 120 using a wired connection, with or without a corresponding wireless connection. Source device 120 may be a smartphone, tablet computer, laptop computer, digital camera, or other user device that connects with sound processing and playback device 110. As shown, source device 120 can connect to network 130 (e.g., the Internet) and can access one or more services on the network. As shown, these services may include one or more cloud services 140.
[0069] In one aspect, source device 120 can access a cloud server in cloud 140 on network 130 using a mobile web browser or a local software application or “app” running on source device 120. In one aspect, the software application or “app” is a local application installed and executed locally on source device 120. In one aspect, a cloud server accessible on cloud 140 includes one or more cloud applications running on the cloud server. The cloud applications may be accessed and executed by source device 120. For example, a cloud application may generate a web page that is rendered by a mobile web browser on source device 120. In one aspect, a mobile software application installed on source device 120 or a cloud application installed on a cloud server may be used, individually or in combination, to implement techniques for low-latency Bluetooth communication between source device 120 and sound processing and playback device 110 according to aspects of the present disclosure. In one aspect, examples of local software applications and cloud applications include gaming applications, audio AR applications, and / or gaming applications with audio AR capabilities. Source device 120 can receive signals (e.g., data and control) from sound processing and reproduction device 110 and send signals to sound processing and reproduction device 110.
[0070] An exemplary sound processing and playback device 110 or partner device 112 may include the components described below (not shown in FIG. 1 ). For example, the sound processing and playback device 110 or partner device 112 may each include one or more processors, memory modules, communications modules, and / or input interfaces for receiving user input. The sound processing and playback device 110 or partner device 112 may each include one or more electroacoustic transducers (or speakers) for outputting audio. The sound processing and playback device 110 also includes a user input interface. The user input interface may include multiple preset indicators, which may be hardware buttons. The preset indicators can provide a user with easy, single-press access to entities assigned to those buttons. The assigned entities can be associated with different digital audio sources, such that a single sound processing and playback device 110 can provide single-press access to a variety of different digital audio sources.
[0071] Sound processing and playback device 110 or partner device 112 may essentially include an acoustic driver or speaker for converting audio signals into acoustic energy via audio hardware. Sound processing and playback device 110 also includes a network interface, at least one processor, the audio hardware, a power supply for powering the various components of sound processing and playback device 110, and memory. In one aspect, the processor, network interface, power supply, and memory are interconnected using various buses, and some of the components may be mounted on a common motherboard or in other manners as needed. In some cases, sound processing and playback device 110 or partner device 112 may include a housing that houses an optional graphical interface (e.g., an OLED display) that can provide a user with information about the music that is currently playing ("Now Playing").
[0072] The network interface may provide communication between the sound processing and playback device 110 and other electronic user devices, such as the source device 120 and the partner device 112, via one or more communication protocols, such as the Bluetooth® Classic protocol, the Bluetooth® Low Energy protocol, etc. Generally, the network interface provides either or both a wireless network interface and a wired interface (optional). The wireless interface allows the sound processing and playback device 110 to communicate wirelessly with other devices according to a wireless communication protocol, such as IEEE 802.11. The wired interface provides network interface functionality over a wired (e.g., Ethernet) connection for reliability and fast transfer speeds, for use, for example, when the sound processing and playback device 110 is not worn by the user.
[0073] In some aspects, the network interface includes a network media processor to support Apple AirPlay® and / or Apple Airplay® 2. For example, if a user connects an AirPlay® or Apple Airplay® 2-enabled device, such as an iPhone or iPad device, to a LAN, the user can then stream music to a network-connected audio playback device via Apple Airplay® or Apple Airplay® 2. In particular, the audio playback device can support audio streaming via AirPlay®, Apple Airplay® 2, and / or DLNA's UPnP protocol, all integrated into one device.
[0074] All other digital audio received as part of a network packet can be passed straight from the network media processor through a USB bridge (not shown) to the processor, to the decoder, to the DSP, and finally to be played back (rendered) via an electroacoustic transducer.
[0075] The network interface may further include Bluetooth circuitry for Bluetooth applications (e.g., for wireless communication with Bluetooth-enabled audio sources such as smartphones or tablets) or other Bluetooth-enabled speaker packages. In some aspects, the Bluetooth circuitry may be the primary network interface due to energy constraints. For example, the network interface may use the Bluetooth circuitry solely for mobile applications when sound processing and playback device 110 or partner device 112 adopts any wearable form factor. For example, BLE technology may be used in sound processing and playback device 110 or partner device 112 to extend battery life, reduce package weight, and provide high-quality performance without other backup or alternative network interfaces.
[0076] In one aspect, the network interface supports communication with other devices using multiple communication protocols simultaneously at one time. For example, sound processing and playback device 110 can support Wi-Fi / Bluetooth coexistence and support simultaneous communication using both Wi-Fi and Bluetooth protocols at one time. For example, sound processing and playback device 110 can receive an audio stream from a smartphone using Bluetooth and further simultaneously rebroadcast the audio stream to one or more other devices over Wi-Fi. In one aspect, the network interface can include only one RF chain capable of communicating using only one communication method (e.g., Wi-Fi or Bluetooth) at a time. In the present context, the network interface can support Wi-Fi communication and Bluetooth communication simultaneously, for example, by time-dividing a single RF chain between Wi-Fi and Bluetooth according to a time division multiplexing (TDM) pattern.
[0077] The streamed data may be passed from the network interface to a processor. The processor may execute instructions (e.g., for performing digital signal processing, decoding, and equalization functions, among others), including instructions stored in memory. The processor may be implemented as a chipset of chips including separate analog and digital processors. The processor may provide audio sound processing and coordination of other components of playback device 110, such as controlling a user interface.
[0078] The memory may store software / firmware related to protocols and their versions used by the sound processing and playback device 110 or partner device 112 to communicate with other networked devices, including the source device 120. For example, the software / firmware manages how the sound processing and playback device 110 communicates with other devices for synchronized audio playback. In one aspect, the software / firmware includes lower-level frame protocols related to control path management and audio path management. Protocols related to control path management generally include protocols used to exchange messages between speakers. Protocols related to audio path management generally include protocols used for clock synchronization, audio distribution / frame synchronization, audio decoder / time alignment, and playback of audio streams. In one aspect, the memory may also store various codecs supported by the speaker package for audio playback of each media format. In one aspect, the software / firmware stored in the memory may be accessible and executable by a processor for synchronized audio playback with other networked speaker packages.
[0079] In aspects, the protocols stored in memory may include, for example, BLE according to Bluetooth Core Specification version 5.2 (BT5.2). Sound processing and playback device 110 or partner device 112, and various components therein, are provided herein to fully conform to or implement aspects of the protocols and associated specifications. For example, BT5.2 includes an enhanced attribute protocol (EATT) that supports concurrent transactions. To support EATT, a new L2CAP mode is defined. Thus, sound processing and playback device 110 includes sufficient hardware and software components to support the BT5.2 specifications and modes of operation, even if not explicitly shown or discussed in this disclosure. For example, sound processing and playback device 110 may utilize LE isochronous channels specified in BT5.2.
[0080] The processor may provide the processed digital audio signals to audio hardware that includes one or more digital-to-analog (D / A) converters for converting the digital audio signals to analog audio signals. The audio hardware also includes one or more amplifiers that provide amplified analog audio signals to electroacoustic transducers for sound output. In addition, the audio hardware may include circuitry for processing analog input signals to provide digital audio signals for sharing with other devices, such as other speaker packages for synchronized output of digital audio.
[0081] The memory may include, for example, non-transitory memory such as flash memory and / or non-volatile random access memory (NVRAM). In some aspects, the instructions (e.g., software) are stored on an information carrier. When executed by one or more processing devices (e.g., processors), the instructions perform one or more processes, such as those described elsewhere herein. The instructions may also be stored by one or more storage devices, such as one or more computer-readable or machine-readable media (e.g., memory, or memory on a processor). The instructions may include instructions for performing decoding (i.e., a software module may include an audio codec for decoding a digital audio stream), as well as digital signal processing and equalization. In some aspects, the memory and processor may cooperate with a microphone on sound processing and playback device 110 or source device 120 in data acquisition and real-time processing.
[0082] Exemplary intelligent dialogue or speech enhancement Aspects of the present disclosure provide techniques for intelligently detecting and enhancing detected speech components in audio signals. For example, an audio device may detect speech components in an original set of audio signals using any number of methods. One example is a trained machine learning network based on mixed categories of audio content. Speech components include sound elements that convey linguistic meaning. The audio device may enhance the detected speech components by transitioning from an original equalization mode to a speech equalization mode that allows a user to better understand the linguistic meaning of the speech. In the absence of detection of speech components (e.g., before or after detection of speech components), the audio device may output the original set of audio signals in the original equalization mode. To correctly detect speech components (e.g., as opposed to singing components), a machine learning network may be trained to recognize what constitutes speech without user intervention. An exemplary intelligent dialogue or speech enhancement according to the present disclosure is provided in FIG. 2.
[0083] 2 is a block diagram 200 illustrating the relationship between audio signals, training data sets, and processing components according to an aspect of the present disclosure. As shown, an original set of audio signals 210, as received by sound processing and reproduction device 110, is provided to machine learning network 220. For example, sound processing and reproduction device 110 may be communicatively coupled to machine learning network 220 via network 130.
[0084] The machine learning network 220 may analyze the original set of audio signals 210 using a deep learning model 260 coupled with the machine learning network 220. While FIG. 2 depicts the deep learning model 260 as separate from the machine learning network 220, in some cases, the deep learning model 260 may be integrated with the machine learning network 220. While FIG. 2 depicts one deep learning model 260 coupled with the machine learning network 220, in some cases, two or more different deep learning models may be coupled or integrated with the machine learning network 220. In some cases, the machine learning network 220 or its interface (e.g., a graphical user interface such as an application on an operating system) may be installed on the source device 120, which may be a smartphone.
[0085] The deep learning model 260 can use various machine learning techniques based on artificial neural networks. For example, the deep learning model 260 can include deep learning architectures such as deep neural networks, deep belief networks, deep reinforcement learning, recurrent neural networks, convolutional neural networks, etc. Similar to speech recognition, the deep learning model 260 can identify sound elements that contain linguistic meaning. Alternatively, the deep learning model 260 can be trained to distinguish sound elements that primarily represent linguistic meaning from sound elements that primarily represent tonal or musical elements other than the linguistic meaning in context.
[0086] For example, the deep learning model 260 may be trained to distinguish speech components from music components or singing components. Speech components may include sounds that convey meaning based on linguistic features. Musical components may include sounds that lack linguistic features. Singing components may include a mixture of sounds that simultaneously contain linguistic and musical components.
[0087] The deep learning model 260 is trained to identify sounds that do not contain musical expressions or lack linguistic features based on the training dataset 250. For example, the training dataset 250 may include various types of singing, music, and dialogue. The deep learning model 260 may be supervised, semi-supervised, or unsupervised to learn whether a sound pattern belongs to one of three categories. For example, the training dataset 250 may include various samples of music, opera, rap music, chorus, conversation, dialogue, speech, etc.
[0088] The machine learning network 220 can then use the deep learning model 260 to calculate the root mean square of the respective energy levels of each category of the mixture categories of audio content. For example, the deep learning model 260 may estimate the respective energy levels of each category of the mixture categories of audio content. The mixture categories of audio content may include a speech component, a music component, and a singing component, which may correspond to samples included in the training dataset 250. Figure 5 and the corresponding description below provide further details of the training dataset 250 below.
[0089] Speech components may be defined or detected by determining the ratio of the energy level of the speech component to the overall energy level above a threshold (example shown in FIG. 3). In some cases, the trained machine learning network 220 is trained to identify categories and thresholds for each of a mixture of categories of audio content based on a known database of movie content (e.g., samples in training dataset 250).
[0090] The machine learning network 220 may include a speech component enhancement module 230. The machine learning network 220 computes in a speech equalization mode when a speech component is detected and provides a computation output 240.
[0091] The speech component enhancement module 230 can be trained to implement speech equalization modes as well as improve speech equalization modes. For example, the speech component enhancement module 230 may increase the magnitude or contrast of speech-related frequency bands to improve speech intelligibility. In some cases, the speech component enhancement module 230 decreases the magnitude of non-speech-related frequency bands or signal channels. In some cases, the speech component enhancement module 230 changes or updates equalization settings or dynamic range compression settings on non-speech-related frequency bands or signal channels. The speech component enhancement module may combine two or more of these operations. An example of a speech equalization mode is provided in FIG. 4 and described below.
[0092] Exemplary Dialogue or Speech Detection 3A and 3B illustrate exemplary determination of different sound categories in an audio signal according to some aspects of the present disclosure. FIG. 3A illustrates a first instance 310 of a video clip 305, and FIG. 3B illustrates a second instance 320. In the first instance, a character in the video clip is speaking. The window includes a category indication frame 312 that monitors the currently detected category of sound (e.g., singing, music, and speech). The window for the first instance 310 highlights frame 314, which indicates a speech component, and the window for the second instance 320 highlights frame 316, which indicates a singing component. Such indications can be useful for supervised learning, allowing a trainer to verify whether the machine learning network 220 correctly classified the sound content.
[0093] 3A and 3B also show a data processing window 320 within each video window. The data processing window 320 plots calculated model outputs, such as energy levels, for each of the sound categories. For example, in the data processing window 320 of the first instance 310, the root mean square of the energy level of the speech component 322 dominates the output. The root mean square of the energy levels of the music component 324 and the singing component 326 are low. In some cases, the ratio of the energy level of the speech component to the overall energy level may be compared to a threshold level (learned based on the dataset). A speech component is detected when the ratio exceeds the threshold. Thus, in the first instance 310, the machine learning network 220 detects the speech component and indicates the detection by highlighting the frame 314 corresponding to singing.
[0094] In the data processing window of the second instance 320, the root mean square of the energy level of the singing component 326 dominates the output, while the root mean square of the energy levels of the music component 324 and the speech component 322 are low. Thus, in the second instance 320, the machine learning network 220 detects the singing component and indicates the detection by highlighting the frame 316 corresponding to the singing.
[0095] 3A and 3B show example energy determinations for three categories of sound, but different sets of categories may be used. For example, the machine learning network 220 may be further trained to identify noises that do not belong to any of the speech, music, or singing categories. According to aspects of the present disclosure, detecting and identifying noises may enable the machine learning network 220 to receive and process captured environmental noise for active noise cancellation. The noise categories may also improve accuracy in detecting other sound categories.
[0096] Exemplary Dialogue or Speech Processing Upon detecting a speech component, the machine learning network 220 may enhance the detected speech component by transitioning from the original equalization mode to a speech equalization mode. FIG. 4 illustrates an example of such an equalization process. The original equalization mode 410 (alternatively, no equalization mode) is transitioned to the speech equalization mode 420 (via process 432) when a speech component is detected and is dominant in the content (e.g., excluding background speech noise). As shown, in the original equalization mode 410, three profiles are used to tune the center channel, the left / right channels, and the left surround / right surround channels. Before transitioning to the speech equalization mode 420, the bass and treble portions of all channels have non-zero amplification to provide a rich surround sound. Dialogue or speech in such an equalization mode may be difficult to understand given other sound components.
[0097] Process 432 implements a transition that minimizes the perceptibility of the change from the original equalization mode 410 (which may not even be an equalization mode) to the speech equalization mode 420. For example, frequency band changes, channel changes, or various other parameters may smoothly transition at different rates or times over different periods to seamlessly blend the two equalization modes 410 and 420. In this way, the user may not perceive the change. Process 432 may be performed by the speech component enhancement module 230 of FIG. 2.
[0098] In the second equalization mode 420, the base and treble frequency spectra of all three channels (center, side, and surround) may be tuned to values close to zero as shown, thus allowing the spectral range corresponding to speech frequencies to be more distinct from other non-speech sound components. Thus, speech is more intelligible when reproduced in the speech equalization mode 420 than in the original equalization mode 410.
[0099] Example training datasets for machine learning FIG. 5 illustrates an exemplary process for generating a training dataset, such as training dataset 250 of FIG. 2, according to some aspects of the present disclosure. As an example, the training dataset may be used by a deep learning model based on a convolutional recurrent neural network (CRNN). The CRNN receives inputs created by mixing multiple clean sound sources with some noise to generate synthetic examples. In some cases, the inputs may take a format such as a Log Mel Spectrogram. As described in FIG. 2, the deep learning model may provide outputs about the energy levels of different categories of sounds. The output may include several frames having a total time length equal to the input duration.
[0100] In some cases, datasets for deep learning models can be configured by a user to update or mix parameters of the datasets and / or parameters for the deep learning model. The training process can include closed-loop feedback (e.g., supervised), where the model being trained receives evaluations based on both sound content detection and equalization (e.g., speech and non-speech mode) output evaluations. For example, as shown in FIG. 5, four source datasets are used to generate synthetic example 570. Speech dataset 510, music dataset 520, singing dataset 530, and noise dataset 540 are mixed in different ways to form first sound category 550 and second sound category 560. First sound category 550 includes all datasets 510-540 and emulates video audio. Second sound category 560 includes singing dataset 530 and noise dataset 540, emulating concert audio, etc. The synthetic example 570 can select one of the first category and the second category 550 or 560 to train the deep learning model.
[0101] Method and process for intelligent dialogue enhancement 6 is a flow diagram illustrating example operations 600 that may be performed by a target device to establish wireless communication with another device. For example, example operations 600 may be performed by sound processing and playback device 110 of FIG.
[0102] The example operations 600 are performed at 602 by analyzing the content of an audio signal during playback and prior to playback of the content to determine whether one or more predefined conditions are met to indicate that the content includes speech.
[0103] At 604, in response to determining that the one or more predefined conditions are met, automatically apply a first playback equalization configured to enhance speech in the content to the audio signal. Automatically applying the first playback equalization may refer to applying the first playback equalization without further input from a user.
[0104] At 606, in response to determining that one or more predefined conditions are not met, either i) no playback equalization or ii) a second playback equalization different from the first playback equalization is applied to the audio signal.
[0105] In particular, in aspects, content is analyzed immediately prior to and during playback of the audio signal. This analysis may occur between 20 milliseconds and 2 seconds before playback of the audio. In aspects, the analysis is more likely to occur in the range of 0.2 to 1 second. In aspects, the analysis occurs 0.5 seconds before playback. By analyzing content before and during playback, aspects of the present disclosure differ from existing techniques that do not have the real-time or near real-time processing challenges associated with the disclosed techniques.
[0106] In an embodiment, the analyzing is performed using a trained machine learning model as described above. In an embodiment, metadata associated with the content is analyzed, the metadata including that the content includes speech. In an embodiment, a voice track of the audio signal is analyzed when one or more predefined conditions exceed a threshold.
[0107] In embodiments, different channels of an audio signal are analyzed, for example, content is correlated between two channels.
[0108] When an audio signal includes different channels, they may be analyzed and / or compared to determine whether speech is likely occurring. For example, such analysis may include comparing the correlation content between two channels (such as the correlation content between the left and right channels of a stereo audio signal) and / or analyzing the center channel (e.g., in a 5.0, 5.1, 7.0, or 7.1 audio signal), since the center channel is typically where the majority of dialogue from movies and television occurs. For example, center channel analysis may include determining when the center channel reproduction exceeds a threshold (e.g., a nominal threshold or a threshold relative to other channels) to determine that the content likely includes speech, and therefore, speech-enhancement equalization should be automatically applied.
[0109] According to an aspect, applying the first playback equalization to the audio signal includes transitioning to the first playback equalization from either i) no playback equalization or ii) a second playback equalization. The transition may include a step change from either i) no playback equalization or ii) the second playback equalization to the first playback equalization.
[0110] In an aspect, applying the first playback equalization to the audio signal includes increasing the volume of speech in the content relative to other content in the audio signal. In an aspect, the speech content is extracted using a machine learning algorithm. As described herein, the speech content is taken from at least one of i) correlation content between two channels, or ii) a center channel. In an aspect, the speech content is taken from a speech component of the audio signal.
[0111] In an aspect, applying the first playback equalization to the audio signal includes reducing the volume of non-speech content in the audio signal. Further, in an aspect, the volume of speech in the content is increased.
[0112] In aspects, the second playback equalization includes low frequency enhancement, music playback enhancement, or a combination of both.
[0113] In an aspect, sounds within an environment in which an audio signal is to be reproduced are analyzed to help determine whether to apply a first reproduction equalization to the audio signal.
[0114] In some aspects, the predefined conditions are user configurable. Additionally or alternatively, in some aspects, the first playback equalization may be user configurable.
[0115] 7 is a flow diagram illustrating example operations 700 that may be performed by a target device to establish wireless communication with another device. For example, the example operations 700 may be performed by the sound processing and playback device 110 of FIG.
[0116] The example operations 700 begin at 702 by detecting speech components in an original set of audio signals using a machine learning network trained on mixed categories of audio content, where the speech components consist of sound elements that convey linguistic meaning.
[0117] At 704, the detected speech components are enhanced by transitioning from the original equalization mode to a speech equalization mode that allows the user to better understand the linguistic meaning therein.
[0118] At 706, the original set of audio signals is output in the original equalization mode without speech detection.
[0119] In operation, the performance of 704 and 706 depends on the instance of content in the audio signal, which may change from time to time. Therefore, the enhancement performed in 704 is dynamic and can be applied automatically to the detected speech components.
[0120] In an aspect, the speech component may be detected using a trained machine learning network. For example, the detection may be based on calculating the root mean square of the respective energy levels of each category of the mixture category of audio content. In some cases, the trained machine learning network may include a deep learning model that estimates the respective energy levels of each category of the mixture category of audio content. The mixture category of audio content may include a speech component, a music component, and a singing component.
[0121] In some cases, speech components may be detected by determining the ratio of the speech component's energy level to an overall energy level above a threshold. For example, speech components may include sounds whose meaning is conveyed based on linguistic features. Musical components may include sounds lacking linguistic features. Vocal components may include a mixture of sounds that simultaneously contain linguistic and musical components. In some cases, a trained machine learning network may be trained to identify categories and thresholds for each of a mixture of categories of audio content based on a known database of movie content.
[0122] In some cases, the speech component may be detected by processing the ongoing audio signal at an advanced time prior to the transition or output operation. For example, the processing or detection operation may be performed in real time or near real time with minimal delay allowed by the computing capabilities of the handling device. In some cases, the processing device and the sound output device may be separate and independent of each other.
[0123] In aspects, enhancing the speech component may include gradually fading from the original equalization mode to the speech equalization mode. In this manner, speech enhancement equalization may be engaged or initiated without being recognized by the user. Similarly, playback may include smoothly fading from the speech equalization mode to the original equalization mode without being recognized by the user. In some cases, the speech equalization mode may include at least one of increasing the magnitude or contrast of speech-related frequency bands, decreasing the magnitude of non-speech-related frequency bands or signal channels, or changing equalization or dynamic range compression settings on non-speech-related frequency bands or signal channels to improve the intelligibility of the speech component.
[0124] In aspects, enhancing the detected speech components may be performed in a first device, and outputting the original set of audio signals may be performed in a second device. For example, the first device may include a sound bar configured to output surround sound (e.g., sound processing and playback device 110 of FIG. 1 ), and the second device may include noise-canceling headphones (e.g., partner device 112 of FIG. 1 ). Thus, when the sound bar receives an audio signal, the sound bar performs calculations (e.g., assisted by machine learning and neural network calculations as described above) to detect the speech components and automatically apply speech mode equalization. Outputting the speech mode enhancement may be performed by the sound bar, the noise-canceling headphones, or both. The first device and the second device may be paired in a short-range wireless communication network.
[0125] In some cases, the first device may be noise-canceling headphones and the second device may be a sound bar configured to output surround sound. The first device and the second device may each include a different device, such as a smartphone or other type of wearable electronic device. Thus, various configurations based on different devices may be configured.
[0126] In aspects, the detected speech components may be extracted and played separately on a third device (e.g., another short-range paired speaker or noise-canceling headphones, such as the second partner device 112 of FIG. 1). In some cases, the extracted speech components may be used to improve an equalization profile, such as by identifying several frequency spectrums for processing in an equalizer. In some cases, the first device, the second device, and the third device are configured to generate a mixed surround sound. For example, different devices may each have an equalization profile for a respective category of sound, such as speech, singing, and background music. In some cases, one or more of the devices may include a microphone or noise sensor for noise cancellation. The devices may be paired with a microphone for measuring ambient noise for cancellation.
[0127] In aspects, the original set of audio signals may be played or output without speech content enhancement based on the determination of speech loss. For example, in a multimedia clip that includes occasional speech or dialog content, speech enhancement processing may be applied only to portions where speech content is detected.
[0128] In aspects, the disclosed method is applicable to wireless earphones, earhooks, or ear-to-ear devices. For example, a host such as a mobile phone may be connected to a bud (e.g., the right bud) via Bluetooth®, and the right bud further connects to the left bud using a Bluetooth® link or other wireless technology such as NFMI or NFEMI. The left bud is initially time-synchronized with the right bud. Audio frames (compressed to mono) are transmitted from the left bud with timestamps (synchronized with the right bud's timestamps) as described in the techniques above. The right bud forwards these encoded mono frames along with its own frames. The right bud does not wait for audio frames from the left bud with the same timestamp. Instead, the right bud transmits any frames that are available and ready to be transmitted using the appropriate packing. It is the responsibility of the receiving application in the host to assemble packets using the timestamps and channel numbers. Depending on how it is configured, the receiving application can choose to merge the decoded mono channel of one bud with the decoded mono channel of the other bud into a stereo track based on the timestamp included in the header of the received encoded frame. This disclosure allows the right bud to simply forward audio frames from the left bud without decoding the frames. This helps conserve battery power in truly wireless audio devices.
[0129] In some aspects, the techniques variously described herein may be used to determine contextual information about a source device and / or a user of the source device. For example, the techniques may be used to help determine aspects of a user's environment (e.g., noisy place, quiet place, indoors, outdoors, on an airplane, in a car, etc.) and / or activity (e.g., commuting, walking, running, sitting, driving, flying, etc.). In some such aspects, sensor data received from the source device may be processed at the target device to determine such contextual information and provide a new or enhanced experience for the user. For example, this may enable customization of playlists or audio content, noise cancellation adjustments, and / or other setting adjustments (e.g., audio equalizer settings, volume settings, notification settings, etc.), to name a few. Because source devices (e.g., headphones or earphones) typically have limited resources (e.g., memory and / or processing resources), using the techniques described herein to offload processing of data from sensors at the source device to the target device while having a system for synchronizing sensor data at the target device offers a variety of applications. In some aspects, the techniques disclosed herein enable a user device to automatically identify an optimized or most preferred configuration or settings for synchronized audio capture operations.
[0130] In some aspects, the techniques variously described herein may be used for numerous audio / video applications. For example, the techniques may be used for stereo or surround sound audio capture from a source device to be synchronized at a target device with video captured from the same source device, another source device, and / or target device. For example, the techniques may be used to synchronize stereo or surround sound audio captured by a microphone on a pair of headphones with video captured from a camera on or connected to the headphones, a separate camera, and / or a smartphone camera, where the smartphone (which is the target device in this example) performs the audio and video synchronization. This may enable real-time playback of stereo or surround sound audio with video (e.g., for live streaming) and capture of recorded video with stereo or surround sound audio (e.g., for posting to a social media platform or a news platform). Additionally, the techniques described herein may enable wirelessly captured audio for audio or video messages without interrupting the user's music or audio playback. Thus, the techniques described herein enable the ability to create immersive and / or noise-free audio for video using a wireless configuration. Furthermore, as can be appreciated based on this disclosure, the described techniques enable schemes that were previously achievable only using wired configurations, and as such, the described techniques free users from the undesirable and uncomfortable experience of being tethered by one or more wires.
[0131] Although the description of the aspects of the present disclosure has been presented above for purposes of illustration, it may be noted that the aspects of the present disclosure are not intended to be limited to any of the disclosed aspects. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described aspects.
[0132] In the above, reference is made to embodiments presented in the present disclosure. However, the scope of the present disclosure is not limited to the particular described embodiments. Aspects of the present disclosure may take the form of entirely hardware embodiments, entirely software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining software and hardware embodiments, all of which may be generally referred to herein as "components," "circuits," "modules," or "systems." Furthermore, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable medium(s) having computer-readable program code embodied thereon.
[0133] Any combination of one or more computer-readable media may be utilized. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media include an electrical connection having one or more wires, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present context, a computer-readable storage medium may be any tangible medium that can contain or store a program.
[0134] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various aspects. In this regard, each block in the flowcharts or block diagrams may correspond to a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions described in the blocks may occur out of the order described in the figures. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or in some cases, the blocks may be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, may be implemented in a dedicated hardware-based system that performs the specified functions or operates a combination of dedicated hardware and computer instructions. [Explanation of symbols]
[0135] 100 systems 110 Playback Devices 112 Partner Devices 120 source devices 130 Network 140 Cloud 200 Block Diagram 210 Audio Signal 220 Machine Learning Network 230 Speech Component Enhancement Module 240 Calculation Output 250 training datasets 260 Deep Learning Models 305 video clips 310 First Instance 312 Category Instruction Frame 314 frames 316 frames 320 Second Instance 322 Speech Components 324 Musical Components 326 Singing component 410 Equalization Mode 420 Second Equalization Mode 510 utterance dataset 520 Music Dataset 530 Singing Dataset 540 Noise Dataset 550 First Sound Category 560 Second Sound Category 600 operations 700 operations
Claims
1. 1. A method for audio signal processing, comprising: During playback of an audio signal, analyzing the content of the audio signal prior to the playback of content to determine whether one or more predefined conditions are met to indicate that the content includes speech; In response to determining that the one or more predefined conditions are met, automatically applying a first playback equalization to the audio signal, the first playback equalization being configured to enhance the speech in the content; and in response to determining that the one or more predefined conditions are not satisfied, applying either i) no playback equalization or ii) a second playback equalization different from the first playback equalization to the audio signal.
2. The method of claim 1 , wherein the analyzing is performed using a trained machine learning model.
3. The method of claim 2 , wherein the trained machine learning model comprises a deep learning model that estimates an energy level of the audio signal.
4. The method of claim 3 , wherein the energy levels of the audio signal include energy levels of any combination of the speech, a musical component of the audio signal, and a singing component of the audio signal.
5. The method of claim 1 , wherein the analyzing comprises analyzing metadata associated with the content, the metadata indicating that the content includes speech.
6. The method of claim 1 , wherein the analyzing comprises analyzing a voice track of the audio signal, and the one or more predefined conditions comprise the voice track exceeding a threshold.
7. The method of claim 1 , wherein the audio signal comprises different channels, and wherein analyzing comprises analyzing the different channels.
8. The method of claim 7 , wherein analyzing the different channels includes comparing correlation content between two of the different channels.
9. The method of claim 1 , wherein the analyzing comprises analyzing a center channel of the audio signal.
10. 2. The method of claim 1, wherein automatically applying the first playback equalization configured to enhance the speech in the content to the audio signal comprises transitioning to the first playback equalization from either i) no playback equalization, or ii) the second playback equalization.
11. 2. The method of claim 1, wherein automatically applying the first playback equalization to the audio signal configured to enhance the speech in the content comprises increasing a volume of the speech in the content relative to other content in the audio signal.
12. 2. The method of claim 1, wherein automatically applying the first playback equalization to the audio signal configured to enhance the speech in the content comprises reducing a volume of non-speech content in the audio signal.
13. The method of claim 12 , further comprising increasing a volume of the speech in the content.
14. The method of claim 1 , wherein the second playback equalization comprises at least one of low frequency enhancement or music playback enhancement.
15. The method of claim 1 , wherein at least one of the first playback equalization or the one or more predefined conditions is user configurable.
16. 10. The method of claim 1, further comprising analyzing sounds in an environment in which the audio signal is to be reproduced to help determine whether to apply the first reproduction equalization to the audio signal.
17. 1. An apparatus for audio signal processing, comprising: Memory and a processor coupled to the memory, wherein the processor and the memory: During playback of an audio signal, analyzing the content of the audio signal prior to the playback of content to determine whether one or more predefined conditions are met to indicate that the content includes speech; In response to determining that the one or more predefined conditions are met, automatically apply a first playback equalization to the audio signal, the first playback equalization being configured to enhance the speech in the content; In response to determining that the one or more predefined conditions are not satisfied, the apparatus is configured to: i) apply no playback equalization, or ii) apply a second playback equalization to the audio signal, the second playback equalization being different from the first playback equalization.
18. 20. The apparatus of claim 17, wherein the processor and the memory are configured to perform detection using a trained machine learning model.
19. 20. The apparatus of claim 17, wherein the audio signal comprises different channels, and the memory and the processor are configured to detect by analyzing the different channels, and analyzing the different channels includes comparing correlation content between two of the different channels.
20. 20. The apparatus of claim 17, wherein the processor and the memory are configured to detect by analyzing a center channel of the audio signal.
21. 18. The apparatus of claim 17, wherein the processor and the memory are configured to automatically apply the first playback equalization configured to enhance the speech in the content to the audio signal by transitioning to the first playback equalization from either i) no playback equalization or ii) the second playback equalization.
22. 18. The apparatus of claim 17, wherein the second playback equalization comprises at least one of low frequency enhancement or music playback enhancement.
23. 20. The apparatus of claim 17, wherein at least one of the first playback equalization or the one or more predefined conditions is user configurable.
24. 20. The apparatus of claim 17, wherein the processor and the memory are further configured to analyze sounds in an environment in which the audio signal is to be reproduced to help determine whether to apply the first reproduction equalization to the audio signal.
25. 1. A non-transitory computer-readable medium storing instructions that, when executed by a device for processing and generating an audio signal, cause the device to: analyzing the content of the audio signal during playback of the audio signal and prior to the playback of content to determine whether one or more predefined conditions are met to indicate that the content includes speech; automatically applying a first playback equalization to the audio signal, the first playback equalization being configured to enhance the speech in the content, in response to determining that the one or more predefined conditions are met; In response to determining that the one or more predefined conditions are not met, the non-transitory computer-readable medium applies either i) no playback equalization or ii) a second playback equalization different from the first playback equalization to the audio signal.
Citation Information
Patent Citations
Karaoke device
JP1997044171A
Sound quality correction apparatus and speech correction method
JP2012063726A
Device and method for audio classification and processing
JP2019194742A
Audio Device with Speech-Based Audio Signal Processing
US20210201926A1